An open API service indexing awesome lists of open source software.

https://github.com/kreuzberg-dev/kreuzberg-crewai

Extract text and metadata from 88+ document formats — PDF, DOCX, XLSX, HTML, images with OCR, and more — directly from your CrewAI agents.
https://github.com/kreuzberg-dev/kreuzberg-crewai

agents ai crewai document-intelligence document-processing kreuzberg llm python rag

Last synced: about 2 months ago
JSON representation

Extract text and metadata from 88+ document formats — PDF, DOCX, XLSX, HTML, images with OCR, and more — directly from your CrewAI agents.

Awesome Lists containing this project

README

          

# kreuzberg-crewai


PyPI version
Python versions
License
Docs
CI

Kreuzberg Banner



Discord

[Kreuzberg](https://github.com/kreuzberg-dev/kreuzberg) document extraction tools for [CrewAI](https://www.crewai.com/) agents.

Extract text and metadata from 90+ document formats — PDF, DOCX, XLSX, HTML, images with OCR, and more — directly from your CrewAI agents.

## Installation

```bash
pip install kreuzberg-crewai
```

## Quick Start

```python
from crewai import Agent, Crew, Task

from kreuzberg_crewai import KreuzbergExtractTool

tool = KreuzbergExtractTool()

agent = Agent(
role="Document Analyst",
goal="Extract and analyze document content",
backstory="You are an expert at reading and understanding documents.",
tools=[tool],
)

task = Task(
description="Extract the content from report.pdf and summarize the key findings.",
expected_output="A summary of the key findings in the report.",
agent=agent,
)

crew = Crew(agents=[agent], tasks=[task])
result = crew.kickoff()
```

## Tools

### KreuzbergExtractTool

Extracts text content from a document file.

**Parameters:**

| Parameter | Type | Default | Description |
|---|---|---|---|
| `file_path` | `str` | required | Path to the document file |
| `output_format` | `"plain" \| "markdown" \| "html"` | `"markdown"` | Output format |

```python
from kreuzberg_crewai import KreuzbergExtractTool

tool = KreuzbergExtractTool()

# The agent calls this automatically, but you can also call it directly:
content = tool._run(file_path="report.pdf", output_format="markdown")
```

### KreuzbergExtractMetadataTool

Extracts metadata (title, authors, dates, page count, format-specific details) from a document file.

**Parameters:**

| Parameter | Type | Default | Description |
|---|---|---|---|
| `file_path` | `str` | required | Path to the document file |

```python
from kreuzberg_crewai import KreuzbergExtractMetadataTool

tool = KreuzbergExtractMetadataTool()

metadata = tool._run(file_path="report.pdf")
# title: Annual Report 2025
# authors: ['John Doe']
# page_count: 42
# pdf_version: 1.7
```

## Agent Example

Using both tools together:

```python
from crewai import Agent, Crew, Task

from kreuzberg_crewai import KreuzbergExtractMetadataTool, KreuzbergExtractTool

extract_tool = KreuzbergExtractTool()
metadata_tool = KreuzbergExtractMetadataTool()

agent = Agent(
role="Research Assistant",
goal="Read documents and extract useful information",
backstory="You help researchers by reading and analyzing documents.",
tools=[extract_tool, metadata_tool],
)

task = Task(
description=(
"First, check the metadata of research-paper.pdf to find the authors and date. "
"Then extract the full content in markdown format and list the key conclusions."
),
expected_output="Authors, date, and key conclusions from the paper.",
agent=agent,
)

crew = Crew(agents=[agent], tasks=[task])
result = crew.kickoff()
```

## Supported Formats

Kreuzberg supports 90+ file formats:

- **Documents:** PDF, DOCX, DOC, XLSX, XLS, PPTX, PPT, ODT, ODS, ODP, RTF, and more
- **Text/Markup:** TXT, MD, HTML, XML, JSON, YAML, LaTeX, Jupyter notebooks
- **Images (OCR):** PNG, JPEG, TIFF, GIF, BMP, WEBP, SVG
- **Email:** EML, MSG (with attachment extraction)
- **eBooks:** EPUB
- **Archives:** ZIP, RAR, 7Z, TAR, GZIP
- **Data:** CSV, DBF

## Development

```bash
# Install dependencies
uv sync

# Run tests
uv run pytest

# Run linting
uv run ruff check src/ tests/
uv run ruff format --check src/ tests/

# Run type checking
uv run mypy src/
```

## License

MIT