{"id":50584860,"url":"https://github.com/kreuzberg-dev/kreuzberg-crewai","last_synced_at":"2026-06-05T05:02:53.506Z","repository":{"id":352459217,"uuid":"1187818568","full_name":"kreuzberg-dev/kreuzberg-crewai","owner":"kreuzberg-dev","description":"Extract text and metadata from 88+ document formats — PDF, DOCX, XLSX, HTML, images with OCR, and more — directly from your CrewAI agents.","archived":false,"fork":false,"pushed_at":"2026-05-20T16:29:58.000Z","size":55,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-05-20T21:52:25.537Z","etag":null,"topics":["agents","ai","crewai","document-intelligence","document-processing","kreuzberg","llm","python","rag"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kreuzberg-dev.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":".github/CODEOWNERS","security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-03-21T07:55:44.000Z","updated_at":"2026-05-20T16:31:28.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/kreuzberg-dev/kreuzberg-crewai","commit_stats":null,"previous_names":["kreuzberg-dev/kreuzberg-crewai"],"tags_count":1,"template":false,"template_full_name":null,"purl":"pkg:github/kreuzberg-dev/kreuzberg-crewai","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kreuzberg-dev%2Fkreuzberg-crewai","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kreuzberg-dev%2Fkreuzberg-crewai/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kreuzberg-dev%2Fkreuzberg-crewai/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kreuzberg-dev%2Fkreuzberg-crewai/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kreuzberg-dev","download_url":"https://codeload.github.com/kreuzberg-dev/kreuzberg-crewai/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kreuzberg-dev%2Fkreuzberg-crewai/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33930311,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-05T02:00:06.157Z","response_time":120,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["agents","ai","crewai","document-intelligence","document-processing","kreuzberg","llm","python","rag"],"created_at":"2026-06-05T05:02:52.688Z","updated_at":"2026-06-05T05:02:53.500Z","avatar_url":"https://github.com/kreuzberg-dev.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# kreuzberg-crewai\n\n\u003cdiv align=\"center\" style=\"display: flex; flex-wrap: wrap; gap: 8px; justify-content: center; margin: 20px 0;\"\u003e\n  \u003ca href=\"https://pypi.org/project/kreuzberg-crewai/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/v/kreuzberg-crewai?label=kreuzberg-crewai\u0026color=007ec6\" alt=\"PyPI version\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://pypi.org/project/kreuzberg-crewai/\"\u003e\u003cimg src=\"https://img.shields.io/pypi/pyversions/kreuzberg-crewai?color=007ec6\" alt=\"Python versions\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/kreuzberg-dev/kreuzberg-crewai/blob/main/LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/badge/License-MIT-blue.svg\" alt=\"License\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://docs.kreuzberg.dev\"\u003e\u003cimg src=\"https://img.shields.io/badge/docs-kreuzberg.dev-blue\" alt=\"Docs\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://github.com/kreuzberg-dev/kreuzberg-crewai/actions/workflows/ci.yaml\"\u003e\u003cimg src=\"https://github.com/kreuzberg-dev/kreuzberg-crewai/actions/workflows/ci.yaml/badge.svg\" alt=\"CI\"\u003e\u003c/a\u003e\n\u003c/div\u003e\n\n\u003cimg width=\"3384\" height=\"573\" alt=\"Kreuzberg Banner\" src=\"https://github.com/user-attachments/assets/1b6c6ad7-3b6d-4171-b1c9-f2026cc9deb8\" /\u003e\n\n\u003cdiv align=\"center\" style=\"margin-top: 20px;\"\u003e\n  \u003ca href=\"https://discord.gg/xt9WY3GnKR\"\u003e\n    \u003cimg height=\"22\" src=\"https://img.shields.io/badge/Discord-Join%20our%20community-7289da?logo=discord\u0026logoColor=white\" alt=\"Discord\"\u003e\n  \u003c/a\u003e\n\u003c/div\u003e\n\n[Kreuzberg](https://github.com/kreuzberg-dev/kreuzberg) document extraction tools for [CrewAI](https://www.crewai.com/) agents.\n\nExtract text and metadata from 90+ document formats — PDF, DOCX, XLSX, HTML, images with OCR, and more — directly from your CrewAI agents.\n\n## Installation\n\n```bash\npip install kreuzberg-crewai\n```\n\n## Quick Start\n\n```python\nfrom crewai import Agent, Crew, Task\n\nfrom kreuzberg_crewai import KreuzbergExtractTool\n\ntool = KreuzbergExtractTool()\n\nagent = Agent(\n    role=\"Document Analyst\",\n    goal=\"Extract and analyze document content\",\n    backstory=\"You are an expert at reading and understanding documents.\",\n    tools=[tool],\n)\n\ntask = Task(\n    description=\"Extract the content from report.pdf and summarize the key findings.\",\n    expected_output=\"A summary of the key findings in the report.\",\n    agent=agent,\n)\n\ncrew = Crew(agents=[agent], tasks=[task])\nresult = crew.kickoff()\n```\n\n## Tools\n\n### KreuzbergExtractTool\n\nExtracts text content from a document file.\n\n**Parameters:**\n\n| Parameter | Type | Default | Description |\n|---|---|---|---|\n| `file_path` | `str` | required | Path to the document file |\n| `output_format` | `\"plain\" \\| \"markdown\" \\| \"html\"` | `\"markdown\"` | Output format |\n\n```python\nfrom kreuzberg_crewai import KreuzbergExtractTool\n\ntool = KreuzbergExtractTool()\n\n# The agent calls this automatically, but you can also call it directly:\ncontent = tool._run(file_path=\"report.pdf\", output_format=\"markdown\")\n```\n\n### KreuzbergExtractMetadataTool\n\nExtracts metadata (title, authors, dates, page count, format-specific details) from a document file.\n\n**Parameters:**\n\n| Parameter | Type | Default | Description |\n|---|---|---|---|\n| `file_path` | `str` | required | Path to the document file |\n\n```python\nfrom kreuzberg_crewai import KreuzbergExtractMetadataTool\n\ntool = KreuzbergExtractMetadataTool()\n\nmetadata = tool._run(file_path=\"report.pdf\")\n# title: Annual Report 2025\n# authors: ['John Doe']\n# page_count: 42\n# pdf_version: 1.7\n```\n\n## Agent Example\n\nUsing both tools together:\n\n```python\nfrom crewai import Agent, Crew, Task\n\nfrom kreuzberg_crewai import KreuzbergExtractMetadataTool, KreuzbergExtractTool\n\nextract_tool = KreuzbergExtractTool()\nmetadata_tool = KreuzbergExtractMetadataTool()\n\nagent = Agent(\n    role=\"Research Assistant\",\n    goal=\"Read documents and extract useful information\",\n    backstory=\"You help researchers by reading and analyzing documents.\",\n    tools=[extract_tool, metadata_tool],\n)\n\ntask = Task(\n    description=(\n        \"First, check the metadata of research-paper.pdf to find the authors and date. \"\n        \"Then extract the full content in markdown format and list the key conclusions.\"\n    ),\n    expected_output=\"Authors, date, and key conclusions from the paper.\",\n    agent=agent,\n)\n\ncrew = Crew(agents=[agent], tasks=[task])\nresult = crew.kickoff()\n```\n\n## Supported Formats\n\nKreuzberg supports 90+ file formats:\n\n- **Documents:** PDF, DOCX, DOC, XLSX, XLS, PPTX, PPT, ODT, ODS, ODP, RTF, and more\n- **Text/Markup:** TXT, MD, HTML, XML, JSON, YAML, LaTeX, Jupyter notebooks\n- **Images (OCR):** PNG, JPEG, TIFF, GIF, BMP, WEBP, SVG\n- **Email:** EML, MSG (with attachment extraction)\n- **eBooks:** EPUB\n- **Archives:** ZIP, RAR, 7Z, TAR, GZIP\n- **Data:** CSV, DBF\n\n## Development\n\n```bash\n# Install dependencies\nuv sync\n\n# Run tests\nuv run pytest\n\n# Run linting\nuv run ruff check src/ tests/\nuv run ruff format --check src/ tests/\n\n# Run type checking\nuv run mypy src/\n```\n\n## License\n\nMIT\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkreuzberg-dev%2Fkreuzberg-crewai","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkreuzberg-dev%2Fkreuzberg-crewai","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkreuzberg-dev%2Fkreuzberg-crewai/lists"}