{"id":16665235,"url":"https://github.com/lumina-ai-inc/chunkr","last_synced_at":"2026-01-03T09:07:41.022Z","repository":{"id":254623049,"uuid":"847078421","full_name":"lumina-ai-inc/chunkr","owner":"lumina-ai-inc","description":"Vision model based PDF chunking","archived":false,"fork":false,"pushed_at":"2024-10-25T04:32:20.000Z","size":1465903,"stargazers_count":1133,"open_issues_count":28,"forks_count":53,"subscribers_count":8,"default_branch":"main","last_synced_at":"2024-10-25T04:51:28.175Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"https://chunkr.ai","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"agpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/lumina-ai-inc.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-08-24T19:32:47.000Z","updated_at":"2024-10-25T04:32:24.000Z","dependencies_parsed_at":"2024-10-28T03:56:03.903Z","dependency_job_id":null,"html_url":"https://github.com/lumina-ai-inc/chunkr","commit_stats":null,"previous_names":["lumina-ai-inc/chunk-my-docs"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lumina-ai-inc%2Fchunkr","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lumina-ai-inc%2Fchunkr/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lumina-ai-inc%2Fchunkr/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lumina-ai-inc%2Fchunkr/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/lumina-ai-inc","download_url":"https://codeload.github.com/lumina-ai-inc/chunkr/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":239011374,"owners_count":19567654,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-10-12T11:02:02.846Z","updated_at":"2025-10-30T17:31:23.761Z","avatar_url":"https://github.com/lumina-ai-inc.png","language":"Python","funding_links":[],"categories":["Rust","Python","Parsers, OCR and extraction"],"sub_categories":[],"readme":"\u003cbr /\u003e\n\u003cdiv align=\"center\"\u003e\n  \u003ca href=\"https://github.com/lumina-ai-inc/chunkr\"\u003e\n    \u003cimg src=\"images/logo.svg\" alt=\"Chunkr Logo\" width=\"80\" height=\"80\"\u003e\n  \u003c/a\u003e\n\n\u003ch3 align=\"center\"\u003eChunkr | Open Source Document Intelligence API\u003c/h3\u003e\n\n  \u003cp align=\"center\"\u003e\n    Production-ready service for document layout analysis, OCR, and semantic chunking.\u003cbr /\u003e\n    Convert PDFs, PPTs, Word docs \u0026 images into RAG/LLM-ready chunks.\n    \u003cbr /\u003e\u003cbr /\u003e\n    \u003cb\u003eLayout Analysis\u003c/b\u003e | \u003cb\u003eOCR + Bounding Boxes\u003c/b\u003e | \u003cb\u003eStructured HTML \u0026 Markdown\u003c/b\u003e | \u003cb\u003eVision-Language Model Processing\u003c/b\u003e\n    \u003cbr /\u003e\u003cbr /\u003e\n    👉 \u003cb\u003eNote:\u003c/b\u003e The \u003ca href=\"https://github.com/lumina-ai-inc/chunkr\"\u003eopen-source AGPL version\u003c/a\u003e is **different** from our fully managed \u003ca href=\"https://www.chunkr.ai\"\u003eCloud API\u003c/a\u003e.  \n    The open-source release uses community/open-source models, while the Cloud API runs **proprietary in-house models** for higher accuracy, speed, and enterprise reliability.\n    \u003cbr /\u003e\u003cbr /\u003e\n    \u003ca href=\"https://www.chunkr.ai\"\u003e\u003cimg src=\"https://img.shields.io/badge/Try_it_out-chunkr.ai-blue?style=flat\u0026logo=rocket\u0026height=20\" alt=\"Try it out\" height=\"20\"\u003e\u003c/a\u003e\n    \u0026nbsp;\u0026nbsp;\u0026nbsp;\n    \u003ca href=\"https://github.com/lumina-ai-inc/chunkr/issues/new\"\u003e\u003cimg src=\"https://img.shields.io/badge/Report_Bug-GitHub_Issues-red?style=flat\u0026logo=github\u0026height=20\" alt=\"Report Bug\" height=\"20\"\u003e\u003c/a\u003e\n    \u0026nbsp;\u0026nbsp;\u0026nbsp;\n    \u003ca href=\"#connect-with-us\"\u003e\u003cimg src=\"https://img.shields.io/badge/Contact-Get_in_Touch-green?style=flat\u0026logo=mail\u0026height=20\" alt=\"Contact\" height=\"20\"\u003e\u003c/a\u003e\n    \u0026nbsp;\u0026nbsp;\u0026nbsp;\n    \u003ca href=\"https://discord.gg/XzKWFByKzW\"\u003e\u003cimg src=\"https://img.shields.io/badge/Discord-Join_Community-5865F2?style=flat\u0026logo=discord\u0026logoColor=white\u0026height=20\" alt=\"Discord\" height=\"20\"\u003e\u003c/a\u003e\n    \u0026nbsp;\u0026nbsp;\u0026nbsp;\n    \u003ca href=\"https://deepwiki.com/lumina-ai-inc/chunkr\"\u003e\u003cimg src=\"https://deepwiki.com/badge.svg\" alt=\"Ask DeepWiki\"\u003e\u003c/a\u003e\n  \u003c/p\u003e\n\u003c/div\u003e\n\n\u003cdiv align=\"center\"\u003e\n  \u003ca href=\"https://www.chunkr.ai\" width=\"1200\" height=\"630\"\u003e\n    \u003cimg src=\"https://chunkr.ai/og-image.png\" alt=\"Chunkr Cloud API\"\u003e\n  \u003c/a\u003e\n\u003c/div\u003e\n\n## Table of Contents\n- [Table of Contents](#table-of-contents)\n- [(Super) Quick Start](#super-quick-start)\n- [Documentation](#documentation)\n- [Open Source vs Cloud API vs Enterprise](#open-source-vs-cloud-api-vs-enterprise)\n- [Quick Start with Docker Compose](#quick-start-with-docker-compose)\n- [LLM Configuration](#llm-configuration)\n  - [Using models.yaml (Recommended)](#using-modelsyaml-recommended)\n  - [Using environment variables (Basic)](#using-environment-variables-basic)\n  - [Common LLM API Providers](#common-llm-api-providers)\n- [Licensing](#licensing)\n- [Connect With Us](#connect-with-us)\n\n## Open Source vs Cloud API vs Enterprise\n\n| Feature | Open Source Repo (good) | Cloud API - chunkr.ai (best) | Enterprise |\n|---------|--------------------|------------------------|------------|\n| **Perfect for** | Development \u0026 testing | Production workloads | Large-scale / High-security |\n| **Layout Analysis** | Uses open-source models | Proprietary in-house models | In-house + custom-tuned |\n| **OCR Accuracy** | Community OCR engines | Optimized OCR stack | Optimized + domain-tuned |\n| **VLM Processing** | Basic open VLMs | Enhanced proprietary VLMs | Custom fine-tunes |\n| **Excel Support** | ❌ | ✅ Native parser | ✅ Native parser |\n| **Document Types** | PDF, PPT, Word, Images | PDF, PPT, Word, Images, Excel | PDF, PPT, Word, Images, Excel |\n| **Infrastructure** | Self-hosted | Fully managed cloud | Managed / On-prem |\n| **Support** | Discord community | Dedicated support | Dedicated founding team |\n| **Migration Support** | Community-driven | Docs + email | Dedicated migration team |\n\n---\n\nThe **open-source release** is ideal if you want transparency, local hosting, or to experiment with Chunkr’s pipeline.  \nFor **best performance, production reliability, and access to in-house models**, we recommend the \u003ca href=\"https://www.chunkr.ai\"\u003eChunkr Cloud API\u003c/a\u003e.  \nFor **high-security or regulated industries**, our **Enterprise edition** offers on-prem or VPC deployments.\n\n\n## Quick Start with Docker Compose\n\n1. Prerequisites:\n   - [Docker and Docker Compose](https://docs.docker.com/get-docker/)\n   - [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html) (for GPU support, optional)\n\n2. Clone the repo:\n```bash\ngit clone https://github.com/lumina-ai-inc/chunkr\ncd chunkr\n```\n\n3. Set up environment variables:\n```bash\n# Copy the example environment file\ncp .env.example .env\n\n# Configure your llm models\ncp models.example.yaml models.yaml\n```\n\nFor more information on how to set up LLMs, see [here](#llm-configuration).\n\n4. Start the services:\n```bash\n# For GPU deployment:\ndocker compose up -d\n\n# For CPU-only deployment:\ndocker compose -f compose.yaml -f compose.cpu.yaml up -d\n\n# For Mac ARM architecture (M1, M2, M3, etc.):\ndocker compose -f compose.yaml -f compose.cpu.yaml -f compose.mac.yaml up -d\n```\n\n5. Access the services:\n   - Web UI: `http://localhost:5173`\n   - API: `http://localhost:8000`\n\n6. Stop the services when done:\n```bash\n# For GPU deployment:\ndocker compose down\n\n# For CPU-only deployment:\ndocker compose -f compose.yaml -f compose.cpu.yaml down\n\n# For Mac ARM architecture (M1, M2, M3, etc.):\ndocker compose -f compose.yaml -f compose.cpu.yaml -f compose.mac.yaml down\n```\n## LLM Configuration\n\nChunkr supports two ways to configure LLMs:\n\n1. **models.yaml file**: Advanced configuration for multiple LLMs with additional options\n2. **Environment variables**: Simple configuration for a single LLM\n\n### Using models.yaml (Recommended)\n\nFor more flexible configuration with multiple models, default/fallback options, and rate limits:\n\n1. Copy the example file to create your configuration:\n```bash\ncp models.example.yaml models.yaml\n```\n\n2. Edit the models.yaml file with your configuration. Example:\n```yaml\nmodels:\n  - id: gpt-4o\n    model: gpt-4o\n    provider_url: https://api.openai.com/v1/chat/completions\n    api_key: \"your_openai_api_key_here\"\n    default: true\n    rate-limit: 200 # requests per minute - optional\n```\n\nBenefits of using models.yaml:\n- Configure multiple LLM providers simultaneously\n- Set default and fallback models\n- Add distributed rate limits per model\n- Reference models by ID in API requests (see docs for more info)\n\n\u003eRead the `models.example.yaml` file for more information on the available options.\n\n### Using environment variables (Basic)\n\nYou can use any OpenAI API compatible endpoint by setting the following variables in your .env file:\n``` \nLLM__KEY:\nLLM__MODEL:\nLLM__URL:\n```\n\n### Common LLM API Providers\n\nBelow is a table of common LLM providers and their configuration details to get you started:\n\n| Provider         | API URL                                                                  | Documentation                                                                                                                          |\n| ---------------- | ------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------- |\n| OpenAI           | https://api.openai.com/v1/chat/completions                               | [OpenAI Docs](https://platform.openai.com/docs)                                                                                        |\n| Google AI Studio | https://generativelanguage.googleapis.com/v1beta/openai/chat/completions | [Google AI Docs](https://ai.google.dev/gemini-api/docs/openai)                                                                         |\n| OpenRouter       | https://openrouter.ai/api/v1/chat/completions                            | [OpenRouter Models](https://openrouter.ai/models)                                                                                      |\n| Self-Hosted      | http://localhost:8000/v1                                                 | [VLLM](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html) or [Ollama](https://ollama.com/blog/openai-compatibility) |\n\n## Licensing\n\nThe core of this project is dual-licensed:\n\n1. [GNU Affero General Public License v3.0 (AGPL-3.0)](LICENSE)\n2. Commercial License\n\nTo use Chunkr without complying with the AGPL-3.0 license terms you can [contact us](mailto:mehul@chunkr.ai) or visit our [website](https://chunkr.ai).\n\n## Connect With Us\n- 📧 Email: [mehul@chunkr.ai](mailto:mehul@chunkr.ai)\n- 📅 Schedule a call: [Book a 30-minute meeting](https://cal.com/mehulc/30min)\n- 🌐 Visit our website: [chunkr.ai](https://chunkr.ai)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flumina-ai-inc%2Fchunkr","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Flumina-ai-inc%2Fchunkr","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flumina-ai-inc%2Fchunkr/lists"}