{"id":51561726,"url":"https://github.com/feyninc/chonkie","last_synced_at":"2026-07-14T01:00:48.279Z","repository":{"id":286521830,"uuid":"956908445","full_name":"feyninc/chonkie","owner":"feyninc","description":"🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines","archived":false,"fork":false,"pushed_at":"2026-06-29T13:18:34.000Z","size":25610,"stargazers_count":4320,"open_issues_count":34,"forks_count":296,"subscribers_count":13,"default_branch":"main","last_synced_at":"2026-07-01T15:34:03.072Z","etag":null,"topics":["ai","chonkie","chunker","chunking-algorithm","llms","rag","retrieval-systems","semantic-chunker","similarity-search","splitting-algorithms","text-splitter"],"latest_commit_sha":null,"homepage":"https://docs.chonkie.ai","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/feyninc.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-03-29T05:10:03.000Z","updated_at":"2026-07-01T12:30:58.000Z","dependencies_parsed_at":null,"dependency_job_id":"2e469d5e-646c-4925-865d-1781e91125ee","html_url":"https://github.com/feyninc/chonkie","commit_stats":null,"previous_names":["chonkie-inc/chonkie","feyninc/chonkie"],"tags_count":44,"template":false,"template_full_name":null,"purl":"pkg:github/feyninc/chonkie","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/feyninc%2Fchonkie","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/feyninc%2Fchonkie/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/feyninc%2Fchonkie/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/feyninc%2Fchonkie/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/feyninc","download_url":"https://codeload.github.com/feyninc/chonkie/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/feyninc%2Fchonkie/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35441637,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-13T02:00:06.543Z","response_time":119,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","chonkie","chunker","chunking-algorithm","llms","rag","retrieval-systems","semantic-chunker","similarity-search","splitting-algorithms","text-splitter"],"created_at":"2026-07-10T11:00:39.408Z","updated_at":"2026-07-14T01:00:48.273Z","avatar_url":"https://github.com/feyninc.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"\u003cdiv align='center'\u003e\n\n![Chonkie Logo](https://github.com/chonkie-inc/chonkie/blob/main/assets/chonkie_logo_br_transparent_bg.png?raw=true)\n\n# 🦛 Chonkie ✨\n\n[![PyPI version](https://img.shields.io/pypi/v/chonkie.svg)](https://pypi.org/project/chonkie/)\n[![License](https://img.shields.io/github/license/chonkie-inc/chonkie.svg)](https://github.com/chonkie-inc/chonkie/blob/main/LICENSE)\n[![Documentation](https://img.shields.io/badge/docs-chonkie.ai-blue.svg)](https://docs.chonkie.ai)\n[![Package size](https://img.shields.io/badge/size-505KB-blue)](https://github.com/chonkie-inc/chonkie/blob/main/README.md#installation)\n[![codecov](https://codecov.io/gh/chonkie-inc/chonkie/graph/badge.svg?token=V4EWIJWREZ)](https://codecov.io/gh/chonkie-inc/chonkie)\n[![Downloads](https://static.pepy.tech/badge/chonkie)](https://pepy.tech/project/chonkie)\n[![Discord](https://dcbadge.limes.pink/api/server/https://discord.gg/vH3SkRqmUz?style=flat)](https://discord.gg/vH3SkRqmUz)\n[![GitHub stars](https://img.shields.io/github/stars/chonkie-inc/chonkie.svg)](https://github.com/chonkie-inc/chonkie/stargazers)\n\n_The lightweight ingestion library for fast, efficient and robust RAG pipelines_\n\n[Installation](#📦-installation) •\n[Usage](#🚀-usage) •\n[Chunkers](#✂️-chunkers) •\n[Integrations](#🔌-integrations) •\n[Benchmarks](#📊-benchmarks)\n\n\u003c/div\u003e\n\nTired of making your gazillionth chunker? Sick of the overhead of large libraries? Want to chunk your texts quickly and efficiently? Chonkie the mighty hippo is here to help!\n\n**🚀 Feature-rich**: All the CHONKs you'd ever need \u003c/br\u003e\n**🔄 End-to-end**: Fetch, CHONK, refine, embed and ship straight to your vector DB! \u003c/br\u003e\n**✨ Easy to use**: Install, Import, CHONK \u003c/br\u003e\n**⚡ Fast**: CHONK at the speed of light! zooooom \u003c/br\u003e\n**🪶 Light-weight**: No bloat, just CHONK \u003c/br\u003e\n**🔌 32+ [integrations](#integrations)**: Works with your favorite tools and vector DBs out of the box! \u003c/br\u003e\n**💬 ️Multilingual**: Out-of-the-box support for 56 languages \u003c/br\u003e\n**☁️ Cloud-Friendly**: CHONK locally or in the [Cloud](https://labs.chonkie.ai) \u003c/br\u003e\n**🦛 Cute CHONK mascot**: psst it's a pygmy hippo btw \u003c/br\u003e\n**❤️ [Moto Moto](#acknowledgements)'s favorite python library** \u003c/br\u003e\n\n**Chonkie** is a chunking library that \"**just works**\" ✨\n\n## 📦 Installation\n\n### Basic Installation\n\nUsing pip:\n\n```bash\npip install chonkie\n```\n\nOr using [uv](https://docs.astral.sh/uv/) (faster):\n\n```bash\nuv pip install chonkie\n```\n\n### Full Installation\n\nChonkie follows the rule of minimum installs.\nHave a favorite chunker? Read our [docs](https://docs.chonkie.ai) to install only what you need.\nDon't want to think about it? Simply install `all` (Not recommended for production environments).\n\nUsing pip:\n\n```bash\npip install \"chonkie[all]\"\n```\n\nOr using [uv](https://docs.astral.sh/uv/):\n\n```bash\nuv pip install \"chonkie[all]\"\n```\n\n## 🚀 Usage\n\n### Basic Usage\n\nHere's a basic example to get you started:\n\n```python\n# First import the chunker you want from Chonkie\nfrom chonkie import RecursiveChunker\n\n# Initialize the chunker\nchunker = RecursiveChunker()\n\n# Chunk some text\nchunks = chunker(\"Chonkie is the goodest boi! My favorite chunking hippo hehe.\")\n\n# Access chunks\nfor chunk in chunks:\n    print(f\"Chunk: {chunk.text}\")\n    print(f\"Tokens: {chunk.token_count}\")\n```\n\n### Pipeline Usage\n\nYou can also use the `chonkie.Pipeline` to chain components together and handle complex workflows. Read more about pipelines in the [docs](https://docs.chonkie.ai/oss/pipelines)!\n\n```python\nfrom chonkie import Pipeline\n\n# Create a pipeline with multiple chunking and refinement steps\npipe = (\n    Pipeline()\n    .chunk_with(\"recursive\", tokenizer=\"gpt2\", chunk_size=2048, recipe=\"markdown\")\n    .chunk_with(\"semantic\", chunk_size=512)\n    .refine_with(\"overlap\", context_size=128)\n    .refine_with(\"embeddings\", embedding_model=\"sentence-transformers/all-MiniLM-L6-v2\")\n)\n\n# CHONK some Texts!\ndoc = pipe.run(texts=\"Chonkie is the goodest boi! My favorite chunking hippo hehe.\")\n\n# Access the processed chunks in the `doc` object\nfor chunk in doc.chunks:\n    print(chunk.text)\n\n# Run asynchronously for high-throughput applications\nimport asyncio\n\nasync def main():\n    doc = await pipe.arun(texts=\"Chonkie runs fast!\")\n    print(len(doc.chunks))\n\nasyncio.run(main())\n```\n\nCheck out more usage examples in the [docs](https://docs.chonkie.ai)!\n\n## 🌐 API Server\n\nRun Chonkie as a self-hosted REST API for easy integration into any application:\n\n```bash\n# Install with API dependencies (includes catsu for multi-provider embeddings)\npip install \"chonkie[api,semantic,code,catsu]\"\n\n# Start the server using the CLI\nchonkie serve\n\n# Or with custom options\nchonkie serve --port 3000 --reload --log-level debug\n\n# Or directly with uvicorn\nuvicorn chonkie.api.main:app --host 0.0.0.0 --port 8000\n```\n\nOr use Docker:\n\n```bash\ndocker compose up\n```\n\nThe API provides endpoints for all chunkers, refineries, and **pipelines** — reusable workflow configurations stored in a local SQLite database.\n\n```bash\n# Create a reusable pipeline\ncurl -X POST http://localhost:8000/v1/pipelines \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"name\": \"rag-chunker\",\n    \"steps\": [\n      {\"type\": \"chunk\", \"chunker\": \"semantic\", \"config\": {\"chunk_size\": 512}},\n      {\"type\": \"refine\", \"refinery\": \"embeddings\", \"config\": {\"embedding_model\": \"text-embedding-3-small\"}}\n    ]\n  }'\n\n# List your pipelines\ncurl http://localhost:8000/v1/pipelines\n```\n\nInteractive documentation is available at `/docs` when the server is running.\n\n## ✂️ Chunkers\n\nChonkie provides several chunkers to help you split your text efficiently for RAG applications. Here's a quick overview of the available chunkers:\n\n| Name               | Alias       | Description                                                                                                                |\n| ------------------ | ----------- | -------------------------------------------------------------------------------------------------------------------------- |\n| `TokenChunker`     | `token`     | Splits text into fixed-size token chunks.                                                                                  |\n| `FastChunker`      | `fast`      | SIMD-accelerated byte-based chunking at 100+ GB/s. Included in the default install.                                        |\n| `SentenceChunker`  | `sentence`  | Splits text into chunks based on sentences.                                                                                |\n| `RecursiveChunker` | `recursive` | Splits text hierarchically using customizable rules to create semantically meaningful chunks.                              |\n| `SemanticChunker`  | `semantic`  | Splits text into chunks based on semantic similarity. Inspired by the work of [Greg Kamradt](https://github.com/gkamradt). |\n| `LateChunker`      | `late`      | Embeds text and then splits it to have better chunk embeddings.                                                            |\n| `CodeChunker`      | `code`      | Splits code into structurally meaningful chunks.                                                                           |\n| `NeuralChunker`    | `neural`    | Splits text using a neural model.                                                                                          |\n| `SlumberChunker`   | `slumber`   | Splits text using an LLM to find semantically meaningful chunks. Also known as _\"AgenticChunker\"_.                         |\n| `TableChunker`     | `table`     | Chunks markdown tables by rows or character count.                                                                          |\n| `TeraflopAIChunker`| `teraflopai`| Splits text using the TeraflopAI Segmentation API for domain-specific segmentation.                                        |\n\nMore on these methods and the approaches taken inside the [docs](https://docs.chonkie.ai)\n\n## 🔌 Integrations\n\nChonkie boasts 45+ integrations across tokenizers, embedding providers, LLMs, refineries, porters, vector databases, and utilities, ensuring it fits seamlessly into your existing workflow.\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e👨‍🍳 Chefs \u0026 📁 Fetchers! Text preprocessing and data loading!\u003c/strong\u003e\u003c/summary\u003e\n\nChefs handle text preprocessing, while Fetchers load data from various sources.\n\n| Component | Class          | Description                                        | Optional Install  |\n| --------- | -------------- | -------------------------------------------------- | ----------------- |\n| `chef`    | `TextChef`     | Text preprocessing and cleaning.                   | `default`         |\n| `chef`    | `MarkdownChef` | Parse markdown into structured MarkdownDocuments.   | `default`         |\n| `chef`    | `TableChef`    | Process CSV/Excel files into MarkdownDocuments.     | `chonkie[table]`  |\n| `chef`    | `MistralOCR`   | Extract text from images/PDFs via Mistral OCR API. | `chonkie[mistral]` |\n| `fetcher` | `FileFetcher`  | Load text from files and directories.              | `default`         |\n\n\u003c/details\u003e\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🏭 Refine your CHONKs with Context and Embeddings! Chonkie supports 2+ refineries!\u003c/strong\u003e\u003c/summary\u003e\n\nRefineries help you post-process and enhance your chunks after initial chunking.\n\n| Refinery Name | Class                | Description                                   | Optional Install    |\n| ------------- | -------------------- | --------------------------------------------- | ------------------- |\n| `overlap`     | `OverlapRefinery`    | Merge overlapping chunks based on similarity. | `default`           |\n| `embeddings`  | `EmbeddingsRefinery` | Add embeddings to chunks using any provider.  | `chonkie[semantic]` |\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🐴 Exporting CHONKs! Chonkie supports 2+ Porters!\u003c/strong\u003e\u003c/summary\u003e\n\nPorters help you save your chunks easily.\n\n| Porter Name | Class            | Description                            | Optional Install    |\n| ----------- | ---------------- | -------------------------------------- | ------------------- |\n| `json`      | `JSONPorter`     | Export chunks to a JSON file.          | `default`           |\n| `datasets`  | `DatasetsPorter` | Export chunks to HuggingFace datasets. | `chonkie[datasets]` |\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🤝 Shake hands with your DB! Chonkie connects with 10+ vector stores!\u003c/strong\u003e\u003c/summary\u003e\n\nHandshakes provide a unified interface to ingest chunks directly into your favorite vector databases.\n\n| Handshake Name | Class                  | Description                                  | Optional Install    |\n| -------------- | ---------------------- | -------------------------------------------- | ------------------- |\n| `chroma`       | `ChromaHandshake`      | Ingest chunks into ChromaDB.                 | `chonkie[chroma]`   |\n| `elastic`      | `ElasticHandshake`     | Ingest chunks into Elasticsearch.            | `chonkie[elastic]`  |\n| `mongodb`      | `MongoDBHandshake`     | Ingest chunks into MongoDB.                  | `chonkie[mongodb]`  |\n| `pgvector`     | `PgvectorHandshake`    | Ingest chunks into PostgreSQL with pgvector. | `chonkie[pgvector]` |\n| `pinecone`     | `PineconeHandshake`    | Ingest chunks into Pinecone.                 | `chonkie[pinecone]` |\n| `qdrant`       | `QdrantHandshake`      | Ingest chunks into Qdrant.                   | `chonkie[qdrant]`   |\n| `turbopuffer`  | `TurbopufferHandshake` | Ingest chunks into Turbopuffer.              | `chonkie[tpuf]`     |\n| `weaviate`     | `WeaviateHandshake`    | Ingest chunks into Weaviate.                 | `chonkie[weaviate]` |\n| `lancedb`      | `LanceDBHandshake`     | Ingest chunks into LanceDB.                  | `chonkie[lancedb]`  |\n| `milvus`       | `MilvusHandshake`      | Ingest chunks into Milvus.                   | `chonkie[milvus]`   |\n\n\u003c/details\u003e\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🪓 Slice 'n' Dice! Chonkie supports 5+ ways to tokenize! \u003c/strong\u003e\u003c/summary\u003e\n\nChoose from supported tokenizers or provide your own custom token counting function. Flexibility first!\n\n| Name           | Description                                                    | Optional Install      |\n| -------------- | -------------------------------------------------------------- | --------------------- |\n| `character`    | Basic character-level tokenizer. **Default tokenizer.**        | `default`             |\n| `word`         | Basic word-level tokenizer.                                    | `default`             |\n| `byte`         | Byte-level tokenizer operating on UTF-8 encoded bytes.         | `default`             |\n| `tokenizers`   | Load any tokenizer from the Hugging Face `tokenizers` library. | `chonkie[tokenizers]` |\n| `tiktoken`     | Use OpenAI's `tiktoken` library (e.g., for `gpt-4`).           | `chonkie[tiktoken]`   |\n| `transformers` | Load tokenizers via `AutoTokenizer` from HF `transformers`.    | `chonkie[neural]`     |\n\n`default` indicates that the feature is available with the default `pip install chonkie`.\n\nTo use a custom token counter, you can pass in any function that takes a string and returns an integer! Something like this:\n\n```python\ndef custom_token_counter(text: str) -\u003e int:\n    return len(text)\n\nchunker = RecursiveChunker(tokenizer=custom_token_counter)\n```\n\nYou can use this to extend Chonkie to support any tokenization scheme you want!\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🧠 Embed like a boss! Chonkie links up with 16+ embedding pals!\u003c/strong\u003e\u003c/summary\u003e\n\nSeamlessly works with various embedding model providers. Bring your favorite embeddings to the CHONK party! Use `AutoEmbeddings` to load models easily.\n\n| Provider / Alias        | Class                           | Description                            | Optional Install        |\n| ----------------------- | ------------------------------- | -------------------------------------- | ----------------------- |\n| `model2vec`             | `Model2VecEmbeddings`           | Use `Model2Vec` models.                | `chonkie[model2vec]`    |\n| `sentence-transformers` | `SentenceTransformerEmbeddings` | Use any `sentence-transformers` model. | `chonkie[st]`           |\n| `openai`                | `OpenAIEmbeddings`              | Use OpenAI's embedding API.            | `chonkie[openai]`       |\n| `azure-openai`          | `AzureOpenAIEmbeddings`         | Use Azure OpenAI embedding service.    | `chonkie[azure-openai]` |\n| `cohere`                | `CohereEmbeddings`              | Use Cohere's embedding API.            | `chonkie[cohere]`       |\n| `gemini`                | `GeminiEmbeddings`              | Use Google's Gemini embedding API.     | `chonkie[gemini]`       |\n| `jina`                  | `JinaEmbeddings`                | Use Jina AI's embedding API.           | `chonkie[jina]`         |\n| `voyageai`              | `VoyageAIEmbeddings`            | Use Voyage AI's embedding API.         | `chonkie[voyageai]`     |\n| `litellm`               | `LiteLLMEmbeddings`             | Use LiteLLM for 100+ embedding models. | `chonkie[litellm]`      |\n| `catsu`                 | `CatsuEmbeddings`               | Unified adapter for 11+ providers.     | `chonkie[catsu]`        |\n| `mistral`               | `MistralEmbeddings`             | Use Mistral's embedding API.           | `chonkie[catsu]`        |\n| `together`              | `TogetherEmbeddings`            | Use Together AI's embedding API.       | `chonkie[catsu]`        |\n| `mixedbread`            | `MixedbreadEmbeddings`          | Use Mixedbread's embedding API.        | `chonkie[catsu]`        |\n| `nomic`                 | `NomicEmbeddings`               | Use Nomic's embedding API.             | `chonkie[catsu]`        |\n| `deepinfra`             | `DeepInfraEmbeddings`           | Use DeepInfra's embedding API.         | `chonkie[catsu]`        |\n| `cloudflare`            | `CloudflareEmbeddings`          | Use Cloudflare Workers AI embeddings.  | `chonkie[catsu]`        |\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🧞‍♂️ Power Up with Genies! Chonkie supports 5+ LLM providers!\u003c/strong\u003e\u003c/summary\u003e\n\nGenies provide interfaces to interact with Large Language Models (LLMs) for advanced chunking strategies or other tasks within the pipeline.\n\n| Genie Name     | Class              | Description                                | Optional Install        |\n| -------------- | ------------------ | ------------------------------------------ | ----------------------- |\n| `gemini`       | `GeminiGenie`      | Interact with Google Gemini APIs.          | `chonkie[gemini]`       |\n| `openai`       | `OpenAIGenie`      | Interact with OpenAI APIs.                 | `chonkie[openai]`       |\n| `azure-openai` | `AzureOpenAIGenie` | Interact with Azure OpenAI APIs.           | `chonkie[azure-openai]` |\n| `groq`         | `GroqGenie`        | Fast inference on Groq hardware.           | `chonkie[groq]`         |\n| `cerebras`     | `CerebrasGenie`    | Fastest inference on Cerebras hardware.    | `chonkie[cerebras]`     |\n\nYou can also use the `OpenAIGenie` to interact with any LLM provider that supports the OpenAI API format, by simply changing the `model`, `base_url`, and `api_key` parameters. For example, here's how to use the `OpenAIGenie` to interact with the `Llama-4-Maverick` model via OpenRouter:\n\n```python\nfrom chonkie import OpenAIGenie\n\ngenie = OpenAIGenie(model=\"meta-llama/llama-4-maverick\",\n                    base_url=\"https://openrouter.ai/api/v1\",\n                    api_key=\"your_api_key\")\n```\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cstrong\u003e🛠️ Utilities \u0026 Helpers! Chonkie includes handy tools!\u003c/strong\u003e\u003c/summary\u003e\n\nAdditional utilities to enhance your chunking workflow.\n\n| Utility Name | Class        | Description                                    | Optional Install |\n| ------------ | ------------ | ---------------------------------------------- | ---------------- |\n| `hub`        | `Hubbie`     | Simple wrapper for HuggingFace Hub operations. | `chonkie[hub]`   |\n| `viz`        | `Visualizer` | Rich console visualizations for chunks.        | `chonkie[viz]`   |\n\n\u003c/details\u003e\n\nWith Chonkie's wide range of integrations, you can easily plug it into your existing infrastructure and start CHONKING!\n\n## 🤖 AI Agent Skills \u0026 Plugins\n\nChonkie provides an official skill and plugin for AI coding agents, giving them deep knowledge of Chonkie's API, chunking strategies, and pipeline patterns — so they can help you build RAG pipelines faster.\n\n**Supported agents:** Claude Code, Cursor, Gemini CLI, and more.\n\n```bash\n# Via skills.sh (works with Claude Code, Cursor, Copilot, and 20+ agents)\nnpx skills add chonkie-inc/skills\n\n# Claude Code only\n/plugin marketplace add chonkie-inc/skills\n```\n\nOnce installed, your agent gains knowledge of all chunkers, the Pipeline API, tokenizer selection, embeddings refineries, vector DB handshakes, the REST API server, recipes, and async/batch processing patterns.\n\nLearn more at [github.com/chonkie-inc/skills](https://github.com/chonkie-inc/skills).\n\n## 📊 Benchmarks\n\n\u003e \"I may be smol hippo, but I pack a big punch!\" 🦛\n\nChonkie is not just cute, it's also fast and efficient! Here's how it stacks up against the competition:\n\n**Size**📦\n\n- **Wheel Size:** 505KB (vs 1-12MB for alternatives)\n- **Installed Size:** 49MB (vs 80-171MB for alternatives)\n- **With Semantic:** Still 10x lighter than the closest competition!\n\n**Speed**⚡\n\n- **Token Chunking:** 33x faster than the slowest alternative\n- **Sentence Chunking:** Almost 2x faster than competitors\n- **Semantic Chunking:** Up to 2.5x faster than others\n\nCheck out our detailed [benchmarks](BENCHMARKS.md) to see how Chonkie races past the competition! 🏃‍♂️💨\n\n## 🤝 Contributing\n\nWant to help grow Chonkie? Check out [CONTRIBUTING.md](CONTRIBUTING.md) to get started! Whether you're fixing bugs, adding features, or improving docs, every contribution helps make Chonkie a better CHONK for everyone.\n\nRemember: No contribution is too small for this tiny hippo! 🦛\n\n## 🙏 Acknowledgements\n\nChonkie would like to CHONK its way through a special thanks to all the users and contributors who have helped make this library what it is today! Your feedback, issue reports, and improvements have helped make Chonkie the CHONKIEST it can be.\n\nAnd of course, special thanks to [Moto Moto](https://www.youtube.com/watch?v=I0zZC4wtqDQ\u0026t=5s) for endorsing Chonkie with his famous quote:\n\n\u003e \"I like them big, I like them chonkie.\" ~ Moto Moto\n\n## 📝 Citation\n\nIf you use Chonkie in your research, please cite it as follows:\n\n```bibtex\n@software{chonkie2025,\n  author = {Minhas, Bhavnick AND Nigam, Shreyash},\n  title = {Chonkie: The lightweight ingestion library for fast, efficient and robust RAG pipelines},\n  year = {2025},\n  publisher = {GitHub},\n  howpublished = {\\url{https://github.com/chonkie-inc/chonkie}},\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffeyninc%2Fchonkie","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ffeyninc%2Fchonkie","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffeyninc%2Fchonkie/lists"}