{"id":22836527,"url":"https://github.com/NVIDIA/nv-ingest","last_synced_at":"2025-08-10T21:32:22.832Z","repository":{"id":254923530,"uuid":"846145923","full_name":"NVIDIA/nv-ingest","owner":"NVIDIA","description":"NVIDIA Ingest is an early access set of microservices for parsing hundreds of thousands of complex, messy unstructured PDFs and other enterprise documents into metadata and text to embed into retrieval systems.","archived":false,"fork":false,"pushed_at":"2024-12-12T20:49:21.000Z","size":4321,"stargazers_count":206,"open_issues_count":55,"forks_count":55,"subscribers_count":12,"default_branch":"main","last_synced_at":"2024-12-12T21:30:43.751Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/NVIDIA.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":"CITATION.md","codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-08-22T16:09:05.000Z","updated_at":"2024-12-12T18:41:35.000Z","dependencies_parsed_at":"2024-10-25T20:43:04.576Z","dependency_job_id":"0c6c07ca-5076-4ea2-84fc-4d39275ead90","html_url":"https://github.com/NVIDIA/nv-ingest","commit_stats":null,"previous_names":["nvidia/nv-ingest"],"tags_count":2,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NVIDIA%2Fnv-ingest","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NVIDIA%2Fnv-ingest/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NVIDIA%2Fnv-ingest/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NVIDIA%2Fnv-ingest/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/NVIDIA","download_url":"https://codeload.github.com/NVIDIA/nv-ingest/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":229464314,"owners_count":18077035,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-12-12T23:02:12.338Z","updated_at":"2025-08-10T21:32:22.807Z","avatar_url":"https://github.com/NVIDIA.png","language":"Python","funding_links":[],"categories":["Python","Repos","🔥LLM Extraction / Parsing"],"sub_categories":[],"readme":"\u003c!--\nSPDX-FileCopyrightText: Copyright (c) 2024, NVIDIA CORPORATION \u0026 AFFILIATES.\nAll rights reserved.\nSPDX-License-Identifier: Apache-2.0\n--\u003e\n\n# What is NeMo Retriever Extraction?\n\nNeMo Retriever extraction is a scalable, performance-oriented document content and metadata extraction microservice. \nNeMo Retriever extraction uses specialized NVIDIA NIM microservices \nto find, contextualize, and extract text, tables, charts and images that you can use in downstream generative applications.\n\n\u003e [!Note]\n\u003e NeMo Retriever extraction is also known as NVIDIA Ingest and nv-ingest.\n\nNeMo Retriever extraction enables parallelization of splitting documents into pages where artifacts are classified (such as text, tables, charts, and images), extracted, and further contextualized through optical character recognition (OCR) into a well defined JSON schema. \nFrom there, NeMo Retriever extraction can optionally manage computation of embeddings for the extracted content, \nand optionally manage storing into a vector database [Milvus](https://milvus.io/).\n\n\u003e [!Note]\n\u003e Cached and Deplot are deprecated. Instead, NeMo Retriever extraction now uses the yolox-graphic-elements NIM. With this change, you should now be able to run NeMo Retriever Extraction on a single 24GB A10G or better GPU. If you want to use the old pipeline, with Cached and Deplot, use the [NeMo Retriever Extraction 24.12.1 release](https://github.com/NVIDIA/nv-ingest/tree/24.12.1).\n\n\nThe following diagram shows the Nemo Retriever extraction pipeline.\n\n![Pipeline Overview](https://docs.nvidia.com/nemo/retriever/extraction/images/overview-extraction.png)\n\n## Table of Contents\n1. [What NeMo Retriever Extraction Is](#what-nvidia-ingest-is)\n2. [Prerequisites](#prerequisites)\n3. [Quickstart](#library-mode-quickstart)\n4. [GitHub Repository Structure](#nv-ingest-repository-structure)\n5. [Notices](#notices)\n\n\n## What NeMo Retriever Extraction Is\n\nNeMo Retriever Extraction is a library and microservice service that does the following:\n\n- Accept a job specification that contains a document payload and a set of ingestion tasks to perform on that payload.\n- Store the result of each job to retrieve later. The result is a dictionary that contains a list of metadata that describes the objects extracted from the base document, and processing annotations and timing/trace data.\n- Support multiple methods of extraction for each document type to balance trade-offs between throughput and accuracy. For example, for .pdf documents, extraction is performed by using pdfium, [nemoretriever-parse](https://build.nvidia.com/nvidia/nemoretriever-parse), Unstructured.io, and Adobe Content Extraction Services.\n- Support various types of before and after processing operations, including text splitting and chunking, transform and filtering, embedding generation, and image offloading to storage.\n\n\nNeMo Retriever Extraction supports the following file types:\n\n- `bmp`\n- `docx`\n- `html` (treated as text)\n- `jpeg`\n- `json` (treated as text)\n- `md` (treated as text)\n- `pdf`\n- `png`\n- `pptx`\n- `sh` (treated as text)\n- `tiff`\n- `txt`\n\n\n### What NeMo Retriever Extraction Isn't\n\nNeMo Retriever extraction does not do the following:\n\n- Run a static pipeline or fixed set of operations on every submitted document.\n- Act as a wrapper for any specific document parsing library.\n\n\nFor more information, see the [full NeMo Retriever Extraction documentation](https://docs.nvidia.com/nemo/retriever/extraction/overview/).\n\n\n## Prerequisites\n\nFor production-level performance and scalability, we recommend that you deploy the pipeline and supporting NIMs by using Docker Compose or Kubernetes ([helm charts](helm)). For more information, refer to [prerequisites](https://docs.nvidia.com/nv-ingest/user-guide/getting-started/prerequisites).\n\n\n## Library Mode Quickstart\n\nFor small-scale workloads, such as workloads of fewer than 100 PDFs, you can use library mode setup. Library mode set up depends on NIMs that are already self-hosted, or, by default, NIMs that are hosted on build.nvidia.com.\n\nLibrary mode deployment of nv-ingest requires:\n\n- Linux operating systems (Ubuntu 22.04 or later recommended)\n- Python 3.12\n- We strongly advise using an isolated Python virtual env, such as provided by [uv](https://docs.astral.sh/uv/getting-started/installation/) or [conda](https://github.com/conda-forge/miniforge)\n\n### Step 1: Prepare Your Environment\n\nCreate a fresh Python environment to install nv-ingest and dependencies.\n\n```shell\nuv venv --python 3.12 nvingest \u0026\u0026 \\\n  source nvingest/bin/activate \u0026\u0026 \\\n  uv pip install nv-ingest==25.6.2 nv-ingest-api==25.6.2 nv-ingest-client==25.6.2\n```\n\nSet your NVIDIA_API_KEY. If you don't have a key, you can get one on [build.nvidia.com](https://org.ngc.nvidia.com/setup/api-keys). For instructions, refer to [Generate Your NGC Keys](/docs/docs/extraction/ngc-api-key.md).\n\n```\nexport NVIDIA_API_KEY=nvapi-...\n```\n\n### Step 2: Ingest Documents\n\nYou can submit jobs programmatically in Python.\n\nTo confirm that you have activated your Python environment, run `which python` and confirm that you see `nvingest` in the result. You can do this before any python command that you run.\n\n```\nwhich python\n/home/dev/projects/nv-ingest/nvingest/bin/python\n```\n\nIf you have a very high number of CPUs, and see the process hang without progress, we recommend that you use `taskset` to limit the number of CPUs visible to the process. Use the following code.\n\n```\ntaskset -c 0-3 python your_ingestion_script.py\n```\n\nOn a 4 CPU core low end laptop, the following code should take about 10 seconds.\n\n```python\nimport logging, os, time\n\nfrom nv_ingest.framework.orchestration.ray.util.pipeline.pipeline_runners import run_pipeline\nfrom nv_ingest.framework.orchestration.ray.util.pipeline.pipeline_runners import PipelineCreationSchema\nfrom nv_ingest_api.util.logging.configuration import configure_logging as configure_local_logging\nfrom nv_ingest_client.client import Ingestor, NvIngestClient\nfrom nv_ingest_api.util.message_brokers.simple_message_broker import SimpleClient\nfrom nv_ingest_client.util.process_json_files import ingest_json_results_to_blob\n\n# Start the pipeline subprocess for library mode\nconfig = PipelineCreationSchema()\n\nrun_pipeline(config, block=False, disable_dynamic_scaling=True, run_in_subprocess=True)\n\nclient = NvIngestClient(\n    message_client_allocator=SimpleClient,\n    message_client_port=7671,\n    message_client_hostname=\"localhost\"\n)\n\n# gpu_cagra accelerated indexing is not available in milvus-lite\n# Provide a filename for milvus_uri to use milvus-lite\nmilvus_uri = \"milvus.db\"\ncollection_name = \"test\"\nsparse = False\n\n# do content extraction from files                                \ningestor = (\n    Ingestor(client=client)\n    .files(\"data/multimodal_test.pdf\")\n    .extract(\n        extract_text=True,\n        extract_tables=True,\n        extract_charts=True,\n        extract_images=True,\n        paddle_output_format=\"markdown\",\n        extract_infographics=True,\n        # extract_method=\"nemoretriever_parse\", #Slower, but maximally accurate, especially for PDFs with pages that are scanned images\n        text_depth=\"page\"\n    ).embed()\n    .vdb_upload(\n        collection_name=collection_name,\n        milvus_uri=milvus_uri,\n        sparse=sparse,\n        # for llama-3.2 embedder, use 1024 for e5-v5\n        dense_dim=2048\n    )\n)\n\nprint(\"Starting ingestion..\")\nt0 = time.time()\nresults = ingestor.ingest(show_progress=True)\nt1 = time.time()\nprint(f\"Time taken: {t1 - t0} seconds\")\n\n# results blob is directly inspectable\nprint(ingest_json_results_to_blob(results[0]))\n```\n\nYou can see the extracted text that represents the content of the ingested test document.\n\n```shell\nStarting ingestion..\nTime taken: 9.243880033493042 seconds\n\nTestingDocument\nA sample document with headings and placeholder text\nIntroduction\nThis is a placeholder document that can be used for any purpose. It contains some \nheadings and some placeholder text to fill the space. The text is not important and contains \nno real value, but it is useful for testing. Below, we will have some simple tables and charts \nthat we can use to confirm Ingest is working as expected.\nTable 1\nThis table describes some animals, and some activities they might be doing in specific \nlocations.\nAnimal Activity Place\nGira@e Driving a car At the beach\nLion Putting on sunscreen At the park\nCat Jumping onto a laptop In a home o@ice\nDog Chasing a squirrel In the front yard\nChart 1\nThis chart shows some gadgets, and some very fictitious costs.\n... document extract continues ...\n```\n\n### Step 3: Query Ingested Content\n\nTo query for relevant snippets of the ingested content, and use them with an LLM to generate answers, use the following code.\n\n```python\nfrom openai import OpenAI\nfrom nv_ingest_client.util.milvus import nvingest_retrieval\nimport os\n\nmilvus_uri = \"milvus.db\"\ncollection_name = \"test\"\nsparse=False\n\nqueries = [\"Which animal is responsible for the typos?\"]\n\nretrieved_docs = nvingest_retrieval(\n    queries,\n    collection_name,\n    milvus_uri=milvus_uri,\n    hybrid=sparse,\n    top_k=1,\n)\n\n# simple generation example\nextract = retrieved_docs[0][0][\"entity\"][\"text\"]\nclient = OpenAI(\n  base_url = \"https://integrate.api.nvidia.com/v1\",\n  api_key = os.environ[\"NVIDIA_API_KEY\"]\n)\n\nprompt = f\"Using the following content: {extract}\\n\\n Answer the user query: {queries[0]}\"\nprint(f\"Prompt: {prompt}\")\ncompletion = client.chat.completions.create(\n  model=\"nvidia/llama-3.1-nemotron-70b-instruct\",\n  messages=[{\"role\":\"user\",\"content\": prompt}],\n)\nresponse = completion.choices[0].message.content\n\nprint(f\"Answer: {response}\")\n```\n\n```shell\nPrompt: Using the following content: TestingDocument\nA sample document with headings and placeholder text\nIntroduction\nThis is a placeholder document that can be used for any purpose. It contains some \nheadings and some placeholder text to fill the space. The text is not important and contains \nno real value, but it is useful for testing. Below, we will have some simple tables and charts \nthat we can use to confirm Ingest is working as expected.\nTable 1\nThis table describes some animals, and some activities they might be doing in specific \nlocations.\nAnimal Activity Place\nGira@e Driving a car At the beach\nLion Putting on sunscreen At the park\nCat Jumping onto a laptop In a home o@ice\nDog Chasing a squirrel In the front yard\nChart 1\nThis chart shows some gadgets, and some very fictitious costs.\n\n Answer the user query: Which animal is responsible for the typos?\nAnswer: A clever query!\n\nAfter carefully examining the provided content, I'd like to point out the potential \"typos\" (assuming you're referring to the unusual or intentionally incorrect text) and attempt to playfully \"assign blame\" to an animal based on the context:\n\n1. **Gira@e** (instead of Giraffe) - **Animal blamed: Giraffe** (Table 1, first row)\n\t* The \"@\" symbol in \"Gira@e\" suggests a possible typo or placeholder character, which we'll humorously attribute to the Giraffe's alleged carelessness.\n2. **o@ice** (instead of Office) - **Animal blamed: Cat**\n\t* The same \"@\" symbol appears in \"o@ice\", which is related to the Cat's activity in the same table. Perhaps the Cat was in a hurry while typing and introduced the error?\n\nSo, according to this whimsical analysis, both the **Giraffe** and the **Cat** are \"responsible\" for the typos, with the Giraffe possibly being the more egregious offender given the more blatant character substitution in its name.\n```\n\n\u003e [!TIP]\n\u003e Beyond inspecting the results, you can read them into things like [llama-index](examples/llama_index_multimodal_rag.ipynb) or [langchain](examples/langchain_multimodal_rag.ipynb) retrieval pipelines.\n\u003e\n\u003e Please also checkout our [demo using a retrieval pipeline on build.nvidia.com](https://build.nvidia.com/nvidia/multimodal-pdf-data-extraction-for-enterprise-rag) to query over document content pre-extracted w/ NVIDIA Ingest.\n\n\n## GitHub Repository Structure\n\nThe following is a description of the folders in the GitHub repository.\n\n- [.devcontainer](https://github.com/NVIDIA/nv-ingest/tree/main/.devcontainer) — VSCode containers for local development\n- [.github](https://github.com/NVIDIA/nv-ingest/tree/main/.github) — GitHub repo configuration files\n- [api](https://github.com/NVIDIA/nv-ingest/tree/main/api) — Core API logic shared across python modules\n- [ci](https://github.com/NVIDIA/nv-ingest/tree/main/ci) — Scripts used to build the nv-ingest container and other packages\n- [client](https://github.com/NVIDIA/nv-ingest/tree/main/client) — Readme, examples, and source code for the nv-ingest-cli utility\n- [conda](https://github.com/NVIDIA/nv-ingest/tree/main/conda) — Conda environment and packaging definitions\n- [config](https://github.com/NVIDIA/nv-ingest/tree/main/config) — Various .yaml files defining configuration for OTEL, Prometheus\n- [data](https://github.com/NVIDIA/nv-ingest/tree/main/data) — Sample PDFs for testing\n- [deploy](https://github.com/NVIDIA/nv-ingest/tree/main/deploy) — Brev.dev-hosted launchable\n- [docker](https://github.com/NVIDIA/nv-ingest/tree/main/docker) — Scripts used by the nv-ingest docker container\n- [docs](https://github.com/NVIDIA/nv-ingest/tree/main/docs/docs) — Documentation for NV Ingest\n- [evaluation](https://github.com/NVIDIA/nv-ingest/tree/main/evaluation) — Notebooks that demonstrate how to test recall accuracy\n- [examples](https://github.com/NVIDIA/nv-ingest/tree/main/examples) — Notebooks, scripts, and tutorial content\n- [helm](https://github.com/NVIDIA/nv-ingest/tree/main/helm) — Documentation for deploying nv-ingest to a Kubernetes cluster via Helm chart\n- [skaffold](https://github.com/NVIDIA/nv-ingest/tree/main/skaffold) — Skaffold configuration\n- [src](https://github.com/NVIDIA/nv-ingest/tree/main/src) — Source code for the nv-ingest pipelines and service\n- [tests](https://github.com/NVIDIA/nv-ingest/tree/main/tests) — Unit tests for nv-ingest\n\n\n## Notices\n\n### Third Party License Notice:\n\nIf configured to do so, this project will download and install additional third-party open source software projects.\nReview the license terms of these open source projects before use:\n\nhttps://pypi.org/project/pdfservices-sdk/\n\n- **`INSTALL_ADOBE_SDK`**:\n  - **Description**: If set to `true`, the Adobe SDK will be installed in the container at launch time. This is\n    required if you want to use the Adobe extraction service for PDF decomposition. Please review the\n    [license agreement](https://github.com/adobe/pdfservices-python-sdk?tab=License-1-ov-file) for the\n    pdfservices-sdk before enabling this option.\n- **`DOWNLOAD_LLAMA_TOKENIZER` (Built With Llama):**:\n  - **Description**: The Split task uses the `meta-llama/Llama-3.2-1B` tokenizer, which will be downloaded\n    from HuggingFace at build time if `DOWNLOAD_LLAMA_TOKENIZER` is set to `True`. Please review the\n    [license agreement](https://huggingface.co/meta-llama/Llama-3.2-1B) for Llama 3.2 materials before using this.\n    This is a gated model so you'll need to [request access](https://huggingface.co/meta-llama/Llama-3.2-1B) and\n    set `HF_ACCESS_TOKEN` to your HuggingFace access token in order to use it.\n\n\n### Contributing\n\nWe require that all contributors \"sign-off\" on their commits. This certifies that the contribution is your original\nwork, or you have rights to submit it under the same license, or a compatible license.\n\nAny contribution which contains commits that are not signed off are not accepted.\n\nTo sign off on a commit, use the --signoff (or -s) option when you commit your changes as shown following.\n\n```\n$ git commit --signoff --message \"Add cool feature.\"\n```\n\nThis appends the following text to your commit message.\n\n```\nSigned-off-by: Your Name \u003cyour@email.com\u003e\n```\n\n#### Developer Certificate of Origin (DCO)\n\nThe following is the full text of the Developer Certificate of Origin (DCO)\n\n```\n  Developer Certificate of Origin\n  Version 1.1\n\n  Copyright (C) 2004, 2006 The Linux Foundation and its contributors.\n  1 Letterman Drive\n  Suite D4700\n  San Francisco, CA, 94129\n\n  Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed.\n```\n\n```\n  Developer's Certificate of Origin 1.1\n\n  By making a contribution to this project, I certify that:\n\n  (a) The contribution was created in whole or in part by me and I have the right to submit it under the open source license indicated in the file; or\n\n  (b) The contribution is based upon previous work that, to the best of my knowledge, is covered under an appropriate open source license and I have the right under that license to submit that work with modifications, whether created in whole or in part by me, under the same open source license (unless I am permitted to submit under a different license), as indicated in the file; or\n\n  (c) The contribution was provided directly to me by some other person who certified (a), (b) or (c) and I have not modified it.\n\n  (d) I understand and agree that this project and the contribution are public and that a record of the contribution (including all personal information I submit with it, including my sign-off) is maintained indefinitely and may be redistributed consistent with this project or the open source license(s) involved.\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FNVIDIA%2Fnv-ingest","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FNVIDIA%2Fnv-ingest","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FNVIDIA%2Fnv-ingest/lists"}