{"id":29526293,"url":"https://github.com/dantetemplar/pdf-extraction-agenda","last_synced_at":"2025-10-16T11:17:14.346Z","repository":{"id":279961391,"uuid":"940576838","full_name":"dantetemplar/pdf-extraction-agenda","owner":"dantetemplar","description":"Overview of pipelines related to PDF to Markdown document processing.","archived":false,"fork":false,"pushed_at":"2025-07-10T12:58:11.000Z","size":3383,"stargazers_count":78,"open_issues_count":25,"forks_count":0,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-07-10T19:54:47.510Z","etag":null,"topics":["benchmark","ocr","pdf","pdf-document","pdf2md","pipeline"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/dantetemplar.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-02-28T12:25:07.000Z","updated_at":"2025-07-10T16:52:24.000Z","dependencies_parsed_at":"2025-06-04T15:37:50.529Z","dependency_job_id":"c62c686a-9fd3-4ea4-a72a-16bb1f2c8b2f","html_url":"https://github.com/dantetemplar/pdf-extraction-agenda","commit_stats":null,"previous_names":["dantetemplar/pdf-extraction-agenda"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/dantetemplar/pdf-extraction-agenda","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dantetemplar%2Fpdf-extraction-agenda","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dantetemplar%2Fpdf-extraction-agenda/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dantetemplar%2Fpdf-extraction-agenda/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dantetemplar%2Fpdf-extraction-agenda/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/dantetemplar","download_url":"https://codeload.github.com/dantetemplar/pdf-extraction-agenda/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dantetemplar%2Fpdf-extraction-agenda/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279182726,"owners_count":26121242,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-16T02:00:06.019Z","response_time":53,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["benchmark","ocr","pdf","pdf-document","pdf2md","pipeline"],"created_at":"2025-07-16T20:04:52.793Z","updated_at":"2025-10-16T11:17:14.339Z","avatar_url":"https://github.com/dantetemplar.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# PDF extraction pipelines and benchmarks agenda\n\n\u003e [!CAUTION]\n\u003e Part of text in this repo written by ChatGPT. Also, I haven't yet run all pipelines because of lack of compute power.\n\nThis repository provides an overview of notable **pipelines** and **benchmarks** related to PDF/OCR document\nprocessing. Each entry includes a brief description, and useful data.\n\n## Table of contents\n\nDid you know that GitHub supports table of\ncontents [by default](https://github.blog/changelog/2021-04-13-table-of-contents-support-in-markdown-files/) 🤔\n\n## Comparison\n\n\u003e [!IMPORTANT]\n\u003e Open [README.md in separate page](https://github.com/dantetemplar/pdf-extraction-agenda/blob/main/README.md), not in repository preview! It will look better.\n\n| Pipeline                                  | [OmniDocBench](#omnidocbench) Overall ↓ | [olmOCR](#olmoocr-eval) Overall ↑ | [Omni OCR](#omni-ocr-benchmark) Accuracy ↑ | [Marker](#marker-benchmarks) Overall ↓ | [Mistral](#mistral-ocr-benchmarks) Overall ↑ | [dp-bench](#dp-bench) NID ↑ | [READoc](#readoc) Overall ↑ | [Actualize.pro](#actualize-pro) Overall ↑ |\n| ----------------------------------------- | --------------------------------------- | --------------------------------- | :----------------------------------------- | -------------------------------------- | :------------------------------------------- | --------------------------- | --------------------------- | ----------------------------------------- |\n| [MinerU](#MinerU)                         | 0.150 \u003csup\u003e[3]\u003c/sup\u003e ⚠️                 | 61.5                              |                                            |                                        |                                              |                             | 60.17                       | **8**                                     |\n| [Marker](#Marker)                         | 0.336                                   | 70.1                              |                                            | **4.24** ⚠️                            |                                              |                             | 63.57                       | 6.5                                       |\n| [MonkeyOCR (pro-3B)](#MonkeyOCR)          | **0.138 \u003csup\u003e[1]\u003c/sup\u003e** ⚠️             | **75.8 \u003csup\u003e[1]\u003c/sup\u003e** ⚠️        |                                            |                                        |                                              |                             |                             |                                           |\n| [olmOCR](#olmOCR)                         | 0.326                                   | 75.5 \u003csup\u003e[2]\u003c/sup\u003e ⚠️            |                                            |                                        |                                              |                             |                             |                                           |\n| [DocLing](#DocLing)                       | 0.589                                   |                                   |                                            | 3.70                                   |                                              |                             |                             | 7.3                                       |\n| [MarkItDown](#MarkItDown)                 |                                         |                                   |                                            |                                        |                                              |                             |                             | 7.78                                      |\n| [Zerox (OmniAI)](#Zerox)                  |                                         |                                   | **91.7 \u003csup\u003e[1]\u003c/sup\u003e** ⚠️                 |                                        |                                              |                             |                             | 7.9                                       |\n| [Unstructured](#Unstructured)             | 0.586                                   |                                   | 50.8                                       |                                        |                                              | 91.18                       |                             | 6.2                                       |\n| [Pix2Text](#Pix2Text)                     | 0.32                                    |                                   |                                            |                                        |                                              |                             | 64.39                       |                                           |\n| [open-parse](#open-parse)                 | 0.646                                   |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| [Markdrop](#markdrop)                     |                                         |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| [Vision Parse](#Vision-Parse)             |                                         |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| _↓ Proprietary pipelines_                 |                                         |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| [Mistral OCR](#MistralOCR)                | 0.268                                   | 72.0 \u003csup\u003e[3]\u003c/sup\u003e               |                                            |                                        | **94.89 ⚠️**                                 |                             |                             |                                           |\n| [Google Document AI](#Google-Document-AI) |                                         |                                   | 67.8                                       |                                        | 83.42                                        | 90.86                       |                             |                                           |\n| [Azure OCR](#Azure-OCR)                   |                                         |                                   | 85.1                                       |                                        | 89.52                                        | 87.69                       |                             |                                           |\n| [Amazon Textract](#Amazon-Textract)       |                                         |                                   | 74.3                                       |                                        |                                              | 96.71                       |                             |                                           |\n| [LlamaParse](#LlamaParse)                 |                                         |                                   |                                            | 3.98                                   |                                              | 92.82                       |                             | 7.1                                       |\n| [Mathpix](#Mathpix)                       | 0.191                                   |                                   |                                            | 4.16                                   |                                              |                             |                             |                                           |\n| [upstage](#upstage-ai)                    |                                         |                                   |                                            |                                        |                                              | **97.02**  ⚠️               |                             |                                           |\n| [doc2x](#doc2x)                           |                                         |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| _↓ Expert VLMs_                           |                                         |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| [Nougat](#Nougat)                         | 0.452                                   |                                   |                                            |                                        |                                              |                             | **81.42**                   |                                           |\n| [GOT-OCR](#GOT-OCR)                       | 0.287                                   | 48.3                              |                                            |                                        |                                              |                             |                             |                                           |\n| [SmolDocling](#SmolDocling)               | 0.493                                   |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| Nanonets-OCR                              |                                         | 64.5                              |                                            |                                        |                                              |                             |                             |                                           |\n| _↓ General VLMs_                          |                                         |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| Gemini-1.5 Flash                          |                                         |                                   |                                            |                                        | 90.23                                        |                             |                             |                                           |\n| Gemini-1.5 Pro                            |                                         |                                   |                                            |                                        | 89.92                                        |                             |                             |                                           |\n| Gemini-2.0 Flash                          | 0.191                                   | 63.8                              | 86.1 \u003csup\u003e[2]\u003c/sup\u003e                        |                                        | 88.69                                        |                             |                             |                                           |\n| Gemini-2.5 Pro                            | 0.148 \u003csup\u003e[2]\u003c/sup\u003e                    |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| GPT4o                                     | 0.233                                   | 69.9                              | 75.5                                       |                                        | 89.77                                        |                             |                             |                                           |\n| Claude Sonnet 3.5                         |                                         |                                   | 69.3                                       |                                        |                                              |                             |                             |                                           |\n| Qwen2-VL-72B                              | 0.252                                   |                                   |                                            |                                        |                                              |                             |                             |                                           |\n| Qwen2.5-VL-72B                            | 0.214                                   | 65.5                              |                                            |                                        |                                              |                             |                             |                                           |\n| InternVL2-76B                             | 0.44                                    |                                   |                                            |                                        |                                              |                             |                             |                                           |\n\n- **Bold** indicates the best result for a given metric, and \u003csup\u003e[2]\u003c/sup\u003e indicates 2nd place in that benchmark.\n- \" \" means the pipeline was not evaluated in that benchmark.\n- ⚠️ means the pipeline authors are the ones who did the benchmark.\n- `Overall ↑` in column name means higher value is better, when `Overall ↓` - lower value is better.\n\n\u003e [!NOTE]\n\u003e I'm working on implementing an easy-to-repeat benchmarking (just run notebook on colab to repeat results, or extend\n\u003e them), but for now I'm struggling with finding suitable dataset.\n\n## Pipelines\n\n### [MinerU](https://github.com/opendatalab/MinerU)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/7)\n[![GitHub last commit](https://img.shields.io/github/last-commit/opendatalab/MinerU?label=GitHub\u0026logo=github)](https://github.com/opendatalab/MinerU)\n![License](https://img.shields.io/badge/License-AGPL--3.0-orange)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://huggingface.co/spaces/opendatalab/MinerU)\n\n**Primary Language:** Python\n\n**License:** AGPL-3.0\n\n**Description:** MinerU is an open-source tool designed to convert PDFs into machine-readable formats, such as Markdown\nand JSON, facilitating seamless data extraction and further processing. Developed during the pre-training phase of\nInternLM, MinerU addresses symbol conversion challenges in scientific literature, making it invaluable for research and\ndevelopment in large language models. Key features include:\n\n- **Content Cleaning**: Removes headers, footers, footnotes, and page numbers to ensure semantic coherence.\n- **Structure Preservation**: Maintains the original document structure, including titles, paragraphs, and lists.\n- **Multimodal Extraction**: Accurately extracts images, image descriptions, tables, and table captions.\n- **Formula Recognition**: Converts recognized formulas into LaTeX format.\n- **Table Conversion**: Transforms tables into LaTeX or HTML formats.\n- **OCR Capabilities**: Detects scanned or corrupted PDFs and enables OCR functionality, supporting text recognition in\n  84 languages.\n- **Cross-Platform Compatibility**: Operates on Windows, Linux, and Mac platforms, supporting both CPU and GPU\n  environments.\n\n### [Marker](https://github.com/VikParuchuri/marker)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/8)\n[![GitHub last commit](https://img.shields.io/github/last-commit/VikParuchuri/marker?label=GitHub\u0026logo=github)](https://github.com/VikParuchuri/marker)\n![License](https://img.shields.io/badge/License-GPL--3.0-yellow)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://www.datalab.to/)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://www.datalab.to/)\n\n**Primary Language:** Python\n\n**License:** GPL-3.0\n\n**Description:** Marker “converts PDFs and images to markdown, JSON, and HTML quickly and accurately.” It is designed to\nhandle a wide range of document types in all languages and produce structured outputs.\n\n**Benchmark Results:** https://github.com/VikParuchuri/marker?tab=readme-ov-file#performance\n\n**API Details:**\n\n- **API URL:** https://www.datalab.to/\n- **Pricing:** https://www.datalab.to/plans\n- **Average Price:** $3 per 1000 pages, at least $25 per month\n\n**Additional Notes:**\n**Demo available after registration on https://www.datalab.to/**\n\n### [MarkItDown](https://github.com/microsoft/markitdown)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/9)\n[![GitHub last commit](https://img.shields.io/github/last-commit/microsoft/markitdown?label=GitHub\u0026logo=github)](https://github.com/microsoft/markitdown)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://mitdown.ca/)\n\n**Primary Language:** Python\n\n**License:** MIT\n\n**Description:** MarkItDown is a Python-based utility developed by Microsoft for converting various file formats into\nMarkdown. It supports a wide range of file types, including:\n\n- **Office Documents**: Word (.docx), PowerPoint (.pptx), Excel (.xlsx)\n- **Media Files**: Images (with EXIF metadata and OCR capabilities), Audio (with speech transcription)\n- **Web and Data Formats**: HTML, CSV, JSON, XML\n- **Archives**: ZIP files (with recursive content parsing)\n- **URLs**: YouTube links\n\nThis versatility makes MarkItDown a valuable tool for tasks such as indexing, text analysis, and preparing content for\nLarge Language Model (LLM) training. The utility offers both command-line and Python API interfaces, providing\nflexibility for various use cases. Additionally, MarkItDown features a plugin-based architecture, allowing for easy\nintegration of third-party extensions to enhance its functionality.\n\n### [olmOCR](https://olmocr.allenai.org/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/10)\n[![GitHub last commit](https://img.shields.io/github/last-commit/allenai/olmocr?label=GitHub\u0026logo=github)](https://github.com/allenai/olmocr)\n![License](https://img.shields.io/badge/License-Apache--2.0-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://olmocr.allenai.org/)\n\n**Primary Language:** Python\n\n**License:** Apache-2.0\n\n**Description:** olmOCR is an open-source toolkit developed by the Allen Institute for AI, designed to convert PDFs and\ndocument images into clean, plain text suitable for large language model (LLM) training and other applications. Key\nfeatures include:\n\n- **High Accuracy**: Preserves reading order and supports complex elements such as tables, equations, and handwriting.\n- **Document Anchoring**: Combines text and visual information to enhance extraction accuracy.\n- **Structured Content Representation**: Utilizes Markdown to represent structured content, including sections, lists,\n  equations, and tables.\n- **Optimized Pipeline**: Compatible with SGLang and vLLM inference engines, enabling efficient scaling from single to\n  multiple GPUs.\n\n### [MonkeyOCR](https://github.com/Yuliang-Liu/MonkeyOCR)\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/27)\n[![GitHub last commit](https://img.shields.io/github/last-commit/Yuliang-Liu/MonkeyOCR?label=GitHub\u0026logo=github)](https://github.com/Yuliang-Liu/MonkeyOCR)\n![License](https://img.shields.io/badge/License-Apache--2.0-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](http://vlrlabmonkey.xyz:7685)\n\n**License:** Apache-2.0\n\n**Description:** MonkeyOCR is an open‑source, **layout‑aware document parsing system** developed by Yuliang‑Liu and collaborators that implements a novel **Structure‑Recognition‑Relation (SRR)** \ntriplet paradigm. It decomposes document analysis into three phases—block structure detection (“Where is it?”), \ncontent recognition (“What is it?”), and reading‑order relation modeling (“How is it organized?”)—delivering both high \naccuracy and inference speed by avoiding heavy end‑to‑end models or brittle modular pipelines. \nTrained on the extensive **MonkeyDoc dataset** (nearly 3.9 million instances across English and Chinese, \ncovering 10+ document types), MonkeyOCR achieves state‑of‑the‑art performance, including significant gains in table \n(+8.6%) and formula (+15.0%) recognition, and outperforms much larger models like Qwen2.5‑VL (72B) and Gemini 2.5 Pro. \nRemarkably, the 3B‑parameter variant runs efficiently—approximately 0.84 pages per second on multi‑page input using a single \nNVIDIA 3090 GPU—making it practical for real‑world document workloads.\n\n\n**Benchmark Results:** https://github.com/Yuliang-Liu/MonkeyOCR?tab=readme-ov-file#benchmark-results\n\n### [MistralOCR](https://mistral.ai/news/mistral-ocr)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/20)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://colab.research.google.com/github/mistralai/cookbook/blob/main/mistral/ocr/structured_ocr.ipynb)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://docs.mistral.ai/capabilities/document/)\n\n**License:** Proprietary\n\n**API Details:**\n\n- **API URL:** https://docs.mistral.ai/capabilities/document/\n- **Pricing:** https://mistral.ai/products/la-plateforme#pricing\n- **Average Price:** 1$ per 1000 pages\n\n### [Google Document AI](https://cloud.google.com/document-ai)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/23)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://console.cloud.google.com/ai/document-ai)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://cloud.google.com/document-ai/docs/reference/rest)\n\n**License:** Proprietary\n\n**Description:** Google Document AI is a cloud-based document processing service that uses machine learning to\nautomatically extract structured data from documents. It supports various document types, including invoices, receipts,\nforms, and identity documents. Key features include:\n\n- **Optical Character Recognition (OCR)**: Converts scanned images and PDFs into editable text.\n- **Data Extraction**: Identifies and extracts key-value pairs, tables, and other structured data.\n- **Document Understanding** Classifies and understands the content of documents.\n- **Customization**: Allows users to train custom models for specific document types.\n\n**API Details:**\n\n- **API URL:** https://cloud.google.com/document-ai/docs/reference/rest\n- **Pricing:** https://cloud.google.com/document-ai/pricing\n- **Average Price:** $1.50 per 1000 pages\n\n### [Azure OCR](https://azure.microsoft.com/en-us/products/ai-services/ai-vision)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/24)\n[![GitHub last commit](https://img.shields.io/github/last-commit/Azure/azure-sdk-for-python?label=GitHub\u0026logo=github)](https://github.com/Azure/azure-sdk-for-python)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://learn.microsoft.com/en-us/azure/cognitive-services/computer-vision/ocr)\n\n**License:** Proprietary\n\n**Description:** Azure AI Vision OCR is a cloud-based service that employs advanced machine-learning algorithms to\nextract printed and handwritten text from images and documents. It supports a wide array of languages and can process\nvarious content types, including posters, street signs, product labels, and business documents. The service is designed\nto detect text lines, words, and paragraphs, providing structured output suitable for integration into applications\nrequiring text extraction capabilities.\n\n**API Details:**\n\n- **API URL:** https://learn.microsoft.com/en-us/azure/cognitive-services/computer-vision/ocr\n- **Pricing:** https://azure.microsoft.com/en-us/pricing/details/cognitive-services/computer-vision/\n- **Average Price:** $1 per 1,000 transactions\n\n### [Amazon Textract](https://aws.amazon.com/textract/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/25)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://docs.aws.amazon.com/textract/latest/dg/API_Reference.html)\n\n**License:** Proprietary\n\n**Description:** Amazon Textract is a machine learning service that automatically extracts text, handwriting, and data\nfrom scanned documents. It goes beyond simple optical character recognition (OCR) by also identifying the contents of\nfields in forms, information stored in tables, and the presence of selection elements such as checkboxes. This enables\nthe conversion of unstructured content into structured data, facilitating integration into various applications and\nworkflows.\n\n**API Details:**\n\n- **API URL:** https://docs.aws.amazon.com/textract/latest/dg/API_Reference.html\n- **Pricing:** https://aws.amazon.com/textract/pricing/\n- **Average Price:** $1.50 per 1000 pages\n\n### [LlamaParse](https://www.llamaindex.ai/llamaparse)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/6)\n[![GitHub last commit](https://img.shields.io/github/last-commit/run-llama/llama_parse?label=GitHub\u0026logo=github)](https://github.com/run-llama/llama_parse)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://api.cloud.llamaindex.ai/api/parsing/upload)\n\n**Primary Language:** Python\n\n**License:** Proprietary\n\n**Description:** LlamaParse is a GenAI-native document parsing platform developed by LlamaIndex. It transforms complex\ndocuments—including PDFs, PowerPoint presentations, Word documents, and spreadsheets—into structured, LLM-ready formats.\nLlamaParse excels in accurately extracting and formatting tables, images, and other non-standard layouts, ensuring\nhigh-quality data for downstream applications such as Retrieval-Augmented Generation (RAG) and data processing. The\nplatform supports over 10 file types and offers features like natural language parsing instructions, JSON output, and\nmultilingual support.\n\n**API Details:**\n\n- **API URL:** https://api.cloud.llamaindex.ai/api/parsing/upload\n- **Pricing:** https://docs.cloud.llamaindex.ai/llamaparse/usage_data\n- **Average Price:** **Free Plan**: 1,000 pages per day; **Paid Plan**: 7,000 pages per week, with additional pages at $\n  3 per 1,000 pages\n\n### [Mathpix](https://mathpix.com/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/5)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://docs.mathpix.com/)\n\n**Primary Language:** Not publicly available\n\n**License:** Proprietary\n\n**Description:** Mathpix offers advanced Optical Character Recognition (OCR) technology tailored for STEM content. Their\nservices include the Convert API, which accurately digitizes images and PDFs containing complex elements such as\nmathematical equations, chemical diagrams, tables, and handwritten notes. The platform supports multiple output formats,\nincluding LaTeX, MathML, HTML, and Markdown, facilitating seamless integration into various applications and workflows.\nAdditionally, Mathpix provides the Snipping Tool, a desktop application that allows users to capture and convert content\nfrom their screens into editable formats with a single keyboard shortcut.\n\n**API Details:**\n\n- **API URL:** https://docs.mathpix.com/\n- **Pricing:** https://mathpix.com/pricing\n- **Average Price:** $5 per 1000 pages\n\n### [Upstage AI](https://upstage.ai/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/22)\n[![GitHub last commit](https://img.shields.io/github/last-commit/UpstageAI/cookbook?label=GitHub\u0026logo=github)](https://github.com/UpstageAI/cookbook)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://console.upstage.ai/docs/getting-started)\n\n**License:** Proprietary\n\n**Description:** The Upstage AI is a comprehensive suite of artificial intelligence solutions designed to enhance\nbusiness operations across various industries. It encompasses advanced large language models (LLMs) and document\nprocessing engines to streamline workflows and improve efficiency.\n\n**Benchmark Results:** https://www.upstage.ai/blog/en/icdar-win-interview\n\n**API Details:**\n\n- **API URL:** https://console.upstage.ai/docs/getting-started\n- **Pricing:** https://upstage.ai/pricing\n- **Average Price:** $10 per 1000 pages\n\n### [Nougat](https://facebookresearch.github.io/nougat/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/4)\n[![GitHub last commit](https://img.shields.io/github/last-commit/facebookresearch/nougat?label=GitHub\u0026logo=github)](https://github.com/facebookresearch/nougat)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n\n**Primary Language:** Python\n\n**License:** MIT\n\n**Description:** Nougat (Neural Optical Understanding for Academic Documents) is an open-source Visual Transformer model\ndeveloped by Meta AI Research. It is designed to perform Optical Character Recognition (OCR) on scientific documents,\nconverting PDFs into a machine-readable markup language. Nougat simplifies the extraction of complex elements such as\nmathematical expressions and tables, enhancing the accessibility of scientific knowledge. The model processes raw pixel\ndata from document images and outputs structured markdown text, bridging the gap between human-readable content and\nmachine-readable formats.\n\n### [GOT-OCR](https://github.com/Ucas-HaoranWei/GOT-OCR2.0)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/3)\n[![GitHub last commit](https://img.shields.io/github/last-commit/Ucas-HaoranWei/GOT-OCR2.0?label=GitHub\u0026logo=github)](https://github.com/Ucas-HaoranWei/GOT-OCR2.0)\n![License](https://img.shields.io/badge/License-Apache--2.0-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://huggingface.co/spaces/ucaslcl/GOT_online)\n\n**Primary Language:** Python\n\n**License:** Apache-2.0\n\n**Description:** GOT-OCR (General OCR Theory) is an open-source, unified end-to-end model designed to advance OCR to\nversion 2.0. It supports a wide range of tasks, including plain document OCR, scene text OCR, formatted document OCR,\nand OCR for tables, charts, mathematical formulas, geometric shapes, molecular formulas, and sheet music. The model is\nhighly versatile, supporting various input types and producing structured outputs, making it well-suited for complex OCR\ntasks.\n\n**Benchmark Results:** https://github.com/Ucas-HaoranWei/GOT-OCR2.0#benchmarks\n\n### [DocLing](https://github.com/DS4SD/docling)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/2)\n[![GitHub last commit](https://img.shields.io/github/last-commit/DS4SD/docling?label=GitHub\u0026logo=github)](https://github.com/DS4SD/docling)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n\n**Primary Language:** Python\n\n**License:** MIT\n\n**Description:** DocLing is an open-source document processing pipeline developed by IBM Research. It simplifies the\nparsing of diverse document formats—including PDF, DOCX, PPTX, HTML, and images—and provides seamless integrations with\nthe generative AI ecosystem. Key features include advanced PDF understanding, optical character recognition (OCR)\nsupport, and plug-and-play integrations with frameworks like LangChain and LlamaIndex.\n\n### [SmolDocling](https://huggingface.co/ds4sd/SmolDocling-256M-preview)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/21)\n![License](https://img.shields.io/badge/License-Apache--2.0-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://huggingface.co/spaces/ds4sd/SmolDocling-256M-Demo)\n\n**License:** Apache-2.0\n\n**Description:** SmolDocling is a multimodal Image-Text-to-Text model designed for efficient document conversion,\ndeveloped by Docling team. It retains Docling's most popular features while ensuring full compatibility with Docling\nthrough seamless support for DoclingDocuments.\n\n### [Zerox](https://getomni.ai/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/12)\n[![GitHub last commit](https://img.shields.io/github/last-commit/getomni-ai/zerox?label=GitHub\u0026logo=github)](https://github.com/getomni-ai/zerox)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://getomni.ai/ocr-demo)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://getomni.ai/)\n\n**Primary Language:** TypeScript\n\n**License:** MIT\n\n**Description:** Zerox is an OCR and document extraction tool that leverages vision models to convert PDFs and images\ninto structured Markdown format. It excels in handling complex layouts, including tables and charts, making it ideal for\nAI ingestion and further text analysis.\n\n**Benchmark Results:** https://getomni.ai/ocr-benchmark\n\n**API Details:**\n\n- **API URL:** https://getomni.ai/\n- **Pricing:** https://getomni.ai/pricing\n- **Average Price:** Extract structured data: 'Startup' plan at $225 per month with 5000 pages included, after that $2\n  per 1000 pages\n\n### [Unstructured](https://unstructured.io/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/13)\n[![GitHub last commit](https://img.shields.io/github/last-commit/Unstructured-IO/unstructured?label=GitHub\u0026logo=github)](https://github.com/Unstructured-IO/unstructured)\n![License](https://img.shields.io/badge/License-Apache--2.0-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://demo.unstructured.io/)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://docs.unstructured.io/platform-api/overview)\n\n**Primary Language:** Python\n\n**License:** Apache-2.0\n\n**Description:** Unstructured is an open-source library that provides components for ingesting and pre-processing\nunstructured data, including images and text documents such as PDFs, HTML, and Word documents. It transforms complex\ndata into structured formats suitable for large language models and AI applications. The platform offers\nenterprise-grade connectors to seamlessly integrate various data sources, making it easier to extract and transform data\nfor analysis and processing.\n\n**API Details:**\n\n- **API URL:** https://docs.unstructured.io/platform-api/overview\n- **Pricing:** https://unstructured.io/developers\n- **Average Price:** **Basic Strategy\n  **: $2 per 1,000 pages, suitable for simple, text-only documents. **Advanced Strategy**: $20 per 1,000 pages, ideal\n  for PDFs, images, and complex file types. **Platinum/VLM Strategy**: $30 per 1,000 pages, designed for challenging\n  documents, including scanned and handwritten content with VLM API integration.\n\n### [Pix2Text](https://p2t.breezedeus.com/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/14)\n[![GitHub last commit](https://img.shields.io/github/last-commit/breezedeus/Pix2Text?label=GitHub\u0026logo=github)](https://github.com/breezedeus/Pix2Text)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://p2t.breezedeus.com/)\n\n**Primary Language:** Python\n\n**License:** MIT\n\n**Description:** Pix2Text (P2T) is an open-source Python3 tool designed to recognize layouts, tables, mathematical\nformulas (LaTeX), and text in images, converting them into Markdown format. It serves as a free alternative to Mathpix,\nsupporting over 80 languages, including English, Simplified Chinese, Traditional Chinese, and Vietnamese. P2T can also\nprocess entire PDF files, extracting content into structured Markdown, facilitating seamless conversion of visual\ncontent into text-based representations.\n\n### [Open-Parse](https://filimoa.github.io/open-parse/)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/15)\n[![GitHub last commit](https://img.shields.io/github/last-commit/Filimoa/open-parse?label=GitHub\u0026logo=github)](https://github.com/Filimoa/open-parse)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n\n**Primary Language:** Python\n\n**License:** MIT\n\n**Description:** Open Parse is a flexible, open-source library designed to enhance document chunking for\nRetrieval-Augmented Generation (RAG) systems. It visually analyzes document layouts to effectively group related\ncontent, surpassing traditional text-splitting methods. Key features include:\n\n- **Visually-Driven Analysis**: Understands complex layouts for superior chunking.\n- **Markdown Support**: Extracts headings, bold, and italic text into Markdown format.\n- **High-Precision Table Extraction**: Converts tables into clean Markdown with high accuracy.\n- **Extensibility**: Allows implementation of custom post-processing steps.\n- **Intuitive Design**: Offers robust editor support for seamless integration.\n\n### [Extractous](https://github.com/yobix-ai/extractous)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/16)\n[![GitHub last commit](https://img.shields.io/github/last-commit/yobix-ai/extractous?label=GitHub\u0026logo=github)](https://github.com/yobix-ai/extractous)\n![License](https://img.shields.io/badge/License-Apache--2.0-brightgreen)\n\n**Primary Language:** Rust\n\n**License:** Apache-2.0\n\n**Description:** Extractous is a high-performance, open-source library designed for efficient extraction of content and\nmetadata from various document types, including PDF, Word, HTML, and more. Developed in Rust, it offers bindings for\nmultiple programming languages, starting with Python. Extractous aims to provide a comprehensive solution for\nunstructured data extraction, enabling local and efficient processing without relying on external services or APIs. Key\nfeatures include:\n\n- **High Performance**: Leveraging Rust's capabilities, Extractous achieves faster processing speeds and lower memory\n  utilization compared to traditional extraction libraries.\n- **Multi-Language Support**: While the core is written in Rust, bindings are available for Python, with plans to\n  support additional languages like JavaScript/TypeScript.\n- **Extensive Format Support**: Through integration with Apache Tika, Extractous supports a wide range of file formats,\n  ensuring versatility in data extraction tasks.\n- **OCR Integration**: Incorporates Tesseract OCR to extract text from images and scanned documents, enhancing its\n  ability to handle diverse content types.\n\n**Benchmark Results:** https://github.com/yobix-ai/extractous-benchmarks\n\n### [Markdrop](https://github.com/shoryasethia/markdrop)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/18)\n[![GitHub last commit](https://img.shields.io/github/last-commit/shoryasethia/markdrop?label=GitHub\u0026logo=github)](https://github.com/shoryasethia/markdrop)\n![License](https://img.shields.io/badge/License-GPL--3.0-yellow)\n\n**Primary Language:** Python\n\n**License:** GPL-3.0\n\n**Description:** A Python package for converting PDFs to markdown while extracting images and tables, generate\ndescriptive text descriptions for extracted tables/images using several LLM clients. And many more functionalities.\nMarkdrop is available on PyPI.\n\n### [Vision Parse](https://github.com/iamarunbrahma/vision-parse)\n\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/19)\n[![GitHub last commit](https://img.shields.io/github/last-commit/iamarunbrahma/vision-parse?label=GitHub\u0026logo=github)](https://github.com/iamarunbrahma/vision-parse)\n![License](https://img.shields.io/badge/License-MIT-brightgreen)\n\n**Primary Language:** Python\n\n**License:** MIT\n\n**Description:** Parse PDFs into markdown using Vision LLMs\n\n\n### [doc2x](https://noedgeai.com/)\n[✏️](https://github.com/dantetemplar/pdf-extraction-agenda/issues/28)\n[![GitHub last commit](https://img.shields.io/github/last-commit/NoEdgeAI/pdfdeal?label=GitHub\u0026logo=github)](https://github.com/NoEdgeAI/pdfdeal)\n![License](https://img.shields.io/badge/License-Proprietary-red)\n[![Demo](https://img.shields.io/badge/DEMO-black?logo=awwwards)](https://doc2x.noedgeai.com/)\n[![API](https://img.shields.io/badge/API-Available-blue?logo=swagger\u0026logoColor=85EA2D)](https://noedgeai.github.io/pdfdeal-docs/)\n\n**License:** Proprietary\n\n**Description:** NoEdgeAI is an open‑source technology initiative focused on enhancing document processing in Retrieval-Augmented Generation (RAG) workflows. Their flagship library, pdfdeal, is a Python wrapper for the Doc2X API that facilitates high‑fidelity PDF-to-text conversion. It extends Doc2X’s capabilities by offering local text preprocessing, Markdown and LaTeX extraction, file splitting, image uploading, and enhancements for better recall when integrating PDFs into knowledge‑base tools like Graphrag, Dify, or FastGPT\n\n**API Details:**\n- **API URL:** https://noedgeai.github.io/pdfdeal-docs/\n\n\n## Benchmarks\n\n### [OmniDocBench](https://github.com/opendatalab/OmniDocBench)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/opendatalab/OmniDocBench?label=GitHub\u0026logo=github)](https://github.com/opendatalab/OmniDocBench)\n![GitHub License](https://img.shields.io/github/license/opendatalab/OmniDocBench)\n\u003c!--- \nLicense: Apache 2.0 \nPrimary language: Python\n--\u003e\n\nOmniDocBench is *“a benchmark for evaluating diverse document parsing in real-world scenarios”* by MinerU devs. It\nestablishes a\ncomprehensive evaluation standard for document content extraction methods.\n\n**Notable features:** OmniDocBench covers a wide variety of document types and layouts, comprising **981 PDF pages\nacross 9 document types, 4 layout styles, and 3 languages**. It provides **rich annotations**: over 20k block-level\nelements (paragraphs, headings, tables, etc.) and 80k+ span-level elements (lines, formulas, etc.), including reading\norder and various attribute tags for pages, text, and tables. The dataset undergoes strict quality control (combining\nmanual annotation, intelligent assistance, and expert review for high accuracy). OmniDocBench also comes with *\n*evaluation code** for fair, end-to-end comparisons of document parsing methods. It supports multiple evaluation tasks (\noverall extraction, layout detection, table recognition, formula recognition, OCR text recognition) and standard\nmetrics (Normalized Edit Distance, BLEU, METEOR, TEDS, COCO mAP/mAR, etc.) to benchmark performance across different\naspects of document parsing.\n\n**End-to-End Evaluation**\n\nEnd-to-end evaluation assesses the model's accuracy in parsing PDF page content. The evaluation uses the model's\nMarkdown output of the entire PDF page parsing results as the prediction.\n\n\u003ctable style=\"width: 92%; margin: auto; border-collapse: collapse;\"\u003e\n  \u003cthead\u003e\n    \u003ctr\u003e\n      \u003cth rowspan=\"2\"\u003eMethod Type\u003c/th\u003e\n      \u003cth rowspan=\"2\"\u003eMethods\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eOverall\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eText\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eFormula\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eFormula\u003csup\u003eCDM\u003c/sup\u003e↑\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eTable\u003csup\u003eTEDS\u003c/sup\u003e↑\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eTable\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eRead Order\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n    \u003c/tr\u003e\n  \u003c/thead\u003e\n  \u003ctbody\u003e\n    \u003ctr\u003e\n      \u003ctd rowspan=\"7\"\u003ePipeline Tools\u003c/td\u003e\n      \u003ctd\u003eMinerU-0.9.3\u003c/td\u003e\n      \u003ctd\u003e0.15\u003c/td\u003e\n      \u003ctd\u003e0.357\u003c/td\u003e\n      \u003ctd\u003e0.061\u003c/td\u003e\n      \u003ctd\u003e0.215\u003c/td\u003e\n      \u003ctd\u003e0.278\u003c/td\u003e\n      \u003ctd\u003e0.577\u003c/td\u003e\n      \u003ctd\u003e57.3\u003c/td\u003e\n      \u003ctd\u003e42.9\u003c/td\u003e\n      \u003ctd\u003e78.6\u003c/td\u003e\n      \u003ctd\u003e62.1\u003c/td\u003e\n      \u003ctd\u003e0.18\u003c/td\u003e\n      \u003ctd\u003e0.344\u003c/td\u003e\n      \u003ctd\u003e0.079\u003c/td\u003e\n      \u003ctd\u003e0.292\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eMarker-1.2.3\u003c/td\u003e\n      \u003ctd\u003e0.336\u003c/td\u003e\n      \u003ctd\u003e0.556\u003c/td\u003e\n      \u003ctd\u003e0.08\u003c/td\u003e\n      \u003ctd\u003e0.315\u003c/td\u003e\n      \u003ctd\u003e0.53\u003c/td\u003e\n      \u003ctd\u003e0.883\u003c/td\u003e\n      \u003ctd\u003e17.6\u003c/td\u003e\n      \u003ctd\u003e11.7\u003c/td\u003e\n      \u003ctd\u003e67.6\u003c/td\u003e\n      \u003ctd\u003e49.2\u003c/td\u003e\n      \u003ctd\u003e0.619\u003c/td\u003e\n      \u003ctd\u003e0.685\u003c/td\u003e\n      \u003ctd\u003e0.114\u003c/td\u003e\n      \u003ctd\u003e0.34\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eMathpix\u003c/td\u003e\n      \u003ctd\u003e0.191\u003c/td\u003e\n      \u003ctd\u003e0.365\u003c/td\u003e\n      \u003ctd\u003e0.105\u003c/td\u003e\n      \u003ctd\u003e0.384\u003c/td\u003e\n      \u003ctd\u003e0.306\u003c/td\u003e\n      \u003ctd\u003e0.454\u003c/td\u003e\n      \u003ctd\u003e62.7\u003c/td\u003e\n      \u003ctd\u003e62.1\u003c/td\u003e\n      \u003ctd\u003e77.0\u003c/td\u003e\n      \u003ctd\u003e67.1\u003c/td\u003e\n      \u003ctd\u003e0.243\u003c/td\u003e\n      \u003ctd\u003e0.32\u003c/td\u003e\n      \u003ctd\u003e0.108\u003c/td\u003e\n      \u003ctd\u003e0.304\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eDocling-2.14.0\u003c/td\u003e\n      \u003ctd\u003e0.589\u003c/td\u003e\n      \u003ctd\u003e0.909\u003c/td\u003e\n      \u003ctd\u003e0.416\u003c/td\u003e\n      \u003ctd\u003e0.987\u003c/td\u003e\n      \u003ctd\u003e0.999\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e61.3\u003c/td\u003e\n      \u003ctd\u003e25.0\u003c/td\u003e\n      \u003ctd\u003e0.627\u003c/td\u003e\n      \u003ctd\u003e0.810\u003c/td\u003e\n      \u003ctd\u003e0.313\u003c/td\u003e\n      \u003ctd\u003e0.837\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003ePix2Text-1.1.2.3\u003c/td\u003e\n      \u003ctd\u003e0.32\u003c/td\u003e\n      \u003ctd\u003e0.528\u003c/td\u003e\n      \u003ctd\u003e0.138\u003c/td\u003e\n      \u003ctd\u003e0.356\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.276\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e0.611\u003c/td\u003e\n      \u003ctd\u003e78.4\u003c/td\u003e\n      \u003ctd\u003e39.6\u003c/td\u003e\n      \u003ctd\u003e73.6\u003c/td\u003e\n      \u003ctd\u003e66.2\u003c/td\u003e\n      \u003ctd\u003e0.584\u003c/td\u003e\n      \u003ctd\u003e0.645\u003c/td\u003e\n      \u003ctd\u003e0.281\u003c/td\u003e\n      \u003ctd\u003e0.499\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eUnstructured-0.17.2\u003c/td\u003e\n      \u003ctd\u003e0.586\u003c/td\u003e\n      \u003ctd\u003e0.716\u003c/td\u003e\n      \u003ctd\u003e0.198\u003c/td\u003e\n      \u003ctd\u003e0.481\u003c/td\u003e\n      \u003ctd\u003e0.999\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e0\u003c/td\u003e\n      \u003ctd\u003e0.064\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e0.998\u003c/td\u003e\n      \u003ctd\u003e0.145\u003c/td\u003e\n      \u003ctd\u003e0.387\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eOpenParse-0.7.0\u003c/td\u003e\n      \u003ctd\u003e0.646\u003c/td\u003e\n      \u003ctd\u003e0.814\u003c/td\u003e\n      \u003ctd\u003e0.681\u003c/td\u003e\n      \u003ctd\u003e0.974\u003c/td\u003e\n      \u003ctd\u003e0.996\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e0.106\u003c/td\u003e\n      \u003ctd\u003e0\u003c/td\u003e\n      \u003ctd\u003e64.8\u003c/td\u003e\n      \u003ctd\u003e27.5\u003c/td\u003e\n      \u003ctd\u003e0.284\u003c/td\u003e\n      \u003ctd\u003e0.639\u003c/td\u003e\n      \u003ctd\u003e0.595\u003c/td\u003e\n      \u003ctd\u003e0.641\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd rowspan=\"5\"\u003eExpert VLMs\u003c/td\u003e\n      \u003ctd\u003eGOT-OCR\u003c/td\u003e\n      \u003ctd\u003e0.287\u003c/td\u003e\n      \u003ctd\u003e0.411\u003c/td\u003e\n      \u003ctd\u003e0.189\u003c/td\u003e\n      \u003ctd\u003e0.315\u003c/td\u003e\n      \u003ctd\u003e0.360\u003c/td\u003e\n      \u003ctd\u003e0.528\u003c/td\u003e\n      \u003ctd\u003e74.3\u003c/td\u003e\n      \u003ctd\u003e45.3\u003c/td\u003e\n      \u003ctd\u003e53.2\u003c/td\u003e\n      \u003ctd\u003e47.2\u003c/td\u003e\n      \u003ctd\u003e0.459\u003c/td\u003e\n      \u003ctd\u003e0.52\u003c/td\u003e\n      \u003ctd\u003e0.141\u003c/td\u003e\n      \u003ctd\u003e0.28\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eNougat\u003c/td\u003e\n      \u003ctd\u003e0.452\u003c/td\u003e\n      \u003ctd\u003e0.973\u003c/td\u003e\n      \u003ctd\u003e0.365\u003c/td\u003e\n      \u003ctd\u003e0.998\u003c/td\u003e\n      \u003ctd\u003e0.488\u003c/td\u003e\n      \u003ctd\u003e0.941\u003c/td\u003e\n      \u003ctd\u003e15.1\u003c/td\u003e\n      \u003ctd\u003e16.8\u003c/td\u003e\n      \u003ctd\u003e39.9\u003c/td\u003e\n      \u003ctd\u003e0.0\u003c/td\u003e\n      \u003ctd\u003e0.572\u003c/td\u003e\n      \u003ctd\u003e1.000\u003c/td\u003e\n      \u003ctd\u003e0.382\u003c/td\u003e\n      \u003ctd\u003e0.954\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eMistral OCR\u003c/td\u003e\n      \u003ctd\u003e0.268\u003c/td\u003e\n      \u003ctd\u003e0.439\u003c/td\u003e\n      \u003ctd\u003e0.072\u003c/td\u003e\n      \u003ctd\u003e0.325\u003c/td\u003e\n      \u003ctd\u003e0.318\u003c/td\u003e\n      \u003ctd\u003e0.495\u003c/td\u003e\n      \u003ctd\u003e64.6\u003c/td\u003e\n      \u003ctd\u003e45.9\u003c/td\u003e\n      \u003ctd\u003e75.8\u003c/td\u003e\n      \u003ctd\u003e63.6\u003c/td\u003e\n      \u003ctd\u003e0.6\u003c/td\u003e\n      \u003ctd\u003e0.65\u003c/td\u003e\n      \u003ctd\u003e0.083\u003c/td\u003e\n      \u003ctd\u003e0.284\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eOLMOCR-sglang\u003c/td\u003e\n      \u003ctd\u003e0.326\u003c/td\u003e\n      \u003ctd\u003e0.469\u003c/td\u003e\n      \u003ctd\u003e0.097\u003c/td\u003e\n      \u003ctd\u003e0.293\u003c/td\u003e\n      \u003ctd\u003e0.455\u003c/td\u003e\n      \u003ctd\u003e0.655\u003c/td\u003e\n      \u003ctd\u003e74.3\u003c/td\u003e\n      \u003ctd\u003e43.2\u003c/td\u003e\n      \u003ctd\u003e68.1\u003c/td\u003e\n      \u003ctd\u003e61.3\u003c/td\u003e\n      \u003ctd\u003e0.608\u003c/td\u003e\n      \u003ctd\u003e0.652\u003c/td\u003e\n      \u003ctd\u003e0.145\u003c/td\u003e\n      \u003ctd\u003e0.277\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eSmolDocling-256M_transformer\u003c/td\u003e\n      \u003ctd\u003e0.493\u003c/td\u003e\n      \u003ctd\u003e0.816\u003c/td\u003e\n      \u003ctd\u003e0.262\u003c/td\u003e\n      \u003ctd\u003e0.838\u003c/td\u003e\n      \u003ctd\u003e0.753\u003c/td\u003e\n      \u003ctd\u003e0.997\u003c/td\u003e\n      \u003ctd\u003e32.1\u003c/td\u003e\n      \u003ctd\u003e0.551\u003c/td\u003e\n      \u003ctd\u003e44.9\u003c/td\u003e\n      \u003ctd\u003e16.5\u003c/td\u003e\n      \u003ctd\u003e0.729\u003c/td\u003e\n      \u003ctd\u003e0.907\u003c/td\u003e\n      \u003ctd\u003e0.227\u003c/td\u003e\n      \u003ctd\u003e0.522\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd rowspan=\"8\"\u003eGeneral VLMs\u003c/td\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eGemini2.0-flash\u003c/td\u003e\n      \u003ctd\u003e0.191\u003c/td\u003e\n      \u003ctd\u003e0.264\u003c/td\u003e\n      \u003ctd\u003e0.091\u003c/td\u003e\n      \u003ctd\u003e0.139\u003c/td\u003e\n      \u003ctd\u003e0.389\u003c/td\u003e\n      \u003ctd\u003e0.584\u003c/td\u003e\n      \u003ctd\u003e77.6\u003c/td\u003e\n      \u003ctd\u003e43.6\u003c/td\u003e\n      \u003ctd\u003e79.7\u003c/td\u003e\n      \u003ctd\u003e78.9\u003c/td\u003e\n      \u003ctd\u003e0.193\u003c/td\u003e\n      \u003ctd\u003e0.206\u003c/td\u003e\n      \u003ctd\u003e0.092\u003c/td\u003e\n      \u003ctd\u003e0.128\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eGemini2.5-Pro\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.148\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.212\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.055\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.168\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e0.356\u003c/td\u003e\n      \u003ctd\u003e0.439\u003c/td\u003e\n      \u003ctd\u003e80.0\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e69.4\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e85.8\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e86.4\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.13\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.119\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.049\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.121\u003c/strong\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eGPT4o\u003c/td\u003e\n      \u003ctd\u003e0.233\u003c/td\u003e\n      \u003ctd\u003e0.399\u003c/td\u003e\n      \u003ctd\u003e0.144\u003c/td\u003e\n      \u003ctd\u003e0.409\u003c/td\u003e\n      \u003ctd\u003e0.425\u003c/td\u003e\n      \u003ctd\u003e0.606\u003c/td\u003e\n      \u003ctd\u003e72.8\u003c/td\u003e\n      \u003ctd\u003e42.8\u003c/td\u003e\n      \u003ctd\u003e72.0\u003c/td\u003e\n      \u003ctd\u003e62.9\u003c/td\u003e\n      \u003ctd\u003e0.234\u003c/td\u003e\n      \u003ctd\u003e0.329\u003c/td\u003e\n      \u003ctd\u003e0.128\u003c/td\u003e\n      \u003ctd\u003e0.251\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eQwen2-VL-72B\u003c/td\u003e\n      \u003ctd\u003e0.252\u003c/td\u003e\n      \u003ctd\u003e0.327\u003c/td\u003e\n      \u003ctd\u003e0.096\u003c/td\u003e\n      \u003ctd\u003e0.218\u003c/td\u003e\n      \u003ctd\u003e0.404\u003c/td\u003e\n      \u003ctd\u003e0.487\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e82.2\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e61.2\u003c/td\u003e\n      \u003ctd\u003e76.8\u003c/td\u003e\n      \u003ctd\u003e76.4\u003c/td\u003e\n      \u003ctd\u003e0.387\u003c/td\u003e\n      \u003ctd\u003e0.408\u003c/td\u003e\n      \u003ctd\u003e0.119\u003c/td\u003e\n      \u003ctd\u003e0.193\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eQwen2.5-VL-72B\u003c/td\u003e\n      \u003ctd\u003e0.214\u003c/td\u003e\n      \u003ctd\u003e0.261\u003c/td\u003e\n      \u003ctd\u003e0.092\u003c/td\u003e\n      \u003ctd\u003e0.18\u003c/td\u003e\n      \u003ctd\u003e0.315\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.434\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e68.8\u003c/td\u003e\n      \u003ctd\u003e62.5\u003c/td\u003e\n      \u003ctd\u003e82.9\u003c/td\u003e\n      \u003ctd\u003e83.9\u003c/td\u003e\n      \u003ctd\u003e0.341\u003c/td\u003e\n      \u003ctd\u003e0.262\u003c/td\u003e\n      \u003ctd\u003e0.106\u003c/td\u003e\n      \u003ctd\u003e0.168\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eInternVL2-76B\u003c/td\u003e\n      \u003ctd\u003e0.44\u003c/td\u003e\n      \u003ctd\u003e0.443\u003c/td\u003e\n      \u003ctd\u003e0.353\u003c/td\u003e\n      \u003ctd\u003e0.290\u003c/td\u003e\n      \u003ctd\u003e0.543\u003c/td\u003e\n      \u003ctd\u003e0.701\u003c/td\u003e\n      \u003ctd\u003e67.4\u003c/td\u003e\n      \u003ctd\u003e44.1\u003c/td\u003e\n      \u003ctd\u003e63.0\u003c/td\u003e\n      \u003ctd\u003e60.2\u003c/td\u003e\n      \u003ctd\u003e0.547\u003c/td\u003e\n      \u003ctd\u003e0.555\u003c/td\u003e\n      \u003ctd\u003e0.317\u003c/td\u003e\n      \u003ctd\u003e0.228\u003c/td\u003e\n    \u003c/tr\u003e\n  \u003c/tbody\u003e\n\u003c/table\u003e\n\n\n\u003ctable style=\"width: 92%; margin: auto; border-collapse: collapse;\"\u003e\n  \u003cthead\u003e\n    \u003ctr\u003e\n      \u003cth rowspan=\"2\"\u003eMethod Type\u003c/th\u003e\n      \u003cth rowspan=\"2\"\u003eMethods\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eOverall\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eText\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eFormula\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eFormula\u003csup\u003eCDM\u003c/sup\u003e↑\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eTable\u003csup\u003eTEDS\u003c/sup\u003e↑\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eTable\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n      \u003cth colspan=\"2\"\u003eRead Order\u003csup\u003eEdit\u003c/sup\u003e↓\u003c/th\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n      \u003cth\u003eEN\u003c/th\u003e\n      \u003cth\u003eZH\u003c/th\u003e\n    \u003c/tr\u003e\n  \u003c/thead\u003e\n  \u003ctbody\u003e\n    \u003ctr\u003e\n      \u003ctd rowspan=\"7\"\u003ePipeline Tools\u003c/td\u003e\n      \u003ctd\u003eMinerU-0.9.3\u003c/td\u003e\n      \u003ctd\u003e0.15\u003c/td\u003e\n      \u003ctd\u003e0.357\u003c/td\u003e\n      \u003ctd\u003e0.061\u003c/td\u003e\n      \u003ctd\u003e0.215\u003c/td\u003e\n      \u003ctd\u003e0.278\u003c/td\u003e\n      \u003ctd\u003e0.577\u003c/td\u003e\n      \u003ctd\u003e57.3\u003c/td\u003e\n      \u003ctd\u003e42.9\u003c/td\u003e\n      \u003ctd\u003e78.6\u003c/td\u003e\n      \u003ctd\u003e62.1\u003c/td\u003e\n      \u003ctd\u003e0.18\u003c/td\u003e\n      \u003ctd\u003e0.344\u003c/td\u003e\n      \u003ctd\u003e0.079\u003c/td\u003e\n      \u003ctd\u003e0.292\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eMarker-1.2.3\u003c/td\u003e\n      \u003ctd\u003e0.336\u003c/td\u003e\n      \u003ctd\u003e0.556\u003c/td\u003e\n      \u003ctd\u003e0.08\u003c/td\u003e\n      \u003ctd\u003e0.315\u003c/td\u003e\n      \u003ctd\u003e0.53\u003c/td\u003e\n      \u003ctd\u003e0.883\u003c/td\u003e\n      \u003ctd\u003e17.6\u003c/td\u003e\n      \u003ctd\u003e11.7\u003c/td\u003e\n      \u003ctd\u003e67.6\u003c/td\u003e\n      \u003ctd\u003e49.2\u003c/td\u003e\n      \u003ctd\u003e0.619\u003c/td\u003e\n      \u003ctd\u003e0.685\u003c/td\u003e\n      \u003ctd\u003e0.114\u003c/td\u003e\n      \u003ctd\u003e0.34\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eMathpix\u003c/td\u003e\n      \u003ctd\u003e0.191\u003c/td\u003e\n      \u003ctd\u003e0.365\u003c/td\u003e\n      \u003ctd\u003e0.105\u003c/td\u003e\n      \u003ctd\u003e0.384\u003c/td\u003e\n      \u003ctd\u003e0.306\u003c/td\u003e\n      \u003ctd\u003e0.454\u003c/td\u003e\n      \u003ctd\u003e62.7\u003c/td\u003e\n      \u003ctd\u003e62.1\u003c/td\u003e\n      \u003ctd\u003e77.0\u003c/td\u003e\n      \u003ctd\u003e67.1\u003c/td\u003e\n      \u003ctd\u003e0.243\u003c/td\u003e\n      \u003ctd\u003e0.32\u003c/td\u003e\n      \u003ctd\u003e0.108\u003c/td\u003e\n      \u003ctd\u003e0.304\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eDocling-2.14.0\u003c/td\u003e\n      \u003ctd\u003e0.589\u003c/td\u003e\n      \u003ctd\u003e0.909\u003c/td\u003e\n      \u003ctd\u003e0.416\u003c/td\u003e\n      \u003ctd\u003e0.987\u003c/td\u003e\n      \u003ctd\u003e0.999\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e61.3\u003c/td\u003e\n      \u003ctd\u003e25.0\u003c/td\u003e\n      \u003ctd\u003e0.627\u003c/td\u003e\n      \u003ctd\u003e0.810\u003c/td\u003e\n      \u003ctd\u003e0.313\u003c/td\u003e\n      \u003ctd\u003e0.837\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003ePix2Text-1.1.2.3\u003c/td\u003e\n      \u003ctd\u003e0.32\u003c/td\u003e\n      \u003ctd\u003e0.528\u003c/td\u003e\n      \u003ctd\u003e0.138\u003c/td\u003e\n      \u003ctd\u003e0.356\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.276\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e0.611\u003c/td\u003e\n      \u003ctd\u003e78.4\u003c/td\u003e\n      \u003ctd\u003e39.6\u003c/td\u003e\n      \u003ctd\u003e73.6\u003c/td\u003e\n      \u003ctd\u003e66.2\u003c/td\u003e\n      \u003ctd\u003e0.584\u003c/td\u003e\n      \u003ctd\u003e0.645\u003c/td\u003e\n      \u003ctd\u003e0.281\u003c/td\u003e\n      \u003ctd\u003e0.499\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eUnstructured-0.17.2\u003c/td\u003e\n      \u003ctd\u003e0.586\u003c/td\u003e\n      \u003ctd\u003e0.716\u003c/td\u003e\n      \u003ctd\u003e0.198\u003c/td\u003e\n      \u003ctd\u003e0.481\u003c/td\u003e\n      \u003ctd\u003e0.999\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e-\u003c/td\u003e\n      \u003ctd\u003e0\u003c/td\u003e\n      \u003ctd\u003e0.064\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e0.998\u003c/td\u003e\n      \u003ctd\u003e0.145\u003c/td\u003e\n      \u003ctd\u003e0.387\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eOpenParse-0.7.0\u003c/td\u003e\n      \u003ctd\u003e0.646\u003c/td\u003e\n      \u003ctd\u003e0.814\u003c/td\u003e\n      \u003ctd\u003e0.681\u003c/td\u003e\n      \u003ctd\u003e0.974\u003c/td\u003e\n      \u003ctd\u003e0.996\u003c/td\u003e\n      \u003ctd\u003e1\u003c/td\u003e\n      \u003ctd\u003e0.106\u003c/td\u003e\n      \u003ctd\u003e0\u003c/td\u003e\n      \u003ctd\u003e64.8\u003c/td\u003e\n      \u003ctd\u003e27.5\u003c/td\u003e\n      \u003ctd\u003e0.284\u003c/td\u003e\n      \u003ctd\u003e0.639\u003c/td\u003e\n      \u003ctd\u003e0.595\u003c/td\u003e\n      \u003ctd\u003e0.641\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd rowspan=\"5\"\u003eExpert VLMs\u003c/td\u003e\n      \u003ctd\u003eGOT-OCR\u003c/td\u003e\n      \u003ctd\u003e0.287\u003c/td\u003e\n      \u003ctd\u003e0.411\u003c/td\u003e\n      \u003ctd\u003e0.189\u003c/td\u003e\n      \u003ctd\u003e0.315\u003c/td\u003e\n      \u003ctd\u003e0.360\u003c/td\u003e\n      \u003ctd\u003e0.528\u003c/td\u003e\n      \u003ctd\u003e74.3\u003c/td\u003e\n      \u003ctd\u003e45.3\u003c/td\u003e\n      \u003ctd\u003e53.2\u003c/td\u003e\n      \u003ctd\u003e47.2\u003c/td\u003e\n      \u003ctd\u003e0.459\u003c/td\u003e\n      \u003ctd\u003e0.52\u003c/td\u003e\n      \u003ctd\u003e0.141\u003c/td\u003e\n      \u003ctd\u003e0.28\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eNougat\u003c/td\u003e\n      \u003ctd\u003e0.452\u003c/td\u003e\n      \u003ctd\u003e0.973\u003c/td\u003e\n      \u003ctd\u003e0.365\u003c/td\u003e\n      \u003ctd\u003e0.998\u003c/td\u003e\n      \u003ctd\u003e0.488\u003c/td\u003e\n      \u003ctd\u003e0.941\u003c/td\u003e\n      \u003ctd\u003e15.1\u003c/td\u003e\n      \u003ctd\u003e16.8\u003c/td\u003e\n      \u003ctd\u003e39.9\u003c/td\u003e\n      \u003ctd\u003e0.0\u003c/td\u003e\n      \u003ctd\u003e0.572\u003c/td\u003e\n      \u003ctd\u003e1.000\u003c/td\u003e\n      \u003ctd\u003e0.382\u003c/td\u003e\n      \u003ctd\u003e0.954\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eMistral OCR\u003c/td\u003e\n      \u003ctd\u003e0.268\u003c/td\u003e\n      \u003ctd\u003e0.439\u003c/td\u003e\n      \u003ctd\u003e0.072\u003c/td\u003e\n      \u003ctd\u003e0.325\u003c/td\u003e\n      \u003ctd\u003e0.318\u003c/td\u003e\n      \u003ctd\u003e0.495\u003c/td\u003e\n      \u003ctd\u003e64.6\u003c/td\u003e\n      \u003ctd\u003e45.9\u003c/td\u003e\n      \u003ctd\u003e75.8\u003c/td\u003e\n      \u003ctd\u003e63.6\u003c/td\u003e\n      \u003ctd\u003e0.6\u003c/td\u003e\n      \u003ctd\u003e0.65\u003c/td\u003e\n      \u003ctd\u003e0.083\u003c/td\u003e\n      \u003ctd\u003e0.284\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eOLMOCR-sglang\u003c/td\u003e\n      \u003ctd\u003e0.326\u003c/td\u003e\n      \u003ctd\u003e0.469\u003c/td\u003e\n      \u003ctd\u003e0.097\u003c/td\u003e\n      \u003ctd\u003e0.293\u003c/td\u003e\n      \u003ctd\u003e0.455\u003c/td\u003e\n      \u003ctd\u003e0.655\u003c/td\u003e\n      \u003ctd\u003e74.3\u003c/td\u003e\n      \u003ctd\u003e43.2\u003c/td\u003e\n      \u003ctd\u003e68.1\u003c/td\u003e\n      \u003ctd\u003e61.3\u003c/td\u003e\n      \u003ctd\u003e0.608\u003c/td\u003e\n      \u003ctd\u003e0.652\u003c/td\u003e\n      \u003ctd\u003e0.145\u003c/td\u003e\n      \u003ctd\u003e0.277\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eSmolDocling-256M_transformer\u003c/td\u003e\n      \u003ctd\u003e0.493\u003c/td\u003e\n      \u003ctd\u003e0.816\u003c/td\u003e\n      \u003ctd\u003e0.262\u003c/td\u003e\n      \u003ctd\u003e0.838\u003c/td\u003e\n      \u003ctd\u003e0.753\u003c/td\u003e\n      \u003ctd\u003e0.997\u003c/td\u003e\n      \u003ctd\u003e32.1\u003c/td\u003e\n      \u003ctd\u003e0.551\u003c/td\u003e\n      \u003ctd\u003e44.9\u003c/td\u003e\n      \u003ctd\u003e16.5\u003c/td\u003e\n      \u003ctd\u003e0.729\u003c/td\u003e\n      \u003ctd\u003e0.907\u003c/td\u003e\n      \u003ctd\u003e0.227\u003c/td\u003e\n      \u003ctd\u003e0.522\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd rowspan=\"8\"\u003eGeneral VLMs\u003c/td\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eGemini2.0-flash\u003c/td\u003e\n      \u003ctd\u003e0.191\u003c/td\u003e\n      \u003ctd\u003e0.264\u003c/td\u003e\n      \u003ctd\u003e0.091\u003c/td\u003e\n      \u003ctd\u003e0.139\u003c/td\u003e\n      \u003ctd\u003e0.389\u003c/td\u003e\n      \u003ctd\u003e0.584\u003c/td\u003e\n      \u003ctd\u003e77.6\u003c/td\u003e\n      \u003ctd\u003e43.6\u003c/td\u003e\n      \u003ctd\u003e79.7\u003c/td\u003e\n      \u003ctd\u003e78.9\u003c/td\u003e\n      \u003ctd\u003e0.193\u003c/td\u003e\n      \u003ctd\u003e0.206\u003c/td\u003e\n      \u003ctd\u003e0.092\u003c/td\u003e\n      \u003ctd\u003e0.128\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eGemini2.5-Pro\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.148\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.212\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.055\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.168\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e0.356\u003c/td\u003e\n      \u003ctd\u003e0.439\u003c/td\u003e\n      \u003ctd\u003e80.0\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e69.4\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e85.8\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e86.4\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.13\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.119\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.049\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.121\u003c/strong\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eGPT4o\u003c/td\u003e\n      \u003ctd\u003e0.233\u003c/td\u003e\n      \u003ctd\u003e0.399\u003c/td\u003e\n      \u003ctd\u003e0.144\u003c/td\u003e\n      \u003ctd\u003e0.409\u003c/td\u003e\n      \u003ctd\u003e0.425\u003c/td\u003e\n      \u003ctd\u003e0.606\u003c/td\u003e\n      \u003ctd\u003e72.8\u003c/td\u003e\n      \u003ctd\u003e42.8\u003c/td\u003e\n      \u003ctd\u003e72.0\u003c/td\u003e\n      \u003ctd\u003e62.9\u003c/td\u003e\n      \u003ctd\u003e0.234\u003c/td\u003e\n      \u003ctd\u003e0.329\u003c/td\u003e\n      \u003ctd\u003e0.128\u003c/td\u003e\n      \u003ctd\u003e0.251\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eQwen2-VL-72B\u003c/td\u003e\n      \u003ctd\u003e0.252\u003c/td\u003e\n      \u003ctd\u003e0.327\u003c/td\u003e\n      \u003ctd\u003e0.096\u003c/td\u003e\n      \u003ctd\u003e0.218\u003c/td\u003e\n      \u003ctd\u003e0.404\u003c/td\u003e\n      \u003ctd\u003e0.487\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e82.2\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e61.2\u003c/td\u003e\n      \u003ctd\u003e76.8\u003c/td\u003e\n      \u003ctd\u003e76.4\u003c/td\u003e\n      \u003ctd\u003e0.387\u003c/td\u003e\n      \u003ctd\u003e0.408\u003c/td\u003e\n      \u003ctd\u003e0.119\u003c/td\u003e\n      \u003ctd\u003e0.193\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eQwen2.5-VL-72B\u003c/td\u003e\n      \u003ctd\u003e0.214\u003c/td\u003e\n      \u003ctd\u003e0.261\u003c/td\u003e\n      \u003ctd\u003e0.092\u003c/td\u003e\n      \u003ctd\u003e0.18\u003c/td\u003e\n      \u003ctd\u003e0.315\u003c/td\u003e\n      \u003ctd\u003e\u003cstrong\u003e0.434\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd\u003e68.8\u003c/td\u003e\n      \u003ctd\u003e62.5\u003c/td\u003e\n      \u003ctd\u003e82.9\u003c/td\u003e\n      \u003ctd\u003e83.9\u003c/td\u003e\n      \u003ctd\u003e0.341\u003c/td\u003e\n      \u003ctd\u003e0.262\u003c/td\u003e\n      \u003ctd\u003e0.106\u003c/td\u003e\n      \u003ctd\u003e0.168\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd\u003eInternVL2-76B\u003c/td\u003e\n      \u003ctd\u003e0.44\u003c/td\u003e\n      \u003ctd\u003e0.443\u003c/td\u003e\n      \u003ctd\u003e0.353\u003c/td\u003e\n      \u003ctd\u003e0.290\u003c/td\u003e\n      \u003ctd\u003e0.543\u003c/td\u003e\n      \u003ctd\u003e0.701\u003c/td\u003e\n      \u003ctd\u003e67.4\u003c/td\u003e\n      \u003ctd\u003e44.1\u003c/td\u003e\n      \u003ctd\u003e63.0\u003c/td\u003e\n      \u003ctd\u003e60.2\u003c/td\u003e\n      \u003ctd\u003e0.547\u003c/td\u003e\n      \u003ctd\u003e0.555\u003c/td\u003e\n      \u003ctd\u003e0.317\u003c/td\u003e\n      \u003ctd\u003e0.228\u003c/td\u003e\n    \u003c/tr\u003e\n  \u003c/tbody\u003e\n\u003c/table\u003e\n\n\u003cp style=\"text-align: center; margin-top: -4pt;\"\u003e\n  Comprehensive evaluation of document parsing algorithms on OmniDocBench: performance metrics for text, formula, table, and reading order extraction, with overall scores derived from ground truth comparisons.\n\u003c/p\u003e\n\n### [olmoOCR eval](https://github.com/allenai/olmocr/tree/main/olmocr/bench)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/allenai/olmocr?label=GitHub\u0026logo=github)](https://github.com/allenai/olmocr)\n![GitHub License](https://img.shields.io/github/license/allenai/olmocr)\n[![Dataset](https://img.shields.io/badge/Dataset-HuggingFace-blue)](https://huggingface.co/datasets/allenai/olmOCR-bench)\n\u003c!--- \nLicense: Apache 2.0 \nPrimary language: Python\n--\u003e\n\nolmOCR-Bench works by testing various \"facts\" about document pages at the PDF-level. Our intention is that each \"fact\" is very simple, \nunambiguous, and machine-checkable, similar to a unit test. For example, once your document has been OCRed, we may check that a\n particular sentence appears exactly somewhere on the page.\n\nDataset Link: https://huggingface.co/datasets/allenai/olmOCR-bench\n\n\n\u003ctable\u003e\n  \u003cthead\u003e\n    \u003ctr\u003e\n      \u003cth align=\"left\"\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/th\u003e\n      \u003cth align=\"center\"\u003eArXiv\u003c/th\u003e\n      \u003cth align=\"center\"\u003eOld Scans Math\u003c/th\u003e\n      \u003cth align=\"center\"\u003eTables\u003c/th\u003e\n      \u003cth align=\"center\"\u003eOld Scans\u003c/th\u003e\n      \u003cth align=\"center\"\u003eHeaders and Footers\u003c/th\u003e\n      \u003cth align=\"center\"\u003eMulti column\u003c/th\u003e\n      \u003cth align=\"center\"\u003eLong tiny text\u003c/th\u003e\n      \u003cth align=\"center\"\u003eBase\u003c/th\u003e\n      \u003cth align=\"center\"\u003eOverall\u003c/th\u003e\n    \u003c/tr\u003e\n  \u003c/thead\u003e\n  \u003ctbody\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eGOT OCR\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e52.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e52.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e0.20\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e22.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e93.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e42.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e29.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e94.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e48.3 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eMarker v1.7.5 (base, force_ocr)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e76.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e57.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e57.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e27.8\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e84.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e72.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e84.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e99.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e70.1 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eMinerU v1.3.10\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e75.4\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e47.4\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e60.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e17.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e96.6\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e59.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e39.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e96.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e61.5 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eMistral OCR API\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e77.2\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e67.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e60.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e29.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e93.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e77.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e99.4\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e72.0 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eNanonets OCR\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e67.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e68.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e77.7\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e39.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e40.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e69.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e53.4\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e99.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e64.5 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eGPT-4o (No Anchor)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e51.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e75.5\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e69.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e40.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e94.2\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e68.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e54.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e96.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e68.9 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eGPT-4o (Anchored)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e53.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e74.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e70.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e40.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e93.8\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e69.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e60.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e96.8\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e69.9 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eGemini Flash 2 (No Anchor)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e32.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e56.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e61.4\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e27.8\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e48.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e58.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e84.4\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e94.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e57.8 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eGemini Flash 2 (Anchored)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e54.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e56.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e72.1\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e34.2\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e64.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e61.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e95.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e63.8 ± 1.2\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eQwen 2 VL (No Anchor)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e19.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e31.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e24.2\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e17.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e88.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e8.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e6.8\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e55.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e31.5 ± 0.9\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eQwen 2.5 VL (No Anchor)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e63.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e65.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e67.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e38.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e73.6\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e68.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e49.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e98.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e65.5 ± 1.2\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eolmOCR v0.1.75 (No Anchor)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.4\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.4\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e42.8\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e94.1\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e77.7\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e97.8\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e74.7 ± 1.1\u003c/td\u003e\n    \u003c/tr\u003e\n    \u003ctr\u003e\n      \u003ctd align=\"left\"\u003eolmOCR v0.1.75 (Anchored)\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e74.9\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.2\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e71.0\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e42.2\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e94.5\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e78.3\u003c/strong\u003e\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e73.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e98.3\u003c/td\u003e\n      \u003ctd align=\"center\"\u003e\u003cstrong\u003e75.5 ± 1.0\u003c/strong\u003e\u003c/td\u003e\n    \u003c/tr\u003e\n  \u003c/tbody\u003e\n\u003c/table\u003e\n\n\nAlso, the olmOCR project provides an **evaluation toolkit** (`runeval.py`) for side-by-side comparison of PDF conversion\npipeline outputs. This tool allows researchers to directly compare text extraction results from different pipeline\nversions against a gold-standard reference. Also olmoOCR authors made some evalutions in\ntheir [technical report](https://olmocr.allenai.org/papers/olmocr.pdf).\n\n\n\u003e We then sampled 2,000 comparison pairs (same PDF, different tool). We asked 11 data researchers and\n\u003e engineers at Ai2 to assess which output was the higher quality representation of the original PDF, focusing on\n\u003e reading order, comprehensiveness of content and representation of structured information. The user interface\n\u003e used is similar to that in Figure 5. Exact participant instructions are listed in Appendix B.\n\n**Bootstrapped Elo Ratings (95% CI)**\n\n| Model   | Elo Rating ± CI | 95% CI Range     |\n| ------- | --------------- | ---------------- |\n| olmoOCR | 1813.0 ± 84.9   | [1605.9, 1930.0] |\n| MinerU  | 1545.2 ± 99.7   | [1336.7, 1714.1] |\n| Marker  | 1429.1 ± 100.7  | [1267.6, 1645.5] |\n| GOTOCOR | 1212.7 ± 82.0   | [1097.3, 1408.3] |\n\n\u003cbr/\u003e\n\n\u003e Table 7: Pairwise Win/Loss Statistics Between Models\n\n| Model Pair         | Wins    | Win Rate (%) |\n| ------------------ | ------- | ------------ |\n| olmOCR vs. Marker  | 49/31   | **61.3**     |\n| olmOCR vs. GOTOCOR | 41/29   | **58.6**     |\n| olmOCR vs. MinerU  | 55/22   | **71.4**     |\n| Marker vs. MinerU  | 53/26   | 67.1         |\n| Marker vs. GOTOCOR | 45/26   | 63.4         |\n| GOTOCOR vs. MinerU | 38/37   | 50.7         |\n| **Total**          | **452** |              |\n\n### [Marker benchmarks](https://github.com/VikParuchuri/marker?tab=readme-ov-file#benchmarks)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/VikParuchuri/marker?label=GitHub\u0026logo=github)](https://github.com/VikParuchuri/marker?tab=readme-ov-file#benchmarks)\n![GitHub License](https://img.shields.io/github/license/VikParuchuri/marker)\n\u003c!--- \nLicense: GPL 3.0\nPrimary language: Python\n--\u003e\n\nThe Marker repository provides benchmark results comparing various PDF processing methods, scored based on a heuristic\nthat aligns text with ground truth text segments, and an LLM as a judge scoring method.\n\n| Method     | Avg Time | Heuristic Score | LLM Score |\n| ---------- | -------- | --------------- | --------- |\n| marker     | 2.83837  | 95.6709         | 4.23916   |\n| llamaparse | 23.348   | 84.2442         | 3.97619   |\n| mathpix    | 6.36223  | 86.4281         | 4.15626   |\n| docling    | 3.69949  | 86.7073         | 3.70429   |\n\n### [READoc](https://arxiv.org/abs/2409.05137)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/icip-cas/READoc?label=GitHub\u0026logo=github)](https://github.com/icip-cas/READoc)\n[![arXiv](https://img.shields.io/badge/arXiv-2409.05137-b31b1b)](https://arxiv.org/abs/2409.05137)\n\n| Methods                    | Text (Concat) | Text (Vocab) | Heading (Concat) | Heading (Tree) | Formula (Embed) | Formula (Isolate) | Table (Concat) | Table (Tree) | Reading Order (Block) | Reading Order (Token) | Average |\n| -------------------------- | ------------- | ------------ | ---------------- | -------------- | --------------- | ----------------- | -------------- | ------------ | --------------------- | --------------------- | ------- |\n| **Baselines**              |               |              |                  |                |                 |                   |                |              |                       |                       |         |\n| PyMuPDF4LLM                | 66.66         | 74.27        | 27.86            | 20.77          | 0.07            | 0.02              | 23.27          | 15.83        | 87.70                 | 89.09                 | 40.55   |\n| Tesseract OCR              | 78.85         | 76.51        | 1.26             | 0.30           | 0.12            | 0.00              | 0.00           | 0.00         | 96.70                 | 97.59                 | 35.13   |\n| **Pipeline Tools**         |               |              |                  |                |                 |                   |                |              |                       |                       |         |\n| MinerU                     | 84.15         | 84.76        | 62.89            | 39.15          | 62.97           | 71.02             | 0.00           | 0.00         | 98.64                 | 97.72                 | 60.17   |\n| Pix2Text                   | 85.85         | 83.72        | 63.23            | 34.53          | 43.18           | 37.45             | 54.08          | 47.35        | 97.68                 | 96.78                 | 64.39   |\n| Marker                     | 83.58         | 81.36        | 68.78            | 54.82          | 5.07            | 56.26             | 47.12          | 43.35        | 98.08                 | 97.26                 | 63.57   |\n| **Expert Visual Models**   |               |              |                  |                |                 |                   |                |              |                       |                       |         |\n| Nougat-small               | 87.35         | 92.00        | 86.40            | 87.88          | 76.52           | 79.39             | 55.63          | 52.35        | 97.97                 | 98.36                 | 81.38   |\n| Nougat-base                | 88.03         | 92.29        | 86.60            | 88.50          | 76.19           | 79.47             | 54.40          | 52.30        | 97.98                 | 98.41                 | 81.42   |\n| **Vision-Language Models** |               |              |                  |                |                 |                   |                |              |                       |                       |         |\n| DeepSeek-VL-7B-Chat        | 31.89         | 39.96        | 23.66            | 12.53          | 17.01           | 16.94             | 22.96          | 16.47        | 88.76                 | 66.75                 | 33.69   |\n| MiniCPM-Llama3-V2.5        | 58.91         | 70.87        | 26.33            | 7.68           | 16.70           | 17.90             | 27.89          | 24.91        | 95.26                 | 93.02                 | 43.95   |\n| LLaVa-1.6-Vicuna-13B       | 27.51         | 37.09        | 8.92             | 6.27           | 17.80           | 11.68             | 23.78          | 16.23        | 76.63                 | 51.68                 | 27.76   |\n| InternVL-Chat-V1.5         | 53.06         | 68.44        | 25.03            | 13.57          | 33.13           | 24.37             | 40.44          | 34.35        | 94.61                 | 91.31                 | 47.83   |\n| GPT-4o-mini                | 79.44         | 84.37        | 31.77            | 18.65          | 42.23           | 41.67             | 47.81          | 39.85        | 97.69                 | 96.35                 | 57.98   |\n\n**Table 3:** Evaluation of various Document Structured Extraction systems on READOC-arXiv.\n\n### [Mistral-OCR benchmarks](https://mistral.ai/news/mistral-ocr)\n\n| Model                | Overall   | Math      | Multilingual | Scanned   | Tables    |\n| -------------------- | --------- | --------- | ------------ | --------- | --------- |\n| Google Document AI   | 83.42     | 80.29     | 86.42        | 92.77     | 78.16     |\n| Azure OCR            | 89.52     | 85.72     | 87.52        | 94.65     | 89.52     |\n| Gemini-1.5-Flash-002 | 90.23     | 89.11     | 86.76        | 94.87     | 90.48     |\n| Gemini-1.5-Pro-002   | 89.92     | 88.48     | 86.33        | 96.15     | 89.71     |\n| Gemini-2.0-Flash-001 | 88.69     | 84.18     | 85.80        | 95.11     | 91.46     |\n| GPT-4o-2024-11-20    | 89.77     | 87.55     | 86.00        | 94.58     | 91.70     |\n| Mistral OCR 2503     | **94.89** | **94.29** | **89.55**    | **98.96** | **96.12** |\n\n### [dp-bench](https://huggingface.co/datasets/upstage/dp-bench)\n\n| Source       | Request date | TEDS ↑ | TEDS-S ↑ | NID ↑ | Avg. Time (secs) ↓ |\n| ------------ | ------------ | ------ | -------- | ----- | ------------------ |\n| upstage      | 2024-10-24   | 93.48  | 94.16    | 97.02 | 3.79               |\n| aws          | 2024-10-24   | 88.05  | 90.79    | 96.71 | 14.47              |\n| llamaparse   | 2024-10-24   | 74.57  | 76.34    | 92.82 | 4.14               |\n| unstructured | 2024-10-24   | 65.56  | 70.00    | 91.18 | 13.14              |\n| google       | 2024-10-24   | 66.13  | 71.58    | 90.86 | 5.85               |\n| microsoft    | 2024-10-24   | 87.19  | 89.75    | 87.69 | 4.44               |\n\n### [Actualize pro](https://www.actualize.pro/recourses/unlocking-insights-from-pdfs-a-comparative-study-of-extraction-tools)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/actualize-ae/pdf-benchmarking?label=GitHub\u0026logo=github)](https://github.com/actualize-ae/pdf-benchmarking)\n\n\u003e In the digital age, PDF documents remain a cornerstone for disseminating and archiving information.\n\u003e However, extracting meaningful data from these structured and unstructured formats continues to challenge modern AI\n\u003e systems.\n\u003e Our recent benchmarking study evaluated seven prominent PDF extraction tools to determine their capabilities across\n\u003e diverse document types and applications.\n\n| PDF Parser   | Overall Score (out of 10) | Text Extraction Accuracy (Score out of 10) | Table Extraction Accuracy (Score out of 10) | Reading Order Accuracy (Score out of 10) | Markdown Conversion Accuracy (Score out of 10) | Code and Math Equations Extraction (Score out of 10) | Image Extraction Accuracy (Score out of 10) |\n| ------------ | ------------------------- | ------------------------------------------ | ------------------------------------------- | ---------------------------------------- | ---------------------------------------------- | ---------------------------------------------------- | ------------------------------------------- |\n| MinerU       | 8                         | 9.3                                        | 7.3                                         | 8.7                                      | 8.3                                            | 6.5                                                  | 7                                           |\n| Xerox        | 7.9                       | 8.7                                        | 7.7                                         | 9                                        | 8.7                                            | 7                                                    | 6                                           |\n| MarkItdown   | 7.78                      | 9                                          | 6.83                                        | 9                                        | 7.67                                           | 7.83                                                 | 5.83                                        |\n| Docling      | 7.3                       | 8.7                                        | 6.3                                         | 9                                        | 8                                              | 6.5                                                  | 5                                           |\n| Llama parse  | 7.1                       | 7.3                                        | 7.7                                         | 8.7                                      | 7.3                                            | 6                                                    | 5.3                                         |\n| Marker       | 6.5                       | 7.3                                        | 5.7                                         | 7.3                                      | 6.7                                            | 4.5                                                  | 6.7                                         |\n| Unstructured | 6.2                       | 7.3                                        | 5                                           | 8.3                                      | 6.7                                            | 5                                                    | 4.7                                         |\n\n### [liduos.com](https://liduos.com/en/ai-develope-tools-series-2-open-source-doucment-parsing.html)\n\n| Function                                          | MinerU | PaddleOCR | Marker | Unstructured | gptpdf | Zerox | Chunkr | pdf-extract-api | Sparrow | LlamaParse | DeepDoc | MegaParse |\n| ------------------------------------------------- | ------ | --------- | ------ | ------------ | ------ | ----- | ------ | --------------- | ------- | ---------- | ------- | --------- |\n| PDF and Image Parsing                             | ✓      | ✓         | ✓      | ✓            | ✓      | ✓     | ✓      | ✓               | ✓       | ✓          | ✓       | ✓         |\n| Parsing of Other Formats (PPT, Excel, DOCX, etc.) | ✓      | -         | -      | ✓            | -      | ✓     | ✓      | -               | ✓       | ✓          | ✓       | ✓         |\n| Layout Analysis                                   | ✓      | ✓         | ✓      | -            | ✓      | -     | ✓      | -               | -       | ✓          | ✓       | -         |\n| Text Recognition                                  | ✓      | ✓         | ✓      | ✓            | ✓      | ✓     | ✓      | ✓               | ✓       | ✓          | ✓       | ✓         |\n| Image Recognition                                 | ✓      | ✓         | ✓      | ✓            | ✓      | ✓     | ✓      | ✓               | ✓       | ✓          | ✓       | ✓         |\n| Simple (Vertical/Horizontal/Hierarchical) Tables  | ✓      | ✓         | ✓      | ✓            | ✓      | ✓     | ✓      | ✓               | ✓       | ✓          | ✓       | ✓         |\n| Complex Tables                                    | -      | -         | -      | -            | -      | -     | -      | -               | -       | -          | -       | -         |\n| Formula Recognition                               | -      | -         | -      | -            | -      | -     | -      | -               | -       | -          | -       | -         |\n| HTML Output                                       | ✓      | -         | ✓      | ✓            | -      | -     | ✓      | -               | -       | -          | ✓       | -         |\n| Markdown Output                                   | ✓      | ✓         | ✓      | -            | ✓      | ✓     | ✓      | ✓               | ✓       | ✓          | -       | ✓         |\n| JSON Output                                       | ✓      | -         | ✓      | ✓            | -      | -     | ✓      | ✓               | -       | ✓          | ✓       | -         |\n\n### [Omni OCR Benchmark](https://getomni.ai/ocr-benchmark)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/yobix-ai/extractous-benchmarks?label=GitHub\u0026logo=github)](https://github.com/getomni-ai/benchmark)\n![GitHub License](https://img.shields.io/github/license/getomni-ai/benchmark)\n\n**JSON Accuracy**\n\n| Model Provider     | JSON Accuracy (%) |\n| ------------------ | ----------------- |\n| OmniAI             | 91.7%             |\n| Gemini 2.0 Flash   | 86.1%             |\n| Azure              | 85.1%             |\n| GPT-4o             | 75.5%             |\n| AWS Textract       | 74.3%             |\n| Claude Sonnet 3.5  | 69.3%             |\n| Google Document AI | 67.8%             |\n| GPT-4o Mini        | 64.8%             |\n| Unstructured       | 50.8%             |\n\n**Cost per 1,000 Pages**\n\n| Model Provider     | Cost per 1,000 Pages ($) |\n| ------------------ | ------------------------ |\n| GPT-4o Mini        | 0.97                     |\n| Gemini 2.0 Flash   | 1.12                     |\n| Google Document AI | 1.50                     |\n| AWS Textract       | 4.00                     |\n| OmniAI             | 10.00                    |\n| Azure              | 10.00                    |\n| GPT-4o             | 18.37                    |\n| Claude Sonnet 3.5  | 19.93                    |\n| Unstructured       | 20.00                    |\n\n**Processing Time per Page**\n\n| Model Provider     | Average Latency (seconds) |\n| ------------------ | ------------------------- |\n| Google Document AI | 3.19                      |\n| Azure              | 4.40                      |\n| AWS Textract       | 4.86                      |\n| Unstructured       | 7.99                      |\n| OmniAI             | 9.69                      |\n| Gemini 2.0 Flash   | 10.71                     |\n| Claude Sonnet 3.5  | 18.42                     |\n| GPT-4o Mini        | 22.73                     |\n| GPT-4o             | 24.85                     |\n\n### [Extractous benchmarks](https://github.com/yobix-ai/extractous-benchmarks/tree/main/docs)\n\n[![GitHub last commit](https://img.shields.io/github/last-commit/yobix-ai/extractous-benchmarks?label=GitHub\u0026logo=github)](https://github.com/yobix-ai/extractous-benchmarks/tree/main/docs)\n![GitHub License](https://img.shields.io/github/license/yobix-ai/extractous-benchmarks)\n\n[`extractous`](https://github.com/yobix-ai/extractous) speedup relative to [\n`unstructured-io`](https://github.com/Unstructured-IO/unstructured)\n\n![image](https://github.com/user-attachments/assets/6d9bc6ba-8e1a-4083-9d6f-864adf854e2f)\n\n[`extractous`](https://github.com/yobix-ai/extractous) memory efficiency relative to [\n`unstructured-io`](https://github.com/Unstructured-IO/unstructured)\n\n![image](https://github.com/user-attachments/assets/e6236232-4fa3-4cd0-8cfa-0bfcd5bc18e3)\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdantetemplar%2Fpdf-extraction-agenda","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdantetemplar%2Fpdf-extraction-agenda","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdantetemplar%2Fpdf-extraction-agenda/lists"}