{"id":25406204,"url":"https://github.com/NoEdgeAI/pdfdeal","last_synced_at":"2025-10-31T02:30:26.762Z","repository":{"id":241686341,"uuid":"807150738","full_name":"NoEdgeAI/pdfdeal","owner":"NoEdgeAI","description":"A python wrapper for the Doc2X API and comes with native texts processing (to improve PDF recall in RAG). | Doc2X API的python封装，同时附带本地的文本处理(提升PDF在RAG中的召回率)。","archived":false,"fork":false,"pushed_at":"2025-02-10T15:49:26.000Z","size":509,"stargazers_count":221,"open_issues_count":1,"forks_count":12,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-02-10T16:33:35.725Z","etag":null,"topics":["doc2x","ocr","pdf","rag"],"latest_commit_sha":null,"homepage":"https://noedgeai.github.io/pdfdeal-docs/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/NoEdgeAI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-05-28T15:11:11.000Z","updated_at":"2025-02-06T12:53:40.000Z","dependencies_parsed_at":"2024-06-13T06:47:11.586Z","dependency_job_id":"35616983-7ceb-4e82-9dac-6fe154dca496","html_url":"https://github.com/NoEdgeAI/pdfdeal","commit_stats":null,"previous_names":["menghuan1918/pdfdeal","noedgeai/pdfdeal"],"tags_count":38,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NoEdgeAI%2Fpdfdeal","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NoEdgeAI%2Fpdfdeal/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NoEdgeAI%2Fpdfdeal/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/NoEdgeAI%2Fpdfdeal/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/NoEdgeAI","download_url":"https://codeload.github.com/NoEdgeAI/pdfdeal/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":239088390,"owners_count":19579435,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["doc2x","ocr","pdf","rag"],"created_at":"2025-02-16T05:08:39.039Z","updated_at":"2025-10-31T02:30:21.382Z","avatar_url":"https://github.com/NoEdgeAI.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"\u003cdiv align=center\u003e\n\u003ch1 aligh=\"center\"\u003e\n\u003cimg src=\"https://github.com/Menghuan1918/pdfdeal/assets/122662527/837cfd7f-4546-4b44-a199-d826d78784fc\" width=\"45\"\u003e  pdfdeal\n\u003c/h1\u003e\n\u003ca href=\"https://github.com/Menghuan1918/pdfdeal/actions/workflows/python-test.yml\"\u003e\n  \u003cimg src=\"https://github.com/Menghuan1918/pdfdeal/actions/workflows/python-test.yml/badge.svg\" \n  alt=\"Package Testing on Python 3.8-3.13 on Win/Linux/macOS\"\u003e\n\u003c/a\u003e\n\u003cbr\u003e\n\u003cbr\u003e\n\n[![Downloads](https://static.pepy.tech/badge/pdfdeal)](https://pepy.tech/project/pdfdeal) ![GitHub License](https://img.shields.io/github/license/Menghuan1918/pdfdeal) ![PyPI - Version](https://img.shields.io/pypi/v/pdfdeal) ![GitHub Repo stars](https://img.shields.io/github/stars/Menghuan1918/pdfdeal)\n\n\u003cbr\u003e\n\n[📄Documentation](https://menghuan1918.github.io/pdfdeal-docs/guide/)\n\n\u003cbr\u003e\n\n🗺️ ENGLISH | [简体中文](README_CN.md)\n\n\u003c/div\u003e\n\nHandle PDF more easily and simply, utilizing Doc2X's powerful document conversion capabilities for retained format file conversion/RAG enhancement.\n\n\u003cdiv align=center\u003e\n\u003cimg src=\"https://github.com/user-attachments/assets/3db3c682-84f1-4712-bd70-47422616f393\" width=\"500px\"\u003e\n\u003c/div\u003e\n\n## Introduction\n\n### Doc2X Support\n\n[Doc2X](https://doc2x.com/) is a new universal document OCR tool that can convert images or PDF files into Markdown/LaTeX text with formulas and text formatting. It performs better than similar tools in most scenarios. `pdfdeal` provides abstract packaged classes to use Doc2X for requests.\n\n### Processing PDFs\n\nUse various OCR or PDF recognition tools to identify images and add them to the original text. You can set the output format to use PDF, which will ensure that the recognized text retains the same page numbers as the original in the new PDF. It also offers various practical file processing tools.\n\nAfter conversion and pre-processing of PDF using Doc2X, you can achieve better recognition rates when used with knowledge base applications such as [graphrag](https://github.com/microsoft/graphrag), [Dify](https://github.com/langgenius/dify), and [FastGPT](https://github.com/labring/FastGPT).\n\n### Markdown Document Processing Features\n\n`pdfdeal` also provides a series of powerful tools to handle Markdown documents:\n\n- **Convert HTML tables to Markdown format**: Allows conversion of HTML formatted tables to Markdown format for easy use in Markdown documents.\n- **Upload images to remote storage services**: Supports uploading local or online images in Markdown documents to remote storage services to ensure image persistence and accessibility.\n- **Convert online images to local images**: Allows downloading and converting online images in Markdown documents to local images for offline use.\n- **Document splitting and separator addition**: Supports splitting Markdown documents by headings or adding separators within documents for better organization and management.\n\nFor detailed feature introduction and usage, please refer to the [documentation link](https://menghuan1918.github.io/pdfdeal-docs/guide/Tools/).\n\n\n## Cases\n\n### graphrag\n\nSee [how to use it with graphrag](https://menghuan1918.github.io/pdfdeal-docs/demo/graphrag.html), [its not supported to recognize pdf](https://github.com/microsoft/graphrag), but you can use the CLI tool `doc2x` to convert it to a txt document for use.\n\n\u003cdiv align=center\u003e\n\u003cimg src=\"https://github.com/user-attachments/assets/f9e8408b-9a4b-42b9-9aee-0d1229065a91\" width=\"600px\"\u003e\n\u003c/div\u003e\n\n### Fastgpt/Dify or other RAG system\n\nOr for knowledge base applications, you can use `pdfdeal`'s built-in variety of enhancements to documents, such as uploading images to remote storage services, adding breaks by paragraph, etc. See [Integration with RAG applications](https://menghuan1918.github.io/pdfdeal-docs/demo/RAG_pre.html).\n\n\u003cdiv align=center\u003e\n\u003cimg src=\"https://github.com/user-attachments/assets/034d3eb0-d77e-4f7d-a707-9be08a092a9a\" width=\"450px\"\u003e\n\u003cimg src=\"https://github.com/user-attachments/assets/6078e585-7c06-485f-bcd3-9fac84eb7301\" width=\"450px\"\u003e\n\u003c/div\u003e\n\n## Documentation\n\nFor details, please refer to the [documentation](https://menghuan1918.github.io/pdfdeal-docs/)\n\nOr check out the [documentation repository pdfdeal-docs](https://github.com/Menghuan1918/pdfdeal-docs).\n\n## Quick Start\n\nFor details, please refer to the [documentation](https://menghuan1918.github.io/pdfdeal-docs/)\n\n### Installation\n\nInstall using pip:\n\n```bash\npip install --upgrade pdfdeal\n```\n\nIf you need [document processing tools](https://menghuan1918.github.io/pdfdeal-docs/guide/Tools/):\n\n```bash\npip install --upgrade \"pdfdeal[rag]\"\n```\n\n### Use the Doc2X PDF API to process all PDF files in a specified folder\n\n```python\nfrom pdfdeal import Doc2X\n\nclient = Doc2X(apikey=\"Your API key\",debug=True)\nsuccess, failed, flag = client.pdf2file(\n    pdf_file=\"tests/pdf\",\n    output_path=\"./Output\",\n    output_format=\"docx\",\n)\nprint(success)\nprint(failed)\nprint(flag)\n```\n\n### Use the Doc2X PDF API to process the specified PDF file and specify the name of the exported file\n\n```python\nfrom pdfdeal import Doc2X\n\nclient = Doc2X(apikey=\"Your API key\",debug=True)\nsuccess, failed, flag = client.pdf2file(\n    pdf_file=\"tests/pdf/sample.pdf\",\n    output_path=\"./Output/test/single/pdf2file\",\n    output_names=[\"sample1.zip\"],\n    output_format=\"md_dollar\",\n)\nprint(success)\nprint(failed)\nprint(flag)\n```\n\nSee the online documentation for details.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FNoEdgeAI%2Fpdfdeal","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FNoEdgeAI%2Fpdfdeal","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FNoEdgeAI%2Fpdfdeal/lists"}