{"id":29482812,"url":"https://github.com/Topdu/OpenOCR","last_synced_at":"2025-07-15T02:02:21.922Z","repository":{"id":242091869,"uuid":"808665467","full_name":"Topdu/OpenOCR","owner":"Topdu","description":"OpenOCR: A general OCR system with accuracy and efficiency. Supporting 24 Scene Text Recognition methods trained from scratch on large-scale real datasets, and will continue to add the latest methods.","archived":false,"fork":false,"pushed_at":"2025-06-08T04:12:00.000Z","size":11699,"stargazers_count":691,"open_issues_count":77,"forks_count":58,"subscribers_count":13,"default_branch":"main","last_synced_at":"2025-07-04T12:15:13.659Z","etag":null,"topics":["chineseocr","ocr","ocr-pytorch","scene-text-detection","scene-text-recognition"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Topdu.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-05-31T14:47:47.000Z","updated_at":"2025-07-04T08:10:02.000Z","dependencies_parsed_at":"2024-06-04T09:28:19.989Z","dependency_job_id":"4aac81f1-affe-4d37-8b7a-0f6dc7380b2d","html_url":"https://github.com/Topdu/OpenOCR","commit_stats":null,"previous_names":["topdu/openocr"],"tags_count":1,"template":false,"template_full_name":null,"purl":"pkg:github/Topdu/OpenOCR","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Topdu%2FOpenOCR","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Topdu%2FOpenOCR/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Topdu%2FOpenOCR/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Topdu%2FOpenOCR/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Topdu","download_url":"https://codeload.github.com/Topdu/OpenOCR/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Topdu%2FOpenOCR/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":265386079,"owners_count":23756747,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["chineseocr","ocr","ocr-pytorch","scene-text-detection","scene-text-recognition"],"created_at":"2025-07-15T02:01:46.416Z","updated_at":"2025-07-15T02:02:21.887Z","avatar_url":"https://github.com/Topdu.png","language":"Python","funding_links":[],"categories":["others","Tools \u0026 Libraries","Python","OCR and Screen Reading"],"sub_categories":["Open Source OCR Systems"],"readme":"\u003cdiv align=\"center\"\u003e\n\n\u003ch1\u003e OpenOCR: A general OCR system with accuracy and efficiency \u003c/h1\u003e\n\n\u003ch5 align=\"center\"\u003e If you find this project useful, please give us a star🌟. \u003c/h5\u003e\n\n\u003ca href=\"https://github.com/Topdu/OpenOCR/blob/main/LICENSE\"\u003e\u003cimg alt=\"license\" src=\"https://img.shields.io/github/license/Topdu/OpenOCR\"\u003e\u003c/a\u003e\n\u003ca href='https://arxiv.org/abs/2411.15858'\u003e\u003cimg src='https://img.shields.io/badge/Paper-Arxiv-red'\u003e\u003c/a\u003e\n\u003ca href=\"https://huggingface.co/spaces/topdu/OpenOCR-Demo\" target=\"_blank\"\u003e\u003cimg src=\"https://img.shields.io/badge/%F0%9F%A4%97-Hugging Face Demo-blue\"\u003e\u003c/a\u003e\n\u003ca href=\"https://modelscope.cn/studios/topdktu/OpenOCR-Demo\" target=\"_blank\"\u003e\u003cimg src=\"https://img.shields.io/badge/魔搭-Demo-blue\"\u003e\u003c/a\u003e\n\u003ca href=\"\"\u003e\u003cimg src=\"https://img.shields.io/badge/OS-Linux%2C%20Win%2C%20Mac-pink.svg\"\u003e\u003c/a\u003e\n\u003ca href=\"https://github.com/Topdu/OpenOCR/graphs/contributors\"\u003e\u003cimg src=\"https://img.shields.io/github/contributors/Topdu/OpenOCR?color=9ea\"\u003e\u003c/a\u003e\n\u003ca href=\"https://pepy.tech/project/openocr\"\u003e\u003cimg src=\"https://static.pepy.tech/personalized-badge/openocr?period=total\u0026units=abbreviation\u0026left_color=grey\u0026right_color=blue\u0026left_text=Clone%20downloads\"\u003e\u003c/a\u003e\n\u003ca href=\"https://github.com/Topdu/OpenOCR/stargazers\"\u003e\u003cimg src=\"https://img.shields.io/github/stars/Topdu/OpenOCR?color=ccf\"\u003e\u003c/a\u003e\n\u003ca href=\"https://pypi.org/project/openocr-python/\"\u003e\u003cimg alt=\"PyPI\" src=\"https://img.shields.io/pypi/v/openocr-python\"\u003e\u003cimg src=\"https://img.shields.io/pypi/dm/openocr-python?label=PyPI%20downloads\"\u003e\u003c/a\u003e\n\n\u003ca href=\"#quick-start\"\u003e 🚀 Quick Start \u003c/a\u003e | English | [简体中文](./README_ch.md)\n\n\u003c/div\u003e\n\n______________________________________________________________________\n\nWe aim to establish a unified benchmark for training and evaluating models in scene text detection and recognition. Building on this benchmark, we introduce a general OCR system with accuracy and efficiency, **OpenOCR**. This repository also serves as the official codebase of the OCR team from the [FVL Laboratory](https://fvl.fudan.edu.cn), Fudan University.\n\nWe sincerely welcome the researcher to recommend OCR or relevant algorithms and point out any potential factual errors or bugs. Upon receiving the suggestions, we will promptly evaluate and critically reproduce them. We look forward to collaborating with you to advance the development of OpenOCR and continuously contribute to the OCR community!\n\n## Features\n\n- 🔥**OpenOCR: A general OCR system with accuracy and efficiency**\n  - ⚡\\[[Quick Start](#quick-start)\\] \\[[Model](https://github.com/Topdu/OpenOCR/releases/tag/develop0.0.1)\\] \\[[ModelScope Demo](https://modelscope.cn/studios/topdktu/OpenOCR-Demo)\\] \\[[Hugging Face Demo](https://huggingface.co/spaces/topdu/OpenOCR-Demo)\\] \\[[Local Demo](#local-demo)\\]  \\[[PaddleOCR Implementation](https://paddlepaddle.github.io/PaddleOCR/latest/algorithm/text_recognition/algorithm_rec_svtrv2.html)\\]\n  - [Introduction](./docs/openocr.md)\n    - A practical OCR system building on SVTRv2.\n    - Outperforms [PP-OCRv4](https://paddlepaddle.github.io/PaddleOCR/latest/ppocr/model_list.html) baseline by 4.5% on the [OCR competition leaderboard](https://aistudio.baidu.com/competition/detail/1131/0/leaderboard) in terms of accuracy, while preserving quite similar inference speed.\n    - [x] Supports Chinese and English text detection and recognition.\n    - [x] Provides server model and mobile model.\n    - [x] Fine-tunes OpenOCR on a custom dataset: [Fine-tuning Det](./docs/finetune_det.md), [Fine-tuning Rec](./docs/finetune_rec.md).\n    - [x] [ONNX model export for wider compatibility](#export-onnx-model).\n- 🔥**SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition (ICCV 2025)**\n  - \\[[Paper](https://arxiv.org/abs/2411.15858)\\] \\[[Doc](./configs/rec/svtrv2/)\\] \\[[Model](./configs/rec/svtrv2/readme.md#11-models-and-results)\\] \\[[Datasets](./docs/svtrv2.md#downloading-datasets)\\] \\[[Config, Training and Inference](./configs/rec/svtrv2/readme.md#3-model-training--evaluation)\\] \\[[Benchmark](./docs/svtrv2.md#results-benchmark--configs--checkpoints)\\]\n  - [Introduction](./docs/svtrv2.md)\n    - A unified training and evaluation benchmark (on top of [Union14M](https://github.com/Mountchicken/Union14M?tab=readme-ov-file#3-union14m-dataset)) for Scene Text Recognition\n    - Supports 24 Scene Text Recognition methods trained from scratch on the large-scale real dataset [Union14M-L-Filter](./docs/svtrv2.md#dataset-details), and will continue to add the latest methods.\n    - Improves accuracy by 20-30% compared to models trained based on synthetic datasets.\n    - Towards Arbitrary-Shaped Text Recognition and Language modeling with a Single Visual Model.\n    - Surpasses Attention-based Encoder-Decoder Methods across challenging scenarios in terms of accuracy and speed\n  - [Get Started](./docs/svtrv2.md#get-started-with-training-a-sota-scene-text-recognition-model-from-scratch) with training a SOTA Scene Text Recognition model from scratch.\n\n## Ours STR algorithms\n\n- [**SVTRv2**](./configs/rec/svtrv2) (*Yongkun Du, Zhineng Chen\\*, Hongtao Xie, Caiyan Jia, Yu-Gang Jiang. SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition,* ICCV 2025. [Doc](./configs/rec/svtrv2/), [Paper](https://arxiv.org/abs/2411.15858))\n- [**IGTR**](./configs/rec/igtr/) (*Yongkun Du, Zhineng Chen\\*, Yuchen Su, Caiyan Jia, Yu-Gang Jiang. Instruction-Guided Scene Text Recognition,* TPAMI 2025. [Doc](./configs/rec/igtr), [Paper](https://ieeexplore.ieee.org/document/10820836))\n- [**CPPD**](./configs/rec/cppd/) (*Yongkun Du, Zhineng Chen\\*, Caiyan Jia, Xiaoting Yin, Chenxia Li, Yuning Du, Yu-Gang Jiang. Context Perception Parallel Decoder for Scene Text Recognition,* TPAMI 2025. [PaddleOCR Doc](https://github.com/PaddlePaddle/PaddleOCR/blob/main/docs/algorithm/text_recognition/algorithm_rec_cppd.en.md), [Paper](https://ieeexplore.ieee.org/document/10902187))\n- [**SMTR\u0026FocalSVTR**](./configs/rec/smtr/) (*Yongkun Du, Zhineng Chen\\*, Caiyan Jia, Xieping Gao, Yu-Gang Jiang. Out of Length Text Recognition with Sub-String Matching,* AAAI 2025. [Doc](./configs/rec/smtr/), [Paper](https://ojs.aaai.org/index.php/AAAI/article/view/32285))\n- [**DPTR**](./configs/rec/dptr/) (*Shuai Zhao, Yongkun Du, Zhineng Chen\\*, Yu-Gang Jiang. Decoder Pre-Training with only Text for Scene Text Recognition,* ACM MM 2024. [Paper](https://dl.acm.org/doi/10.1145/3664647.3681390))\n- [**CDistNet**](./configs/rec/cdistnet/) (*Tianlun Zheng, Zhineng Chen\\*, Shancheng Fang, Hongtao Xie, Yu-Gang Jiang. CDistNet: Perceiving Multi-Domain Character Distance for Robust Text Recognition,* IJCV 2024. [Paper](https://link.springer.com/article/10.1007/s11263-023-01880-0))\n- **MRN** (*Tianlun Zheng, Zhineng Chen\\*, Bingchen Huang, Wei Zhang, Yu-Gang Jiang. MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition,* ICCV 2023. [Paper](https://openaccess.thecvf.com/content/ICCV2023/html/Zheng_MRN_Multiplexed_Routing_Network_for_Incremental_Multilingual_Text_Recognition_ICCV_2023_paper.html), [Code](https://github.com/simplify23/MRN))\n- **TPS++** (*Tianlun Zheng, Zhineng Chen\\*, Jinfeng Bai, Hongtao Xie, Yu-Gang Jiang. TPS++: Attention-Enhanced Thin-Plate Spline for Scene Text Recognition,* IJCAI 2023. [Paper](https://arxiv.org/abs/2305.05322), [Code](https://github.com/simplify23/TPS_PP))\n- [**SVTR**](./configs/rec/svtr/) (*Yongkun Du, Zhineng Chen\\*, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, Yu-Gang Jiang. SVTR: Scene Text Recognition with a Single Visual Model,* IJCAI 2022 (Long). [PaddleOCR Doc](https://github.com/Topdu/PaddleOCR/blob/main/doc/doc_ch/algorithm_rec_svtr.md), [Paper](https://www.ijcai.org/proceedings/2022/124))\n- [**NRTR**](./configs/rec/nrtr/) (*Fenfen Sheng, Zhineng Chen, Bo Xu. NRTR: A No-Recurrence Sequence-to-Sequence Model For Scene Text Recognition,* ICDAR 2019. [Paper](https://arxiv.org/abs/1806.00926))\n\n## Recent Updates\n\n- **2025.07.10**: Our paper [SVTRv2](https://arxiv.org/abs/2411.15858) is accepted by ICCV 2025. Accessible in [Doc](./configs/rec/svtrv2/).\n\n- **2025.03.24**: 🔥 Releasing the feature of fine-tuning OpenOCR on a custom dataset: [Fine-tuning Det](./docs/finetune_det.md), [Fine-tuning Rec](./docs/finetune_rec.md)\n\n- **2025.03.23**: 🔥 Releasing the feature of [ONNX model export for wider compatibility](#export-onnx-model).\n\n- **2025.02.22**: Our paper [CPPD](https://ieeexplore.ieee.org/document/10902187) is accepted by TPAMI. Accessible in [Doc](./configs/rec/cppd/) and [PaddleOCR Doc](https://github.com/PaddlePaddle/PaddleOCR/blob/main/docs/algorithm/text_recognition/algorithm_rec_cppd.en.md).\n\n- **2024.12.31**: Our paper [IGTR](https://ieeexplore.ieee.org/document/10820836) is accepted by TPAMI. Accessible in [Doc](./configs/rec/igtr/).\n\n- **2024.12.16**: Our paper [SMTR](https://ojs.aaai.org/index.php/AAAI/article/view/32285) is accepted by AAAI 2025. Accessible in [Doc](./configs/rec/smtr/).\n\n- **2024.12.03**: The pre-training code for [DPTR](https://dl.acm.org/doi/10.1145/3664647.3681390) is merged.\n\n- **🔥 2024.11.23 release notes**:\n\n  - **OpenOCR: A general OCR system with accuracy and efficiency**\n    - ⚡\\[[Quick Start](#quick-start)\\] \\[[Model](https://github.com/Topdu/OpenOCR/releases/tag/develop0.0.1)\\] \\[[ModelScope Demo](https://modelscope.cn/studios/topdktu/OpenOCR-Demo)\\] \\[[Hugging Face Demo](https://huggingface.co/spaces/topdu/OpenOCR-Demo)\\] \\[[Local Demo](#local-demo)\\]  \\[[PaddleOCR Implementation](https://paddlepaddle.github.io/PaddleOCR/latest/algorithm/text_recognition/algorithm_rec_svtrv2.html)\\]\n    - [Introduction](./docs/openocr.md)\n  - **SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition**\n    - \\[[Paper](https://arxiv.org/abs/2411.15858)\\] \\[[Doc](./configs/rec/svtrv2/)\\] \\[[Model](./configs/rec/svtrv2/readme.md#11-models-and-results)\\] \\[[Datasets](./docs/svtrv2.md#downloading-datasets)\\] \\[[Config, Training and Inference](./configs/rec/svtrv2/readme.md#3-model-training--evaluation)\\] \\[[Benchmark](./docs/svtrv2.md#results--configs--checkpoints)\\]\n    - [Introduction](./docs/svtrv2.md)\n    - [Get Started](./docs/svtrv2.md#get-started-with-training-a-sota-scene-text-recognition-model-from-scratch) with training a SOTA Scene Text Recognition model from scratch.\n\n## Quick Start\n\n**Note**: OpenOCR supports inference using both the ONNX and Torch frameworks, with the dependency environments for the two frameworks being isolated. When using ONNX for inference, there is no need to install Torch, and vice versa.\n\n### 1. ONNX Inference\n\n#### Install OpenOCR and Dependencies:\n\n```shell\npip install openocr-python\npip install onnxruntime\n```\n\n#### Usage:\n\n```python\nfrom openocr import OpenOCR\nonnx_engine = OpenOCR(backend='onnx', device='cpu')\nimg_path = '/path/img_path or /path/img_file'\nresult, elapse = onnx_engine(img_path)\n```\n\n### 2. Pytorch inference\n\n#### Dependencies:\n\n- [PyTorch](http://pytorch.org/) version \u003e= 1.13.0\n- Python version \u003e= 3.7\n\n```shell\nconda create -n openocr python==3.8\nconda activate openocr\n# install gpu version torch\nconda install pytorch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 pytorch-cuda=11.8 -c pytorch -c nvidia\n# or cpu version\nconda install pytorch torchvision torchaudio cpuonly -c pytorch\n```\n\nAfter installing dependencies, the following two installation methods are available. Either one can be chosen.\n\n#### 2.1. Python Modules\n\n**Install OpenOCR**:\n\n```shell\npip install openocr-python\n```\n\n**Usage**:\n\n```python\nfrom openocr import OpenOCR\nengine = OpenOCR()\nimg_path = '/path/img_path or /path/img_file'\nresult, elapse = engine(img_path)\n\n# Server mode\n# engine = OpenOCR(mode='server')\n```\n\n#### 2.2. Clone this repository:\n\n```shell\ngit clone https://github.com/Topdu/OpenOCR.git\ncd OpenOCR\npip install -r requirements.txt\nwget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_det_repvit_ch.pth\nwget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_repsvtr_ch.pth\n# Rec Server model\n# wget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/openocr_svtrv2_ch.pth\n```\n\n**Usage**:\n\n```shell\n# OpenOCR system: Det + Rec model\npython tools/infer_e2e.py --img_path=/path/img_fold or /path/img_file\n# Det model\npython tools/infer_det.py --c ./configs/det/dbnet/repvit_db.yml --o Global.infer_img=/path/img_fold or /path/img_file\n# Rec model\npython tools/infer_rec.py --c ./configs/rec/svtrv2/repsvtr_ch.yml --o Global.infer_img=/path/img_fold or /path/img_file\n```\n\n##### Export ONNX model\n\n```shell\npip install onnx\npython tools/toonnx.py --c configs/rec/svtrv2/repsvtr_ch.yml --o Global.device=cpu\npython tools/toonnx.py --c configs/det/dbnet/repvit_db.yml --o Global.device=cpu\n```\n\n##### Inference with ONNXRuntime\n\n```shell\npip install onnxruntime\n# OpenOCR system: Det + Rec model\npython tools/infer_e2e.py --img_path=/path/img_fold or /path/img_file --backend=onnx --device=cpu\n# Det model\npython tools/infer_det.py --c ./configs/det/dbnet/repvit_db.yml --o Global.backend=onnx Global.device=cpu Global.infer_img=/path/img_fold or /path/img_file\n# Rec model\npython tools/infer_rec.py --c ./configs/rec/svtrv2/repsvtr_ch.yml --o Global.backend=onnx Global.device=cpu Global.infer_img=/path/img_fold or /path/img_file\n```\n\n#### Local Demo\n\n```shell\npip install gradio==4.20.0\nwget https://github.com/Topdu/OpenOCR/releases/download/develop0.0.1/OCR_e2e_img.tar\ntar xf OCR_e2e_img.tar\n# start demo\npython demo_gradio.py\n```\n\n## Reproduction schedule:\n\n### Scene Text Recognition\n\n| Method                                        | Venue                                                                                          | Training | Evaluation | Contributor                                 |\n| --------------------------------------------- | ---------------------------------------------------------------------------------------------- | -------- | ---------- | ------------------------------------------- |\n| [CRNN](./configs/rec/svtrs/)                  | [TPAMI 2016](https://arxiv.org/abs/1507.05717)                                                 | ✅       | ✅         |                                             |\n| [ASTER](./configs/rec/aster/)                 | [TPAMI 2019](https://ieeexplore.ieee.org/document/8395027)                                     | ✅       | ✅         | [pretto0](https://github.com/pretto0)       |\n| [NRTR](./configs/rec/nrtr/)                   | [ICDAR 2019](https://arxiv.org/abs/1806.00926)                                                 | ✅       | ✅         |                                             |\n| [SAR](./configs/rec/sar/)                     | [AAAI 2019](https://aaai.org/papers/08610-show-attend-and-read-a-simple-and-strong-baseline-for-irregular-text-recognition/) | ✅       | ✅         | [pretto0](https://github.com/pretto0)       |\n| [MORAN](./configs/rec/moran/)                 | [PR 2019](https://www.sciencedirect.com/science/article/abs/pii/S0031320319300263)             | ✅       | ✅         |                                             |\n| [DAN](./configs/rec/dan/)                     | [AAAI 2020](https://arxiv.org/pdf/1912.10205)                                                  | ✅       | ✅         |                                             |\n| [RobustScanner](./configs/rec/robustscanner/) | [ECCV 2020](https://www.ecva.net/papers/eccv_2020/papers_ECCV/html/3160_ECCV_2020_paper.php)   | ✅       | ✅         | [pretto0](https://github.com/pretto0)       |\n| [AutoSTR](./configs/rec/autostr/)             | [ECCV 2020](https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123690732.pdf)            | ✅       | ✅         |                                             |\n| [SRN](./configs/rec/srn/)                     | [CVPR 2020](https://openaccess.thecvf.com/content_CVPR_2020/html/Yu_Towards_Accurate_Scene_Text_Recognition_With_Semantic_Reasoning_Networks_CVPR_2020_paper.html) | ✅       | ✅         | [pretto0](https://github.com/pretto0)       |\n| [SEED](./configs/rec/seed/)                   | [CVPR 2020](https://openaccess.thecvf.com/content_CVPR_2020/html/Qiao_SEED_Semantics_Enhanced_Encoder-Decoder_Framework_for_Scene_Text_Recognition_CVPR_2020_paper.html) | ✅       | ✅         |                                             |\n| [ABINet](./configs/rec/abinet/)               | [CVPR 2021](https://openaccess.thecvf.com//content/CVPR2021/html/Fang_Read_Like_Humans_Autonomous_Bidirectional_and_Iterative_Language_Modeling_for_CVPR_2021_paper.html) | ✅       | ✅         | [YesianRohn](https://github.com/YesianRohn) |\n| [VisionLAN](./configs/rec/visionlan/)         | [ICCV 2021](https://openaccess.thecvf.com/content/ICCV2021/html/Wang_From_Two_to_One_A_New_Scene_Text_Recognizer_With_ICCV_2021_paper.html) | ✅       | ✅         | [YesianRohn](https://github.com/YesianRohn) |\n| PIMNet                                        | [ACM MM 2021](https://dl.acm.org/doi/10.1145/3474085.3475238)                                  |          |            | TODO                                        |\n| [SVTR](./configs/rec/svtrs/)                  | [IJCAI 2022](https://www.ijcai.org/proceedings/2022/124)                                       | ✅       | ✅         |                                             |\n| [PARSeq](./configs/rec/parseq/)               | [ECCV 2022](https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136880177.pdf)            | ✅       | ✅         |                                             |\n| [MATRN](./configs/rec/matrn/)                 | [ECCV 2022](https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136880442.pdf)            | ✅       | ✅         |                                             |\n| [MGP-STR](./configs/rec/mgpstr/)              | [ECCV 2022](https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136880336.pdf)            | ✅       | ✅         |                                             |\n| [LPV](./configs/rec/lpv/)                     | [IJCAI 2023](https://www.ijcai.org/proceedings/2023/0189.pdf)                                  | ✅       | ✅         |                                             |\n| [MAERec](./configs/rec/maerec/)(Union14M)     | [ICCV 2023](https://openaccess.thecvf.com/content/ICCV2023/papers/Jiang_Revisiting_Scene_Text_Recognition_A_Data_Perspective_ICCV_2023_paper.pdf) | ✅       | ✅         |                                             |\n| [LISTER](./configs/rec/lister/)               | [ICCV 2023](https://openaccess.thecvf.com/content/ICCV2023/papers/Cheng_LISTER_Neighbor_Decoding_for_Length-Insensitive_Scene_Text_Recognition_ICCV_2023_paper.pdf) | ✅       | ✅         |                                             |\n| [CDistNet](./configs/rec/cdistnet/)           | [IJCV 2024](https://link.springer.com/article/10.1007/s11263-023-01880-0)                      | ✅       | ✅         | [YesianRohn](https://github.com/YesianRohn) |\n| [BUSNet](./configs/rec/busnet/)               | [AAAI 2024](https://ojs.aaai.org/index.php/AAAI/article/view/28402)                            | ✅       | ✅         |                                             |\n| DCTC                                          | [AAAI 2024](https://ojs.aaai.org/index.php/AAAI/article/view/28575)                            |          |            | TODO                                        |\n| [CAM](./configs/rec/cam/)                     | [PR 2024](https://arxiv.org/abs/2402.13643)                                                    | ✅       | ✅         |                                             |\n| [OTE](./configs/rec/ote/)                     | [CVPR 2024](https://openaccess.thecvf.com/content/CVPR2024/html/Xu_OTE_Exploring_Accurate_Scene_Text_Recognition_Using_One_Token_CVPR_2024_paper.html) | ✅       | ✅         |                                             |\n| CFF                                           | [IJCAI 2024](https://arxiv.org/abs/2407.05562)                                                 |          |            | TODO                                        |\n| [DPTR](./configs/rec/dptr/)                   | [ACM MM 2024](https://dl.acm.org/doi/10.1145/3664647.3681390)                                  |          |            | [fd-zs](https://github.com/fd-zs)           |\n| VIPTR                                         | [ACM CIKM 2024](https://arxiv.org/abs/2401.10110)                                              |          |            | TODO                                        |\n| [IGTR](./configs/rec/igtr/)                   | [TPAMI 2025](https://ieeexplore.ieee.org/document/10820836)                                    | ✅       | ✅         |                                             |\n| [SMTR](./configs/rec/smtr/)                   | [AAAI 2025](https://ojs.aaai.org/index.php/AAAI/article/view/32285)                            | ✅       | ✅         |                                             |\n| [CPPD](./configs/rec/cppd/)                   | [TPAMI 2025](https://ieeexplore.ieee.org/document/10902187)                                    | ✅       | ✅         |                                             |\n| [FocalSVTR-CTC](./configs/rec/svtrs/)         | [AAAI 2025](https://ojs.aaai.org/index.php/AAAI/article/view/32285)                            | ✅       | ✅         |                                             |\n| [SVTRv2](./configs/rec/svtrv2/)               | [ICCV 2025](https://arxiv.org/abs/2411.15858)                                                  | ✅       | ✅         |                                             |\n| [ResNet+Trans-CTC](./configs/rec/svtrs/)      |                                                                                                | ✅       | ✅         |                                             |\n| [ViT-CTC](./configs/rec/svtrs/)               |                                                                                                | ✅       | ✅         |                                             |\n\n#### Contributors\n\n______________________________________________________________________\n\nYiming Lei ([pretto0](https://github.com/pretto0)), Xingsong Ye ([YesianRohn](https://github.com/YesianRohn)), and Shuai Zhao ([fd-zs](https://github.com/fd-zs)) from the [FVL Laboratory](https://fvl.fudan.edu.cn), Fudan University, with guidance from Dr. Zhineng Chen ([Homepage](https://zhinchenfd.github.io/)), completed the majority work of the algorithm reproduction. Grateful for their outstanding contributions.\n\n### Scene Text Detection (STD)\n\nTODO\n\n### Text Spotting\n\nTODO\n\n______________________________________________________________________\n\n## Citation\n\nIf you find our method useful for your reserach, please cite:\n\n```bibtex\n@inproceedings{Du2024SVTRv2,\n      title={SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition},\n      author={Yongkun Du and Zhineng Chen and Hongtao Xie and Caiyan Jia and Yu-Gang Jiang},\n      booktitle={ICCV},\n      year={2025}\n}\n```\n\n# Acknowledgement\n\nThis codebase is built based on the [PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR), [PytorchOCR](https://github.com/WenmuZhou/PytorchOCR), and [MMOCR](https://github.com/open-mmlab/mmocr). Thanks for their awesome work!\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FTopdu%2FOpenOCR","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FTopdu%2FOpenOCR","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FTopdu%2FOpenOCR/lists"}