{"id":13411223,"url":"https://github.com/adbar/trafilatura","last_synced_at":"2025-12-24T17:27:53.089Z","repository":{"id":38206633,"uuid":"180136168","full_name":"adbar/trafilatura","owner":"adbar","description":"Python \u0026 command-line tool to gather text on the Web: web crawling/scraping, extraction of text, metadata, comments","archived":false,"fork":false,"pushed_at":"2024-06-11T10:40:24.000Z","size":34084,"stargazers_count":3090,"open_issues_count":61,"forks_count":232,"subscribers_count":30,"default_branch":"master","last_synced_at":"2024-06-11T20:20:20.887Z","etag":null,"topics":["article-extractor","corpus","corpus-builder","corpus-tools","crawler","html-to-markdown","html2text","news","news-aggregator","news-crawler","nlp","readability","rss-feed","scraping","tei","text-cleaning","text-extraction","text-mining","text-preprocessing","web-scraping"],"latest_commit_sha":null,"homepage":"https://trafilatura.readthedocs.io","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/adbar.png","metadata":{"files":{"readme":"README.md","changelog":"HISTORY.md","contributing":"CONTRIBUTING.md","funding":".github/FUNDING.yml","license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null},"funding":{"github":null,"patreon":null,"open_collective":null,"ko_fi":"adbarbaresi","tidelift":null,"community_bridge":null,"liberapay":null,"issuehunt":null,"otechie":null,"custom":null}},"created_at":"2019-04-08T11:38:48.000Z","updated_at":"2024-08-20T12:00:23.517Z","dependencies_parsed_at":"2024-05-28T16:20:13.935Z","dependency_job_id":"094bed18-c30b-4b59-bdcb-3c6842efc9b7","html_url":"https://github.com/adbar/trafilatura","commit_stats":{"total_commits":1395,"total_committers":39,"mean_commits":35.76923076923077,"dds":"0.10179211469534055","last_synced_commit":"85cd3d8aa1d349e931bdb247fc041c67f1a66f2b"},"previous_names":[],"tags_count":37,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adbar%2Ftrafilatura","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adbar%2Ftrafilatura/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adbar%2Ftrafilatura/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/adbar%2Ftrafilatura/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/adbar","download_url":"https://codeload.github.com/adbar/trafilatura/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":242756856,"owners_count":20180206,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["article-extractor","corpus","corpus-builder","corpus-tools","crawler","html-to-markdown","html2text","news","news-aggregator","news-crawler","nlp","readability","rss-feed","scraping","tei","text-cleaning","text-extraction","text-mining","text-preprocessing","web-scraping"],"created_at":"2024-07-30T20:01:12.276Z","updated_at":"2025-12-24T17:27:53.080Z","avatar_url":"https://github.com/adbar.png","language":"Python","funding_links":["https://ko-fi.com/adbarbaresi","https://github.com/sponsors/adbar"],"categories":["Python","📝 Content \u0026 Text Extraction","网络服务","Web Scraping \u0026 Crawling","Building","web-scraping","🕸️ Web Scraping \u0026 Crawling","Web Scraping","转换工具"],"sub_categories":["Ruby","网络爬虫","Tools","开发组件"],"readme":"# Trafilatura: Discover and Extract Text Data on the Web\n\n\u003cbr/\u003e\n\n\u003cimg alt=\"Trafilatura Logo\" src=\"https://raw.githubusercontent.com/adbar/trafilatura/master/docs/trafilatura-logo.png\" align=\"center\" width=\"60%\"/\u003e\n\n\u003cbr/\u003e\n\n[![Python package](https://img.shields.io/pypi/v/trafilatura.svg)](https://pypi.python.org/pypi/trafilatura)\n[![Python versions](https://img.shields.io/pypi/pyversions/trafilatura.svg)](https://pypi.python.org/pypi/trafilatura)\n[![Documentation Status](https://readthedocs.org/projects/trafilatura/badge/?version=latest)](http://trafilatura.readthedocs.org/en/latest/?badge=latest)\n[![Code Coverage](https://img.shields.io/codecov/c/github/adbar/trafilatura.svg)](https://codecov.io/gh/adbar/trafilatura)\n[![Downloads](https://static.pepy.tech/badge/trafilatura/month)](https://pepy.tech/project/trafilatura)\n[![Reference DOI: 10.18653/v1/2021.acl-demo.15](https://img.shields.io/badge/DOI-10.18653%2Fv1%2F2021.acl--demo.15-blue)](https://aclanthology.org/2021.acl-demo.15/)\n\n\u003cbr/\u003e\n\n\u003cimg alt=\"Demo as GIF image\" src=\"https://raw.githubusercontent.com/adbar/trafilatura/master/docs/trafilatura-demo.gif\" align=\"center\" width=\"80%\"/\u003e\n\n\u003cbr/\u003e\n\n\n## Introduction\n\nTrafilatura is a cutting-edge **Python package and command-line tool**\ndesigned to **gather text on the Web and simplify the process of turning\nraw HTML into structured, meaningful data**. It includes all necessary\ndiscovery and text processing components to perform **web crawling,\ndownloads, scraping, and extraction** of main texts, metadata and\ncomments. It aims at staying **handy and modular**: no database is\nrequired, the output can be converted to commonly used formats.\n\nGoing from HTML bulk to essential parts can alleviate many problems\nrelated to text quality, by **focusing on the actual content**,\n**avoiding the noise** caused by recurring elements like headers and footers\nand by **making sense of the data and metadata** with selected information.\nThe extractor strikes a balance between limiting noise (precision) and\nincluding all valid parts (recall). It is **robust and reasonably fast**.\n\nTrafilatura is [widely used](https://trafilatura.readthedocs.io/en/latest/used-by.html)\nand integrated into [thousands of projects](https://github.com/adbar/trafilatura/network/dependents)\nby companies like HuggingFace, IBM, and Microsoft Research as well as institutions like\nthe Allen Institute, Stanford, the Tokyo Institute of Technology, and\nthe University of Munich.\n\n\n### Features\n\n- Advanced web crawling and text discovery:\n   - Support for sitemaps (TXT, XML) and feeds (ATOM, JSON, RSS)\n   - Smart crawling and URL management (filtering and deduplication)\n\n- Parallel processing of online and offline input:\n   - Live URLs, efficient and polite processing of download queues\n   - Previously downloaded HTML files and parsed HTML trees\n\n- Robust and configurable extraction of key elements:\n   - Main text (common patterns and generic algorithms like jusText and readability)\n   - Metadata (title, author, date, site name, categories and tags)\n   - Formatting and structure: paragraphs, titles, lists, quotes, code, line breaks, in-line text formatting\n   - Optional elements: comments, links, images, tables\n\n- Multiple output formats:\n   - TXT and Markdown\n   - CSV\n   - JSON\n   - HTML, XML and [XML-TEI](https://tei-c.org/)\n\n- Optional add-ons:\n   - Language detection on extracted content\n   - Speed optimizations\n\n- Actively maintained with support from the open-source community:\n   - Regular updates, feature additions, and optimizations\n   - Comprehensive documentation\n\n\n### Evaluation and alternatives\n\nTrafilatura consistently outperforms other open-source libraries in text\nextraction benchmarks, showcasing its efficiency and accuracy in\nextracting web content. The extractor tries to strike a balance between\nlimiting noise and including all valid parts.\n\nFor more information see the [benchmark section](https://trafilatura.readthedocs.io/en/latest/evaluation.html)\nand the [evaluation readme](https://github.com/adbar/trafilatura/blob/master/tests/README.rst)\nto run the evaluation with the latest data and packages.\n\n\n#### Other evaluations:\n\n- Most efficient open-source library in *ScrapingHub*'s [article extraction benchmark](https://github.com/scrapinghub/article-extraction-benchmark)\n- Best overall tool according to [Bien choisir son outil d'extraction de contenu à partir du Web](https://hal.archives-ouvertes.fr/hal-02768510v3/document)\n  (Lejeune \u0026 Barbaresi 2020)\n- Best single tool by ROUGE-LSum Mean F1 Page Scores in [An Empirical Comparison of Web Content Extraction Algorithms](https://webis.de/downloads/publications/papers/bevendorff_2023b.pdf)\n  (Bevendorff et al. 2023)\n\n\n## Usage and documentation\n\n[Getting started with Trafilatura](https://trafilatura.readthedocs.io/en/latest/quickstart.html)\nis straightforward. For more information and detailed guides, visit\n[Trafilatura's documentation](https://trafilatura.readthedocs.io/):\n\n- [Installation](https://trafilatura.readthedocs.io/en/latest/installation.html)\n- Usage:\n  [On the command-line](https://trafilatura.readthedocs.io/en/latest/usage-cli.html),\n  [With Python](https://trafilatura.readthedocs.io/en/latest/usage-python.html),\n  [With R](https://trafilatura.readthedocs.io/en/latest/usage-r.html)\n- [Core Python functions](https://trafilatura.readthedocs.io/en/latest/corefunctions.html)\n- Interactive Python Notebook: [Trafilatura Overview](docs/Trafilatura_Overview.ipynb)\n- [Tutorials and use cases](https://trafilatura.readthedocs.io/en/latest/tutorials.html)\n\nYoutube playlist with video tutorials in several languages:\n\n- [Web scraping tutorials and how-tos](https://www.youtube.com/watch?v=8GkiOM17t0Q\u0026list=PL-pKWbySIRGMgxXQOtGIz1-nbfYLvqrci)\n\n\n## License\n\nThis package is distributed under the [Apache 2.0 license](https://www.apache.org/licenses/LICENSE-2.0.html).\n\nVersions prior to v1.8.0 are under GPLv3+ license.\n\n\n### Contributing\n\nContributions of all kinds are welcome. Visit the [Contributing\npage](https://github.com/adbar/trafilatura/blob/master/CONTRIBUTING.md)\nfor more information. Bug reports can be filed on the [dedicated issue\npage](https://github.com/adbar/trafilatura/issues).\n\nMany thanks to the\n[contributors](https://github.com/adbar/trafilatura/graphs/contributors)\nwho extended the docs or submitted bug reports, features and bugfixes!\n\n\n## Context\n\nThis work started as a PhD project at the crossroads of linguistics and\nNLP, this expertise has been instrumental in shaping Trafilatura over\nthe years. Initially launched to create text databases for research purposes\nat the Berlin-Brandenburg Academy of Sciences (DWDS and ZDL units),\nthis package continues to be maintained but its future depends on community support.\n\n**If you value this software or depend on it for your product, consider\nsponsoring it and contributing to its codebase**. Your support\n[on GitHub](https://github.com/sponsors/adbar) or [ko-fi.com](https://ko-fi.com/adbarbaresi)\nwill help maintain and enhance this popular package.\n\n*Trafilatura* is an Italian word for [wire\ndrawing](https://en.wikipedia.org/wiki/Wire_drawing) symbolizing the\nrefinement and conversion process. It is also the way shapes of pasta\nare formed.\n\n### Author\n\nReach out via ia the software repository or the [contact\npage](https://adrien.barbaresi.eu/) for inquiries, collaborations, or\nfeedback. See also social networks for the latest updates.\n\n-   Barbaresi, A. [Trafilatura: A Web Scraping Library and Command-Line\n    Tool for Text Discovery and\n    Extraction](https://aclanthology.org/2021.acl-demo.15/), Proceedings\n    of ACL/IJCNLP 2021: System Demonstrations, 2021, p. 122-131.\n-   Barbaresi, A. \"[Generic Web Content Extraction with Open-Source\n    Software](https://hal.archives-ouvertes.fr/hal-02447264/document)\",\n    Proceedings of KONVENS 2019, Kaleidoscope Abstracts, 2019.\n-   Barbaresi, A. \"[Efficient construction of metadata-enhanced web\n    corpora](https://hal.archives-ouvertes.fr/hal-01371704v2/document)\",\n    Proceedings of the [10th Web as Corpus Workshop\n    (WAC-X)](https://www.sigwac.org.uk/wiki/WAC-X), 2016.\n\n\n### Citing Trafilatura\n\nTrafilatura is widely used in the academic domain, chiefly for data\nacquisition. Here is how to cite it:\n\n[![Reference DOI: 10.18653/v1/2021.acl-demo.15](https://img.shields.io/badge/DOI-10.18653%2Fv1%2F2021.acl--demo.15-blue)](https://aclanthology.org/2021.acl-demo.15/)\n[![Zenodo archive DOI: 10.5281/zenodo.3460969](https://zenodo.org/badge/DOI/10.5281/zenodo.3460969.svg)](https://doi.org/10.5281/zenodo.3460969)\n\n``` shell\n@inproceedings{barbaresi-2021-trafilatura,\n  title = {{Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction}},\n  author = \"Barbaresi, Adrien\",\n  booktitle = \"Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations\",\n  pages = \"122--131\",\n  publisher = \"Association for Computational Linguistics\",\n  url = \"https://aclanthology.org/2021.acl-demo.15\",\n  year = 2021,\n}\n```\n\n\n### Software ecosystem\n\nJointly developed plugins and additional packages also contribute to the\nfield of web data extraction and analysis:\n\n\u003cimg alt=\"Software ecosystem\" src=\"https://raw.githubusercontent.com/adbar/htmldate/master/docs/software-ecosystem.png\" align=\"center\" width=\"65%\"/\u003e\n\nCorresponding posts can be found on [Bits of\nLanguage](https://adrien.barbaresi.eu/blog/tag/trafilatura.html).\n\nImpressive, you have reached the end of the page: Thank you for your\ninterest!\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fadbar%2Ftrafilatura","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fadbar%2Ftrafilatura","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fadbar%2Ftrafilatura/lists"}