{"id":13634519,"url":"https://github.com/seomoz/reppy","last_synced_at":"2025-04-18T15:31:58.930Z","repository":{"id":46562242,"uuid":"2668665","full_name":"seomoz/reppy","owner":"seomoz","description":"Modern robots.txt Parser for Python","archived":false,"fork":false,"pushed_at":"2024-01-12T05:07:07.000Z","size":469,"stargazers_count":194,"open_issues_count":26,"forks_count":41,"subscribers_count":112,"default_branch":"master","last_synced_at":"2025-04-12T21:46:50.502Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/seomoz.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2011-10-29T00:06:46.000Z","updated_at":"2025-04-03T21:21:35.000Z","dependencies_parsed_at":"2024-06-18T18:20:04.227Z","dependency_job_id":"692f46a2-40a5-423f-a7b1-0617a47578cf","html_url":"https://github.com/seomoz/reppy","commit_stats":{"total_commits":201,"total_committers":20,"mean_commits":10.05,"dds":0.7512437810945274,"last_synced_commit":"cb57131799ff2ba7e263a12c82861c8169dc8c70"},"previous_names":[],"tags_count":23,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/seomoz%2Freppy","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/seomoz%2Freppy/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/seomoz%2Freppy/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/seomoz%2Freppy/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/seomoz","download_url":"https://codeload.github.com/seomoz/reppy/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":249514108,"owners_count":21284470,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-01T23:01:06.767Z","updated_at":"2025-04-18T15:31:58.622Z","avatar_url":"https://github.com/seomoz.png","language":"Python","funding_links":[],"categories":["Others"],"sub_categories":[],"readme":"Robots Exclusion Protocol Parser for Python\n===========================================\n\n[![Build Status](https://travis-ci.org/seomoz/reppy.svg?branch=master)](https://travis-ci.org/seomoz/reppy)\n\n`Robots.txt` parsing in Python.\n\nGoals\n=====\n\n- __Fetching__ -- helper utilities for fetching and parsing `robots.txt`s, including\n    checking `cache-control` and `expires` headers\n- __Support for newer features__ -- like `Crawl-Delay` and `Sitemaps`\n- __Wildcard matching__ -- without using regexes, no less\n- __Performance__ -- with \u003e100k parses per second, \u003e1M URL checks per second once parsed\n- __Caching__ -- utilities to help with the caching of `robots.txt` responses\n\nInstallation\n============\n`reppy` is available on `pypi`:\n\n```bash\npip install reppy\n```\n\nWhen installing from source, there are submodule dependencies that must also be fetched:\n\n```bash\ngit submodule update --init --recursive\nmake install\n```\n\nUsage\n=====\n\nChecking when pages are allowed\n-------------------------------\nTwo classes answer questions about whether a URL is allowed: `Robots` and\n`Agent`:\n\n```python\nfrom reppy.robots import Robots\n\n# This utility uses `requests` to fetch the content\nrobots = Robots.fetch('http://example.com/robots.txt')\nrobots.allowed('http://example.com/some/path/', 'my-user-agent')\n\n# Get the rules for a specific agent\nagent = robots.agent('my-user-agent')\nagent.allowed('http://example.com/some/path/')\n```\n\nThe `Robots` class also exposes properties `expired` and `ttl` to describe how\nlong the response should be considered valid. A `reppy.ttl` policy is used to\ndetermine what that should be:\n\n```python\nfrom reppy.ttl import HeaderWithDefaultPolicy\n\n# Use the `cache-control` or `expires` headers, defaulting to a 30 minutes and\n# ensuring it's at least 10 minutes\npolicy = HeaderWithDefaultPolicy(default=1800, minimum=600)\n\nrobots = Robots.fetch('http://example.com/robots.txt', ttl_policy=policy)\n```\n\nCustomizing fetch\n-----------------\nThe `fetch` method accepts `*args` and `**kwargs` that are passed on to `requests.get`,\nallowing you to customize the way the `fetch` is executed:\n\n```python\nrobots = Robots.fetch('http://example.com/robots.txt', headers={...})\n```\n\nMatching Rules and Wildcards\n----------------------------\nBoth `*` and `$` are supported for wildcard matching.\n\nThis library follows the matching [1996 RFC](http://www.robotstxt.org/norobots-rfc.txt)\ndescribes. In the case where multiple rules match a query, the longest rules wins as\nit is presumed to be the most specific.\n\nChecking sitemaps\n-----------------\nThe `Robots` class also lists the sitemaps that are listed in a `robots.txt`\n\n```python\n# This property holds a list of URL strings of all the sitemaps listed\nrobots.sitemaps\n```\n\nDelay\n-----\nThe `Crawl-Delay` directive is per agent and can be accessed through that class. If\nnone was specified, it's `None`:\n\n```python\n# What's the delay my-user-agent should use\nrobots.agent('my-user-agent').delay\n```\n\nDetermining the `robots.txt` URL\n--------------------------------\nGiven a URL, there's a utility to determine the URL of the corresponding `robots.txt`.\nIt preserves the scheme and hostname and the port (if it's not the default port for the\nscheme).\n\n```python\n# Get robots.txt URL for http://userinfo@example.com:8080/path;params?query#fragment\n# It's http://example.com:8080/robots.txt\nRobots.robots_url('http://userinfo@example.com:8080/path;params?query#fragment')\n```\n\nCaching\n=======\nThere are two cache classes provided -- `RobotsCache`, which caches entire `reppy.Robots`\nobjects, and `AgentCache`, which only caches the `reppy.Agent` relevant to a client. These\ncaches duck-type the class that they cache for the purposes of checking if a URL is\nallowed:\n\n```python\nfrom reppy.cache import RobotsCache\ncache = RobotsCache(capacity=100)\ncache.allowed('http://example.com/foo/bar', 'my-user-agent')\n\nfrom reppy.cache import AgentCache\ncache = AgentCache(agent='my-user-agent', capacity=100)\ncache.allowed('http://example.com/foo/bar')\n```\n\nLike `reppy.Robots.fetch`, the cache constructory accepts a `ttl_policy` to inform the\nexpiration of the fetched `Robots` objects, as well as `*args` and `**kwargs` to be passed\nto `reppy.Robots.fetch`.\n\nCaching Failures\n----------------\nThere's a piece of classic caching advice: \"don't cache failures.\" However, this is not\nalways appropriate in certain circumstances. For example, if the failure is a timeout,\nclients may want to cache this result so that every check doesn't take a very long time.\n\nTo this end, the `cache` module provides a notion of a cache policy. It determines what\nto do in the case of an exception. The default is to cache a form of a disallowed response\nfor 10 minutes, but you can configure it as you see fit:\n\n```python\n# Do not cache failures (note the `ttl=0`):\nfrom reppy.cache.policy import ReraiseExceptionPolicy\ncache = AgentCache('my-user-agent', cache_policy=ReraiseExceptionPolicy(ttl=0))\n\n# Cache and reraise failures for 10 minutes (note the `ttl=600`):\ncache = AgentCache('my-user-agent', cache_policy=ReraiseExceptionPolicy(ttl=600))\n\n# Treat failures as being disallowed\ncache = AgentCache(\n    'my-user-agent',\n    cache_policy=DefaultObjectPolicy(ttl=600, lambda _: Agent().disallow('/')))\n```\n\nDevelopment\n===========\nA `Vagrantfile` is provided to bootstrap a development environment:\n\n```bash\nvagrant up\n```\n\nAlternatively, development can be conducted using a `virtualenv`:\n\n```bash\nvirtualenv venv\nsource venv/bin/activate\npip install -r requirements.txt\n```\n\nTests\n=====\nTests may be run in `vagrant`:\n\n```bash\nmake test\n```\n\nDevelopment\n===========\n\nEnvironment\n-----------\nTo launch the `vagrant` image, we only need to\n`vagrant up` (though you may have to provide a `--provider` flag):\n\n```bash\nvagrant up\n```\n\nWith a running `vagrant` instance, you can log in and run tests:\n\n```bash\nvagrant ssh\nmake test\n```\n\nRunning Tests\n-------------\nTests are run with the top-level `Makefile`:\n\n```bash\nmake test\n```\n\nPRs\n===\nThese are not all hard-and-fast rules, but in general PRs have the following expectations:\n\n- __pass Travis__ -- or more generally, whatever CI is used for the particular project\n- __be a complete unit__ -- whether a bug fix or feature, it should appear as a complete\n    unit before consideration.\n- __maintain code coverage__ -- some projects may include code coverage requirements as\n    part of the build as well\n- __maintain the established style__ -- this means the existing style of established\n    projects, the established conventions of the team for a given language on new\n    projects, and the guidelines of the community of the relevant languages and\n    frameworks.\n- __include failing tests__ -- in the case of bugs, failing tests demonstrating the bug\n    should be included as one commit, followed by a commit making the test succeed. This\n    allows us to jump to a world with a bug included, and prove that our test in fact\n    exercises the bug.\n- __be reviewed by one or more developers__ -- not all feedback has to be accepted, but\n    it should all be considered.\n- __avoid 'addressed PR feedback' commits__ -- in general, PR feedback should be rebased\n    back into the appropriate commits that introduced the change. In cases, where this\n    is burdensome, PR feedback commits may be used but should still describe the changed\n    contained therein.\n\nPR reviews consider the design, organization, and functionality of the submitted code.\n\nCommits\n=======\nCertain types of changes should be made in their own commits to improve readability. When\ntoo many different types of changes happen simultaneous to a single commit, the purpose of\neach change is muddled. By giving each commit a single logical purpose, it is implicitly\nclear why changes in that commit took place.\n\n- __updating / upgrading dependencies__ -- this is especially true for invocations like\n    `bundle update` or `berks update`.\n- __introducing a new dependency__ -- often preceeded by a commit updating existing\n    dependencies, this should only include the changes for the new dependency.\n- __refactoring__ -- these commits should preserve all the existing functionality and\n    merely update how it's done.\n- __utility components to be used by a new feature__ -- if introducing an auxiliary class\n    in support of a subsequent commit, add this new class (and its tests) in its own\n    commit.\n- __config changes__ -- when adjusting configuration in isolation\n- __formatting / whitespace commits__ -- when adjusting code only for stylistic purposes.\n\nNew Features\n------------\nSmall new features (where small refers to the size and complexity of the change, not the\nimpact) are often introduced in a single commit. Larger features or components might be\nbuilt up piecewise, with each commit containing a single part of it (and its corresponding\ntests).\n\nBug Fixes\n---------\nIn general, bug fixes should come in two-commit pairs: a commit adding a failing test\ndemonstrating the bug, and a commit making that failing test pass.\n\nTagging and Versioning\n======================\nWhenever the version included in `setup.py` is changed (and it should be changed when\nappropriate using [http://semver.org/](http://semver.org/)), a corresponding tag should\nbe created with the same version number (formatted `v\u003cversion\u003e`).\n\n```bash\ngit tag -a v0.1.0 -m 'Version 0.1.0\n\nThis release contains an initial working version of the `crawl` and `parse`\nutilities.'\n\ngit push --tags origin\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fseomoz%2Freppy","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fseomoz%2Freppy","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fseomoz%2Freppy/lists"}