{"id":24518363,"url":"https://github.com/schafeld/crawler","last_synced_at":"2025-06-10T11:36:45.488Z","repository":{"id":73840530,"uuid":"182534469","full_name":"schafeld/crawler","owner":"schafeld","description":"Website crawler, Python ","archived":false,"fork":false,"pushed_at":"2019-04-21T14:00:04.000Z","size":4,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-06-02T22:34:16.641Z","etag":null,"topics":["python","python3"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/schafeld.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-04-21T13:06:41.000Z","updated_at":"2019-04-21T14:00:05.000Z","dependencies_parsed_at":"2024-01-20T06:45:11.174Z","dependency_job_id":null,"html_url":"https://github.com/schafeld/crawler","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/schafeld%2Fcrawler","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/schafeld%2Fcrawler/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/schafeld%2Fcrawler/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/schafeld%2Fcrawler/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/schafeld","download_url":"https://codeload.github.com/schafeld/crawler/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/schafeld%2Fcrawler/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":259068242,"owners_count":22800470,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["python","python3"],"created_at":"2025-01-22T01:42:04.107Z","updated_at":"2025-06-10T11:36:45.416Z","avatar_url":"https://github.com/schafeld.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# crawler\nWebsite crawler, Python \n\n## Installation\n\nCheck out from Github\n\nInstall dependencies/modules\n\n```  pip install beautifulsoup4 ``` \nor (if like me you have Python 2.x and Python 3.x)\n```  pip3 install bs4 ``` \n```  pip3 install requests ``` \n```  pip3 install lxml ``` \n\nExecution\n```python3 crawler.py``` \n\n\n## Acknowledgment\nThanks for source/inspiration: https://medium.freecodecamp.org/how-to-build-a-url-crawler-to-map-a-website-using-python-6a287be1da11 \n\n## Final words\nUse this script at your own risk. Crawling websites puts load on them and may be seen as aggressive act, i.e. can get you banned, blocked or worse. As always in life: Don't be rude to people!\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fschafeld%2Fcrawler","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fschafeld%2Fcrawler","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fschafeld%2Fcrawler/lists"}