{"id":26700062,"url":"https://github.com/samparsky/web-crawler","last_synced_at":"2025-08-16T16:45:57.440Z","repository":{"id":203498819,"uuid":"84051963","full_name":"samparsky/web-crawler","owner":"samparsky","description":"This a site crawler built with scrapy and stores data generated in mongodb using scrapy","archived":false,"fork":false,"pushed_at":"2017-03-06T10:39:13.000Z","size":13,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-03-26T23:18:37.953Z","etag":null,"topics":["extruct","scrapy","scrapy-demo"],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/samparsky.png","metadata":{"files":{"readme":"readme.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2017-03-06T09:03:04.000Z","updated_at":"2017-03-06T10:45:12.000Z","dependencies_parsed_at":null,"dependency_job_id":"91ff039f-81f2-4ee4-98e6-56d42f54017d","html_url":"https://github.com/samparsky/web-crawler","commit_stats":null,"previous_names":["samparsky/web-crawler"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/samparsky/web-crawler","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/samparsky%2Fweb-crawler","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/samparsky%2Fweb-crawler/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/samparsky%2Fweb-crawler/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/samparsky%2Fweb-crawler/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/samparsky","download_url":"https://codeload.github.com/samparsky/web-crawler/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/samparsky%2Fweb-crawler/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":270741933,"owners_count":24637491,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-08-16T02:00:11.002Z","response_time":91,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["extruct","scrapy","scrapy-demo"],"created_at":"2025-03-26T23:18:40.196Z","updated_at":"2025-08-16T16:45:56.955Z","avatar_url":"https://github.com/samparsky.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"## Simple Web Crawler\n-----------------\n\nThis is a simple web crawler that crawls\n[link](https://mommypoppins.com/events?area%5B%5D=118\u0026field_event_date_value%5B%5D=03-04-2017\u0026event_end=2017-04-07). Parses through the results page. It works based on the (Scrapy)[https://scrapy.org/] crawling engine. Its uses Extruct to parse application/ld+json content of the pages to retrieve basic content and Xpath to query the  \n\n### To start \n------------\n```sh\n\npip install -r requirements.txt\n\n````\n\n### To run the crawler\n\n```sh\n cd \u003cdirectory\u003e\n scrapy crawl wizard\n\n````\n\n### MongoDB\n\nThe mongodb collection schema is as follows\n\n```python\n    event_name  \n    description \n    age_group    \n    location     \n    price        \n    link\t\t \n    event_link \n    date\n```\n\nThe mongodb database is `mommy` and the collection is `crawl`\nTo view the crawled data run the below commands at the mongo shell\n\n```sh\n \u003e use mommy\n \u003e db.crawl.find()\n\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsamparsky%2Fweb-crawler","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsamparsky%2Fweb-crawler","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsamparsky%2Fweb-crawler/lists"}