{"id":13398624,"url":"https://github.com/geekan/scrapy-examples","last_synced_at":"2025-10-10T22:35:32.474Z","repository":{"id":13137888,"uuid":"15820094","full_name":"geekan/scrapy-examples","owner":"geekan","description":"Multifarious Scrapy examples. Spiders for alexa / amazon / douban / douyu / github / linkedin etc.","archived":false,"fork":false,"pushed_at":"2023-11-03T06:14:04.000Z","size":18246,"stargazers_count":3235,"open_issues_count":8,"forks_count":1041,"subscribers_count":232,"default_branch":"master","last_synced_at":"2025-05-08T00:08:41.353Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/geekan.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2014-01-11T09:37:39.000Z","updated_at":"2025-05-07T15:50:44.000Z","dependencies_parsed_at":"2024-04-06T12:46:32.348Z","dependency_job_id":null,"html_url":"https://github.com/geekan/scrapy-examples","commit_stats":null,"previous_names":[],"tags_count":1,"template":false,"template_full_name":null,"purl":"pkg:github/geekan/scrapy-examples","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/geekan%2Fscrapy-examples","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/geekan%2Fscrapy-examples/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/geekan%2Fscrapy-examples/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/geekan%2Fscrapy-examples/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/geekan","download_url":"https://codeload.github.com/geekan/scrapy-examples/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/geekan%2Fscrapy-examples/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279005460,"owners_count":26083902,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-10T02:00:06.843Z","response_time":62,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-30T19:00:29.507Z","updated_at":"2025-10-10T22:35:32.458Z","avatar_url":"https://github.com/geekan.png","language":"Python","funding_links":[],"categories":["Python","\u003ca id=\"cf11bcd58b4ec7a549bfb11297003180\"\u003e\u003c/a\u003e爬虫"],"sub_categories":[],"readme":"scrapy-examples\n==============\n\nMultifarious scrapy examples with integrated proxies and agents, which make you comfy to write a spider.\n\nDon't use it to do anything illegal!\n\n***\n\n## Real spider example: doubanbook\n\n#### Tutorial\n\n    git clone https://github.com/geekan/scrapy-examples\n    cd scrapy-examples/doubanbook\n    scrapy crawl doubanbook\n\n#### Depth\n\nThere are several depths in the spider, and the spider gets\nreal data from depth2.\n\n- Depth0: The entrance is `http://book.douban.com/tag/`\n- Depth1: Urls like `http://book.douban.com/tag/外国文学` from depth0\n- Depth2: Urls like `http://book.douban.com/subject/1770782/` from depth1\n\n#### Example image\n![douban book](https://raw.githubusercontent.com/geekan/scrapy-examples/master/doubanbook/sample.jpg)\n\n***\n\n## Avaiable Spiders\n\n* tutorial\n  * dmoz_item\n  * douban_book\n  * page_recorder\n  * douban_tag_book\n* doubanbook\n* linkedin\n* hrtencent\n* sis\n* zhihu\n* alexa\n  * alexa\n  * alexa.cn\n\n## Advanced\n\n* Use `parse_with_rules` to write a spider quickly.  \n  See dmoz spider for more details.\n\n* Proxies\n  * If you don't want to use proxy, just comment the proxy middleware in settings.  \n  * If you want to custom it, hack `misc/proxy.py` by yourself.  \n\n* Notice\n  * Don't use `parse` as your method name, it's an inner method of CrawlSpider.\n\n### Advanced Usage\n\n* Run `./startproject.sh \u003cPROJECT\u003e` to start a new project.  \n  It will automatically generate most things, the only left things are:\n  * `PROJECT/PROJECT/items.py`\n  * `PROJECT/PROJECT/spider/spider.py`\n\n#### Example to hack `items.py` and `spider.py`\n\nHacked `items.py` with additional fields `url` and `description`:  \n```\nfrom scrapy.item import Item, Field\n\nclass exampleItem(Item):\n    url = Field()\n    name = Field()\n    description = Field()\n```\n\nHacked `spider.py` with start rules and css rules (here only display the class exampleSpider):  \n```\nclass exampleSpider(CommonSpider):\n    name = \"dmoz\"\n    allowed_domains = [\"dmoz.org\"]\n    start_urls = [\n        \"http://www.dmoz.com/\",\n    ]\n    # Crawler would start on start_urls, and follow the valid urls allowed by below rules.\n    rules = [\n        Rule(sle(allow=[\"/Arts/\", \"/Games/\"]), callback='parse', follow=True),\n    ]\n\n    css_rules = {\n        '.directory-url li': {\n            '__use': 'dump', # dump data directly\n            '__list': True, # it's a list\n            'url': 'li \u003e a::attr(href)',\n            'name': 'a::text',\n            'description': 'li::text',\n        }\n    }\n\n    def parse(self, response):\n        info('Parse '+response.url)\n        # parse_with_rules is implemented here:\n        #   https://github.com/geekan/scrapy-examples/blob/master/misc/spider.py\n        self.parse_with_rules(response, self.css_rules, exampleItem)\n```\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgeekan%2Fscrapy-examples","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fgeekan%2Fscrapy-examples","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fgeekan%2Fscrapy-examples/lists"}