{"id":13532502,"url":"https://github.com/dpapathanasiou/recipebook","last_synced_at":"2025-04-13T11:05:55.990Z","repository":{"id":144914189,"uuid":"81031685","full_name":"dpapathanasiou/recipebook","owner":"dpapathanasiou","description":"This is a simple application for scraping and parsing food recipe data found on the web in hRecipe format, producing results in json","archived":false,"fork":false,"pushed_at":"2020-05-09T23:32:50.000Z","size":51,"stargazers_count":106,"open_issues_count":1,"forks_count":23,"subscribers_count":5,"default_branch":"master","last_synced_at":"2025-03-27T02:12:27.947Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/dpapathanasiou.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2017-02-06T00:11:31.000Z","updated_at":"2025-03-25T23:39:57.000Z","dependencies_parsed_at":null,"dependency_job_id":"89982a3a-c6de-4a8b-a3c4-194ba3bb6f8c","html_url":"https://github.com/dpapathanasiou/recipebook","commit_stats":null,"previous_names":[],"tags_count":8,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dpapathanasiou%2Frecipebook","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dpapathanasiou%2Frecipebook/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dpapathanasiou%2Frecipebook/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dpapathanasiou%2Frecipebook/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/dpapathanasiou","download_url":"https://codeload.github.com/dpapathanasiou/recipebook/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248703200,"owners_count":21148118,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-01T07:01:11.396Z","updated_at":"2025-04-13T11:05:55.960Z","avatar_url":"https://github.com/dpapathanasiou.png","language":"Python","funding_links":[],"categories":["Python","others","Tools"],"sub_categories":[],"readme":"# About\n\nThis is a simple application for scraping and parsing food recipe data found on the web in [hRecipe](http://microformats.org/wiki/hrecipe) format, producing results in [json](http://json.org/).\n\nThis project was inspired by [this answer](http://opendata.stackexchange.com/a/4286) to a query for an open database of recipes.\n\n[Contribute](sites/README.md) your favorite site by implementing a [RecipeParser](parser.py) class for it, and make a [pull request](https://help.github.com/articles/about-pull-requests/).\n\n# Data\n\nRecipes collected using the crawler are now available in a [parallel repository](https://github.com/dpapathanasiou/recipes).\n\n# Usage\n\nThe current version uses [Python 3](https://docs.python.org/3/), which is easily installed with [Anaconda](https://www.anaconda.com): use the `Python 3.x` [command line installer](https://www.anaconda.com/products/individual), and optionally, the [PyCharm IDE](https://www.anaconda.com/pycharm) (the community edition is free).\n\n## Individual recipes\n\nImport the class corresponding to the site you want, and use the recipe URL in its constructor.\n\nHere's an example to fetch and parse the [Chocolate, Almond, and Banana Parfaits](http://www.epicurious.com/recipes/food/views/Chocolate-Almond-and-Banana-Parfaits-357369) recipe from [Epicurious](http://www.epicurious.com/):\n\n```python\n\u003e\u003e\u003e import sys; sys.path.append('sites')\n\u003e\u003e\u003e from epicurious import Epicurious\n\u003e\u003e\u003e recipe = Epicurious(\"http://www.epicurious.com/recipes/food/views/Chocolate-Almond-and-Banana-Parfaits-357369\")\n```\n\nUse the \u003ctt\u003esave()\u003c/tt\u003e method to create a file of the recipe in json object.\n\nThe file name is determined from the URL, and the output folder is defined in the [settings.py](settings.py) file as \u003ctt\u003eOUTPUT_FOLDER\u003c/tt\u003e, and can be overridden by creating a local_settings.py file:\n\n```python\n\u003e\u003e\u003e recipe.save()\n```\n\nResults in the creation of \u003ctt\u003e/tmp/chocolate-almond-and-banana-parfaits-357369.json\u003c/tt\u003e with these contents:\n\n```json\n{\n    \"directions\": [\n        \"Heat chocolate chips and 4 tablespoons cream in microwave in 1-cup glass measuring cup at 50 percent power just until chocolate is melted, about 30 to 35 seconds. Stir to blend; cool chocolate sauce to lukewarm. Whisk mascarpone, amaretto, sugar, and remaining 2 tablespoons cream in medium bowl until blended and mixture just starts to thicken.\",\n        \"Using 2 1/2-inch-diameter cookie cutter, cut out round from each angel food cake slice. Place 1 cake round in each of 4 wine goblets or old-fashioned glasses. Top each cake round with 3 banana slices, 1 heaping tablespoon mascarpone mixture, bittersweet chocolate sauce, and sprinkling of almonds. Repeat parfait layering 1 more time and serve.\"\n    ],\n    \"ingredients\": [\n        \"1/2 cup bittersweet chocolate chips\",\n        \"6 tablespoons heavy whipping cream, divided\",\n        \"3/4 cup mascarpone cheese\",\n        \"3 tablespoons amaretto\",\n        \"2 tablespoons sugar\",\n        \"8 1/2-inch-thick angel food cake slices\",\n        \"24 1/3-inch-thick diagonal banana slices (from about 3 bananas)\",\n        \"1/3 cup (about) sliced almonds, toasted\"\n    ],\n    \"language\": \"en-US\",\n    \"source\": \"www.epicurious.com\",\n    \"tags\": [\n        \"Chocolate\",\n        \"Dessert\",\n        \"Quick \u0026 Easy\",\n        \"High Fiber\",\n        \"Banana\",\n        \"Almond\",\n        \"Amaretto\",\n        \"Shower\",\n        \"Party\",\n        \"Vegetarian\",\n        \"Pescatarian\",\n        \"Peanut Free\",\n        \"Soy Free\",\n        \"Kosher\"\n    ],\n    \"title\": \"Chocolate, Almond, and Banana Parfaits\",\n    \"url\": \"http://www.epicurious.com/recipes/food/views/Chocolate-Almond-and-Banana-Parfaits-357369\"\n}\n```\n\n## Crawling\n\nMost sites offer related links within each recipe.\n\nFrom the example above, the \u003ctt\u003egetOtherRecipeLinks()\u003c/tt\u003e method produces more URLs to fetch:\n\n```python\n\u003e\u003e\u003e recipe.getOtherRecipeLinks()\n['http://www.epicurious.com/recipes/food/views/chocolate-amaretto-souffles-104730', 'http://www.epicurious.com/recipes/food/views/coffee-almond-ice-cream-cake-with-dark-chocolate-sauce-11036', 'http://www.epicurious.com/recipes/food/views/toasted-almond-mocha-ice-cream-tart-12550', 'http://www.epicurious.com/recipes/food/views/chocolate-marble-cheesecake-241488', 'http://www.epicurious.com/recipes/food/views/hazelnut-dome-cake-4246']\n```\n\nThe [crawler.py](crawler.py) application takes advantage of this by visiting each related recipe link in parallel, getting even more recipe links, fetching each of those, and so on.\n\nKick it off with a specific site and a file of initial seed links, and it will automatically fetch and parse all the related links it finds, without repeating the same link twice.\n\nFrom the example above, here is how to start the crawler with four parallel worker threads.\n\nThe file \u003ctt\u003e/tmp/epi.link\u003c/tt\u003e passed in the second argument contains the seed URL \u003ctt\u003ehttp://www.epicurious.com/recipes/food/views/Chocolate-Almond-and-Banana-Parfaits-357369\u003c/tt\u003e for this example, though it could contain more links, too.\n\nIt is also a good idea to capture the output into a log file, as shown here, in order to see the full list of parsed recipes, along with any error messages.\n\n```sh\npython crawler.py Epicurious /tmp/epi.link 4 \u003e epicurious.log 2\u003e\u00261\n```\n\nBy default, all the json files are written to the \u003ctt\u003eOUTPUT_FOLDER\u003c/tt\u003e folder specified in [settings.py](settings.py) local_settings.py, but this can be changed by passing a fourth argument: \"False\" or \"F\" (in either upper or lower case) will prevent the individual recipes from being written to to the \u003ctt\u003eOUTPUT_FOLDER\u003c/tt\u003e folder at all.\n\nSimilarly, storing the results to a [ARMS mongo service](https://github.com/dpapathanasiou/ARMS) is off by default, but if the fifth and sixth arguments specify a database and collection, respectively, the crawler will attempt to store them, using the \u003ctt\u003eARMS\u003c/tt\u003e server, api key and seed definitions in [settings.py](settings.py) or local_settings.py.\n\n### Avoiding server blocks\n\nThe crawler can also be configured to pause a random number of seconds in between fetches, to prevent recipe hosts from blocking it for too many requests.\n\nThe pause default configuration is defined in lines 12 and 13 of the [settings.py](settings.py) file, which can be overridden in a local_settings.py definition.\n\nAnother strategy, which can done in conjunction with pausing, is to change the [user agent](https://en.wikipedia.org/wiki/User_agent) from the default defined in line 11 of the [settings.py](settings.py) file to something resembling a human user.\n\n[MDN maintains a list](https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/User-Agent) of current common browser agent strings, which can be used in a local_settings.py definition of the \u003ctt\u003eUA\u003c/tt\u003e variable.\n\n### Usage\nHere is the crawler usage in full:\n\n```sh\npython crawler.py [site: (AllRecipes|Epicurious|FoodNetwork|Saveur|SiroGohan|WilliamsSonoma)] [file of seed urls] [threads] [save() (defaults to True)] [store() database (defaults to None)] [store() collection (defaults to None)]\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdpapathanasiou%2Frecipebook","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdpapathanasiou%2Frecipebook","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdpapathanasiou%2Frecipebook/lists"}