{"id":18146439,"url":"https://github.com/whoisjayd/imdb-scrapper","last_synced_at":"2025-04-06T20:45:18.730Z","repository":{"id":258861136,"uuid":"875795254","full_name":"WhoIsJayD/IMDB-Scrapper","owner":"WhoIsJayD","description":"IMDB Movie Scrapping","archived":false,"fork":false,"pushed_at":"2024-11-05T12:46:01.000Z","size":39,"stargazers_count":1,"open_issues_count":1,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-02-13T02:47:33.902Z","etag":null,"topics":["dataset","imdb","python","scrape","scrapping","scrapping-python","scrappy","selenium","selenium-python","tmdb"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/WhoIsJayD.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.md","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-10-20T20:59:30.000Z","updated_at":"2025-01-05T01:56:49.000Z","dependencies_parsed_at":"2024-12-20T08:28:15.518Z","dependency_job_id":"ebd9ce1e-baf0-49ff-91db-e7ba5b5f4381","html_url":"https://github.com/WhoIsJayD/IMDB-Scrapper","commit_stats":null,"previous_names":["whoisjayd/imdb-scrapper"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/WhoIsJayD%2FIMDB-Scrapper","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/WhoIsJayD%2FIMDB-Scrapper/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/WhoIsJayD%2FIMDB-Scrapper/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/WhoIsJayD%2FIMDB-Scrapper/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/WhoIsJayD","download_url":"https://codeload.github.com/WhoIsJayD/IMDB-Scrapper/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247550642,"owners_count":20956984,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["dataset","imdb","python","scrape","scrapping","scrapping-python","scrappy","selenium","selenium-python","tmdb"],"created_at":"2024-11-01T21:07:42.523Z","updated_at":"2025-04-06T20:45:18.496Z","avatar_url":"https://github.com/WhoIsJayD.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# IMDb-TMDb Scraper\n\n[![License](https://img.shields.io/github/license/WhoIsJayD/IMDB-Scrapper)](https://github.com/WhoIsJayD/IMDB-Scrapper/blob/main/LICENSE.md)\n[![Python](https://img.shields.io/badge/python-3.10%2B-blue)](https://www.python.org/downloads/release/python-310/)\n[![Issues](https://img.shields.io/github/issues/WhoIsJayD/IMDB-Scrapper)](https://github.com/WhoIsJayD/IMDB-Scrapper/issues)\n\nThis project provides two Scrapy spiders for scraping movie data from **IMDb** and **TMDb**: a **basic scraper** and an **advanced scraper** with additional capabilities for concurrent and customizable scraping.\n\n## 📂 Project Structure\n\n```plaintext\n.\n├── imdbscrapper\n│   ├── spiders\n│   │   ├── __init__.py\n│   │   ├── advance_scrapper.py\n│   │   ├── basic_scrapper.py\n│   ├── __init__.py\n│   ├── items.py\n│   ├── middlewares.py\n│   ├── pipelines.py\n│   └── settings.py\n│\n├── LICENSE.md\n├── README.md\n├── requirements.txt\n├── scrapy.cfg\n└── setup.py\n```\n\n## 🚀 Features\n\n### Basic Scraper (`basic_scrapper.py`)\n- **IMDb Search Pages**: Scrapes movie details from IMDb search pages.\n- **TMDb Integration**: Fetches additional data, such as posters and ratings, using the TMDb API.\n- **Pagination**: Supports pagination through the IMDb search \"Show More\" option.\n- **Output**: Saves scraped data as JSON and CSV files.\n\n### Advanced Scraper (`advance_scrapper.py`)\n- **Enhanced Features**: Includes all functionalities of the basic scraper.\n- **Multi-threaded Scraping**: Utilizes concurrent scraping for faster data collection.\n- **Robust Error Handling**: Implements improved retry and error management.\n- **Data Enrichment**: Collects extended movie metadata and applies data cleaning.\n\n## 📋 Requirements\n\n- **Python** 3.10+\n- **Scrapy**\n- **Selenium** (for JavaScript-heavy pages)\n- **Requests** (for TMDb API requests)\n- **Concurrent Futures** (for parallel scraping)\n\nInstall dependencies with:\n```bash\npip install -r requirements.txt\n```\n\n## 🛠️ Installation\n\n1. **Clone the repository**:\n   ```bash\n   git clone https://github.com/WhoIsJayD/IMDB-Scrapper\n   cd IMDB-Scrapper\n   ```\n\n2. **Set up API Key**:\n   Add your TMDb API key in each spider file or pass it as an argument.\n\n3. **Set up Selenium** (for advanced scraping):\n   - Download [ChromeDriver](https://sites.google.com/a/chromium.org/chromedriver/downloads) compatible with your Chrome version.\n   - Add ChromeDriver to your system's PATH.\n\n## ⚙️ Usage\n\n### Basic Scraper\n\nRun the basic scraper with:\n```bash\nscrapy crawl basic_scrapper\n```\n\n### Advanced Scraper\n\nRun the advanced scraper with custom parameters:\n```bash\nscrapy crawl advance_scrapper -a tmdb_api_key=\"your_tmdb_api_key\" -a start_year=2000 -a end_year=2023 -a num_instances=5\n```\n\n### Configuration Options\n\n- **`start_year`**: Start year for the movie range.\n- **`end_year`**: End year for the movie range.\n- **`num_instances`**: Number of concurrent scraping instances.\n\n## 📁 Output\n\nThe scrapers produce the following files:\n- **movies.json**: Contains movie data in JSON format.\n- **movies.csv**: Contains movie data in CSV format.\n\n## 🔧 Customization\n\n- Modify `custom_settings` in each spider to configure scraping behavior.\n- Adjust the `clean_movie_data` method in `advance_scrapper.py` to customize data cleaning.\n\n## ⚠️ Notes\n\n- **Legal Compliance**: Ensure your usage complies with IMDb and TMDb terms of service.\n- **Rate Limiting**: To avoid blocking, set appropriate delays or intervals.\n\n## 📝 License\n\nThis project is licensed under the MIT License. See the [LICENSE.md](LICENSE.md) file for details.\n\n## 📞 Contact\n\nFor any inquiries, please reach out:\n\n- **Name**: Jaydeep Solanki\n- **Email**: jaydeep.solankee@yahoo.com\n- **LinkedIn**: [LinkedIn Profile](https://www.linkedin.com/in/solanki-jaydeep)\n\n## 🙌 Acknowledgments\n\nSpecial thanks to:\n- **IMDb** for the movie data.\n- **TMDb** for their API resources.\n- The **Scrapy** and **Selenium** communities for their robust tools and documentation.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwhoisjayd%2Fimdb-scrapper","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fwhoisjayd%2Fimdb-scrapper","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwhoisjayd%2Fimdb-scrapper/lists"}