https://github.com/rohit1kumar/internsync
A Python web scraper to scrape internships from Internshala , with the option to integrate Google Sheets and sync it with GitHub Actions
https://github.com/rohit1kumar/internsync
python scraping webscraping
Last synced: over 1 year ago
JSON representation
A Python web scraper to scrape internships from Internshala , with the option to integrate Google Sheets and sync it with GitHub Actions
- Host: GitHub
- URL: https://github.com/rohit1kumar/internsync
- Owner: rohit1kumar
- Created: 2023-10-29T20:27:57.000Z (almost 3 years ago)
- Default Branch: main
- Last Pushed: 2024-09-09T04:17:24.000Z (almost 2 years ago)
- Last Synced: 2025-01-22T10:23:14.483Z (over 1 year ago)
- Topics: python, scraping, webscraping
- Language: Python
- Homepage: https://docs.google.com/spreadsheets/d/1vK_4_jWJ4H1HJJr_FWxPNavddi5AU-HLqc3UCRC9hnM/edit?usp=sharing
- Size: 28.3 KB
- Stars: 1
- Watchers: 1
- Forks: 0
- Open Issues: 0
-
Metadata Files:
- Readme: README.md
Awesome Lists containing this project
README
# InternSync
A Playwright based web scraper to scrape internships from [Internshala](https://internshala.com/), written in Python. Data is stored in a CSV file.
## Disclaimer
*This is for educational purposes only. I am not responsible for any misuse of this code.*
## Features
- Store data in a CSV file
- Store data in a Google Sheet (optional)
- Keeps the Google Sheet data synced using GitHub Actions (optional)
## Prerequisites
Make sure you have the following dependencies installed:
- [Playwright](https://playwright.dev/python/docs/library/) library for Python
- Python 3.x
*you can install them using the following commands too:*
```bash
pip install playwright && playwright install chromium
```
## Usage
1. Clone the repository
2. Install the dependencies using `pip install -r requirements.txt`
3. Run the script using `python main.py`
Optional steps (for Google Sheets mode only):
1. Create a new Google Sheet
2. Create a new project in [Google Cloud Platform](https://console.cloud.google.com/)
3. Follow this guide for setting up the [Google Sheets API](https://docs.gspread.org/en/v5.12.0/oauth2.html#for-bots-using-service-account)
4. Download the JSON file and add all the credentials to the `.env` file (refer to `.env.example`)
5. Get the Google Sheet ID from the URL e.g `https://docs.google.com/spreadsheets/d/GOOGLE_SHEET_ID/edit`
6. Add to `GOOGLE_SHEET_ID` to the `.env` file
Optional steps (for syncing Google Sheets using GitHub Actions):
1. GitHub Actions are already setup in the repository
2. Download GitHub CLI or add secrets manually to the repository from `.env` file
3. With GitHub CLI run `gh secret set -R -f .env`
## Options
- `--headful`: Run the script in non-headless mode (show the browser)
```bash
python main.py --headful
```
- `--gs`: Run the script in Google Sheets mode (store data in Google Sheets)
```bash
python main.py --gs
```
[//]: # (- `--limit`: Limit the number of internships to scrape (default: 10))
[//]: # (- `--output`: Output file name (default: internships.csv) )