{"id":26748240,"url":"https://github.com/mohidex/data-pipeline-on-gcp","last_synced_at":"2025-04-14T22:14:34.800Z","repository":{"id":185918487,"uuid":"673416074","full_name":"mohidex/data-pipeline-on-gcp","owner":"mohidex","description":"The Real-time Ecommerce Data Collection and Processing project empowers businesses with real-time insights by efficiently extracting, processing, and storing ecommerce data from multiple sources. Combining Golang and Python, this cutting-edge solution streamlines data handling from diverse ecommerce websites.","archived":false,"fork":false,"pushed_at":"2025-01-18T18:04:56.000Z","size":1037,"stargazers_count":6,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-04-14T22:14:30.396Z","etag":null,"topics":["beautifulsoup","data-engineer","data-pipeline","data-science","database","datastore","dependency-injection","firebase","firestore","gcp","go","golang","google","google-cloud","pubsub","python","solid-principles","storage","web-scraping"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/mohidex.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2023-08-01T15:14:11.000Z","updated_at":"2025-01-18T18:04:58.000Z","dependencies_parsed_at":"2023-09-17T02:45:29.343Z","dependency_job_id":null,"html_url":"https://github.com/mohidex/data-pipeline-on-gcp","commit_stats":null,"previous_names":["mohidex/data-pipeline","mohidex/data-pipeline-on-gcp"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mohidex%2Fdata-pipeline-on-gcp","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mohidex%2Fdata-pipeline-on-gcp/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mohidex%2Fdata-pipeline-on-gcp/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mohidex%2Fdata-pipeline-on-gcp/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/mohidex","download_url":"https://codeload.github.com/mohidex/data-pipeline-on-gcp/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248968917,"owners_count":21191162,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["beautifulsoup","data-engineer","data-pipeline","data-science","database","datastore","dependency-injection","firebase","firestore","gcp","go","golang","google","google-cloud","pubsub","python","solid-principles","storage","web-scraping"],"created_at":"2025-03-28T10:17:03.829Z","updated_at":"2025-04-14T22:14:34.785Z","avatar_url":"https://github.com/mohidex.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"The Real-time Ecommerce Data Collection and Processing project offers a robust and efficient solution for gathering data from various ecommerce websites, processing it in real-time, and storing it in Google Cloud Datastore. The project comprises two parts: a Golang-based data processing pipeline named \"Topic2Warehouse\" and a Python-based web scraper and publisher.\n\n![hight-level-design](system-scatch.png)\n\n**Golang Data Processing Pipeline (topic2storage):**\n\nThe Golang pipeline is responsible for real-time data processing and storage. It subscribes to Google Cloud Pub/Sub, receiving data sent from the Python scraper. The received data is then stored securely in Google Cloud Datastore, providing reliable and scalable data storage capabilities. The pipeline leverages Golang's concurrency features, ensuring high throughput and seamless handling of incoming data from multiple sources.\n\n**Python Web Scraper and Publisher (source2topic):**\n\nThe Python application excels in web scraping, collecting valuable ecommerce data from various websites. The scraper uses Python's BeautifulSoup library for HTML parsing and efficiently extracts relevant product details. After scraping, the data is published to Google Cloud Pub/Sub, enabling real-time data transfer to the Golang data processing pipeline.\n\n\n```mermaid\nflowchart LR\n    Sources((Sources))\n    Scheduler[Scheduler - GHA]\n    Blobs[Blobs - Google Drive]\n    Client((Client))\n    APIService[API Service - Manage Curves]\n    PSQL[(PSQL)]\n    Dashboard[Dashboard]\n    InfluxDB[(Influx DB)]\n    Telegraf[Telegraf]\n    Consumer[Consumer]\n    Compose[Compose]\n    RabbitMQ[[RabbitMQ]]\n    ForecastingModel[Forecasting Model]\n\n    Sources --\u003e|write row data| Scheduler\n    Scheduler --\u003e Blobs\n    Blobs --\u003e|Get row file| Compose\n    Compose --\u003e|Publish processed data| RabbitMQ\n    RabbitMQ --\u003e|Read Msg| Telegraf\n    Telegraf --\u003e|Write TS| InfluxDB\n    APIService --\u003e|Read TS| InfluxDB\n    InfluxDB --\u003e Dashboard\n    Client --\u003e APIService\n    APIService --\u003e|Update Metadata| PSQL\n    PSQL --\u003e|Update Metadata| Consumer\n    Consumer --\u003e|Update Metadata| PSQL\n    RabbitMQ --\u003e ForecastingModel\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmohidex%2Fdata-pipeline-on-gcp","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmohidex%2Fdata-pipeline-on-gcp","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmohidex%2Fdata-pipeline-on-gcp/lists"}