{"id":51602724,"url":"https://github.com/jacktheprogrammer/time-series-forecasting-mlops","last_synced_at":"2026-07-11T23:01:29.981Z","repository":{"id":354282044,"uuid":"1222768061","full_name":"JackTheProgrammer/time-series-forecasting-mlops","owner":"JackTheProgrammer","description":"MLOps driven time series forecasting using DL with PyTorch","archived":false,"fork":false,"pushed_at":"2026-06-08T16:03:55.000Z","size":747,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-06-08T17:29:30.526Z","etag":null,"topics":["clusterized-ml-delivery","containerized-ml-deployment","docker","dvc","github-actions","gold-price-prediction","gold-stock-forecasting","kubernetes","mlops","stock-market","stock-price-prediction","timeseries-forecasting","yfinance"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/JackTheProgrammer.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-04-27T17:30:45.000Z","updated_at":"2026-06-08T16:07:06.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/JackTheProgrammer/time-series-forecasting-mlops","commit_stats":null,"previous_names":["jacktheprogrammer/time-series-forecasting-mlops"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/JackTheProgrammer/time-series-forecasting-mlops","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackTheProgrammer%2Ftime-series-forecasting-mlops","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackTheProgrammer%2Ftime-series-forecasting-mlops/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackTheProgrammer%2Ftime-series-forecasting-mlops/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackTheProgrammer%2Ftime-series-forecasting-mlops/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/JackTheProgrammer","download_url":"https://codeload.github.com/JackTheProgrammer/time-series-forecasting-mlops/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackTheProgrammer%2Ftime-series-forecasting-mlops/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35377013,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-11T02:00:05.354Z","response_time":104,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["clusterized-ml-delivery","containerized-ml-deployment","docker","dvc","github-actions","gold-price-prediction","gold-stock-forecasting","kubernetes","mlops","stock-market","stock-price-prediction","timeseries-forecasting","yfinance"],"created_at":"2026-07-11T23:01:28.913Z","updated_at":"2026-07-11T23:01:29.973Z","avatar_url":"https://github.com/JackTheProgrammer.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# 🪙 Automated MLOps Gold Price Forecasting Pipeline\n\n[![Python 3.13](https://img.shields.io/badge/Python-3.13-blue.svg?style=flat-square\u0026logo=python)](https://www.python.org/)\n[![PyTorch](https://img.shields.io/badge/PyTorch-2.6.0-EE4C2C.svg?style=flat-square\u0026logo=pytorch)](https://pytorch.org/)\n[![DVC](https://img.shields.io/badge/DVC-Data_Version_Control-9CF.svg?style=flat-square\u0026logo=data-version-control)](https://dvc.org/)\n[![Docker](https://img.shields.io/badge/Docker-Containerized-2496ED.svg?style=flat-square\u0026logo=docker)](https://www.docker.com/)\n[![Kubernetes](https://img.shields.io/badge/Kubernetes-Orchestrated-326CE5.svg?style=flat-square\u0026logo=kubernetes)](https://kubernetes.io/)\n[![GitHub Actions](https://img.shields.io/badge/CI/CD-GitHub_Actions-2088FF.svg?style=flat-square\u0026logo=github-actions)](https://github.com/features/actions)\n\nAn enterprise-grade, fully automated end-to-end MLOps pipeline for time-series forecasting. This system abstracts heavy data and deep learning model weights using content-addressable storage (DVC), orchestrates decentralized compute layers via GitHub Actions, and handles immutable production deliveries using Docker and Kubernetes.\n\n---\n\n## 📖 Introduction\n\nThis project implements a self-correcting, scheduled deep learning retraining mechanism to forecast gold stock prices. By treating data, preprocessing parameters, and neural network weights as versioned assets distinct from code, the system mitigates data drift, handles continuous delivery, and provides a scalable runtime environment with zero human intervention.\n\n## 🛠️ Tech Stack \u0026 Ecosystem\n\n* **Core Engine:** Python 3.13, PyTorch 2.6 (CPU-optimized compilation)\n* **Data Tracking:** Data Version Control (DVC) backed by Remote Cloud Cache Storage\n* **CI/CD Orchestration:** GitHub Actions Engine (Decoupled Runner Architectures)\n* **Container Delivery:** Docker, Multi-stage Buildx, Docker Hub Registry\n* **Production Cluster:** Kubernetes (Minikube Cluster Simulation, LoadBalancer Ingress)\n* **Presentation Layer:** FastAPI (Backend Prediction Engine) \u0026 Streamlit (Frontend Client UI)\n\n---\n\n## 🔬 Methodology \u0026 System Operations\n\n1. **Automated Ingestion \u0026 Preprocessing (`ingestion.py`, `preprocessing.py`)**\n   * Dynamically tracks real-time historical market data from 2010 to the present moment using `yfinance`.\n   * Standardizes raw values, manages historical sequence lengths, and packages scaling parameters into isolated data serialization formats (`scaled_transform.pkl`).\n\n2. **Decoupled Model Development (`dl_pipeline.py`, `architectures/`)**\n   * Parallel training of multiple deep learning architectures: Recurrent neural paths (**LSTM**, **GRU**) and convolutional temporal matchers (**Conv1D**).\n   * Models are benchmarked via automated validation scripts (`evaluate.py`) tracking Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE) to dynamically flag the absolute \"Winner Model.\"\n\n3. **Continuous Deployment \u0026 Storage Decoupling**\n   * Heavy binaries are offloaded to an abstract content-addressable storage remote via an automated cloud credential handshake.\n   * Runtime images are compiled with strict metadata layers and deployed into simulated container orchestration layers for target scale verification.\n\n---\n\n## ⚙️ Architecture Blueprints\n\n### Overall System \u0026 Data Flow\n\nThe system keeps the application runtime lightweight by cleanly separating code logic (tracked in Git) from heavy binaries (managed by DVC).\n\n![Overall System Flow](diagrams/overall_generic_diag.jpg)\n\n### CI/CD Scheduled Orchestration\n\nThe retraining pipeline operates on an independent serverless virtual machine context, checking out code pointers, parsing delta transformations, writing back metadata hashes, and delivering production-ready containers.\n\n![CI/CD Workflow Architecture](diagrams/mlops_ci_cd_workflow.jpg)\n\n---\n\n## 🔄 Local Workstation Synchronization (Windows OS 🪟)\n\nTo sync your local environment with the cloud infrastructure context, a dedicated batch file (`sync_pipeline.bat`) executes safe upstream rebase pulls and hydrates your local data cache.\n\n### Method 1: Execution via Command Line\n\nOpen a standard terminal and trigger the Task Scheduler engine immediately:\n\n```bat\nschtasks /run /tn \"MLOps_TimeSeries_Sync\"\n```\n\n### Method 2: Execution via Task Scheduler GUI\n\n1. Press `Win + R`, type `taskschd.msc`, and press **Enter**.\n2. Navigate to the **Task Scheduler Library** in the left sidebar directory.\n3. Locate **`MLOps_TimeSeries_Sync`** in the middle console pane.\n4. Right-click the task entry and select **Run**.\n\n### Task Orchestration \u0026 Deployment Setup\n\n1. Open **Command Prompt as Administrator** and map the script execution to run precisely at 17:00 (5:00 PM PKT) on Day 1 of every 4th month to stay perfectly aligned with the cloud cron architecture (`0 17 1 */4 *`):\n\n```bat\nschtasks /create /tn \"MLOps_TimeSeries_Sync\" /tr \"'D:\\path\\to\\your\\sync_pipeline.bat'\" /sc monthly /mo 4 /d 1 /st 17:00\n```\n\n1. **Workstation Deployment References:**\n\n*Local Command Prompt initialization verification.*\n\n*Active operational task mapped inside the native Windows Task Scheduler environment.*\n\n\u003e ⚠️ **Configuration Note:** The local [Synchronization Script](https://www.google.com/search?q=sync_pipeline.bat) contains hardcoded routing parameters pointing to `cd \"D:\\ml projects\\mlops_time_series_modeling\"`. Ensure you update this path to reflect the absolute workspace path on your local system before deploying the scheduler task.\n\n---\n\n## 🚀 Key Automation Features\n\n* **Data Drift Safeguards:** Automated periodic cron execution fetches real-time updates and retrains models down to the latest decimal without manual workflow initialization.\n* **Smart Delta Synchronization:** The DVC delivery layer evaluates system state hashes at the block level; if retraining yields identical artifacts, network serialization uploads are safely skipped (`2 files pushed`).\n* **Self-Healing Git Loops:** The CI/CD runner checks changes using `git status --porcelain`. If tracking hashes diverge, it signs, updates, commits, and pushes pointer metadata back to `main` while injecting `[skip ci]` flags to completely block recursive execution loops.\n\n---\n\n## 💻 Application Client Interface\n\n![Alt demo](demos/streamlit_app_running_screenshot.png)\n\n---\n\n## 📂 Project Directory Structure\n\n```text\ntime-series-forecasting-mlops/\n├── .dockerignore\n├── .dvc/\n│   └── .gitignore\n├── .dvcignore\n├── .github/\n│   └── workflows/\n│       └── actions.yaml          # Scheduled Retraining \u0026 CD Infrastructure Blueprint\n├── .gitignore\n├── compose.yml                   # Multi-Container Application Stack Topology\n├── data.dvc                      # DVC Data Vector Tracking Pointer\n├── demos/                        # Visual Execution Logs \u0026 Demos\n├── diagrams/                     # System Infrastructure Flow Charts\n├── Dockerfile                    # Multi-Stage Complete Architectural Image Build\n├── forecasts.Dockerfile          # Lightweight Production Target Forecasting Runtime\n├── forecasts.Dockerfile.dockerignore\n├── k8s/                          # Production Deployment Declarative Contexts\n│   ├── deployment.yaml\n│   └── service.yaml\n├── LICENSE\n├── models.dvc                    # DVC Trained Weights Tracking Pointer\n├── notebooks/\n│   └── timeseries_modelling.ipynb\n├── README.md\n├── requirements.txt\n├── scaled_transform.dvc          # DVC Data Processing State Pointer\n├── scripts/\n│   ├── __init__.py\n│   ├── api/                      # FastAPI Backend Core Rest Layer\n│   │   ├── main.py\n│   │   └── schemas.py\n│   ├── app/                      # Streamlit Frontend Client Rendering Interface\n│   │   ├── home/\n│   │   │   ├── daily_forecasts.py\n│   │   │   └── home.py\n│   │   └── server/\n│   │       ├── forecasting.py\n│   │       ├── latest_model_forecasting.py\n│   │       └── server.py\n│   ├── architectures/            # Deep Learning Layer Paradigms (LSTM, GRU, Conv1D)\n│   │   ├── conv1d.py\n│   │   ├── gru.py\n│   │   └── lstm.py\n│   ├── dl_pipeline.py            # Deep Learning Core Training Loop Execution Script\n│   ├── evaluate.py               # Model Metric Evaluation Matrix Block\n│   ├── ingestion.py              # Real-time API Asset Pipeline\n│   └── preprocessing.py          # Vector Normalization \u0026 Sequential Splitting\n├── sync_pipeline.bat             # Native Local Windows Task Scheduler Automation Target\n└── winner_models.dvc             # DVC Production Model Artifact Tracker\n\n```\n\n---\n\n## 🚀 Application Deployment Guide\n\n### Method 1: Local Native Installation\n\n```bash\n# Clone the repository code base\ngit clone https://github.com/JackTheProgrammer/time-series-forecasting-mlops.git\ncd time-series-forecasting-mlops\n\n# Build the base CPU-optimized computing environment and requirements\npip install --no-cache-dir --extra-index-url [https://download.pytorch.org/whl/cpu](https://download.pytorch.org/whl/cpu) torch==2.6.0\npip install --no-cache-dir -r requirements.txt\n\n# Initialize the DVC\n# step 1: install DVC via pip\npip install dvc --no-cache-dir\n# step 2: run the preprocessing and dl_pipeline scripts to generate the required DVC tracked assets (data, models, winner_models)\npython scripts/preprocessing.py\npython scripts/dl_pipeline.py\n# step 3: Add the generated artifacts to the DVC tracking, no need to craft .dvcignore though\ndvc add data\ndvc add models\ndvc add winner_models\n\n# Launch Backend Engine \u0026 Frontend Interface simultaneously so that you can also generate the scaled_transform folder as well, containing the scalers as pickled files\npython scripts/api/main.py \u0026\u0026 streamlit run scripts/app/home/home.py\n\n# adding the scaled_transform folder to DVC artifacts\ndvc add scaled_transform\ndvc commit -f\ndvc remote add -d myremote \u003cLINK_TO_YOUR_REMOTE_STORAGE\u003e # I used google drive as remote storage cause this method was free entirely\n# you may need to do the necessary configurations related to the remote storage you are using, for example in case of google storage,\n# I followed this tutorial: https://youtu.be/e3GuonR1r-0?si=1f7f4dniQwH-6uCi\ndvc push -r myremote\n```\n\n### Method 2: Isolated Docker Runtimes\n\n#### Option A: Run Full Ingestion, Training, \u0026 Prediction Ecosystem\n\n```bash\ndocker pull fawadawan143/gold_stock_prediction:latest\ndocker run -p 8501:8501 -p 5050:5050 fawadawan143/gold_stock_prediction:latest\n```\n\n#### Option B: Run Optimized Production Serving UI (Latest Assets Only)\n\n```bash\ndocker pull fawadawan143/daily-forecasting:latest\ndocker run -p 8501:8501 -p 5050:5050 fawadawan143/daily-forecasting:latest\n```\n\n### Method 3: Multi-Container Orchestration via Docker Compose\n\n```bash\ngit clone https://github.com/JackTheProgrammer/time-series-forecasting-mlops.git\ncd time-series-forecasting-mlops\n\n# Launch system instances in safe daemon modes\ndocker compose up --build -d\n\n# Spin up isolated processing nodes or production endpoints independently\ndocker compose up -d ml-pipeline       # Executes full ingestion, retraining, and testing\ndocker compose up -d forecasting-app    # Serves prediction features using the latest assets\n\n# Spin down the active stack environment safely\ndocker compose down\n```\n\n### Method 4: Production Cluster Deployment (Kubernetes)\n\n```bash\ngit clone https://github.com/JackTheProgrammer/time-series-forecasting-mlops.git\ncd time-series-forecasting-mlops\n\n# Start your localized cluster management context\nminikube start --driver=docker\n# Apply declarative state manifests into active pods and nodes\nkubectl apply -f k8s/\n# Inject local image assets directly into Minikube container runtime scopes\nminikube image load fawadawan143/daily-forecasting:latest\n# Expose your cluster services directly to external clients via LoadBalancer emulation\nkubectl expose deployment gold-forecasting-deployment --type=LoadBalancer --port=80 --target-port=80\n# In a separate terminal session, initiate the network ingress bridge tunnel\nminikube tunnel\n```\n\n*Your application endpoint route paths are now accessible at `http://localhost:80/forecast` or `http://localhost:80/latest-forecast` mapping directly into your cluster load-balancer matrix.*\n\n#### Cluster Cleanup Routine\n\n```bash\nkubectl delete service gold-forecasting-deployment\nkubectl delete deployment gold-forecasting-deployment\nminikube stop\n```\n\n---\n\n## 🔗 Engineering References \u0026 Resources\n\n* **[Core Analytical Blueprint Repository](https://github.com/JackTheProgrammer/Time-Series-Forecasting-and-Analysis)**\n* [Docker for Machine Learning | Docker Crash Course | CampusX](https://youtu.be/GToyQTGDOS4?si=Lkb3lNbxzPKwPf2I)\n* [Docker Simply Explained with a Machine Learning Project for Beginners](https://youtu.be/-l7YocEQtA0?si=GQ425cnC0SJnHLG9)\n* [Data Version Control | DVC | Remote Target Implementations | CodeKamikaze](https://youtu.be/e3GuonR1r-0?si=1f7f4dniQwH-6uCi)\n* [Automating Data Pipelines with Python \u0026 GitHub Actions Engine](https://youtu.be/wJ794jLP2Tw?si=mXsWGGRQGxmI1rfl)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjacktheprogrammer%2Ftime-series-forecasting-mlops","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjacktheprogrammer%2Ftime-series-forecasting-mlops","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjacktheprogrammer%2Ftime-series-forecasting-mlops/lists"}