{"id":50384533,"url":"https://github.com/nitish2773/food-delivery-delay-analyzer","last_synced_at":"2026-05-30T14:01:32.932Z","repository":{"id":343680109,"uuid":"1178723221","full_name":"Nitish2773/food-delivery-delay-analyzer","owner":"Nitish2773","description":null,"archived":false,"fork":false,"pushed_at":"2026-03-11T09:56:49.000Z","size":9969,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-03-11T16:44:58.599Z","etag":null,"topics":["matplotlib-pyplot","numpy","pandas","python3","scipy-stats","seaborn"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Nitish2773.png","metadata":{"files":{"readme":"Readme.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-03-11T09:54:55.000Z","updated_at":"2026-03-11T09:57:51.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/Nitish2773/food-delivery-delay-analyzer","commit_stats":null,"previous_names":["nitish2773/food-delivery-delay-analyzer"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/Nitish2773/food-delivery-delay-analyzer","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Nitish2773%2Ffood-delivery-delay-analyzer","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Nitish2773%2Ffood-delivery-delay-analyzer/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Nitish2773%2Ffood-delivery-delay-analyzer/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Nitish2773%2Ffood-delivery-delay-analyzer/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Nitish2773","download_url":"https://codeload.github.com/Nitish2773/food-delivery-delay-analyzer/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Nitish2773%2Ffood-delivery-delay-analyzer/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33694714,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-05-30T02:00:06.278Z","response_time":92,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["matplotlib-pyplot","numpy","pandas","python3","scipy-stats","seaborn"],"created_at":"2026-05-30T14:01:31.909Z","updated_at":"2026-05-30T14:01:32.925Z","avatar_url":"https://github.com/Nitish2773.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n\n# 📦 Food Delivery Delay Analyzer\n### A Data-Driven Analysis of What Causes Late Deliveries — Operations, HR Analytics \u0026 Fleet Management\n\n[![Python](https://img.shields.io/badge/Python-3.9+-3776AB?style=for-the-badge\u0026logo=python\u0026logoColor=white)](https://python.org)\n[![Pandas](https://img.shields.io/badge/Pandas-2.x-150458?style=for-the-badge\u0026logo=pandas\u0026logoColor=white)](https://pandas.pydata.org)\n[![Seaborn](https://img.shields.io/badge/Seaborn-0.12+-4C72B0?style=for-the-badge\u0026logo=python\u0026logoColor=white)](https://seaborn.pydata.org)\n[![SciPy](https://img.shields.io/badge/SciPy-1.x-8CAAE6?style=for-the-badge\u0026logo=scipy\u0026logoColor=white)](https://scipy.org)\n[![Jupyter](https://img.shields.io/badge/Jupyter-Notebook-F37626?style=for-the-badge\u0026logo=jupyter\u0026logoColor=white)](https://jupyter.org)\n\n\u003cbr\u003e\n\n\u003e **Only 30% of delivery orders arrive on time during peak traffic + bad weather — yet Zomato and Swiggy promise 30-minute delivery to every customer, every time.**  \n\u003e This project maps the exact operational, HR, and fleet factors driving that gap.\n\n\u003cbr\u003e\n\n![Project Banner](outputs/correlation_target.png)\n\n\u003c/div\u003e\n\n---\n\n## 📎 Project Presentation\n\n| Format | Link |\n|--------|------|\n| 📥 Download PPTX | [FoodDelivery_Delay_Analyzer_Presentation.pptx](presentation/FoodDelivery_Delay_Analyzer_Presentation.pptx) |\n| 📓 Jupyter Notebook | [analysis.ipynb](analysis.ipynb) |\n\n---\n\n## 📌 Table of Contents\n\n- [Project Overview](#-project-overview)\n- [Key Findings](#-key-findings)\n- [Dataset](#-dataset)\n- [Project Structure](#-project-structure)\n- [Methodology](#-methodology)\n- [Visualisations](#-visualisations)\n- [Statistical Validation](#-statistical-validation)\n- [Recommendations](#-recommendations)\n- [How to Run](#-how-to-run)\n- [Tech Stack](#-tech-stack)\n- [Limitations](#-limitations)\n- [Author](#-author)\n\n---\n\n## 🎯 Project Overview\n\nIndia's food delivery market is worth **₹38,000 Crore** and growing. Platforms like Zomato and Swiggy serve millions of orders daily — but late deliveries remain the **#1 reason for customer churn and negative reviews**. Yet most companies don't systematically analyze *what* causes those delays.\n\nThis project analyzes **45,000+ real food delivery orders** from Kaggle using Python and statistical methods to identify root causes of delays across **Operations, HR Analytics, and Fleet Management** — and provides **7 data-driven business recommendations**.\n\n### Business Problems Solved\n\n| # | Problem | Question |\n|---|---------|----------|\n| 1 | **Distribution \u0026 Consistency** | What is the actual delivery time spread and how consistent is service? |\n| 2 | **Weather Impact** | Which weather conditions cause the most delays? |\n| 3 | **Partner Rating vs Speed** | Do higher-rated partners actually deliver faster? |\n| 4 | **Outlier Detection** | Which orders are extreme delays — and what causes them? |\n| 5 | **City Type Performance** | Does delivery performance differ across Metro, Urban, and Semi-Urban cities? |\n| 6 | **Partner Age vs Speed** | Does delivery partner age affect speed? Which age group is fastest? |\n| 7 | **Vehicle Type Efficiency** | Which vehicle type is most efficient — and does it change with distance? |\n\n---\n\n## 🔍 Key Findings\n\n\u003ctable\u003e\n\u003ctr\u003e\n\u003ctd width=\"50%\"\u003e\n\n### 🔴 Finding 1 — Delivery Inconsistency\nMean delivery time ≈ **26 minutes**, but Std Dev ≈ **5 minutes** with positive skew — meaning a small percentage of orders experience extreme delays pulling the average up. Inconsistency damages customer trust more than just being slow.\n\n\u003c/td\u003e\n\u003ctd width=\"50%\"\u003e\n\n### 🟠 Finding 2 — Weather Impact\nRainy and foggy conditions add **25–35% extra delivery time** vs sunny baseline. Apps still show the same ETA regardless of weather — creating a **guaranteed disappointment** for customers ordering during bad weather.\n\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"50%\"\u003e\n\n### 🟡 Finding 3 — Rating-Speed Link\nDelivery partner rating has a **negative correlation** with delivery time — higher-rated partners deliver faster on average. Statistically significant finding that directly informs HR incentive program design.\n\n\u003c/td\u003e\n\u003ctd width=\"50%\"\u003e\n\n### 🩵 Finding 4 — Extreme Delay Outliers\n**~5–8% of all orders** are statistical outliers — extreme delays identified using the IQR method. These orders generate the majority of customer complaints and refund requests despite being a small percentage of volume.\n\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/table\u003e\n\n### ⭐ Finding 5 — City Type Performance Gap\n\n| Rank | City Type | Avg Delivery Time | Std Deviation | Opportunity |\n|------|-----------|------------------|---------------|-------------|\n| 🥇 Best | Urban | 22.98 mins | 8.87 mins | Maintain SLA |\n| 🥈 Mid | Metropolitian | 27.32 mins | 9.18 mins | Optimise routing and traffic handling |\n| 🥉 Worst | Semi-Urban | 49.73 mins | 2.69 mins | Priority operational intervention |\n\n\u003e *City-type-specific SLA framework needed — one delivery promise cannot work for Metro, Urban, and Semi-Urban equally.*\n\u003e *Semi-Urban deliveries take more than double the time of Urban deliveries, suggesting potential infrastructure, routing, or fleet allocation issues*\n\n---\n\n## 📊 Dataset\n\n| Dataset | Source | Rows | Key Columns Used |\n|---------|--------|------|-----------------|\n| Food Delivery Dataset | [Kaggle — Gaurav Malik](https://www.kaggle.com/datasets/gauravmalik26/food-delivery-dataset) | ~45,000 | `Time_taken(min)`, `Weatherconditions`, `Road_traffic_density`, `Delivery_person_ratings`, `Delivery_person_Age`, `City`, `Type_of_vehicle`, `distance(km)` |\n\n**Download the dataset and place it in the `data/` folder as `food_delivery.csv` before running.**\n\n---\n\n## 📁 Project Structure\n\n```\nfood_delivery_analysis/\n│\n├── 📂 data/\n│   └── food_delivery.csv           ← Raw dataset (download from Kaggle)\n│\n├── 📓 analysis.ipynb               ← Main Jupyter Notebook (all 7 problems)\n│\n└── 📂 outputs/                     ← Auto-generated charts\n    ├── P1_delivery_distribution.png\n    ├── P2_weather_impact.png\n    ├── P3_rating_vs_speed.png\n    ├── P4_outlier_detection.png\n    ├── P5_city_performance.png\n    ├── P6_age_vs_speed.png\n    ├── P7_vehicle_efficiency.png\n    └── P_bonus_heatmap.png\n```\n\n---\n\n## 🔬 Methodology\n\n### Feature Engineering — Age Buckets \u0026 Distance Ranges\n\nTwo key engineered features power the HR and Fleet analyses:\n\n```python\n# Age buckets — for HR analytics (Problem 6)\ndf['age_bucket'] = pd.cut(\n    df['Delivery_person_Age'],\n    bins=[18, 25, 30, 35, 50],\n    labels=['18-25', '25-30', '30-35', '35+']\n)\n\n# Distance ranges — for vehicle efficiency cross-analysis (Problem 7)\ndf['distance_range'] = pd.cut(\n    df['distance(km)'],\n    bins=[0, 5, 10, 100],\n    labels=['Short (\u003c5km)', 'Medium (5-10km)', 'Long (\u003e10km)']\n)\n\n# Vehicle × Distance cross-analysis\nvehicle_distance = df.groupby(\n    ['Type_of_vehicle', 'distance_range'], observed=True\n)['Time_taken(min)'].mean().round(2).unstack()\n```\n\n### Data Cleaning Steps\n\n1. Stripped whitespace from all column names using `.str.strip()`\n2. Fixed `Time_taken(min)` column — values stored as `'(min) 24'` extracted using `str.extract(r'(\\d+)')`\n3. Dropped rows where target variable `Time_taken(min)` was null using `dropna(subset=...)`\n4. Converted `Delivery_person_ratings` and `Delivery_person_Age` to numeric using `pd.to_numeric(errors='coerce')`\n5. Stripped whitespace from all string columns using `.str.strip()`\n\n### Outlier Detection — IQR Method\n\n```python\nq1  = df['Time_taken(min)'].quantile(0.25)\nq3  = df['Time_taken(min)'].quantile(0.75)\niqr = q3 - q1\n\nupper_fence = q3 + 1.5 * iqr   # Extreme delay threshold\nlower_fence = q1 - 1.5 * iqr\n\noutliers = df[\n    (df['Time_taken(min)'] \u003e upper_fence) |\n    (df['Time_taken(min)'] \u003c lower_fence)\n]\n```\n\n---\n\n## 📈 Visualisations\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eClick to view all 8 charts\u003c/b\u003e\u003c/summary\u003e\n\n### Chart 1 — Delivery Time Distribution (Histogram + KDE)\n![Distribution](outputs/delivery_distribution.png)\n\n### Chart 2 — Weather Impact (Bar Chart)\n![Weather](outputs/weather_impact.png)\n\n### Chart 3 — Partner Rating vs Speed (Scatter + Trend Line)\n![Rating](outputs/rating_vs_speed.png)\n\n### Chart 4 — Outlier Detection (Box Plot + Strip Plot)\n![Outliers](outputs/outlier_detection.png)\n\n### Chart 5 — City Type Performance (Box Plot)\n![City](outputs/city_performance.png)\n\n### Chart 6 — Partner Age vs Speed (Scatter + Bar)\n![Age](outputs/age_vs_delivery.png)\n\n### Chart 7 — Vehicle Efficiency by Distance (Grouped Bar)\n![Vehicle](outputs/vehicle_efficiency.png)\n\n### Chart 8 — Full Correlation Heatmap\n![Heatmap](outputs/correlation_target.png)\n\n\u003c/details\u003e\n\n---\n\n## 📐 Statistical Validation\n\nAll key findings are backed by statistical methods — not just visual observation.\n\n| Method | Applied To | Result | Interpretation |\n|--------|-----------|--------|----------------|\n| **Mean vs Median comparison** | Delivery time distribution | Mean \u003e Median | Indicates right-skewed distribution with some extreme delays |\n| **Standard Deviation** | Delivery time spread | ~9.4 mins | High inconsistency — some very fast, some very late |\n| **IQR Outlier Detection** | Extreme delay orders | ~0.6% of orders | Statistically proven outlier orders identified |\n| **Pearson Correlation** | Rating vs delivery time | Negative r=-0.339 | Higher rating = faster delivery — statistically linked |\n| **Pearson Correlation** | Age vs delivery time | r = -0.299| Older Partners tend to take slightly longer deliveries |\n| **Skewness** | Delivery time shape | \u003e 0 (right-skewed) | Confirms presence of extreme delay tail |\n| **GroupBy Aggregation** | Weather, City, Vehicle | Per-group means | Prevents large categories distorting raw counts |\n\n\u003e **Why IQR over Standard Deviation for outlier detection?**  \n\u003e Delivery time data is right-skewed — the std deviation method assumes normality and would misclassify many valid orders. IQR is distribution-agnostic and more robust for real-world operational data.\n\n---\n\n## 💡 Recommendations\n\n### 1️⃣ Weather-Aware Dynamic ETA System\nIntegrate a weather API into the ETA calculation. When rain or fog is detected, automatically add the measured extra time (25–35%) to the delivery estimate shown to customers. This costs nothing to implement but directly reduces disappointment and complaint rates.\n\n### 2️⃣ Partner Performance Incentive Program\nCreate a performance score combining rating AND average delivery speed. Introduce bonuses for top-performing partners and targeted training for low-rated or consistently slow partners. Data shows rating correlates with speed — rewarding it drives improvement.\n\n### 3️⃣ Outlier Early Alert System\nBuild a real-time risk score at order placement: Bad Weather + Jam Traffic + Long Distance = high-risk order. Flag these automatically and send a proactive \"your order may take longer\" message before the customer starts waiting and complaining.\n\n### 4️⃣ City-Type-Specific SLA Framework\nSet different ETA promises per city type. Metro cities with high traffic density need more partners per km and real-time routing. Semi-urban zones need different vehicle types and routing tools. One 30-minute promise across all city types will always fail somewhere.\n\n### 5️⃣ Age-Targeted Partner Training\nDesign training programs based on age group performance data — younger partners may need customer service and safety training, older partners may benefit from route optimisation tools and app navigation support.\n\n### 6️⃣ Vehicle-Distance Matching Rules\nAssign vehicle types by delivery zone distance: bicycles for under 3km urban deliveries, bikes and scooters for 3–10km, larger vehicles for 10km+. This operational change improves delivery speed without adding new resources.\n\n### 7️⃣ Data-Driven SLA Targets\nReplace guesswork delivery promises with targets based on actual distribution data — set promises at the 75th percentile delivery time per city type, not a universal \"30 minutes.\" This creates honest, achievable promises.\n\n---\n\n## ▶️ How to Run\n\n### Prerequisites\n```bash\npip install pandas numpy matplotlib seaborn scipy jupyter\n```\n\n### Steps\n\n```bash\n# 1. Clone this repository\ngit clone https://github.com/Nitish2773/food-delivery-delay-analyzer.git\ncd food-delivery-delay-analyzer\n\n# 2. Download dataset from Kaggle and place in data/ folder\n#    - food_delivery.csv → https://www.kaggle.com/datasets/gauravmalik26/food-delivery-dataset\n\n# 3. Launch Jupyter\njupyter notebook\n\n# 4. Open and run analysis.ipynb from top to bottom\n#    All 7 problems run sequentially — charts auto-save to outputs/\n```\n\n\u003e ⚠️ **Important:** Run all cells from top to bottom. Each problem section depends on the cleaning and feature engineering steps above it.\n\n### Requirements\n\n```\npandas\u003e=1.5.0\nnumpy\u003e=1.23.0\nmatplotlib\u003e=3.6.0\nseaborn\u003e=0.12.0\nscipy\u003e=1.9.0\njupyter\u003e=1.0.0\n```\n\n---\n\n## 🛠️ Tech Stack\n\n| Tool | Version | Purpose |\n|------|---------|---------|\n| **Python** | 3.9+ | Core language |\n| **Pandas** | 2.x | Data loading, cleaning, manipulation, GroupBy |\n| **NumPy** | 1.x | Numerical operations, correlation, polyfit |\n| **Matplotlib** | 3.x | Base plotting, figure layout, saving charts |\n| **Seaborn** | 0.12+ | Statistical visualisations (histplot, boxplot, barplot, heatmap, stripplot) |\n| **SciPy** | 1.x | Skewness, kurtosis (`scipy.stats`) |\n| **Jupyter Notebook** | Latest | Interactive development in VS Code |\n\n---\n\n## ⚠️ Limitations\n\n- **No real-time data:** The Kaggle dataset is static — live operational patterns may differ. However, structural findings (weather impact, traffic delays, vehicle efficiency gaps) are expected to persist.\n- **Partner-level, not order-level:** Vehicle and age analyses are at partner level — individual order complexity and route difficulty are not captured.\n- **No actual order volume:** Order count data would be a more precise demand signal than inferring from dataset size alone.\n- **City type classification:** The Metro/Urban/Semi-Urban classification is dataset-provided — ground-truth city infrastructure data would improve city-level analysis.\n- **Correlation, not causation:** All correlation findings (rating vs speed, age vs speed) show statistical relationships — controlled experiments would be needed to establish causation.\n\n---\n\n## 🚀 Future Scope\n\n- [ ] Time-series analysis: peak hour and day-of-week delivery pattern identification\n- [ ] Predictive model: logistic regression to flag orders likely to become outlier delays at placement time\n- [ ] Multi-city replication: Mumbai, Chennai, Hyderabad, Delhi comparison\n- [ ] Dashboard: interactive Power BI or Streamlit dashboard for operations teams\n- [ ] Swiggy comparison: does the same delay pattern exist on a competing platform?\n\n---\n\n## 👤 Author\n\n**Sri Nitish Kamisetti**  \nB.Tech Computer Science Engineering | Batch of 2025  \nGodavari Institute of Engineering and Technology, Rajamahendravaram\n\n[![LinkedIn](https://img.shields.io/badge/LinkedIn-Connect-0A66C2?style=for-the-badge\u0026logo=linkedin\u0026logoColor=white)](https://www.linkedin.com/in/sri-nitish-kamisetti/)\n[![GitHub](https://img.shields.io/badge/GitHub-Follow-181717?style=for-the-badge\u0026logo=github\u0026logoColor=white)](https://github.com/Nitish2773)\n[![Email](https://img.shields.io/badge/Email-Contact-EA4335?style=for-the-badge\u0026logo=gmail\u0026logoColor=white)](mailto:nitishkamisetti123@gmail.com)\n\n---\n\n## 📄 License\n\nThis project is open source and available under the [MIT License](LICENSE).\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\n**If this project helped you, please consider giving it a ⭐**\n\n*Built with curiosity, cleaned with Pandas, validated with statistics.*\n\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fnitish2773%2Ffood-delivery-delay-analyzer","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fnitish2773%2Ffood-delivery-delay-analyzer","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fnitish2773%2Ffood-delivery-delay-analyzer/lists"}