{"id":29534418,"url":"https://github.com/amr-yasser226/machine-learning-for-network-intrusion-detection","last_synced_at":"2025-07-17T00:07:46.750Z","repository":{"id":300403193,"uuid":"948722167","full_name":"amr-yasser226/machine-learning-for-network-intrusion-detection","owner":"amr-yasser226","description":"A complete pipeline for network intrusion detection comparing label encoding and one‑hot encoding, with SMOTE resampling, feature selection, and ensemble modeling using scikit‑learn and XGBoost, also this was phase one of our University's \"CSAI 253- Machine Learning\" course.","archived":false,"fork":false,"pushed_at":"2025-06-21T13:27:46.000Z","size":6785,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-06-21T14:28:05.686Z","etag":null,"topics":["csai-253","cybersecurity","cybersecurity-training","ensamble-methods","feature-engineering","imbalanced-learning","machine-learning","machine-learning-algorithms","network-intrusion-detection","one-hot-encoding","sckit-learn","smote","tree-based-model","xgboost","zewailcity"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/amr-yasser226.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-03-14T20:55:39.000Z","updated_at":"2025-06-21T13:29:59.000Z","dependencies_parsed_at":"2025-06-21T14:38:14.684Z","dependency_job_id":null,"html_url":"https://github.com/amr-yasser226/machine-learning-for-network-intrusion-detection","commit_stats":null,"previous_names":["amr-yasser226/machine-learning-for-network-intrusion-detection"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/amr-yasser226/machine-learning-for-network-intrusion-detection","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amr-yasser226%2Fmachine-learning-for-network-intrusion-detection","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amr-yasser226%2Fmachine-learning-for-network-intrusion-detection/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amr-yasser226%2Fmachine-learning-for-network-intrusion-detection/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amr-yasser226%2Fmachine-learning-for-network-intrusion-detection/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/amr-yasser226","download_url":"https://codeload.github.com/amr-yasser226/machine-learning-for-network-intrusion-detection/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amr-yasser226%2Fmachine-learning-for-network-intrusion-detection/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":265553486,"owners_count":23787083,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["csai-253","cybersecurity","cybersecurity-training","ensamble-methods","feature-engineering","imbalanced-learning","machine-learning","machine-learning-algorithms","network-intrusion-detection","one-hot-encoding","sckit-learn","smote","tree-based-model","xgboost","zewailcity"],"created_at":"2025-07-17T00:07:46.102Z","updated_at":"2025-07-17T00:07:46.744Z","avatar_url":"https://github.com/amr-yasser226.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Machine Learning for Network Intrusion Detection\r\n\r\nThis repository implements a reproducible pipeline to detect network intrusions, comparing **Label Encoding** vs. **One‑Hot Encoding** and culminating in ensemble methods. The notebooks walk through data ingestion, cleaning, feature engineering, model training, and evaluation.\r\n\r\n## Repository Structure\r\n\r\n```\r\n\r\n.\r\n├── Data\r\n│   └── Project\\_Phase1\\_before\\_cleaning.csv\r\n├── LICENSE\r\n├── .gitattributes\r\n├── .gitignore\r\n├── model\r\n│   ├── Final\\_(ADDED\\_LABEL\\_ENCODING).ipynb\r\n│   └── Final\\_(One\\_Hot\\_Encoding).ipynb\r\n├── REPORT.pdf\r\n└── README.md\r\n\r\n````\r\n\r\n- **Data/**  \r\n  Raw CSV dataset (pre‑cleaning).\r\n\r\n- **model/Final_(ADDED_LABEL_ENCODING).ipynb**  \r\n  Full pipeline using **Label Encoding**  \r\n  1. Mount \u0026 load data  \r\n  2. Missing‑value analysis \u0026 outlier handling (IQR + winsorization)  \r\n  3. Data‑type corrections \u0026 new feature creation  \r\n  4. Label encoding + mutual information for feature selection  \r\n  5. SMOTE resampling for class imbalance  \r\n  6. Training and comparing 5 classifiers (RF, KNN, SVM, Logistic Regression, Decision Tree)  \r\n  7. Hyperparameter tuning, feature‑importance filtering, stacking ensembles  \r\n  8. Final metrics \u0026 recommendation  \r\n\r\n- **model/Final_(One_Hot_Encoding).ipynb**  \r\n  Identical pipeline, but uses **One‑Hot Encoding** (drop‑first) instead of label encoding. Facilitates direct comparison of encoding strategies.\r\n\r\n- **REPORT.pdf**  \r\n  Narrative report with tables, charts, and a concise recommendation.\r\n\r\n## Key Results\r\n\r\n| Encoding       | Best Model     | Accuracy | False Negatives | Notes                                  |\r\n|----------------|----------------|---------:|----------------:|----------------------------------------|\r\n| Label Encoding | Random Forest  |   99.75% |               0 | Chosen for zero FN in test set         |\r\n| One‑Hot        | XGBoost        |   99.82% |               1 | Slightly higher accuracy but 1 FN      |\r\n\r\n- **Random Forest (Label Encoding)** achieved 99.75% accuracy with **0 false negatives**, critical for intrusion detection.\r\n- **XGBoost (One‑Hot Encoding)** delivered 99.82% accuracy but incurred 1 false negative.\r\n- All other models (KNN, SVM, Logistic Regression, Decision Tree) performed competitively but with higher FN rates.\r\n- **Stacking** ensembles (RF/DT/SVM) did not improve upon a single Random Forest for zero-FN performance.\r\n\r\n## Quickstart\r\n\r\n```bash\r\ngit clone https://github.com/amr-yasser226/machine-learning-for-network-intrusion-detection.git\r\ncd machine-learning-for-network-intrusion-detection\r\n\r\npython3 -m venv venv\r\nsource venv/bin/activate\r\n\r\njupyter lab\r\n````\r\n\r\nOpen the two notebooks under `model/` and run end-to-end.\r\n\r\n## Dependencies\r\n\r\n* Python 3.8+\r\n* pandas, numpy, matplotlib, seaborn\r\n* scikit‑learn, imbalanced‑learn\r\n* xgboost\r\n\r\n## What This Solves\r\n\r\n* Demonstrates best practices in **EDA**, **feature engineering**, and **model evaluation**\r\n* Compares two encoding strategies for categorical data\r\n* Addresses class imbalance with **SMOTE**\r\n* Benchmarks multiple classifiers and stacking ensembles\r\n* Prioritizes **zero false negatives**—paramount in intrusion detection\r\n\r\n## License\r\n\r\nThis project is released under the MIT License. See [LICENSE](LICENSE) for details.\r\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Famr-yasser226%2Fmachine-learning-for-network-intrusion-detection","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Famr-yasser226%2Fmachine-learning-for-network-intrusion-detection","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Famr-yasser226%2Fmachine-learning-for-network-intrusion-detection/lists"}