{"id":25105178,"url":"https://github.com/soroushesnaashari/customer-clustering","last_synced_at":"2026-03-03T21:01:27.935Z","repository":{"id":264665181,"uuid":"894002489","full_name":"soroushesnaashari/Customer-Clustering","owner":"soroushesnaashari","description":"An Unsupervised Machine Learning project","archived":false,"fork":false,"pushed_at":"2024-12-16T13:05:48.000Z","size":6708,"stargazers_count":5,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-06-17T05:04:23.346Z","etag":null,"topics":["dbscan","gmm","kmeans","machine-learning"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/soroushesnaashari.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2024-11-25T15:23:40.000Z","updated_at":"2025-02-10T09:06:46.000Z","dependencies_parsed_at":"2025-04-20T15:00:18.863Z","dependency_job_id":null,"html_url":"https://github.com/soroushesnaashari/Customer-Clustering","commit_stats":null,"previous_names":["soroushesnaashari/customer-clustering"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/soroushesnaashari/Customer-Clustering","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soroushesnaashari%2FCustomer-Clustering","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soroushesnaashari%2FCustomer-Clustering/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soroushesnaashari%2FCustomer-Clustering/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soroushesnaashari%2FCustomer-Clustering/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/soroushesnaashari","download_url":"https://codeload.github.com/soroushesnaashari/Customer-Clustering/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/soroushesnaashari%2FCustomer-Clustering/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":30060626,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-03T18:21:05.932Z","status":"ssl_error","status_checked_at":"2026-03-03T18:20:59.341Z","response_time":61,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["dbscan","gmm","kmeans","machine-learning"],"created_at":"2025-02-07T22:42:40.325Z","updated_at":"2026-03-03T21:01:27.909Z","avatar_url":"https://github.com/soroushesnaashari.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"## Customer Clustering Project\n[![](Image.jpg)](https://unsplash.com/photos/group-of-people-standing-in-front-of-food-stall-counter-66RxrYlPShI)\n\n### Overview\nThis project applies **Unsupervised Machine Learning** techniques to a customer dataset to identify meaningful clusters. The goal is to segment customers based on their features, using some popular clustering algorithms such as **K-Means**, **DBSCAN** and **Gaussian Mixture Models (GMM)**. The clustering results are evaluated and visualized to extract actionable insights.\n\n\u003cbr\u003e\n\n### Process Workflow\nThe project follows a structured approach for clustering:\n\n1. **Data Cleaning:** \u003cbr\u003e\nImported the dataset and handled missing values using *dropna()*.\n\n2. **Data Visualization:** \u003cbr\u003e\nExplored the dataset through detailed visualizations (scatter plots, pair plots, bar charts, etc.) to understand feature relationships.\n\n3. **Data Scaling:** \u003cbr\u003e\nApplied *MinMaxScaler* to normalize the dataset to the range [0, 1] for better performance with clustering algorithms.\n\n4. **Clustering Algorithms:** \u003cbr\u003e\nPerformed clustering using the following methods:\n\n  - *K-Means Clustering:* Identified optimal clusters using Silhouette Score Analysis.\n  - *DBSCAN:* Explored density-based clustering with parameter tuning (eps and min_samples).\n  - *Gaussian Mixture Models (GMM):* Modeled clusters using probability distributions.\n\n5. **Evaluation:** \u003cbr\u003e\nUsed *Silhouette Score* to evaluate the clustering performance for all algorithms.\n\n6. **Parameter Optimization:** \u003cbr\u003e\nConducted parameter tuning for *K-Means*, *DBSCAN* and *GMM* to find the optimal number of clusters and density thresholds.\n\n7. **Visualization:** \u003cbr\u003e\nCreated various plots to interpret and compare clustering results:\n\n  - Scatter plots with PCA reduction.\n  - Heatmaps of cluster feature averages.\n  - Barplot to illustrate the cluster's size.\n\n\u003cbr\u003e\n\n### Results\nAfter evaluation and parameter tuning:\n\n  - **K-Means** provided the best clustering results with:\n    - Optimal Number of Clusters: 2\n    - Silhouette Score: 0.3906\n  - **DBSCAN** struggled with the dataset, forming a single cluster and noise.\n  - **GMM** gave reasonable results but was outperformed by K-Means in this case.\n\n\u003cbr\u003e\n\n### Repository Contents\n\n- **Data:** Contains the [Original Dataset](https://www.kaggle.com/datasets/mahnazarjmand/customer-segmentation) and you can see the cleaned dataset in notebook.\n\n- **Notebook:** Jupyter notebook detailing the entire process, including data cleaning, visualization, clustering, and evaluation.\n\n- **README.md:** Project documentation.\n\n\u003cbr\u003e\n\n### How to Contribute\nContributions are welcome! If you'd like to improve the project or add new features:\n\n1. **Fork the repository.**\n2. **Create a new branch.**\n3. **Make your changes and submit a pull request.**\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsoroushesnaashari%2Fcustomer-clustering","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsoroushesnaashari%2Fcustomer-clustering","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsoroushesnaashari%2Fcustomer-clustering/lists"}