{"id":19721212,"url":"https://github.com/niteshchawla/clustering-ml","last_synced_at":"2026-04-17T03:35:16.086Z","repository":{"id":248735993,"uuid":"829556895","full_name":"Niteshchawla/Clustering-ML","owner":"Niteshchawla","description":"Analyzing the vast data of learners can uncover patterns in their professional backgrounds and preferences. Allowing Scaler to make tailored content recommendations and provide specialized mentorship.","archived":false,"fork":false,"pushed_at":"2024-10-01T10:54:22.000Z","size":13327,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-02-28T00:16:10.635Z","etag":null,"topics":["cluster-analysis","clustering","hierarchical-clustering","k-means-clustering","machine-learning","numpy","pca-analysis","visualisation"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Niteshchawla.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-07-16T17:09:05.000Z","updated_at":"2024-10-01T10:54:26.000Z","dependencies_parsed_at":"2025-02-27T18:34:24.393Z","dependency_job_id":"fd898872-8e34-4f14-ac41-e2184c97917e","html_url":"https://github.com/Niteshchawla/Clustering-ML","commit_stats":null,"previous_names":["niteshchawla/clustering-ml"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/Niteshchawla/Clustering-ML","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Niteshchawla%2FClustering-ML","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Niteshchawla%2FClustering-ML/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Niteshchawla%2FClustering-ML/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Niteshchawla%2FClustering-ML/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Niteshchawla","download_url":"https://codeload.github.com/Niteshchawla/Clustering-ML/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Niteshchawla%2FClustering-ML/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31913771,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-16T18:22:33.417Z","status":"online","status_checked_at":"2026-04-17T02:00:06.879Z","response_time":62,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["cluster-analysis","clustering","hierarchical-clustering","k-means-clustering","machine-learning","numpy","pca-analysis","visualisation"],"created_at":"2024-11-11T23:13:41.197Z","updated_at":"2026-04-17T03:35:16.063Z","avatar_url":"https://github.com/Niteshchawla.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Clustering-ML\n**Problem Statement**\n\nScaler is an online tech-versity offering intensive computer science \u0026 Data Science courses through live classes delivered by tech leaders and subject matter experts. The meticulously structured program enhances the skills of software professionals by offering a modern curriculum with exposure to the latest technologies. It is a product by InterviewBit.\n\nYou are working as a data scientist with the analytics vertical of Scaler, focused on profiling the best companies and job positions to work for from the Scaler database. You are provided with the information for a segment of learners and tasked to cluster them on the basis of their job profile, company, and other features. Ideally, these clusters should have similar characteristics.\n\n\n**Dataset:**\n\nDataset Link: scaler_kmeans.csv\n\n\n**Data Dictionary:**\n\n‘Unnamed 0’ - Index of the dataset\n\nEmail_hash - Anonymised Personal Identifiable Information (PII)\n\nCompany_hash - This represents an anonymized identifier for the company, which is the current employer of the learner.\n\norgyear - Employment start date\n\nCTC - Current CTC\n\nJob_position - Job profile in the company\n\nCTC_updated_year - Year in which CTC got updated (Yearly increments, Promotions)\n\n\n**Concept Used:**\n\nManual Clustering\n\nUnsupervised Clustering - K- means, Hierarchical Clustering\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fniteshchawla%2Fclustering-ml","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fniteshchawla%2Fclustering-ml","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fniteshchawla%2Fclustering-ml/lists"}