{"id":24086396,"url":"https://github.com/zuzann18/customerbehavioranalysis-sas","last_synced_at":"2026-03-03T02:36:09.403Z","repository":{"id":233983260,"uuid":"767194700","full_name":"zuzann18/CustomerBehaviorAnalysis-SAS","owner":"zuzann18","description":"The aim of this project was to identify and analyze predictors capable of forecasting customers credit defaults based on historical data. A key element involved analyzing the impact of various predictors on the target variable, which indicates whether a client defaults after 12 months.","archived":false,"fork":false,"pushed_at":"2024-04-21T21:18:19.000Z","size":4719,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-01-10T01:57:38.673Z","etag":null,"topics":["customerbehaviour","sas"],"latest_commit_sha":null,"homepage":"","language":"SAS","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/zuzann18.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2024-03-04T21:37:23.000Z","updated_at":"2024-04-21T21:18:22.000Z","dependencies_parsed_at":"2024-04-17T22:05:10.923Z","dependency_job_id":null,"html_url":"https://github.com/zuzann18/CustomerBehaviorAnalysis-SAS","commit_stats":null,"previous_names":["zuzann18/customerbehavioranalysis-sas"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zuzann18%2FCustomerBehaviorAnalysis-SAS","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zuzann18%2FCustomerBehaviorAnalysis-SAS/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zuzann18%2FCustomerBehaviorAnalysis-SAS/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/zuzann18%2FCustomerBehaviorAnalysis-SAS/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/zuzann18","download_url":"https://codeload.github.com/zuzann18/CustomerBehaviorAnalysis-SAS/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":240972315,"owners_count":19886937,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["customerbehaviour","sas"],"created_at":"2025-01-10T01:57:51.066Z","updated_at":"2026-03-03T02:36:09.373Z","avatar_url":"https://github.com/zuzann18.png","language":"SAS","funding_links":[],"categories":[],"sub_categories":[],"readme":"# CustomerBehaviorAnalysis-SAS\n## Forecasting customers credit defaults based on historical data\n\n### Project Objective\nThe aim of this project was to identify and analyze predictors capable of forecasting customers credit defaults based on historical data. A key element involved analyzing the impact of various predictors on the target variable, default_cus12, which indicates whether a client defaults after 12 months.\n\n## Statistical Methods\nData Missingness Analysis:\nAn analysis of data missingness was conducted to understand data stability over time and improve data quality by:\n\nRemoving observations with missing values in the key target variable.\nCleaning data, marking missing data, and deriving distinct periods and customer default statuses.\nOutlier Analysis\nOutlying observations were examined to eliminate errors and anomalies in the data, using quartile methods to identify outliers.\n\n## Exploratory Data Analysis (EDA) Techniques\n**Frequency Analysis:** Used to get an overview of the distribution of categorical data, which aids in understanding the diversity and prevalence of categorical features.\n**Univariate Analysis**: Histograms and summary statistics (mean, standard deviation) were employed to understand the distribution of continuous variables and to check for normality.\n**Descriptive Statistics**: Insights into central tendencies and variability were provided\nKolmogorov-Smirnov Test: Used to assess the consistency of distributions, which is important for verifying data homogeneity across datasets.\n## Predictor Analysis\nGini Coefficient and V-Cramer's Coefficient: Applied to evaluate the strength of the relationship between individual variables and the target variable, crucial for building an effective predictive model.\nDependent Variable\nThe dependent variable, default_cus12, is a categorical/binary (qualitative) variable that indicates whether a client defaults on financial obligations after 12 months. This classification is essential in credit risk modeling.\n\nApplication of the Gini Coefficient\nThe Gini coefficient was used to assess the discriminatory power of predictors within the context of credit risk modeling:\n\n## Model Evaluation\nSingle-Variable Logistic Regression Models: Each model used one predictor at a time with default_cus12 as the dependent variable to evaluate performance.\n## Predictive Power Assessment\nCoefficient Range: Values ranged from 0 to 1, where a higher Gini coefficient indicates a stronger ability to differentiate between the outcomes (default or no default).\nThreshold Determination\n## Cutoff Point: Established at a Gini coefficient of 0.4 to select the best predictors from the dataset for further analysis.\n### Comparative Analysis Ranking Predictors: \nThe Gini coefficient was calculated for different predictors to rank and prioritize variables based on their effectiveness in predicting customer defaults.\nImportance of Gini Coefficient in Credit Risk Modeling\nModel Quality: Indicates the model's ability to correctly classify individuals into default or no default categories.\nDecision Making: Financial institutions use these insights to make informed credit granting decisions, reducing the risk of defaults.\n\n\nThe findings from the project based on V-Cramer's coefficient were specifically aimed at assessing the strength of the relationship between textual (categorical) variables and the target variable default_cus12. Here’s a summary of the findings derived from using V-Cramer's coefficient:\n\n### Strength of Association:\n\nV-Cramer's coefficient was used to measure the association strength between various categorical variables and the likelihood of default (default_cus12). This included variables such as employment type, marital status, and residential status among others.\nSelection of Predictive Variables:\n\nVariables that exhibited higher V-Cramer's coefficients were considered strong predictors and were thus prioritized for inclusion in the predictive modeling. This approach helped in refining the model by focusing on variables that have a more substantial impact on the prediction of default.\nEnhanced Model Accuracy:\n\nBy identifying and including variables with significant V-Cramer's coefficients, the project aimed to enhance the accuracy of the predictive models. This is because these variables have a demonstrated strong relationship with the outcome, making them valuable for improving model performance.\nDecision-Making Insights:\n\nThe analysis provided insights into which customer characteristics are most predictive of default, aiding in decision-making processes related to credit risk management. For instance, if a variable like marital status showed a high V-Cramer's coefficient, it indicated a significant impact on default likelihood, which could influence credit policies and risk assessments.\nThe use of V-Cramer's coefficient thus played a crucial role in identifying and validating the importance of various categorical variables in the context of credit default risk, leading to more focused and effective predictive analytics\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzuzann18%2Fcustomerbehavioranalysis-sas","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fzuzann18%2Fcustomerbehavioranalysis-sas","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fzuzann18%2Fcustomerbehavioranalysis-sas/lists"}