{"id":22117452,"url":"https://github.com/benitomartin/mlops-databricks-credit-default","last_synced_at":"2025-04-10T05:07:53.882Z","repository":{"id":264692886,"uuid":"894011114","full_name":"benitomartin/mlops-databricks-credit-default","owner":"benitomartin","description":"End-to-end MLOps Credit Default Project using DABs","archived":false,"fork":false,"pushed_at":"2024-11-29T10:43:47.000Z","size":1132,"stargazers_count":17,"open_issues_count":0,"forks_count":16,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-04-10T05:07:48.829Z","etag":null,"topics":["aws","continuous-deployment","continuous-integration","databricks","mlflow","precommit-hooks","pydantic","python","ruff","uv"],"latest_commit_sha":null,"homepage":"https://medium.com/marvelous-mlops/building-an-end-to-end-mlops-project-with-databricks-8cd9a85cc3c0","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/benitomartin.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-11-25T15:40:38.000Z","updated_at":"2025-03-12T14:43:38.000Z","dependencies_parsed_at":"2024-12-01T23:01:40.350Z","dependency_job_id":null,"html_url":"https://github.com/benitomartin/mlops-databricks-credit-default","commit_stats":null,"previous_names":["benitomartin/mlops-databricks-credit-default"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benitomartin%2Fmlops-databricks-credit-default","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benitomartin%2Fmlops-databricks-credit-default/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benitomartin%2Fmlops-databricks-credit-default/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benitomartin%2Fmlops-databricks-credit-default/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/benitomartin","download_url":"https://codeload.github.com/benitomartin/mlops-databricks-credit-default/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248161273,"owners_count":21057555,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["aws","continuous-deployment","continuous-integration","databricks","mlflow","precommit-hooks","pydantic","python","ruff","uv"],"created_at":"2024-12-01T13:34:06.265Z","updated_at":"2025-04-10T05:07:53.861Z","avatar_url":"https://github.com/benitomartin.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"## MLOps Credit Default ✈️\n\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"737\" alt=\"cover\" src=\"https://github.com/user-attachments/assets/a1c18fba-9e39-45b5-8fcd-bceb1f5f5af9\"\u003e\n\u003c/p\u003e\n\nThis is a personal MLOps project based on a [Kaggle](https://www.kaggle.com/datasets/uciml/default-of-credit-card-clients-dataset/data) dataset for credit default predictions.\n\nIt was developed as part of the this [End-to-end MLOps with Databricks](https://maven.com/marvelousmlops/mlops-with-databricks) course and you can walk through it together with this [Medium](https://medium.com/@benitomartin/8cd9a85cc3c0) publication.\n\nFeel free to ⭐ and clone this repo 😉\n\n## Tech Stack\n\n![Visual Studio Code](https://img.shields.io/badge/Visual%20Studio%20Code-0078d7.svg?style=for-the-badge\u0026logo=visual-studio-code\u0026logoColor=white)\n![Jupyter Notebook](https://img.shields.io/badge/jupyter-%23FA0F00.svg?style=for-the-badge\u0026logo=jupyter\u0026logoColor=white)\n![Python](https://img.shields.io/badge/python-3670A0?style=for-the-badge\u0026logo=python\u0026logoColor=ffdd54)\n![Pandas](https://img.shields.io/badge/pandas-%23150458.svg?style=for-the-badge\u0026logo=pandas\u0026logoColor=white)\n![NumPy](https://img.shields.io/badge/numpy-%23013243.svg?style=for-the-badge\u0026logo=numpy\u0026logoColor=white)\n![Databricks](https://img.shields.io/badge/Databricks-FF3621?style=for-the-badge\u0026logo=Databricks\u0026logoColor=white)\n![scikit-learn](https://img.shields.io/badge/scikit--learn-%23F7931E.svg?style=for-the-badge\u0026logo=scikit-learn\u0026logoColor=white)\n![Spark](https://img.shields.io/badge/Apache_Spark-FFFFFF?style=for-the-badge\u0026logo=apachespark\u0026logoColor=#E35A16)\n![MLflow](https://img.shields.io/badge/MLflow-0194E2.svg?style=for-the-badge\u0026logo=MLflow\u0026logoColor=white)\n![Anaconda](https://img.shields.io/badge/Anaconda-%2344A833.svg?style=for-the-badge\u0026logo=anaconda\u0026logoColor=white)\n![Linux](https://img.shields.io/badge/Linux-FCC624?style=for-the-badge\u0026logo=linux\u0026logoColor=white)\n![AWS](https://img.shields.io/badge/AWS-%23FF9900.svg?style=for-the-badge\u0026logo=amazon-aws\u0026logoColor=white)\n![Git](https://img.shields.io/badge/git-%23F05033.svg?style=for-the-badge\u0026logo=git\u0026logoColor=white)\n\n## Project Structure\n\nThe project has been structured with the following folders and files:\n\n- `.github/workflows`: CI/CD configuration files\n  - `cd.yml`\n  - `ci.yml`\n- `data`: raw data\n  - `data.csv`\n- `notebooks`: notebooks for various stages of the project\n  - `create_source_data`: notebook for generating synthetic data\n    - `create_source_data_notebook.py`\n  - `feature_engineering`: feature engineering and MLflow experiments\n    - `basic_mlflow_experiment_notebook.py`\n    - `combined_mlflow_experiment_notebook.py`\n    - `custom_mlflow_experiment_notebook.py`\n    - `prepare_data_notebook.py`\n  - `model_feature_serving`: notebooks for serving models and features\n    - `AB_test_model_serving_notebbok.py`\n    - `feature_serving_notebook.py`\n    - `model_serving_feat_lookup_notebook.py`\n    - `model_serving_notebook.py`\n  - `monitoring`: monitoring and alerts setup\n    - `create_alert.py`\n    - `create_inference_data.py`\n    - `lakehouse_monitoring.py`\n    - `send_request_to_endpoint.py`\n- `src`: source code for the project\n  - `credit_default`\n    - `data_cleaning.py`\n    - `data_cleaning_spark.py`\n    - `data_preprocessing.py`\n    - `data_preprocessing_spark.py`\n    - `utils.py`\n- `tests`: unit tests for the project\n  - `test_data_cleaning.py`\n  - `test_data_preprocessor.py`\n- `workflows`: workflows for Databricks asset bundle\n  - `deploy_model.py`\n  - `evaluate_model.py`\n  - `preprocess.py`\n  - `refresh_monitor.py`\n  - `train_model.py`\n- `.pre-commit-config.yaml`: configuration for pre-commit hooks\n- `Makefile`: helper commands for installing requirements, formatting, testing, linting, and cleaning\n- `project_config.yml`: configuration settings for the project\n- `databricks.yml`: Databricks asset bundle configuration\n- `bundle_monitoring.yml`: monitoring settings for Databricks asset bundle\n\n## Project Set Up\n\nThe Python version used for this project is Python 3.11.\n\n1. Clone the repo:\n\n   ```bash\n   git clone https://github.com/benitomartin/mlops-databricks-credit-default.git\n   ```\n\n2. Create the virtual environment using `uv` with Python version 3.11 and install the requirements:\n\n   ```bash\n    uv venv -p 3.11.0 .venv\n    source .venv/bin/activate\n    uv pip install -r pyproject.toml --all-extras\n    uv lock\n    ```\n\n3. Build the wheel package:\n\n    ```bash\n    # Build\n    uv build\n    ```\n\n4. Install the Databricks extension for VS Code and Databricks CLI:\n\n   ```bash\n   curl -fsSL https://raw.githubusercontent.com/databricks/setup-cli/main/install.sh | sh\n   ```\n\n5. Authenticate on Databricks:\n\n   ```bash\n   # Authentication\n   databricks auth login --configure-cluster --host \u003cworkspace-url\u003e\n\n   # Profiles\n   databricks auth profiles\n   cat ~/.databrickscfg\n   ```\n\nAfter entering your information, the CLI will prompt you to save it under a Databricks configuration profile `~/.databrickscfg`\n\n\n## Catalog Set Up\n\nOnce the project is set up, you need to create the volumes to store the data and the wheel package that will you have to install in the cluster:\n\n- **catalog name**: *credit*\n- **schema_name**: *default*\n- **volume name**: *data* and *packages*\n\n  ```bash\n  # Create volumes\n  databricks volumes create credit default data MANAGED\n  databricks volumes create credit default packages MANAGED\n\n  # Push volumes\n  databricks fs cp data/data.csv dbfs:/Volumes/credit/default/data/data.csv\n  databricks fs cp dist/credit_default_databricks-0.0.1-py3-none-any.whl dbfs:/Volumes/credit/default/packages\n\n  # Show volumes\n  databricks fs ls dbfs:/Volumes/credit/default/data\n  databricks fs ls dbfs:/Volumes/credit/default/packages\n  ```\n\n## Token Creation\n\nSome project files require a Databricks authentication token. This token allows secure access to Databricks resources and APIs:\n\n1. Create a token in the Databricks UI:\n\n   - Navigate to `Settings` --\u003e `User` --\u003e `Developer` --\u003e `Access tokens`\n\n   - Generate a new personal access token\n\n2. Create a secret scope for securely storing the token:\n\n    ```bash\n    # Create Scope\n    databricks secrets create-scope secret-scope\n\n    # Add secret after running command\n    databricks secrets put-secret secret-scope databricks-token\n\n    # List secrets\n    databricks secrets list-secrets secret-scope\n    ```\n\n**Note**: For GitHub Actions (in `cd.yml`), the token must also be added as a GitHub Secret in your repository settings.\n\nNow you can follow the code along the [Medium](https://medium.com/@benitomartin/8cd9a85cc3c0) publication or use it as supporting material if you enroll in the [course](https://maven.com/marvelousmlops/mlops-with-databricks). The blog does not contain an explanation of all files. Just the main ones used for the final deployment, but you can test out other files as well 🙂.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbenitomartin%2Fmlops-databricks-credit-default","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fbenitomartin%2Fmlops-databricks-credit-default","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbenitomartin%2Fmlops-databricks-credit-default/lists"}