{"id":16085721,"url":"https://github.com/iamraphson/de-2024-project-book-recommendation","last_synced_at":"2025-10-08T08:45:55.281Z","repository":{"id":229739267,"uuid":"777520870","full_name":"iamraphson/DE-2024-project-book-recommendation","owner":"iamraphson","description":null,"archived":false,"fork":false,"pushed_at":"2024-03-31T03:04:30.000Z","size":1894,"stargazers_count":22,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-06-09T07:04:06.126Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/iamraphson.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-03-26T02:09:08.000Z","updated_at":"2025-04-08T12:15:45.000Z","dependencies_parsed_at":"2024-04-03T12:30:38.598Z","dependency_job_id":null,"html_url":"https://github.com/iamraphson/DE-2024-project-book-recommendation","commit_stats":null,"previous_names":["iamraphson/de-2024-project-book-recommendation"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/iamraphson/DE-2024-project-book-recommendation","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/iamraphson%2FDE-2024-project-book-recommendation","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/iamraphson%2FDE-2024-project-book-recommendation/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/iamraphson%2FDE-2024-project-book-recommendation/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/iamraphson%2FDE-2024-project-book-recommendation/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/iamraphson","download_url":"https://codeload.github.com/iamraphson/DE-2024-project-book-recommendation/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/iamraphson%2FDE-2024-project-book-recommendation/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":278916442,"owners_count":26068090,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-08T02:00:06.501Z","response_time":56,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-10-09T13:09:09.025Z","updated_at":"2025-10-08T08:45:55.212Z","avatar_url":"https://github.com/iamraphson.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Data Pipeline Project for Book Recommendation\n\n\u003cdetails\u003e\n    \u003csummary\u003eTable of Contents\u003c/summary\u003e\n    \u003col\u003e\n        \u003cli\u003e\n            \u003ca href=\"#introduction\"\u003eIntroduction\u003c/a\u003e\n            \u003cul\u003e\n                \u003cli\u003e\u003ca href=\"#built-with\"\u003eBuilt With\u003c/a\u003e\u003c/li\u003e\n            \u003c/ul\u003e\n        \u003c/li\u003e\n        \u003cli\u003e\n            \u003ca href=\"#project-architecture\"\u003eProject Architecture\u003c/a\u003e\n        \u003c/li\u003e\n        \u003cli\u003e\n             \u003ca href=\"#getting-started\"\u003eGetting Started\u003c/a\u003e\n             \u003cul\u003e\n                \u003cli\u003e\n                \u003ca href=\"#create-a-google-cloud-project\"\u003eCreate a Google Cloud Project\u003ca\u003e\n                \u003c/li\u003e\n                \u003cli\u003e\n                \u003ca href=\"#set-up-kaggle\"\u003eSet up Kaggle\u003ca\u003e\n                \u003c/li\u003e\n                \u003cli\u003e\n                    \u003ca href=\"#set-up-the-infrastructure-on-GCP-with-terraform\"\u003eSet up the infrastructure on GCP with Terraform\u003c/a\u003e\n                \u003c/li\u003e\n                \u003cli\u003e\n                    \u003ca href=\"#set-up-airflow-and-metabase\"\u003eSet up Airflow and Metabase\u003c/a\u003e\n                \u003c/li\u003e\n            \u003c/ul\u003e\n        \u003c/li\u003e\n        \u003cli\u003e\n            \u003ca href=\"#data-ingestion\"\u003eData Ingestion\u003c/a\u003e\n        \u003c/li\u003e\n        \u003cli\u003e\n            \u003ca href=\"#data-transformation\"\u003eData Transformation\u003c/a\u003e\n        \u003c/li\u003e\n        \u003cli\u003e\n            \u003ca href=\"#data-visualization\"\u003eData Visualization\u003c/a\u003e\n        \u003c/li\u003e\n        \u003cli\u003e\n            \u003ca href=\"#contact\"\u003eContact\u003c/a\u003e\n        \u003c/li\u003e\n         \u003cli\u003e\n            \u003ca href=\"#acknowledgments\"\u003eAcknowledgments\u003c/a\u003e\n        \u003c/li\u003e\n    \u003c/ol\u003e\n\u003c/details\u003e\n\n## Introduction\n\nThis project is part of the [Data Engineering Zoomcamp](https://github.com/DataTalksClub/data-engineering-zoomcamp). As part of the project, I developed a data pipeline to load and process data from a Kaggle dataset containing bookstore information for a book recommendation system. The dataset can be accessed on [this Kaggle](https://www.kaggle.com/datasets/arashnic/book-recommendation-dataset/).\n\nThis dataset offers book ratings from users of various ages and geographical locations. It comprises three files: User.csv, which includes age and location data of bookstore users; Books.csv, containing information such as authors, titles, and ISBNs; and Ratings.csv, which details the ratings given by users for each book. Additional information about the dataset is available on Kaggle.\n\nThe primary objective of this project is to establish a streamlined data pipeline for obtaining, storing, cleansing, and visualizing data automatically. This pipeline aims to address various queries, such as identifying top-rated publishers and authors, and analyzing ratings based on geographical location.\n\nGiven that the data is static, the data pipeline operates as a one-time process.\n\n### Built With\n\n- Dataset repo: [Kaggle](https://www.kaggle.com)\n- Infrastructure as Code: [Terraform](https://www.terraform.io/)\n- Workflow Orchestration: [Airflow](https://airflow.apache.org)\n- Data Lake: [Google Cloud Storage](https://cloud.google.com/storage)\n- Data Warehouse: [Google BigQuery](https://cloud.google.com/bigquery)\n- Transformation: [DBT](https://www.getdbt.com/)\n- Visualisation: [Metabase](https://www.metabase.com/)\n- Programming Language: Python and SQL\n\n## Project Architecture\n\n![architecture](./screenshots/architecture.png)\nCloud infrastructure is set up with Terraform.\n\nAirflow is run on a local docker container.\n\n## Getting Started\n\n### Prerequisites\n\n1. A [Google Cloud Platform](https://cloud.google.com/) account.\n2. A [kaggle](https://www.kaggle.com/) account.\n3. Install VSCode or [Zed](https://zed.dev/) or any other IDE that works for you.\n4. [Install Terraform](https://www.terraform.io/downloads)\n5. [Install Docker Desktop](https://docs.docker.com/get-docker/)\n6. [Install Google Cloud SDK](https://cloud.google.com/sdk)\n7. Clone this repository onto your local machine.\n\n### Create a Google Cloud Project\n\n- Go to [Google Cloud](https://console.cloud.google.com/) and create a new project.\n- Get the project ID and define the environment variables `GCP_PROJECT_ID` in the .env file located in the root directory\n- Create a [Service account](https://cloud.google.com/iam/docs/service-account-overview) with the following roles:\n  - `BigQuery Admin`\n  - `Storage Admin`\n  - `Storage Object Admin`\n  - `Viewer`\n- Download the Service Account credentials and store it in `$HOME/.google/credentials/`.\n- You need to activate the following APIs [here](https://console.cloud.google.com/apis/library/browse)\n  - Cloud Storage API\n  - BigQuery API\n- Assign the `GOOGLE_APPLICATION_CREDENTIALS` environment variable to the path of your JSON credentials file, such that `GOOGLE_APPLICATION_CREDENTIALS` will be $HOME/.google/credentials/\u003cauthkeys_filename\u003e.json\n  - add this line to the end of the `.bashrc` file\n  ```bash\n  export GOOGLE_APPLICATION_CREDENTIALS=${HOME}/.google/google_credentials.json\n  ```\n  - Activate the enviroment variable by runing `source .bashrc`\n\n### Set up kaggle\n\n- A detailed description on how to authenicate is found [here](https://www.kaggle.com/docs/api)\n- Define the environment variables `KAGGLE_USER` and `KAGGLE_TOKEN` in the .env file located in the root directory. Note: `KAGGLE_TOKEN` is the same as `KAGGLE_KEY`\n\n### Set up the infrastructure on GCP with Terraform\n\n- Using Zed or VSCode, open the cloned project `DE-2024-project-bookrecommendation`.\n- To customize the default values of `variable \"project\"` and `variable \"region\"` to your preferred project ID and region, you have two options: either edit the variables.tf file in Terraform directly and modify the values, or set the environment variables `TF_VAR_project` and `TF_VAR_region`.\n- Open the terminal to the root project.\n- Navigate to the root directory of the project in the terminal and then change the directory to the terraform folder using the command `cd terraform`.\n- Set an alias `alias tf='terraform'`\n- Initialise Terraform: `tf init`\n- Plan the infrastructure: `tf plan`\n- Apply the changes: `tf apply`\n\n### Set up Airflow and Metabase\n\n- Please confirm that the following environment variables are configured in `.env` in the root directory of the project.\n  - `AIRFLOW_UID`. The default value is 50000\n  - `KAGGLE_USERNAME`. This should be set from [Set up kaggle](#set-up-kaggle) section.\n  - `KAGGLE_TOKEN`. This should be set from [Set up kaggle](#set-up-kaggle) section too\n  - `GCP_PROJECT_ID`. This should be set from [Create a Google Cloud Project](#create-a-google-cloud-project) section\n  - `GCP_BOOK_RECOMMENDATION_BUCKET=book_recommendation_datalake_\u003cGCP project id\u003e`\n  - `GCP_BOOK_RECOMMENDATION_WH_DATASET=book_recommendation_analytics`\n  - `GCP_BOOK_RECOMMENDATION_WH_EXT_DATASET=book_recommendataion_wh`\n- Run `docker-compose up`.\n- Access the Airflow dashboard by visiting `http://localhost:8080/` in your web browser. The interface will resemble the following. Use the username and password airflow to log in.\n\n![Airflow](./screenshots/airflow_home.png)\n\n- Visit `http://localhost:1460` in your web browser to access the Metabase dashboard. The interface will resemble the following. You will need to sign up to use the UI.\n\n![Metabase](./screenshots/metabase_home.png)\n\n## Data Ingestion\n\nOnce you've completed all the steps outlined in the previous section, you should now be able to view the Airflow dashboard in your web browser. Below will display as list of DAGs\n![DAGS](./screenshots/dags_index.png)\nBelow is the DAG's graph.\n![DAG Graph](./screenshots/dag_graph.png)\nTo run the DAG, Click on the play button(Figure 1)\n![Run Graph](./screenshots/run_dag.png)\n\n## Data Transformation\n\n- Navigate to the root directory of the project in the terminal and then change the directory to the terraform folder using the command `cd data_dbt`.\n- Generate a profiles.yml file within `${HOME}/.dbt`, followed by defining a profile for this project as instructed below.\n\n```yaml\ndata_dbt_book_recommendation:\n  outputs:\n    dev:\n      dataset: book_recommendation_analytics\n      fixed_retries: 1\n      keyfile: \u003clocation_google_auth_key\u003e\n      location: \u003cpreferred project region\u003e\n      method: service-account\n      priority: interactive\n      project: \u003cpreferred project id\u003e\n      threads: 6\n      timeout_seconds: 300\n      type: bigquery\n  target: dev\n```\n\n- To run all models, run `dbt run -t dev`\n- Navigate to your Google [BigQuery](https://console.cloud.google.com/bigquery) project by clicking on this link. There, you'll find all the tables and views created by DBT.\n  ![Big Query](./screenshots/bigquery_schema_1.png)\n\n## Data Visualization\n\nPlease watch the [provided video tutorial](https://youtu.be/BnLkrA7a6gM\u0026) to configure your Metabase database connection with BigQuery.You have the flexibility to customize your dashboard according to your preferences. Additionally, this [PDF](./screenshots/DE_2024_Dashboard.pdf) linked below contains the complete screenshot of the dashboard I created.\n\n![Dashboard](./screenshots/DE_2024_Dashboard.png)\n\n## Contact\n\nTwitter: [@iamraphson](https://twitter.com/iamraphson)\n\n## Acknowledgments\n\nI would like to extend my heartfelt gratitude to the organizers of the [Data Engineering Zoomcamp](https://github.com/DataTalksClub/data-engineering-zoomcamp) for providing such a valuable course. The insights I gained have been instrumental in broadening my understanding of the field of Data Engineering. Additionally, I want to express my appreciation to my fellow colleague with whom I took the course. Thank you all for your support and collaboration throughout this journey.\n\n🦅\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fiamraphson%2Fde-2024-project-book-recommendation","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fiamraphson%2Fde-2024-project-book-recommendation","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fiamraphson%2Fde-2024-project-book-recommendation/lists"}