{"id":102406,"url":"https://github.com/brandonhimpfen/awesome-data-science","name":"awesome-data-science","description":"A curated list of tools, libraries, platforms, datasets, workflows, and learning resources for data science.","projects_count":83,"last_synced_at":"2026-10-11T08:00:21.898Z","repository":{"id":328409306,"uuid":"1115462597","full_name":"brandonhimpfen/awesome-data-science","owner":"brandonhimpfen","description":"A curated list of tools, libraries, platforms, datasets, workflows, and learning resources for data science.","archived":false,"fork":false,"pushed_at":"2026-09-06T19:40:47.000Z","size":48,"stargazers_count":17,"open_issues_count":0,"forks_count":2,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-10-04T02:44:28.620Z","etag":null,"topics":["awesome","awesome-list","awesome-lists","big-data","data","data-analysis","data-collection","data-science","machine-learning","open-data"],"latest_commit_sha":null,"homepage":"https://lnktr.net/awesome","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/brandonhimpfen.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":null,"code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"claude":null,"gemini":null,"cursor":null,"copilot":null,"dco":null,"cla":null,"disclosure":null},"funding":{"github":"brandonhimpfen","ko_fi":"brandonhimpfen","buy_me_a_coffee":"brandonhimpfen","custom":["https://paypal.me/brandonhimpfen","https://www.brandonhimpfen.com/#/portal/support"]}},"created_at":"2025-12-12T22:55:35.000Z","updated_at":"2026-09-26T16:29:28.000Z","dependencies_parsed_at":"2026-09-01T20:01:33.005Z","dependency_job_id":"4255018f-cf2f-4827-89f0-4bfd4dd14400","html_url":"https://github.com/brandonhimpfen/awesome-data-science","commit_stats":null,"previous_names":["awesomelistsio/awesome-data-science","brandonhimpfen/awesome-data-science"],"tags_count":2,"template":false,"template_full_name":"brandonhimpfen/awesome-lists-template","purl":"pkg:github/brandonhimpfen/awesome-data-science","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/brandonhimpfen%2Fawesome-data-science","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/brandonhimpfen%2Fawesome-data-science/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/brandonhimpfen%2Fawesome-data-science/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/brandonhimpfen%2Fawesome-data-science/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/brandonhimpfen","download_url":"https://codeload.github.com/brandonhimpfen/awesome-data-science/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/brandonhimpfen%2Fawesome-data-science/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":343843146,"owners_count":38178528,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-10-03T21:59:58.778Z","status":"online","status_checked_at":"2026-10-11T02:00:05.703Z","response_time":113,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-01-02T00:00:35.853Z","updated_at":"2026-10-11T08:00:21.898Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Databases \u0026 Storage","Programming Languages","Visualization","Notebooks \u0026 Experimentation","Data Wrangling \u0026 ETL","Machine Learning","Deep Learning","Datasets \u0026 Open Data","MLOps \u0026 Production","Big Data \u0026 Distributed Computing","Foundations \u0026 References","License","Statistics \u0026 Probability","Learning Resources","Time Series \u0026 Forecasting","Exploratory Data Analysis","Related Awesome Lists"],"sub_categories":["Courses","Tutorials","Guides"],"readme":"# Awesome Data Science [![Awesome Lists](https://srv-cdn.himpfen.io/badges/awesome-lists/awesomelists-flat.svg)](https://github.com/brandonhimpfen/awesome-lists)\n\n[![DOI](https://zenodo.org/badge/1115462597.svg)](https://doi.org/10.5281/zenodo.19673261) \u003cbr /\u003e\n[![Support Open Work](https://img.shields.io/badge/Support-Open%20Work-0A0A0A?style=flat\u0026logo=github)](https://github.com/brandonhimpfen/support) \n[![GitHub Sponsor](https://srv-cdn.himpfen.io/badges/github/github-flat.svg)](https://github.com/sponsors/brandonhimpfen) \n[![Buy Me a Coffee](https://srv-cdn.himpfen.io/badges/buymeacoffee/buymeacoffee-flat.svg)](https://buymeacoffee.com/brandonhimpfen) \n[![Ko-Fi](https://srv-cdn.himpfen.io/badges/kofi/kofi-flat.svg)](https://ko-fi.com/brandonhimpfen) \n[![PayPal](https://srv-cdn.himpfen.io/badges/paypal/paypal-flat.svg)](https://www.paypal.com/donate/?hosted_button_id=3LLKRXJU44EJJ) \u003cbr /\u003e\n[![X](https://srv-cdn.himpfen.io/badges/twitter/twitter-flat.svg)](https://x.com/ListsAwesome) \n[![Facebook](https://srv-cdn.himpfen.io/badges/facebook-pages/facebook-pages-flat.svg)](https://www.facebook.com/awesomelists)\n\n\u003e A curated list of tools, libraries, platforms, datasets, workflows, and learning resources for **data science**, spanning data collection, analysis, visualization, machine learning, and production analytics.\n\n_Support ongoing maintenance and curation via [GitHub Sponsors](https://github.com/sponsors/brandonhimpfen)._\n\n## Contents\n\n- [Foundations \u0026 References](#foundations--references)\n- [Programming Languages](#programming-languages)\n- [Data Wrangling \u0026 ETL](#data-wrangling--etl)\n- [Exploratory Data Analysis](#exploratory-data-analysis)\n- [Visualization](#visualization)\n- [Statistics \u0026 Probability](#statistics--probability)\n- [Machine Learning](#machine-learning)\n- [Deep Learning](#deep-learning)\n- [Time Series \u0026 Forecasting](#time-series--forecasting)\n- [Big Data \u0026 Distributed Computing](#big-data--distributed-computing)\n- [Databases \u0026 Storage](#databases--storage)\n- [MLOps \u0026 Production](#mlops--production)\n- [Notebooks \u0026 Experimentation](#notebooks--experimentation)\n- [Datasets \u0026 Open Data](#datasets--open-data)\n- [Learning Resources](#learning-resources)\n- [Related Awesome Lists](#related-awesome-lists)\n\n## Foundations \u0026 References\n\n- [Data Science Stack Exchange](https://datascience.stackexchange.com/) – Community Q\u0026A covering theory, tools, and practice.\n- [arXiv Data Science](https://arxiv.org/) – Open-access preprints across statistics, ML, and data analysis.\n- [CRISP-DM](https://www.ibm.com/topics/crisp-dm) – Widely used process model for data mining projects.\n- [KDnuggets](https://www.kdnuggets.com/) – News, tutorials, and opinions in data science and analytics.\n- [Towards Data Science](https://towardsdatascience.com/) – Popular publication with practical data science articles.\n\n## Programming Languages\n\n- [Python](https://www.python.org/) – Dominant language for data science, ML, and scientific computing.\n- [R](https://www.r-project.org/) – Statistical computing language with rich analysis packages.\n- [Julia](https://julialang.org/) – High-performance language for numerical and scientific computing.\n- [Scala](https://www.scala-lang.org/) – JVM language commonly used with Apache Spark.\n- [SQL](https://en.wikipedia.org/wiki/SQL) – Query language essential for data extraction and analysis.\n\n## Data Wrangling \u0026 ETL\n\n- [pandas](https://pandas.pydata.org/) – Core Python library for data manipulation and analysis.\n- [Polars](https://www.pola.rs/) – Fast DataFrame library optimized for performance.\n- [Apache Airflow](https://airflow.apache.org/) – Workflow orchestration platform for data pipelines.\n- [dbt](https://www.getdbt.com/) – Transformation tool for analytics engineering workflows.\n- [Apache NiFi](https://nifi.apache.org/) – Visual tool for automating data flows between systems.\n- [Talend Open Studio](https://www.talend.com/products/talend-open-studio/) – Open-source ETL and integration platform.\n\n## Exploratory Data Analysis\n\n- [ydata-profiling](https://github.com/ydataai/ydata-profiling) – Automated EDA reports for pandas DataFrames.\n- [Sweetviz](https://github.com/fbdesignpro/sweetviz) – Visual EDA tool for quick dataset inspection.\n- [PandasGUI](https://github.com/adamerose/PandasGUI) – Interactive GUI for exploring pandas DataFrames.\n- [D-Tale](https://github.com/man-group/dtale) – Interactive visualizer for pandas and NumPy data.\n\n## Visualization\n\n- [Matplotlib](https://matplotlib.org/) – Foundational plotting library for Python.\n- [Seaborn](https://seaborn.pydata.org/) – Statistical data visualization built on Matplotlib.\n- [Plotly](https://plotly.com/python/) – Interactive plotting library for dashboards and notebooks.\n- [Altair](https://altair-viz.github.io/) – Declarative statistical visualization library for Python.\n- [Tableau](https://www.tableau.com/) – Business intelligence and data visualization platform.\n- [Power BI](https://powerbi.microsoft.com/) – Analytics and visualization service by Microsoft.\n\n## Statistics \u0026 Probability\n\n- [SciPy Stats](https://docs.scipy.org/doc/scipy/reference/stats.html) – Statistical functions for scientific computing.\n- [Statsmodels](https://www.statsmodels.org/) – Statistical modeling and hypothesis testing in Python.\n- [Stan](https://mc-stan.org/) – Probabilistic programming for Bayesian inference.\n- [PyMC](https://www.pymc.io/) – Bayesian statistical modeling and probabilistic ML.\n- [scikit-posthocs](https://github.com/maximtrp/scikit-posthocs) – Post-hoc statistical tests for Python.\n\n## Machine Learning\n\n- [scikit-learn](https://scikit-learn.org/) – Core ML library for classical algorithms in Python.\n- [XGBoost](https://xgboost.ai/) – Gradient boosting library for structured data.\n- [LightGBM](https://lightgbm.readthedocs.io/) – Fast gradient boosting framework by Microsoft.\n- [CatBoost](https://catboost.ai/) – Gradient boosting with strong categorical feature support.\n- [MLflow](https://mlflow.org/) – Platform for tracking experiments and managing ML lifecycles.\n\n## Deep Learning\n\n- [TensorFlow](https://www.tensorflow.org/) – End-to-end deep learning framework.\n- [PyTorch](https://pytorch.org/) – Popular deep learning library with dynamic computation graphs.\n- [Keras](https://keras.io/) – High-level neural networks API.\n- [Hugging Face Transformers](https://huggingface.co/transformers/) – Pretrained models for NLP and multimodal tasks.\n- [FastAI](https://www.fast.ai/) – High-level library simplifying deep learning workflows.\n\n## Time Series \u0026 Forecasting\n\n- [statsforecast](https://github.com/Nixtla/statsforecast) – Fast statistical forecasting library.\n- [Prophet](https://facebook.github.io/prophet/) – Time series forecasting tool for business use cases.\n- [GluonTS](https://github.com/awslabs/gluonts) – Probabilistic time series modeling by AWS.\n- [Darts](https://github.com/unit8co/darts) – Python library for easy time series forecasting.\n- [tslearn](https://tslearn.readthedocs.io/) – Machine learning toolkit for time series data.\n\n## Big Data \u0026 Distributed Computing\n\n- [Apache Spark](https://spark.apache.org/) – Distributed data processing engine.\n- [Apache Hadoop](https://hadoop.apache.org/) – Framework for distributed storage and processing.\n- [Dask](https://www.dask.org/) – Parallel computing library that scales Python workflows.\n- [Ray](https://www.ray.io/) – Distributed computing framework for ML and Python apps.\n- [Flink](https://flink.apache.org/) – Stream and batch processing framework.\n\n## Databases \u0026 Storage\n\n- [PostgreSQL](https://www.postgresql.org/) – Advanced open-source relational database.\n- [MySQL](https://www.mysql.com/) – Popular relational database system.\n- [MongoDB](https://www.mongodb.com/) – NoSQL document-oriented database.\n- [DuckDB](https://duckdb.org/) – In-process analytical SQL database.\n- [BigQuery](https://cloud.google.com/bigquery) – Serverless data warehouse on GCP.\n- [Snowflake](https://www.snowflake.com/) – Cloud-native data warehouse platform.\n\n## MLOps \u0026 Production\n\n- [Kubeflow](https://www.kubeflow.org/) – Kubernetes-native ML workflows.\n- [Seldon Core](https://www.seldon.io/) – Model deployment and inference on Kubernetes.\n- [BentoML](https://bentoml.com/) – Framework for serving ML models in production.\n- [Weights \u0026 Biases](https://wandb.ai/) – Experiment tracking and model monitoring.\n- [Evidently](https://www.evidentlyai.com/) – Data and model monitoring for ML systems.\n\n## Notebooks \u0026 Experimentation\n\n- [Jupyter](https://jupyter.org/) – Interactive notebooks for data analysis and ML.\n- [Google Colab](https://colab.research.google.com/) – Cloud-hosted Jupyter notebooks with free GPUs.\n- [Kaggle Notebooks](https://www.kaggle.com/notebooks) – Collaborative notebooks with datasets and competitions.\n- [VS Code Notebooks](https://code.visualstudio.com/docs/datascience/jupyter-notebooks) – Notebook support inside VS Code.\n\n## Datasets \u0026 Open Data\n\n- [Kaggle Datasets](https://www.kaggle.com/datasets) – Public datasets for data science projects.\n- [UCI ML Repository](https://archive.ics.uci.edu/ml/) – Classic datasets for ML research.\n- [OpenML](https://www.openml.org/) – Open datasets and benchmarks for ML experiments.\n- [Google Dataset Search](https://datasetsearch.research.google.com/) – Search engine for public datasets.\n- [World Bank Open Data](https://data.worldbank.org/) – Global development and economic data.\n\n## Learning Resources\n\n### Tutorials\n- [Kaggle Learn](https://www.kaggle.com/learn) – Hands-on micro-courses in data science.\n- [DataCamp](https://www.datacamp.com/) – Interactive courses for data science skills.\n- [Coursera Data Science](https://www.coursera.org/) – University-backed data science programs.\n\n### Guides\n- [Python Data Science Handbook](https://jakevdp.github.io/PythonDataScienceHandbook/) – Comprehensive guide to Python data tools.\n- [Google ML Crash Course](https://developers.google.com/machine-learning/crash-course) – Practical intro to ML concepts.\n- [The Data Science Lifecycle](https://www.ibm.com/topics/data-science) – Overview of end-to-end data science workflows.\n\n### Courses\n- *Applied Data Science* – End-to-end data analysis and modeling.\n- *Machine Learning Engineering* – Production ML systems and MLOps.\n- *Statistics for Data Science* – Probability and inference foundations.\n\n## Related Awesome Lists\n\n- [Awesome Machine Learning](https://github.com/brandonhimpfen/awesome-machine-learning)\n- [Awesome AI](https://github.com/brandonhimpfen/awesome-ai)\n- [Awesome Python](https://github.com/brandonhimpfen/awesome-python)\n- [Awesome Big Data](https://github.com/brandonhimpfen/awesome-big-data)\n- [Awesome Scientific Computing](https://github.com/brandonhimpfen/awesome-scientific-computing)\n\n## Contribute\n\nContributions are welcome. Please ensure your submission fully follows the requirements outlined in [`CONTRIBUTING.md`](CONTRIBUTING.md), including formatting, scope alignment, and category placement.\n\nPull requests that do not adhere to the contribution guidelines may be closed.\n\n## License\n\n[![CC0](https://mirrors.creativecommons.org/presskit/buttons/88x31/svg/by-sa.svg)](http://creativecommons.org/licenses/by-sa/4.0/)\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/brandonhimpfen%2Fawesome-data-science/projects"}