awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
https://github.com/academic/awesome-datascience
Last synced: 2 days ago
JSON representation
-
The Data Science Toolbox
-
Deep Learning Packages
- NetworkX
- Redash
- C3
- geomap
- Netron
- PyTorch
- torchvision
- torchtext
- torchaudio
- ignite
- PyToune
- skorch
- PyVarInf
- pytorch_geometric
- GPyTorch
- pyro
- Catalyst
- pytorch_tabular
- Yolov3
- Yolov5
- Yolov8
- TensorFlow
- TensorLayer
- TFLearn
- tensorpack
- Polyaxon
- NeuPy
- tfdeploy
- TensorFlow Fold
- tensorlm
- Mesh TensorFlow
- Ludwig
- TF-Agents
- TensorForce
- keras-contrib
- Hyperas
- Elephas
- Hera
- Spektral
- qkeras
- keras-rl
- Talos
- cartodb
- Cube
- Resseract Lite
- vizzu
- Sonnet
- TRFL
- tensorflow-upstream
- Glue
- Wrangler
- r2d3
- TensorWatch
- Dash
- TensorLight
- PyTorchNet
- pytorch_tabular
- Metabase
- MetaReview - Free online meta-analysis platform with 11 interactive D3.js statistical charts (forest plot, funnel plot, Galbraith, L'Abbé, Baujat, etc.), 5 effect size measures, AI literature screening, and publication-ready report export. [github.com](https://github.com/TerryFYL/metareview)
- torchvista - Interactive notebook-based tool to visualize the forward pass of any PyTorch model.
- plot.ly
- raw
- geomap
-
General Machine Learning Packages
- scikit-learn
- Shogun
- scikit-survival
- scikit-multilearn
- sklearn-expertsys
- scikit-feature
- scikit-rebate
- seqlearn
- sklearn-bayes
- sklearn-crfsuite
- sklearn-deap
- sigopt_sklearn
- sklearn-evaluation
- scikit-image
- scikit-opt
- scikit-posthocs
- pystruct
- xLearn
- cuML
- causalml
- mlpack
- MLxtend
- modAL
- Sparkit-learn
- dlib
- imodels
- RuleFit
- pyGAM
- Deepchecks
- XGBoost
- LightGBM
- CatBoost
- interpretable
- feature-engine
- PerpetualBooster
- hyperlearn
- jSciPy - A Java port of SciPy's signal processing module, offering filters, transformations, and other scientific computing utilities.
- JAX
- feature-engine
- scikit-survival
-
Miscellaneous Tools
- Hortonworks Sandbox
- R
- Tidyverse
- Scikit-Learn
- NumPy - dimensional arrays and matrices and includes an assortment of high-level mathematical functions to operate on these arrays. |
- Vaex
- SciPy
- Data Science Toolbox
- Datadog - scale data science. |
- Variance
- Kite Development Kit
- Apache Flink - purpose data processing. |
- Apache Hama - Level open source project, allowing you to do advanced analytics beyond MapReduce. |
- Apache Spark - fast cluster computing |
- Data Mechanics - friendly and cost-effective. |
- Caffe
- Torch
- Aerosolve
- Datawrapper
- Tensor Flow
- Natural Language Toolkit
- nlp-toolkit for node.js
- Apache Zeppelin - based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more |
- LightTag
- Amazon Rekognition
- Amazon Textract
- Amazon Lookout for Vision
- Amazon CodeGuru - powered recommendations.|
- Statsmodels - based inferential statistics, hypothesis testing and regression framework |
- Gensim - source library for topic modeling of natural language text |
- spaCy
- PyMC3
- Explore Data Science Libraries
- MLflow
- Arize AI - causing issues such as data quality and performance drift. |
- Aureo.io - code platform that focuses on building artificial intelligence. It provides users with the capability to create pipelines, automations and integrate them with artificial intelligence models – all with their basic data. |
- ERD Lab
- Arize-Phoenix - uncover insights, surface problems, monitor, and fine tune your models. |
- Synthical - powered collaborative environment for research. Find relevant papers, create collections to manage bibliography, and summarize content — all in one place |
- The Data Science Lifecycle Process
- Data Science Lifecycle Template Repo
- RexMex
- ChemicalX
- PyTorch Geometric Temporal
- Little Ball of Fur - Learn like API. |
- Karate Club - Learn like API. |
- ML Workspace - in-one web-based IDE for machine learning and data science. The workspace is deployed as a Docker container and is preloaded with a variety of popular data science libraries (e.g., Tensorflow, PyTorch) and dev tools (e.g., Jupyter, VS Code) |
- steppy
- steppy-toolkit
- Data Science Toolbox
- Kite Development Kit
- Weka
- Octave - level interpreted language, primarily intended for numerical computations.(Free Matlab) |
- Hydrosphere Mist
- Torch
- Nervana's python based Deep Learning Framework
- Skale
- Intel framework
- IJulia - language backend combined with the Jupyter interactive environment |
- Featuretools
- Optimus - processing, feature engineering, exploratory data analysis and easy ML with PySpark backend. |
- Albumentations
- Lambdo
- Feast
- Hopsworks - source data-intensive machine learning platform with a feature store. Ingest and manage features for both online (MySQL Cluster) and offline (Apache Hive) access, train and serve models at scale. |
- MindsDB
- Lightwood
- AWS Data Wrangler - source Python package that extends the power of Pandas library to AWS connecting DataFrames and AWS data related services (Amazon Redshift, AWS Glue, Amazon Athena, Amazon EMR, etc). |
- CML - like environments with GitHub Actions & GitLab CI, and autogenerate visual reports on pull/merge requests. |
- Grid Studio - based spreadsheet application with full integration of the Python programming language. |
- Python Data Science Handbook
- Shapley - driven framework to quantify the value of classifiers in a machine learning ensemble. |
- DAGsHub
- Nimblebox - stack MLOps platform designed to help data scientists and machine learning practitioners around the world discover, create, and launch multi-cloud apps from their web browser. |
- Towhee
- LineaPy
- envd
- Explore Data Science Libraries
- MLEM
- cleanlab - centric AI and automatically detecting various issues in ML datasets |
- AutoGluon - series, and multi-modal data |
- Arize-Phoenix - uncover insights, surface problems, monitor, and fine tune your models. |
- Comet
- Opik
- Synthical - powered collaborative environment for research. Find relevant papers, create collections to manage bibliography, and summarize content — all in one place |
- teeplot
- Streamlit
- Gradio
- Weights & Biases
- Optuna
- Ray Tune
- Apache Airflow
- Prefect
- Kedro - source Python framework for creating reproducible, maintainable data science code |
- InterpretML - source package also provides visualization tools for EBMs, other glass-box models, and black-box explanations |
- LIME
- flyte
-
Programming Languages
Categories
Sub Categories
Miscellaneous Tools
144
Bloggers
125
Books
97
Deep Learning Packages
92
Datasets
88
Comparison
78
Twitter Accounts
70
YouTube Videos & Channels
59
MOOC's
50
Comics
50
Journals, Publications and Magazines
42
Facebook Accounts
41
General Machine Learning Packages
40
Podcasts
36
Tutorials
29
Algorithms
28
Free Courses
24
Colleges
23
Infographics
15
Presentations
11
Data Science Competitions
5
Tools
5
Newsletters
4
Research & Knowledge Retrieval
4
Telegram Channels
3
Intensive Programs
2
Frameworks
2
Workflow
2
Mailing lists
1
Hobby
1
Slack Communities
1
GitHub Groups
1
Disaster
1
Keywords
machine-learning
86
python
60
deep-learning
58
data-science
50
pytorch
26
tensorflow
21
scikit-learn
13
data-analysis
13
keras
13
ml
12
neural-network
11
reinforcement-learning
11
mlops
10
artificial-intelligence
10
data-visualization
9
ai
9
computer-vision
8
numpy
8
llm
7
hyperparameter-optimization
7
neural-networks
7
gradient-boosting
7
awesome-list
7
object-detection
7
jupyter-notebook
6
dataset
6
data-mining
6
r
6
jupyter
6
pandas
6
explainable-ml
5
image-processing
5
explainable-ai
5
data
5
spark
5
big-data
5
awesome
5
workflow
5
nlp
5
statistics
5
pipeline
5
reproducibility
4
data-engineering
4
feature-engineering
4
scientific-computing
4
gpu
4
machine-learning-algorithms
4
cli
4
optimization
4
open-source
4