An open API service indexing awesome lists of open source software.

awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
https://github.com/EthicalML/awesome-production-machine-learning

Last synced: 18 days ago
JSON representation

  • Adversarial Robustness

    • Nicolas Carlini’s Adversarial ML reading list - not a library, but a curated list of the most important adversarial papers by one of the leading minds in Adversarial ML, Nicholas Carlini. If you want to discover the 10 papers that matter the most - I would start here.
    • Robust ML - another robustness resource maintained by some of the leading names in adversarial ML. They specifically focus on defenses, and ones that have published code available next to papers. Practical and useful.
    • Robust ML - another robustness resource maintained by some of the leading names in adversarial ML. They specifically focus on defenses, and ones that have published code available next to papers. Practical and useful.
    • Robust ML - another robustness resource maintained by some of the leading names in adversarial ML. They specifically focus on defenses, and ones that have published code available next to papers. Practical and useful.
    • Robust ML - another robustness resource maintained by some of the leading names in adversarial ML. They specifically focus on defenses, and ones that have published code available next to papers. Practical and useful.
    • AdvBox - A toolbox to generate adversarial examples that fool neural networks in PaddlePaddle, PyTorch, Caffe2, MxNet, Keras, TensorFlow, and Advbox can benchmark the robustness of machine learning models.
    • Adversarial DNN Playground - Playground.svg?style=social) - think [TensorFlow Playground](https://playground.tensorflow.org), but for Adversarial Examples! A visualization tool designed for learning and teaching - the attack library is limited in size, but it has a nice front-end to it with buttons you can press!
    • AdverTorch - library for adversarial attacks / defenses specifically for PyTorch.
    • Artificial Adversary - adversary.svg?style=social) AirBnB's library to generate text that reads the same to a human but passes adversarial classifiers.
    • Counterfit - Counterfit is a command-line tool and generic automation layer for assessing the security of machine learning systems.
    • Foolbox - Foolbox is a Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX.
    • MIA - epfl/mia.svg?style=social) - A library for running membership inference attacks (MIA) against machine learning models.
    • OpenAttack - OpenAttack is a Python-based textual adversarial attack toolkit, which handles the whole process of textual adversarial attacking, including preprocessing text, accessing the victim model, generating adversarial examples and evaluation.
    • TextFool - kulynych/textfool.svg?style=social) - plausible looking adversarial examples for text generation.
    • Trickster - epfl/trickster.svg?style=social) - Library and experiments for attacking machine learning in discrete domains using graph search.
    • Factool - NLP/factool.svg?style=social) - Factool is a tool augmented framework for detecting factual errors of texts generated by large language models.
  • Agentic Framework

    • Agents - Agents allows users to build AI-driven server programs that can see, hear, and speak in realtime.
    • Chidori - Chidori is a reactive runtime that supports building robust AI agents using languages like Node.js, Python, and Rust, with a focus on reactivity and observability in agent workflows.
    • Modelscope-Agent - agent.svg?style=social) - Modelscope-Agent is a customizable and scalable agent framework.
    • OpenAGI - OpenAGI is used as the agent creation package to build agents for AIOS.
    • Swarm - Swarm is an educational framework exploring ergonomic, lightweight multi-agent orchestration.
    • Swarms - Swarms is an enterprise grade and production ready multi-agent collaboration framework that enables you to orchestrate many agents to work collaboratively at scale to automate real-world activities.
    • AutoGen - AutoGen is an open-source framework for building AI agent systems.
    • CrewAI - CrewAI is a cutting-edge framework for orchestrating role-playing, autonomous AI agents.
    • LangGraph - ai/langgraph.svg?style=social) - LangGraph is a library for building stateful, multi-actor applications with LLMs, used to create agent and multi-agent workflows.
    • Eko - Eko is a production-ready JavaScript framework that enables developers to create reliable agents, from simple commands to complex workflows.
    • Composio - Composio equip's your AI agents & LLMs with 100+ high-quality integrations via function calling.
    • AgentOps - AI/agentops.svg?style=social) - AgentOps helps developers build, evaluate, and monitor AI agents from prototype to production.
    • AIOpsLab - AIOpsLab is a holistic framework to enable the design, development, and evaluation of autonomous AIOps agents..
    • PydanticAI - ai.svg?style=social) - PydanticAI is a Python agent framework designed to make it less painful to build production grade applications with Generative AI.
    • IntellAgent - ai/intellagent.svg?style=social) - IntellAgent is an advanced multi-agent framework that transforms the evaluation and optimization of conversational agents.
    • AgentStack - AI/AgentStack.svg?style=social) - AgentStack scaffolds your agent stack.
    • TensorZero - TensorZero is an open-source framework for building production-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluations, and experimentation.
  • AutoML

    • Colombus - A scalable framework to perform exploratory feature selection implemented in R.
    • AutoML-GS - yonder/tsfresh.svg?style=social) - Automatic feature and model search with code generation in Python, on top of common data science libraries (tensorflow, sklearn, etc.).
    • auto-sklearn - sklearn.svg?style=social) - Framework to automate algorithm and hyperparameter tuning for sklearn.
    • ENAS via Parameter Sharing - Efficient Neural Architecture Search via Parameter Sharing by [authors of paper](https://arxiv.org/abs/1802.03268).
    • ENAS-PyTorch - pytorch.svg?style=social) - Efficient Neural Architecture Search (ENAS) in PyTorch based [on this paper](https://arxiv.org/abs/1802.03268).
    • ENAS-Tensorflow - Tensorflow.svg?style=social) - Efficient Neural Architecture search via parameter sharing(ENAS) micro search Tensorflow code for windows user.
    • Feature Engine - engine/feature_engine.svg?style=social) - Feature-engine is a Python library that contains several transformers to engineer features for use in machine learning models.
    • Featuretools - An open source framework for automated feature engineering.
    • FLAML - FLAML is a fast library for automated machine learning & tuning.
    • go-featureprocessing - featureprocessing.svg?style=social) - A feature pre-processing framework in Go that matches functionality of sklearn.
    • HEBO - noah/HEBO.svg?style=social) - Set of open-source hyperparameter optimization frameworks, including the winning submission to the [NeurIPS 2020 Black-Box Optimisation Challenge](https://bbochallenge.com/leaderboard) tested on hyperparameter tuning tasks.
    • Katib - A Kubernetes-based system for Hyperparameter Tuning and Neural Architecture Search.
    • keras-tuner - team/keras-tuner.svg?style=social) - Keras Tuner is an easy-to-use, distributable hyperparameter optimisation framework that solves the pain points of performing a hyperparameter search. Keras Tuner makes it easy to define a search space and leverage included algorithms to find the best hyperparameter values.
    • Maggy - Asynchronous, directed Hyperparameter search and parallel ablation studies on Apache Spark - [(Video)](https://www.youtube.com/watch?v=0Hd1iYEL03w).
    • Neural Architecture Search with Controller RNN - architecture-search.svg?style=social) - Basic implementation of Controller RNN from [Neural Architecture Search with Reinforcement Learning](https://arxiv.org/abs/1611.01578) and [Learning Transferable Architectures for Scalable Image Recognition](https://arxiv.org/abs/1707.07012).
    • Neural Network Intelligence - NNI (Neural Network Intelligence) is a toolkit to help users run automated machine learning (AutoML) experiments.
    • Optuna - Optuna is an automatic hyperparameter optimisation software framework, particularly designed for machine learning.
    • OSS Vizier - OSS Vizier is a Python-based service for black-box optimisation and research, one of the first hyperparameter tuning services designed to work at scale.
    • sklearn-deap - deap.svg?style=social) Use evolutionary algorithms instead of gridsearch in scikit-learn.
    • TPOT - Automation of sklearn pipeline creation (including feature selection, pre-processor, etc.).
    • tsfresh - yonder/tsfresh.svg?style=social) - Automatic extraction of relevant features from time series.
    • Upgini - Free automated data & feature enrichment library for machine learning: automatically searches through thousands of ready-to-use features from public and community shared data sources and enriches your training dataset with only the accuracy improving features.
    • AutoGluon - Automated feature, model, and hyperparameter selection for tabular, image, and text data on top of popular machine learning libraries (Scikit-Learn, LightGBM, CatBoost, PyTorch, MXNet).
    • Autokeras - team/autokeras.svg?style=social) - AutoML library for Keras based on ["Auto-Keras: Efficient Neural Architecture Search with Network Morphism"](https://arxiv.org/abs/1806.10282).
    • EvalML - EvalML is an AutoML library which builds, optimizes, and evaluates machine learning pipelines using domain-specific objective functions.
    • Ax - Ax is an accessible, general-purpose platform for understanding, managing, deploying, and automating adaptive experiments.
    • BoTorch - pytorch/botorch.svg?cacheSeconds=86400) - BoTorch is a library for Bayesian Optimization built on PyTorch.
    • Perpetual - ml/perpetual.svg?cacheSeconds=86400) - A gradient boosting machine that doesn't need hyperparameter optimization, with a simple budget parameter to control model complexity.
    • AIDE - AIDE is an open-source ML engineering agent that uses a tree search algorithm to autonomously explore, implement, and evaluate solution strategies for machine learning tasks.
    • HEBO - noah/hebo.svg?cacheSeconds=172800) - Set of open-source hyperparameter optimization frameworks, including the winning submission to the [NeurIPS 2020 Black-Box Optimisation Challenge](https://bbochallenge.com/leaderboard) tested on hyperparameter tuning tasks.
  • Commercial Platform

    • Amazon Web Services - AWS (Amazon Web Services) is a comprehensive, evolving cloud computing platform provided by Amazon that includes a mixture of infrastructure-as-a-service (IaaS), platform-as-a-service (PaaS) and packaged-software-as-a-service (SaaS) offerings, including: [Amazon Augmented AI](https://aws.amazon.com/augmented-ai/), [Amazon Rekognition](https://aws.amazon.com/rekognition/), [Amazon SageMaker](https://aws.amazon.com/sagemaker/).
    • Anthropic - Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
    • Anyscale - Anyscale is a unified compute platform that makes it easy to develop, deploy, and manage scalable AI and Python applications using Ray.
    • Apheris - A platform for federated and privacy-preserving data science that lets you securely collaborate on AI with partners without sharing any data.
    • Arize - ML observability and automated model monitoring to help ML practitioners understand how their models perform in production, troubleshoot issues, and improve model performance. ML teams can upload offline (training or validation) baselines into an evaluation/inference store alongside online production data for model validation, drift detection, data quality checks, and model performance management.
    • Arthur - Authur is a platform that measures, monitors, and improves machine learning models to deliver better results.
    • Azure Machine Learning - Azure Machine Learning empowers data scientists and developers to build, deploy, and manage high-quality models faster and with confidence.
    • BigML - A consumable, programmable, and scalable Machine Learning platform that makes it easy to solve and automate classification, regression, time series, etc..
    • Censius - Censius is an AI Observability Platform that assists enterprises in continuously monitoring, analyzing, and explaining their production models. It combines monitoring, accountability, and explainability into one Observability Platform.
    • Comet - Machine learning experiment management. Free for open source and students - [(Video)](https://www.youtube.com/watch?v=xaybRkapeNE).
    • D2iQ Kaptain - An end-to-end machine learning platform built for security, scale, and speed, that allows enterprises to develop and deploy machine learning models that runs in the cloud, on premises (incl. air-gapped), in hybrid environments, or on the edge; based on Kubeflow and open-source [Kubernetes Universal Declarative Operators](https://kudo.dev) (KUDO).
    • DAGsHub - Community platform for Open Source ML – Manage experiments, data & models and create collaborative ML projects easily.
    • Databricks - An integrated end-to-end machine learning environment incorporating managed services for experiment tracking, model training, feature development and management, and feature and model serving.
    • Dataiku - Collaborative data science platform powering both self-service analytics and the operationalization of machine learning models in production.
    • DataRobot - Automated machine learning platform which enables users to build and deploy machine learning models.
    • Datatron - Machine Learning Model Governance Platform for all your AI models in production for large Enterprises.
    • Deep Cognition Deep Learning Studio - E2E platform for deep learning.
    • deepsense.ai - deepsense.ai helps companies gain competitive advantage by providing customized AI-powered end-to-end solutions, with the main focus on AI software, team augmentation and AI advisory.
    • Diffgram - Training Data First platform. Database & Training Data Pipelines for Supervised AI. Integrated with GCP, AWS, Azure and top Annotation Supervision UIs (or use built-in Diffgram UI, or build your own). Plus a growing list of integrated service providers! For Computer Vision, NLP, and Supervised Deep Learning / Machine Learning.
    • Fennel - Realtime feature engineering platform for fast moving machine learning teams. Python / Pandas native, built in Rust. Easy to install/use/run, builds upon best practices for reducing data/feature quality issues, and keeps cloud spend low. Fully managed, zero ops.
    • Fiddler - Fiddler is a model performance management platform that offers model monitoring, observability, explainability & fairness.
    • Gemesys - GEMESYS aims to design a chip that emulates the human brain, overcoming computing bottlenecks and shaping a better future for everyone.
    • Graphsignal - Machine learning profiler that helps make model training and inference faster and more efficient.
    • H2O Driverless AI - Automates key machine learning tasks, delivering automatic feature engineering, model validation, model tuning, model selection and deployment, machine learning interpretability, bring your own recipe, time-series and automatic pipeline generation for model scoring - [(Video)](https://www.youtube.com/watch?v=ZqCoFp3-rGc).
    • Hugging Face - Hugging Face is a platform that allows users to share machine learning models and datasets.
    • IBM Watson Studio - Build and scale trusted AI on any cloud. Automate the AI lifecycle for ModelOps.
    • InnerEye - InnerEye combines human intelligence with artificial intelligence. By capitalizing on the merging of human neural processing and deep artificial neural networks, InnerEye allows fast and accurate visual inspection, real-time AI training and validation, and establishes a unique human-machine interface for connected user applications.
    • Iguazio Data Science Platform - Bring your Data Science to life by automating MLOps with end-to-end machine learning pipelines, transforming AI projects into real-world business outcomes, and supporting real-time performance at enterprise scale.
    • Iterative Studio - Seamless data and model management, experiment tracking, visualization and automation, with Git as the single source of truth.
    • Katonic.ai - Automate your cycle of Intelligence with Katonic MLOps Platform.
    • Kern AI - Kern AI builds the self-service development environment for NLP training data, used by data scientists to quickly build high-quality, large-scale labeled datasets.
    • Labelbox - Image labelling service with support for semantic segmentation (brush & superpixels), bounding boxes and nested classifications.
    • ModelOp - An enterprise MLOps platform that automates the governance, management and monitoring of deployed AI, ML models across platforms and teams, resulting in reliable, compliant and scalable AI initiatives.
    • Modelplace - Modelplace provides a directory of tested and benchmarked AI models from around the world curated by OpenCV.
    • MLJAR - Platform for rapid prototyping, developing and deploying machine learning models.
    • OpenAI - OpenAI aims to promote and develop friendly AI in a way that benefits humanity as a whole.
    • Pinecone - Pinecone vector database makes it easy to build high-performance vector search applications
    • Prodigy - Prodigy is a scriptable annotation tool so efficient that data scientists can do the annotation themselves, enabling a new level of rapid iteration.
    • Replicate - Replicate lets you run machine learning models with a cloud API, without having to understand the intricacies of machine learning or manage your own infrastructure.
    • Robust Intelligence - Robust Intelligence is an end-to-end ML integrity solution that proactively eliminates failure at every stage of the model lifecycle. From pre-deployment vulnerability detection and validation to post-deployment monitoring and protection, Robust Intelligence gives teams the confidence to scale models in production across a variety of use cases and modalities.
    • SambaNova - SambaNova Systems is a company that specializes in generative AI. They offer a full-stack platform that allows users to build powerful AI models, customized with their data, and owned by them.
    • Scale - Scale AI turns raw data into high-quality training data by combining machine learning powered pre-labeling and active tooling with varying levels and types of human review.
    • Scribble Enrich - Customizable, auditable, privacy-aware feature store. It is designed to help mid-sized data teams gain trust in the data that they use for training and analysis, and support emerging needs such drift computation and bias assessment.
    • SigOpt - SigOpt is a model development platform that makes it easy to track runs, visualize training, and scale hyperparameter optimisation for any type of model built with any library on any infrastructure.
    • Skytree - End to end machine learning platform - [(Video)](https://www.youtube.com/watch?v=XuCwpnU-F1k).
    • SuperAnnotate - A complete set of solutions for image and video annotation and an annotation service with integrated tooling, on-demand narrow expertise in various fields, and a custom neural network, automation, and training models powered by AI.
    • Syndicai - Easy-to-use cloud agnostic platform that deploys, manages, and scales any trained AI model in minutes with no configuration & infrastructure setup.
    • Talend Studio - Data integration platform that provides various software and services for data integration, data management, enterprise application integration, data quality, cloud storage and Big Data.
    • Tecton - Tecton is an all-in-one system to build, automate, and centralize feature workflows for production ML.
    • Valohai - Machine orchestration, version control and pipeline management for deep learning.
    • Vertex AI - Vertex AI Workbench is the single environment for data scientists to complete all of their ML work, from experimentation, to deployment, to managing and monitoring models. It is a Jupyter-based fully managed, scalable, enterprise-ready compute infrastructure with security controls and user management capabilities.
    • Ultralytics
    • WhyLabs - Enable observability to detect data and ML issues faster, deliver continuous improvements, and avoid costly incidents.
    • Zilliz - Zilliz builds vector database to accelerate development of next generation data fabric.
    • Skymind - Software distribution designed to help enterprise IT teams manage, deploy, and retrain machine learning models at scale.
    • Wallaroo.AI - Production AI platform for deploying, managing and observing any model at scale across any enviornment from cloud to edge. Go from python notebook to inferencing in minutes. [Community edition available](https://portal.wallaroo.community/).
    • OpenAI - OpenAI aims to promote and develop friendly AI in a way that benefits humanity as a whole.
    • Zeno - Zeno is a platform for evaluating AI systems.
  • Computation and Communication Optimisation

    • Accelerate - Accelerate abstracts exactly and only the boilerplate code related to multi-GPU/TPU/mixed-precision and leaves the rest of your code unchanged.
    • NVIDIA TensorRT - TensorRT is a C++ library for high-performance inference on NVIDIA GPUs and deep learning accelerators.
    • Colossal-AI - A unified deep learning system for big model era, which helps users to efficiently and quickly deploy large AI model training and inference.
    • Dask - Distributed parallel processing framework for Pandas and NumPy computations.
    • DEAP - A novel evolutionary computation framework for rapid prototyping and testing of ideas. It seeks to make algorithms explicit and data structures transparent. It works in perfect harmony with parallelisation mechanisms such as multiprocessing and SCOOP.
    • einops - Flexible and powerful tensor operations for readable and reliable code.
    • Flashlight - A fast, flexible machine learning library written entirely in C++ from the Facebook AI Research and the creators of Torch, TensorFlow, Eigen and Deep Speech.
    • Hivemind - at-home/hivemind.svg?style=social) - Decentralized deep learning in PyTorch.
    • Horovod - Uber's distributed training framework for TensorFlow, Keras, and PyTorch.
    • LightGBM - LightGBM is a gradient boosting framework that uses tree based learning algorithms.
    • PaddlePaddle - PaddlePaddle is a framework to perform large-scale deep network training, using data sources distributed across hundreds of nodes.
    • PyTorch Lightning - AI/pytorch-lightning.svg?style=social) - PyTorch Lightning pretrains, finetunes and deploys AI models on multiple GPUs, TPUs with zero code changes.
    • Ray - project/ray.svg?style=social) - Ray is a flexible, high-performance distributed execution framework for machine learning.
    • veScale - veScale is a PyTorch native LLM training framework.
    • Composer - Composer is a PyTorch library that enables you to train neural networks faster, at lower cost, and to higher accuracy.
    • CuDF - Built based on the Apache Arrow columnar memory format, cuDF is a GPU DataFrame library for loading, joining, aggregating, filtering, and otherwise manipulating data.
    • CuML - cuML is a suite of libraries that implement machine learning algorithms and mathematical primitives functions that share compatible APIs with other RAPIDS projects.
    • CuPy - An implementation of NumPy-compatible multi-dimensional array on CUDA. CuPy consists of the core multi-dimensional array class, cupy.ndarray, and many functions on it.
    • Flax - A neural network library and ecosystem for JAX designed for flexibility.
    • Lava - Lava is an open source framework to develop applications for neuromorphic hardware architectures.
    • MLX - explore/mlx.svg?style=social) - MLX is an array framework for machine learning on Apple silicon.
    • Modin - project/modin.svg?style=social) - Speed up your Pandas workflows by changing a single line of code.
    • Nevergrad - Nevergrad is a gradient-free optimisation platform.
    • Norse - Norse aims to exploit the advantages of bio-inspired neural components, which are sparse and event-driven - a fundamental difference from artificial neural networks.
    • Numba - A compiler for Python array and numerical functions.
    • Optimum - Optimum is an extension of Transformers and Diffusers, providing a set of optimization tools enabling maximum efficiency to train and run models on targeted hardware while keeping things easy to use.
    • PEFT - Parameter-Efficient Fine-Tuning (PEFT) methods enable efficient adaptation of pre-trained language models (PLMs) to various downstream applications without fine-tuning all the model's parameters.
    • PyTorch - PyTorch is a library to develop and train neural network based deep learning models.
    • scikit-learn - learn/scikit-learn.svg?style=social) - Scikit-learn is a powerful machine learning library that provides a wide variety of modules for data access, data preparation and statistical model building.
    • snnTorch - snnTorch is a deep and online learning library with spiking neural networks.
    • Sonnet - deepmind/sonnet.svg?style=social) - Sonnet is a library built on top of TensorFlow 2 designed to provide simple, composable abstractions for machine learning research.
    • TensorFlow - TensorFlow is a leading library designed for developing and deploying state-of-the-art machine learning applications.
    • ThunderKittens
    • torchkeras - ov-file.svg?style=social) The torchkeras library is a simple tool for training neural network in pytorch jusk in a keras style.
    • TorchOpt - TorchOpt is an efficient library for differentiable optimization built upon PyTorch.
    • Vaex - of-Core DataFrames (similar to Pandas), to visualize and explore big tabular datasets. Vaex uses memory mapping, zero memory copy policy and lazy computations for best performance (no memory wasted).
    • Vowpal Wabbit
    • XGBoost - XGBoost is an optimized distributed gradient boosting library designed to be highly efficient, flexible and portable.
    • yellowbrick - yellowbrick is a matplotlib-based model evaluation plots for scikit-learn and other machine learning libraries.
    • bitsandbytes - foundation/bitsandbytes.svg?style=social) - Bitsandbytes library is a lightweight Python wrapper around CUDA custom functions, in particular 8-bit optimizers, matrix multiplication (LLM.int8()), and 8 & 4-bit quantization functions.
    • BitBLAS - BitBLAS is a library to support mixed-precision BLAS operations on GPUs
    • Liger Kernel - Kernel.svg?style=social) - Liger Kernel is a collection of Triton kernels designed specifically for LLM training.
    • DLRover - machine-learning/dlrover.svg?style=social) - DLRover makes the distributed training of large AI models easy, stable, fast and green.
    • torchdistill - matsubara/torchdistill.svg?style=social) - torchdistill offers various state-of-the-art knowledge distillation methods and enables you to design (new) experiments simply by editing a declarative yaml config file instead of Python code.
    • Adapters - hub/adapters.svg?style=social) - Adapters is a unified library for parameter-efficient and modular transfer learning.
    • SetFit - SetFit is an efficient and prompt-free framework for few-shot fine-tuning of Sentence Transformers.
    • DeepSpeed - DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
    • Jax - ml/jax.svg?style=social) - Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more.
    • YDF - decision-forests.svg?style=social) - YDF (Yggdrasil Decision Forests) is a library to train, evaluate, interpret, and serve Random Forest, Gradient Boosted Decision Trees, CART and Isolation forest models.
    • DGL - DGL is an easy-to-use, high performance and scalable Python package for deep learning on graphs.
    • PyG - team/pytorch_geometric.svg?style=social) - PyG (PyTorch Geometric) is a library built upon PyTorch to easily write and train Graph Neural Networks (GNNs) for a wide range of applications related to structured data.
    • Kompute - nc/lava.svg?style=social) - Blazing fast, lightweight and mobile phone-enabled Vulkan compute framework optimized for advanced GPU data processing usecases.
    • DeepEP - ai/DeepEP.svg?style=social) - DeepEP is a communication library tailored for Mixture-of-Experts (MoE) and expert parallelism (EP). It provides high-throughput and low-latency all-to-all GPU kernels, which are also known as MoE dispatch and combine. The library also supports low-precision operations, including FP8.
    • FlagGems - FlagGems is a high-performance general operator library implemented in OpenAI Triton. It builds on a collection of backend neutral kernels that aims to accelerate LLM training and inference across diverse hardware platforms.
    • Triton - lang/triton.svg?style=social) - Triton is a language and compiler for writing highly efficient custom Deep-Learning primitives. The aim of Triton is to provide an open-source environment to write fast code at higher productivity than CUDA, but also with higher flexibility than other existing DSLs.
    • GPUStack - GPUStack is an open-source GPU cluster manager for running AI models.
    • Cache-DiT - dit.svg?cacheSeconds=86400) - Cache-DiT is built on top of Diffusers and supports nearly all DiTs, providing hybrid cache acceleration (DBCache, TaylorSeer, SCM, etc.) and comprehensive parallelism optimizations including Context Parallelism, Tensor Parallelism, and hybrid 2D/3D parallelism, with compatibility for compilation, CPU offloading, and quantization.
  • Computation Load Distribution

    • Apache Spark MLlib - Apache Spark's scalable machine learning library in Java, Scala, Python and R.
    • Bagua - Bagua is a performant and flexible distributed training framework for PyTorch, providing a faster alternative to PyTorch DDP and Horovod. It supports advanced distributed training algorithms such as quantization and decentralization.
    • Fiber - Distributed computing library for modern computer clusters from Uber.
    • PyWren - Answer the question of the "cloud button" for python function execution. It's a framework that abstracts AWS Lambda to enable data scientists to execute any Python function - [(Video)](https://www.youtube.com/watch?v=OskQytBBdJU).
    • TensorFlowOnSpark - TensorFlowOnSpark brings TensorFlow programs to Apache Spark clusters.
  • Data Annotation and Synthesis

    • Argilla - io/argilla.svg?style=social) - Argilla helps domain experts and data teams to build better NLP datasets in less time.
    • cleanlab - Python library for data-centric AI. Can automatically: find mislabeled data, detect outliers, estimate consensus + annotator-quality for multi-annotator datasets, suggest which data is best to (re)label next.
    • COCO Annotator - annotator.svg?style=social) - Web-based image segmentation tool for object detection, localization and keypoints
    • refinery - kern-ai/refinery.svg?style=social) - The data scientist's open-source choice to scale, assess and maintain natural language data.
    • SDV - dev/SDV.svg?style=social) - Synthetic Data Vault (SDV) is a Synthetic Data Generation ecosystem of libraries that allows users to easily learn single-table, multi-table and timeseries datasets to later on generate new Synthetic Data that has the same format and statistical properties as the original dataset.
    • Semantic Segmentation Editor - Automotive-And-Industry-Lab/semantic-segmentation-editor.svg?style=social) - Hitachi's Open source tool for labelling camera and LIDAR data.
    • NeMo Curator - Curator.svg?style=social) - NeMo Curator is a GPU-accelerated framework for efficient large language model data curation.
    • CVAT - ai/cvat.svg?style=social) - CVAT (Computer Vision Annotation Tool) is OpenCV's web-based annotation tool for both videos and images for computer algorithms.
    • Doccano - Open source text annotation tools for humans, providing functionality for sentiment analysis, named entity recognition, and machine translation.
    • Label Studio - studio.svg?style=social) - Multi-domain data labeling and annotation tool with standardized output format.
    • YData Synthetic - synthetic.svg?style=social) - YData Synthetic is a package to generate synthetic tabular and time-series data leveraging the state of the art generative models.
    • Gretel Synthetics - synthetics.svg?style=social) - Gretel Synthetics is a synthetic data generators for structured and unstructured text, featuring differentially private learning.
    • synthcity - synthcity is a library for generating and evaluating synthetic tabular data.
    • NeMo Curator - Curator.svg?style=social) - NeMo Curator is a GPU-accelerated framework for efficient large language model data curation.
    • ViPE - tlabs/vipe.svg?style=social) - ViPE is a spatial AI tool for annotating camera poses and dense depth maps from raw videos.
    • LightlyStudio - ai/lightly-studio.svg?cacheSeconds=86400) - An open source tool to curate, annotate, and manage vision datasets (images and videos). Supports embedding-based auto-selection, annotation, and auto-labeling for bounding boxes and segmentation.
    • TabGAN - data-generation.svg?cacheSeconds=86400) - Synthetic tabular data generation using GANs (CTGAN), Diffusion Models, and LLMs with adversarial filtering, privacy metrics, and sklearn integration.
Sub Categories