awesome-llmops
An awesome & curated list of best LLMOps tools for developers
https://github.com/tensorchord/awesome-llmops
Last synced: 13 days ago
JSON representation
-
Model
-
Audio Foundation Model
- whisper - Scale Weak Supervision |  |
-
CV Foundation Model
- midjourney
- disco-diffusion - diffusion.svg?style=flat-square) |
- segment-anything (SAM) - anything.svg?style=flat-square) |
- stable-diffusion - to-image diffusion model |  |
- stable-diffusion v2 - Resolution Image Synthesis with Latent Diffusion Models |  |
-
Large Language Model
- Falcon 40B - 40B-Instruct is a 40B parameters causal decoder-only model built by TII based on Falcon-40B and finetuned on a mixture of Baize. It is made available under the Apache 2.0 license. | |
- Mixtral-8x7B-v0.1 - 8x7B Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts. | |
- Gemma
- Alpaca - lab/stanford_alpaca.svg?style=flat-square) |
- BELLE - tune by 34B Chinese Character Corpus, based on LLaMA and Alpaca. |  |
- Bloom - science Open-access Multilingual Language Model |  |
- dolly - square) |
- FastChat (Vicuna) - T5. |  |
- GLM-6B (ChatGLM) - Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs. |  |
- ChatGLM2-6B - 6B is the second-generation version of the open-source bilingual (Chinese-English) chat model [ChatGLM-6B](https://github.com/THUDM/ChatGLM-6B). |  |
- GLM-130B (ChatGLM) - Trained Model (ICLR 2023) |  |
- GPT-NeoX - neox.svg?style=flat-square) |
- Luotuo - Alpaca-LoRA. |  |
- StableLM - AI/StableLM.svg?style=flat-square) |
-
Robotics Foundation Model
- LeRobot - to-end learning tools, data pipelines, and support for training/deploying VLA models. |  |
- Octo - based generalist robot policy pretrained on 800K+ robot trajectories from the Open X-Embodiment dataset. Supports language instructions, goal images, and fine-tuning to new embodiments. |  |
- OpenPI - source VLA models from Physical Intelligence, including π₀ and π₀.5 — flow-based vision-language-action models pretrained on large-scale robot data with fine-tuning support. |  |
- OpenVLA - parameter open-source Vision-Language-Action model trained on 970K+ robot demonstrations from the Open X-Embodiment dataset for generalist robotic manipulation. |  |
- SmolVLA
-
-
Observability
- PromptHub - Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.
- Prompteams - Prompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).
- Doku - An open-source LLM Observability platform streamlining the monitoring of LLM applications with just two lines of code. It provides valuable insights into token usage and user engagement, tracks API usage for providers like OpenAI, and facilitates easy data export to observability platforms like Grafana and DataDog.
-
Optimizations
-
Profiling
- FeatherCNN - square) |
- Forward - square) |
- NCNN - performance neural network inference framework optimized for the mobile platform. |  |
- PocketFlow - square) |
- TensorFlow Model Optimization - optimization.svg?style=flat-square) |
- TNN - square) |
- optimum-tpu - tpu.svg?style=flat-square) |
- LangWatch - square) |
- agent-opt - driven iterative refinements. |  |
- Entroly - theoretic context optimization proxy. Cuts LLM token costs by 70–95% with zero accuracy loss using greedy submodular knapsack maximization. |  |
- lean-ctx - aware compression, and shell output patterns. [Website](https://leanctx.com) |  |
-
-
Performance
-
ML Compiler
- ONNX-MLIR - mlir.svg?style=flat-square) |
- TVM - square) |
- bitsandbytes - bit quantization for PyTorch. |  |
-
Profiling
- octoml-profile - profile is a python library and cloud service designed to provide the simplest experience for assessing and optimizing the performance of PyTorch models on cloud hardware with state-of-the-art ML acceleration technology. |  |
- scalene - performance, high-precision CPU, GPU, and memory profiler for Python |  |
-
-
Search
-
Hybrid search
- Airweave - ai/airweave.svg?style=flat-square) |
-
Vector search
- Pinecone - performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles. | |
- Vellum - of-box support for OCR, text chunking, embedding model experimentation, metadata filtering, and production-grade APIs. | |
- Awadb - ai/awadb.svg?style=flat-square) |
- Chroma - core/chroma.svg?style=flat-square) |
- Infinity - native database built for LLM applications, providing incredibly fast vector and full-text search |  |
- Lancedb - friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps! |  |
- Marqo - ai/marqo.svg?style=flat-square) |
- Milvus - io/milvus.svg?style=flat-square) |
- pgvector - source vector similarity search for Postgres. |  |
- pgvecto.rs - square) |
- Qdrant - square) |
- txtai - powered semantic search applications |  |
- Vald - square) |
- Vearch - based vector retrieval |  |
- VectorDB - no more, no less. |  |
- Epsilla - cloud/vectordb.svg?style=flat-square) |
- VectorChord - friendly vector search in Postgres, the successor of `pgvecto.rs`. |  |
- ParadeDB - square) |
- AquilaDB - NN search. |  |
- Weaviate - tolerance and scalability of a cloud-native database, all accessible through GraphQL, REST, and various language clients. |  |
- Omnigraph - native, Rust, traversal + vector + BM25 in one runtime. |  |
- Rivestack - in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage. | |
-
-
Security
-
Frameworks for LLM security
- Plexiglass - labs/plexiglass?style=flat-square) |
- Cordum - first agent orchestration platform with pre-dispatch policy evaluation, output scanning (PII, secrets, injection), job scheduling, workflow engine, and full audit trail. |  |
- brood-box - isolated microVMs with snapshot isolation, egress control, and MCP authorization. |  |
- dstack - source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing. |  |
-
Observability
- Azure OpenAI Logger - openai-logger?style=flat-square) |
- Deepchecks - square) |
- Fiddler AI - production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity. |  |
- Giskard - AI/giskard.svg?style=flat-square) |
- Great Expectations - expectations/great_expectations.svg?style=flat-square) |
- whylogs - square) |
- Traceloop OpenLLMetry - based observability and monitoring for LLM and agents workflows. | 
- traceAI - source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows. |  |
- Future AGI - grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows. |  |
- semantic-coverage - agi/futureagi-sdk?style=flat-square) |
- QWED - AI/qwed-verification.svg?style=flat-square) |
- RagTune - square) |
- ClevAgent - restart. Python SDK or HTTP API. | |
- EvalView - call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API. |  |
- onWatch - dev/onwatch.svg?style=flat-square) |
-
-
Serving
-
Frameworks/Servers for Serving
- BentoML - square) |
- Mosec - to-use Python interface. |  |
- TFServing - performance serving system for machine learning models. |  |
- Torchserve - square) |
- Triton Server (TRTIS) - inference-server/server.svg?style=flat-square) |
- langchain-serve - ai/langchain-serve.svg?style=flat-square) |
- lanarky - grade LLM applications |  |
- ray-llm - RayLLM |  |
- Xinference - source language models, speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop. |  |
- KubeAI - to-text. |  |
- Kaito - 3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers. |  |
- Open Responses - source platform for building long-running LLM agents with tool use. |  |
- Open Responses - source platform for building long-running LLM agents with tool use. |  |
- Open Responses - source platform for building long-running LLM agents with tool use. |  |
- Jina - ai/jina.svg?style=flat-square) |
- mcpproxy-go - source MCP proxy with BM25 tool filtering, quarantine security, activity logging, and web UI. Routes multiple MCP servers through single endpoint, reducing context bloat by ~97%. |  |
- KubeStellar Console - powered multi-cluster Kubernetes dashboard for hybrid edge and cloud. GPU monitoring, LLM inference cluster management, benchmark streaming, and 20+ CNCF integrations. CNCF Sandbox (Apache 2.0). |  |
-
Large Model Serving
- CTranslate2 - square) |
- Clip-as-a-service - ai/clip-as-service.svg?style=flat-square) |
- Flowise - square) |
- Infinity - embeddings |  |
- Modelz-LLM - llm.svg?style=flat-square) |
- TensorRT-LLM - LLM.svg?style=flat-square) |
- text-generation-inference - generation-inference.svg?style=flat-square) |
- text-embeddings-inference - embedding models |  |
- vllm - throughput and memory-efficient inference and serving engine for LLMs. |  |
- x-stable-diffusion - time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. |  |
- tokenizers - of-the-Art Tokenizers optimized for Research and Production |  |
- DeepSpeed-MII - latency and high-throughput inference possible, powered by DeepSpeed. |  |
- prima.cpp - square) |
- FlexGen - oriented scenarios. |  |
- Shimmy - free Rust inference server with OpenAI API compatibility and hot model swapping |  |
- llama.cpp - square) |
- Ollama - square) |
- whisper.cpp - square) |
- Rapid-MLX - compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching. |  |
- OneComp - training quantization pipeline for LLMs (QEP, AutoBit, JointQ, rotation) with vLLM plugin (arXiv:2603.28845). |  |
- LLMKube - GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats. |  |
- Off Grid - source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline. |  |
- whisper-ctranslate2 - memory usage drop-in cli replacement that supports word-level timestamps and VAD filter |  |
-
-
Training
-
Experiment Tracking
- Aim - to-use and performant open-source experiment tracker. |  |
- Guild AI - square) |
- Kedro-Viz - Viz is an interactive development tool for building data science pipelines with Kedro. Kedro-Viz also allows users to view and compare different runs in the Kedro project. |  |
- LabNotebook - square) |
- Sacred - square) |
- ClearML - Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-Management |  |
-
Foundation Model Fine Tuning
- Flyflow - devs/flyflow.svg?style=flat-square) |
- alpaca-lora - tune LLaMA on consumer hardware |  |
- finetuning-scheduler - tuning schedules. |  |
- LMFlow - square) |
- Lora - rank adaptation to quickly fine-tune diffusion models. |  |
- peft - of-the-art Parameter-Efficient Fine-Tuning. |  |
- p-tuning-v2 - tuning on small/medium-sized models and sequence tagging challenges. [(ACL 2022)](https://arxiv.org/abs/2110.07602) |  |
- QLoRA - bit finetuning task performance. |  |
- TRL - square) |
-
Frameworks for Training
- Accelerate - GPU, TPU, mixed-precision. |  |
- Apache MXNet - aware Dataflow Dep Scheduler. |  |
- Caffe - square) |
- ColossalAI - scale model training system with efficient parallelization techniques. |  |
- Horovod - square) |
- Kedro - source Python framework for creating reproducible, maintainable and modular data science code. |  |
- Keras - team/keras.svg?style=flat-square) |
- LightGBM - square) |
- MegEngine - to-use deep learning framework, with auto-differentiation. |  |
- metric-learn - learn-contrib/metric-learn.svg?style=flat-square) |
- MindSpore - ai/mindspore.svg?style=flat-square) |
- Oneflow - centered and open-source deep learning framework. |  |
- PaddlePaddle - square) |
- PyTorch - square) |
- XGBoost - square) |
- scikit-learn - learn/scikit-learn.svg?style=flat-square) |
- TensorFlow - square) |
- VectorFlow - square) |
- Candle - square`) |
- DeepSpeed - square) |
- Jax - performance machine learning research. |  |
- axolotl - tuning of various AI models, offering support for multiple configurations and architectures. |  |
-
IDEs and Workspaces
- code server - server.svg?style=flat-square) |
- conda - agnostic, system-level binary package manager and ecosystem. |  |
- Docker - source project created by Docker to enable and accelerate software containerization. |  |
- envd - square) |
- Jupyter Notebooks - based notebook environment for interactive computing. |  |
- Kurtosis - container environments. |  |
-
Model Editing
- FastEdit - square) |
-
Visualization
- Fiddler AI
- Maniford - agnostic visual debugging tool for machine learning. |  |
- netron - square) |
- OpenOps - square) |
- TensorBoard - square) |
- TensorSpace - trained deep learning models from TensorFlow, Keras, TensorFlow.js. |  |
- dtreeviz - square) |
- Zetane Viewer - square) |
- Zeno - ml/zeno.svg?style=flat-square) |
- OpenOps - square) |
-
Programming Languages
Categories
Sub Categories
Observability
90
Profiling
73
Vector search
32
Large Model Serving
23
Frameworks for Training
22
Frameworks/Servers for Serving
17
Large Language Model
14
ML Platforms
14
Workflow
13
Visualization
10
Foundation Model Fine Tuning
9
IDEs and Workspaces
6
Experiment Tracking
6
Data Management
6
CV Foundation Model
5
Scheduling
5
Robotics Foundation Model
5
Model Management
4
Frameworks for LLM security
4
Data/Feature enrichment
4
ML Compiler
3
Data Storage
3
Data Tracking
2
Feature Engineering
2
Audio Foundation Model
2
Hybrid search
1
Model Editing
1
Keywords
machine-learning
103
llm
62
python
60
deep-learning
54
ai
46
mlops
45
pytorch
39
data-science
39
llmops
29
tensorflow
25
kubernetes
25
automl
24
ml
23
inference
18
openai
18
large-language-models
17
hyperparameter-optimization
17
llms
17
rag
16
prompt-engineering
15
developer-tools
15
langchain
15
chatgpt
14
keras
14
vector-database
14
gpu
13
observability
12
gpt
12
vector-search
12
neural-architecture-search
12
evaluation
12
docker
11
neural-network
11
llama
10
golang
10
artificial-intelligence
10
agents
10
generative-ai
10
chatbot
9
analytics
9
ai-agents
9
hyperparameter-tuning
9
scikit-learn
9
open-source
9
go
8
transformers
8
workflow
8
nearest-neighbor-search
8
model-serving
8
language-model
8