An open API service indexing awesome lists of open source software.

awesome-llmops

An awesome & curated list of best LLMOps tools for developers
https://github.com/tensorchord/awesome-llmops

Last synced: 13 days ago
JSON representation

  • Model

    • Audio Foundation Model

      • whisper - Scale Weak Supervision | ![GitHub Badge](https://img.shields.io/github/stars/openai/whisper.svg?style=flat-square) |
    • CV Foundation Model

      • midjourney
      • disco-diffusion - diffusion.svg?style=flat-square) |
      • segment-anything (SAM) - anything.svg?style=flat-square) |
      • stable-diffusion - to-image diffusion model | ![GitHub Badge](https://img.shields.io/github/stars/CompVis/stable-diffusion.svg?style=flat-square) |
      • stable-diffusion v2 - Resolution Image Synthesis with Latent Diffusion Models | ![GitHub Badge](https://img.shields.io/github/stars/Stability-AI/stablediffusion.svg?style=flat-square) |
    • Large Language Model

      • Falcon 40B - 40B-Instruct is a 40B parameters causal decoder-only model built by TII based on Falcon-40B and finetuned on a mixture of Baize. It is made available under the Apache 2.0 license. | |
      • Mixtral-8x7B-v0.1 - 8x7B Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts. | |
      • Gemma
      • Alpaca - lab/stanford_alpaca.svg?style=flat-square) |
      • BELLE - tune by 34B Chinese Character Corpus, based on LLaMA and Alpaca. | ![GitHub Badge](https://img.shields.io/github/stars/LianjiaTech/BELLE.svg?style=flat-square) |
      • Bloom - science Open-access Multilingual Language Model | ![GitHub Badge](https://img.shields.io/github/stars/bigscience-workshop/model_card.svg?style=flat-square) |
      • dolly - square) |
      • FastChat (Vicuna) - T5. | ![GitHub Badge](https://img.shields.io/github/stars/lm-sys/FastChat.svg?style=flat-square) |
      • GLM-6B (ChatGLM) - Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs. | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/ChatGLM-6B.svg?style=flat-square) |
      • ChatGLM2-6B - 6B is the second-generation version of the open-source bilingual (Chinese-English) chat model [ChatGLM-6B](https://github.com/THUDM/ChatGLM-6B). | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/ChatGLM2-6B.svg?style=flat-square) |
      • GLM-130B (ChatGLM) - Trained Model (ICLR 2023) | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/GLM-130B.svg?style=flat-square) |
      • GPT-NeoX - neox.svg?style=flat-square) |
      • Luotuo - Alpaca-LoRA. | ![GitHub Badge](https://img.shields.io/github/stars/LC1332/Luotuo-Chinese-LLM.svg?style=flat-square) |
      • StableLM - AI/StableLM.svg?style=flat-square) |
    • Robotics Foundation Model

      • LeRobot - to-end learning tools, data pipelines, and support for training/deploying VLA models. | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/lerobot.svg?style=flat-square) |
      • Octo - based generalist robot policy pretrained on 800K+ robot trajectories from the Open X-Embodiment dataset. Supports language instructions, goal images, and fine-tuning to new embodiments. | ![GitHub Badge](https://img.shields.io/github/stars/octo-models/octo.svg?style=flat-square) |
      • OpenPI - source VLA models from Physical Intelligence, including π₀ and π₀.5 — flow-based vision-language-action models pretrained on large-scale robot data with fine-tuning support. | ![GitHub Badge](https://img.shields.io/github/stars/Physical-Intelligence/openpi.svg?style=flat-square) |
      • OpenVLA - parameter open-source Vision-Language-Action model trained on 970K+ robot demonstrations from the Open X-Embodiment dataset for generalist robotic manipulation. | ![GitHub Badge](https://img.shields.io/github/stars/openvla/openvla.svg?style=flat-square) |
      • SmolVLA
  • Observability

    • PromptHub - Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.
    • Prompteams - Prompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).
    • Doku - An open-source LLM Observability platform streamlining the monitoring of LLM applications with just two lines of code. It provides valuable insights into token usage and user engagement, tracks API usage for providers like OpenAI, and facilitates easy data export to observability platforms like Grafana and DataDog.
  • Optimizations

    • Profiling

      • FeatherCNN - square) |
      • Forward - square) |
      • NCNN - performance neural network inference framework optimized for the mobile platform. | ![GitHub Badge](https://img.shields.io/github/stars/Tencent/ncnn.svg?style=flat-square) |
      • PocketFlow - square) |
      • TensorFlow Model Optimization - optimization.svg?style=flat-square) |
      • TNN - square) |
      • optimum-tpu - tpu.svg?style=flat-square) |
      • LangWatch - square) |
      • agent-opt - driven iterative refinements. | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/agent-opt?style=flat-square) |
      • Entroly - theoretic context optimization proxy. Cuts LLM token costs by 70–95% with zero accuracy loss using greedy submodular knapsack maximization. | ![GitHub Badge](https://img.shields.io/github/stars/juyterman1000/entroly.svg?style=flat-square) |
      • lean-ctx - aware compression, and shell output patterns. [Website](https://leanctx.com) | ![GitHub Badge](https://img.shields.io/github/stars/yvgude/lean-ctx.svg?style=flat-square) |
  • Performance

    • ML Compiler

      • ONNX-MLIR - mlir.svg?style=flat-square) |
      • TVM - square) |
      • bitsandbytes - bit quantization for PyTorch. | ![GitHub Badge](https://img.shields.io/github/stars/bitsandbytes-foundation/bitsandbytes?style=flat-square) |
    • Profiling

      • octoml-profile - profile is a python library and cloud service designed to provide the simplest experience for assessing and optimizing the performance of PyTorch models on cloud hardware with state-of-the-art ML acceleration technology. | ![GitHub Badge](https://img.shields.io/github/stars/octoml/octoml-profile.svg?style=flat-square) |
      • scalene - performance, high-precision CPU, GPU, and memory profiler for Python | ![GitHub Badge](https://img.shields.io/github/stars/plasma-umass/scalene.svg?style=flat-square) |
      • Airweave - ai/airweave.svg?style=flat-square) |
      • Pinecone - performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles. | |
      • Vellum - of-box support for OCR, text chunking, embedding model experimentation, metadata filtering, and production-grade APIs. | |
      • Awadb - ai/awadb.svg?style=flat-square) |
      • Chroma - core/chroma.svg?style=flat-square) |
      • Infinity - native database built for LLM applications, providing incredibly fast vector and full-text search | ![GitHub Badge](https://img.shields.io/github/stars/infiniflow/infinity.svg?style=flat-square) |
      • Lancedb - friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps! | ![GitHub Badge](https://img.shields.io/github/stars/lancedb/lancedb.svg?style=flat-square) |
      • Marqo - ai/marqo.svg?style=flat-square) |
      • Milvus - io/milvus.svg?style=flat-square) |
      • pgvector - source vector similarity search for Postgres. | ![GitHub Badge](https://img.shields.io/github/stars/pgvector/pgvector.svg?style=flat-square) |
      • pgvecto.rs - square) |
      • Qdrant - square) |
      • txtai - powered semantic search applications | ![GitHub Badge](https://img.shields.io/github/stars/neuml/txtai.svg?style=flat-square) |
      • Vald - square) |
      • Vearch - based vector retrieval | ![GitHub Badge](https://img.shields.io/github/stars/vearch/vearch.svg?style=flat-square) |
      • VectorDB - no more, no less. | ![GitHub Badge](https://img.shields.io/github/stars/jina-ai/vectordb.svg?style=flat-square) |
      • Epsilla - cloud/vectordb.svg?style=flat-square) |
      • VectorChord - friendly vector search in Postgres, the successor of `pgvecto.rs`. | ![GitHub Badge](https://img.shields.io/github/stars/tensorchord/VectorChord.svg?style=flat-square) |
      • ParadeDB - square) |
      • AquilaDB - NN search. | ![GitHub Badge](https://img.shields.io/github/stars/Aquila-Network/AquilaDB.svg?style=flat-square) |
      • Weaviate - tolerance and scalability of a cloud-native database, all accessible through GraphQL, REST, and various language clients. | ![GitHub Badge](https://img.shields.io/github/stars/semi-technologies/weaviate.svg?style=flat-square) |
      • Omnigraph - native, Rust, traversal + vector + BM25 in one runtime. | ![GitHub Badge](https://img.shields.io/github/stars/ModernRelay/omnigraph.svg?style=flat-square) |
      • Rivestack - in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage. | |
  • Security

    • Frameworks for LLM security

      • Plexiglass - labs/plexiglass?style=flat-square) |
      • Cordum - first agent orchestration platform with pre-dispatch policy evaluation, output scanning (PII, secrets, injection), job scheduling, workflow engine, and full audit trail. | ![GitHub Badge](https://img.shields.io/github/stars/cordum-io/cordum.svg?style=flat-square) |
      • brood-box - isolated microVMs with snapshot isolation, egress control, and MCP authorization. | ![GitHub Badge](https://img.shields.io/github/stars/stacklok/brood-box?style=flat-square) |
      • dstack - source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing. | ![GitHub Badge](https://img.shields.io/github/stars/Dstack-TEE/dstack?style=flat-square) |
    • Observability

      • Azure OpenAI Logger - openai-logger?style=flat-square) |
      • Deepchecks - square) |
      • Fiddler AI - production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity. | ![GitHub Badge](https://img.shields.io/github/stars/fiddler-labs/fiddler-auditor.svg?style=flat-square) |
      • Giskard - AI/giskard.svg?style=flat-square) |
      • Great Expectations - expectations/great_expectations.svg?style=flat-square) |
      • whylogs - square) |
      • Traceloop OpenLLMetry - based observability and monitoring for LLM and agents workflows. | ![GitHub Badge](https://img.shields.io/github/stars/traceloop/openllmetry.svg?style=flat-square)
      • traceAI - source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows. | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/traceAI?style=flat-square) |
      • Future AGI - grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows. | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/futureagi-sdk?style=flat-square) |
      • semantic-coverage - agi/futureagi-sdk?style=flat-square) |
      • QWED - AI/qwed-verification.svg?style=flat-square) |
      • RagTune - square) |
      • ClevAgent - restart. Python SDK or HTTP API. | |
      • EvalView - call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API. | ![GitHub Badge](https://img.shields.io/github/stars/hidai25/eval-view.svg?style=flat-square) |
      • onWatch - dev/onwatch.svg?style=flat-square) |
  • Serving

    • Frameworks/Servers for Serving

      • BentoML - square) |
      • Mosec - to-use Python interface. | ![GitHub Badge](https://img.shields.io/github/stars/mosecorg/mosec?style=flat-square) |
      • TFServing - performance serving system for machine learning models. | ![GitHub Badge](https://img.shields.io/github/stars/tensorflow/serving.svg?style=flat-square) |
      • Torchserve - square) |
      • Triton Server (TRTIS) - inference-server/server.svg?style=flat-square) |
      • langchain-serve - ai/langchain-serve.svg?style=flat-square) |
      • lanarky - grade LLM applications | ![GitHub Badge](https://img.shields.io/github/stars/ajndkr/lanarky.svg?style=flat-square) |
      • ray-llm - RayLLM | ![GitHub Badge](https://img.shields.io/github/stars/ray-project/ray-llm.svg?style=flat-square) |
      • Xinference - source language models, speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop. | ![GitHub Badge](https://img.shields.io/github/stars/xorbitsai/inference.svg?style=flat-square) |
      • KubeAI - to-text. | ![GitHub Badge](https://img.shields.io/github/stars/substratusai/kubeai.svg?style=flat-square) |
      • Kaito - 3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers. | ![GitHub Badge](https://img.shields.io/github/stars/kaito-project/kaito.svg?style=flat-square) |
      • Open Responses - source platform for building long-running LLM agents with tool use. | ![GitHub Badge](https://img.shields.io/github/stars/julep-ai/julep.svg?style=flat-square) |
      • Open Responses - source platform for building long-running LLM agents with tool use. | ![GitHub Badge](https://img.shields.io/github/stars/julep-ai/julep.svg?style=flat-square) |
      • Open Responses - source platform for building long-running LLM agents with tool use. | ![GitHub Badge](https://img.shields.io/github/stars/julep-ai/julep.svg?style=flat-square) |
      • Jina - ai/jina.svg?style=flat-square) |
      • mcpproxy-go - source MCP proxy with BM25 tool filtering, quarantine security, activity logging, and web UI. Routes multiple MCP servers through single endpoint, reducing context bloat by ~97%. | ![GitHub Badge](https://img.shields.io/github/stars/smart-mcp-proxy/mcpproxy-go.svg?style=flat-square) |
      • KubeStellar Console - powered multi-cluster Kubernetes dashboard for hybrid edge and cloud. GPU monitoring, LLM inference cluster management, benchmark streaming, and 20+ CNCF integrations. CNCF Sandbox (Apache 2.0). | ![GitHub Badge](https://img.shields.io/github/stars/kubestellar/console.svg?style=flat-square) |
    • Large Model Serving

      • CTranslate2 - square) |
      • Clip-as-a-service - ai/clip-as-service.svg?style=flat-square) |
      • Flowise - square) |
      • Infinity - embeddings | ![GitHub Badge](https://img.shields.io/github/stars/michaelfeil/infinity.svg?style=flat-square) |
      • Modelz-LLM - llm.svg?style=flat-square) |
      • TensorRT-LLM - LLM.svg?style=flat-square) |
      • text-generation-inference - generation-inference.svg?style=flat-square) |
      • text-embeddings-inference - embedding models | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/text-embeddings-inference.svg?style=flat-square) |
      • vllm - throughput and memory-efficient inference and serving engine for LLMs. | ![GitHub stars](https://img.shields.io/github/stars/vllm-project/vllm.svg?style=flat-square) |
      • x-stable-diffusion - time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. | ![GitHub Badge](https://img.shields.io/github/stars/stochasticai/x-stable-diffusion.svg?style=flat-square) |
      • tokenizers - of-the-Art Tokenizers optimized for Research and Production | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/tokenizers.svg?style=flat-square) |
      • DeepSpeed-MII - latency and high-throughput inference possible, powered by DeepSpeed. | ![GitHub Badge](https://img.shields.io/github/stars/microsoft/DeepSpeed-MII.svg?style=flat-square) |
      • prima.cpp - square) |
      • FlexGen - oriented scenarios. | ![GitHub Badge](https://img.shields.io/github/stars/FMInference/FlexGen.svg?style=flat-square) |
      • Shimmy - free Rust inference server with OpenAI API compatibility and hot model swapping | ![GitHub Badge](https://img.shields.io/github/stars/Michael-A-Kuykendall/shimmy.svg?style=flat-square) |
      • llama.cpp - square) |
      • Ollama - square) |
      • whisper.cpp - square) |
      • Rapid-MLX - compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching. | ![GitHub Badge](https://img.shields.io/github/stars/raullenchai/Rapid-MLX.svg?style=flat-square) |
      • OneComp - training quantization pipeline for LLMs (QEP, AutoBit, JointQ, rotation) with vLLM plugin (arXiv:2603.28845). | ![GitHub Badge](https://img.shields.io/github/stars/FujitsuResearch/OneCompression.svg?style=flat-square) |
      • LLMKube - GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats. | ![GitHub Badge](https://img.shields.io/github/stars/defilantech/LLMKube.svg?style=flat-square) |
      • Off Grid - source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline. | ![GitHub Badge](https://img.shields.io/github/stars/alichherawalla/off-grid-mobile-ai.svg?style=flat-square) |
      • whisper-ctranslate2 - memory usage drop-in cli replacement that supports word-level timestamps and VAD filter | ![GitHub Badge](https://img.shields.io/github/stars/Softcatala/whisper-cTranslate2?style=flat-square) |
  • Training

    • Experiment Tracking

      • Aim - to-use and performant open-source experiment tracker. | ![GitHub Badge](https://img.shields.io/github/stars/aimhubio/aim.svg?style=flat-square) |
      • Guild AI - square) |
      • Kedro-Viz - Viz is an interactive development tool for building data science pipelines with Kedro. Kedro-Viz also allows users to view and compare different runs in the Kedro project. | ![GitHub Badge](https://img.shields.io/github/stars/kedro-org/kedro-viz.svg?style=flat-square) |
      • LabNotebook - square) |
      • Sacred - square) |
      • ClearML - Magical CI/CD to streamline your ML workflow. Experiment Manager, MLOps and Data-Management | ![GitHub Badge](https://img.shields.io/github/stars/allegroai/clearml.svg?style=flat-square) |
    • Foundation Model Fine Tuning

      • Flyflow - devs/flyflow.svg?style=flat-square) |
      • alpaca-lora - tune LLaMA on consumer hardware | ![GitHub Badge](https://img.shields.io/github/stars/tloen/alpaca-lora.svg?style=flat-square) |
      • finetuning-scheduler - tuning schedules. | ![GitHub Badge](https://img.shields.io/github/stars/speediedan/finetuning-scheduler.svg?style=flat-square) |
      • LMFlow - square) |
      • Lora - rank adaptation to quickly fine-tune diffusion models. | ![GitHub Badge](https://img.shields.io/github/stars/cloneofsimo/lora.svg?style=flat-square) |
      • peft - of-the-art Parameter-Efficient Fine-Tuning. | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/peft.svg?style=flat-square) |
      • p-tuning-v2 - tuning on small/medium-sized models and sequence tagging challenges. [(ACL 2022)](https://arxiv.org/abs/2110.07602) | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/P-tuning-v2.svg?style=flat-square) |
      • QLoRA - bit finetuning task performance. | ![GitHub Badge](https://img.shields.io/github/stars/artidoro/qlora.svg?style=flat-square) |
      • TRL - square) |
    • Frameworks for Training

      • Accelerate - GPU, TPU, mixed-precision. | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/accelerate.svg?style=flat-square) |
      • Apache MXNet - aware Dataflow Dep Scheduler. | ![GitHub Badge](https://img.shields.io/github/stars/apache/mxnet.svg?style=flat-square) |
      • Caffe - square) |
      • ColossalAI - scale model training system with efficient parallelization techniques. | ![GitHub Badge](https://img.shields.io/github/stars/hpcaitech/ColossalAI.svg?style=flat-square) |
      • Horovod - square) |
      • Kedro - source Python framework for creating reproducible, maintainable and modular data science code. | ![GitHub Badge](https://img.shields.io/github/stars/kedro-org/kedro.svg?style=flat-square) |
      • Keras - team/keras.svg?style=flat-square) |
      • LightGBM - square) |
      • MegEngine - to-use deep learning framework, with auto-differentiation. | ![GitHub Badge](https://img.shields.io/github/stars/MegEngine/MegEngine.svg?style=flat-square) |
      • metric-learn - learn-contrib/metric-learn.svg?style=flat-square) |
      • MindSpore - ai/mindspore.svg?style=flat-square) |
      • Oneflow - centered and open-source deep learning framework. | ![GitHub Badge](https://img.shields.io/github/stars/Oneflow-Inc/oneflow.svg?style=flat-square) |
      • PaddlePaddle - square) |
      • PyTorch - square) |
      • XGBoost - square) |
      • scikit-learn - learn/scikit-learn.svg?style=flat-square) |
      • TensorFlow - square) |
      • VectorFlow - square) |
      • Candle - square`) |
      • DeepSpeed - square) |
      • Jax - performance machine learning research. | ![GitHub Badge](https://img.shields.io/github/stars/google/jax.svg?style=flat-square) |
      • axolotl - tuning of various AI models, offering support for multiple configurations and architectures. | ![GitHub Badge](https://img.shields.io/github/stars/OpenAccess-AI-Collective/axolotl.svg?style=flat-square) |
    • IDEs and Workspaces

      • code server - server.svg?style=flat-square) |
      • conda - agnostic, system-level binary package manager and ecosystem. | ![GitHub Badge](https://img.shields.io/github/stars/conda/conda.svg?style=flat-square) |
      • Docker - source project created by Docker to enable and accelerate software containerization. | ![GitHub Badge](https://img.shields.io/github/stars/moby/moby.svg?style=flat-square) |
      • envd - square) |
      • Jupyter Notebooks - based notebook environment for interactive computing. | ![GitHub Badge](https://img.shields.io/github/stars/jupyter/notebook.svg?style=flat-square) |
      • Kurtosis - container environments. | ![GitHub Badge](https://img.shields.io/github/stars/kurtosis-tech/kurtosis.svg?style=flat-square) |
    • Model Editing

    • Visualization

      • Fiddler AI
      • Maniford - agnostic visual debugging tool for machine learning. | ![GitHub Badge](https://img.shields.io/github/stars/uber/manifold.svg?style=flat-square) |
      • netron - square) |
      • OpenOps - square) |
      • TensorBoard - square) |
      • TensorSpace - trained deep learning models from TensorFlow, Keras, TensorFlow.js. | ![GitHub Badge](https://img.shields.io/github/stars/tensorspace-team/tensorspace.svg?style=flat-square) |
      • dtreeviz - square) |
      • Zetane Viewer - square) |
      • Zeno - ml/zeno.svg?style=flat-square) |
      • OpenOps - square) |