{"id":45999,"url":"https://github.com/tensorchord/awesome-llmops","name":"awesome-llmops","description":"An awesome \u0026 curated list of best LLMOps tools for developers","projects_count":400,"last_synced_at":"2026-08-02T16:00:22.385Z","repository":{"id":40565412,"uuid":"481805419","full_name":"tensorchord/Awesome-LLMOps","owner":"tensorchord","description":"An awesome \u0026 curated list of best LLMOps tools for developers","archived":false,"fork":false,"pushed_at":"2026-05-21T09:12:50.000Z","size":1035,"stargazers_count":5877,"open_issues_count":163,"forks_count":907,"subscribers_count":82,"default_branch":"main","last_synced_at":"2026-07-12T23:03:48.702Z","etag":null,"topics":["ai-development-tools","awesome-list","llmops","mlops"],"latest_commit_sha":null,"homepage":"","language":"Shell","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc0-1.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/tensorchord.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"contributing.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2022-04-15T01:56:44.000Z","updated_at":"2026-07-12T06:52:30.000Z","dependencies_parsed_at":"2024-02-13T21:28:20.934Z","dependency_job_id":"50afcdcd-fa73-4790-ad2b-85e17aea3789","html_url":"https://github.com/tensorchord/Awesome-LLMOps","commit_stats":null,"previous_names":["tensorchord/awesome-open-source-llmops","tensorchord/awesome-open-source-mlops"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/tensorchord/Awesome-LLMOps","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tensorchord%2FAwesome-LLMOps","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tensorchord%2FAwesome-LLMOps/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tensorchord%2FAwesome-LLMOps/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tensorchord%2FAwesome-LLMOps/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/tensorchord","download_url":"https://codeload.github.com/tensorchord/Awesome-LLMOps/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/tensorchord%2FAwesome-LLMOps/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36199567,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-08-02T02:00:06.915Z","response_time":58,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-01-14T02:14:39.008Z","updated_at":"2026-08-02T16:00:22.386Z","primary_language":"Python","list_of_lists":false,"displayable":true,"categories":["Large Scale Deployment","Search","Model","LLMOps","AutoML","Observability","Training","ML Platforms","Awesome Lists","Serving","Data","Optimizations","Code AI","Performance","Security","Federated ML"],"sub_categories":["Workflow","Vector search","CV Foundation Model","Observability","Profiling","Large Language Model","Visualization","Foundation Model Fine Tuning","Frameworks for Training","Frameworks/Servers for Serving","Data Management","IDEs and Workspaces","Audio Foundation Model","Large Model Serving","ML Platforms","Data Storage","ML Compiler","Robotics Foundation Model","Experiment Tracking","Scheduling","Data Tracking","Data/Feature enrichment","Model Management","Feature Engineering","Frameworks for LLM security","Model Editing","Hybrid search"],"readme":"# Awesome LLMOps\n\n\u003ca href=\"https://discord.gg/KqswhpVgdU\"\u003e\u003cimg alt=\"discord invitation link\" src=\"https://img.shields.io/discord/974584200327991326?style=flat\u0026logo=discord\u0026cacheSeconds=60\"\u003e\u003c/a\u003e\n\u003ca href=\"https://awesome.re\"\u003e\u003cimg src=\"https://awesome.re/badge-flat2.svg\"\u003e\u003c/a\u003e\n\nAn awesome \u0026 curated list of the best LLMOps tools for developers.\n\n\u003e [!NOTE]\n\u003e Contributions are most welcome, please adhere to the [contribution guidelines](contributing.md).\n\n## Table of Contents\n\n- [Awesome LLMOps](#awesome-llmops)\n  - [Table of Contents](#table-of-contents)\n  - [Model](#model)\n    - [Large Language Model](#large-language-model)\n    - [CV Foundation Model](#cv-foundation-model)\n    - [Audio Foundation Model](#audio-foundation-model)\n    - [Robotics Foundation Model](#robotics-foundation-model)\n  - [Serving](#serving)\n    - [Large Model Serving](#large-model-serving)\n    - [Frameworks/Servers for Serving](#frameworksservers-for-serving)\n  - [Security](#security)\n    - [Frameworks for LLM security](#frameworks-for-llm-security)\n    - [Observability](#observability)\n  - [LLMOps](#llmops)\n  - [Search](#search)\n    - [Vector search](#vector-search)\n  - [Code AI](#code-ai)\n  - [Training](#training)\n    - [IDEs and Workspaces](#ides-and-workspaces)\n    - [Foundation Model Fine Tuning](#foundation-model-fine-tuning)\n    - [Frameworks for Training](#frameworks-for-training)\n    - [Experiment Tracking](#experiment-tracking)\n    - [Visualization](#visualization)\n    - [Model Editing](#model-editing)\n  - [Data](#data)\n    - [Data Management](#data-management)\n    - [Data Storage](#data-storage)\n    - [Data Tracking](#data-tracking)\n    - [Feature Engineering](#feature-engineering)\n    - [Data/Feature enrichment](#datafeature-enrichment)\n  - [Large Scale Deployment](#large-scale-deployment)\n    - [ML Platforms](#ml-platforms)\n    - [Workflow](#workflow)\n    - [Scheduling](#scheduling)\n    - [Model Management](#model-management)\n  - [Performance](#performance)\n    - [ML Compiler](#ml-compiler)\n    - [Profiling](#profiling)\n  - [AutoML](#automl)\n  - [Optimizations](#optimizations)\n  - [Federated ML](#federated-ml)\n  - [Awesome Lists](#awesome-lists)\n\n\u003c!-- Created by https://github.com/ekalinin/github-markdown-toc --\u003e\n\n## Model\n\n### Large Language Model\n\n| Project                                                                 | Details                                                                                                                                                                                    | Repository                                                                                                |\n| ----------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------- |\n| [Alpaca](https://github.com/tatsu-lab/stanford_alpaca)                  | Code and documentation to train Stanford's Alpaca models, and generate the data.                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/tatsu-lab/stanford_alpaca.svg?style=flat-square)      |\n| [BELLE](https://github.com/LianjiaTech/BELLE)                           | A 7B Large Language Model fine-tune by 34B Chinese Character Corpus, based on LLaMA and Alpaca.                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/LianjiaTech/BELLE.svg?style=flat-square)              |\n| [Bloom](https://github.com/bigscience-workshop/model_card)              | BigScience Large Open-science Open-access Multilingual Language Model                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/bigscience-workshop/model_card.svg?style=flat-square) |\n| [dolly](https://github.com/databrickslabs/dolly)                        | Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform                                                                                              | ![GitHub Badge](https://img.shields.io/github/stars/databrickslabs/dolly.svg?style=flat-square)           |\n| [Falcon 40B](https://huggingface.co/tiiuae/falcon-40b-instruct)         | Falcon-40B-Instruct is a 40B parameters causal decoder-only model built by TII based on Falcon-40B and finetuned on a mixture of Baize. It is made available under the Apache 2.0 license. |                                                                                                           |\n| [FastChat (Vicuna)](https://github.com/lm-sys/FastChat)                 | An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and FastChat-T5.                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/lm-sys/FastChat.svg?style=flat-square)                |\n| [Gemma](https://www.kaggle.com/models/google/gemma)                     | Gemma is a family of lightweight, open models built from the research and technology that Google used to create the Gemini models.                                                         |                                                                                                           |\n| [GLM-6B (ChatGLM)](https://github.com/THUDM/ChatGLM-6B)                 | An Open Bilingual Pre-Trained Model, quantization of ChatGLM-130B, can run on consumer-level GPUs.                                                                                         | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/ChatGLM-6B.svg?style=flat-square)               |\n| [ChatGLM2-6B](https://github.com/THUDM/ChatGLM2-6B)                     | ChatGLM2-6B is the second-generation version of the open-source bilingual (Chinese-English) chat model [ChatGLM-6B](https://github.com/THUDM/ChatGLM-6B).                                  | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/ChatGLM2-6B.svg?style=flat-square)              |\n| [GLM-130B (ChatGLM)](https://github.com/THUDM/GLM-130B)                 | An Open Bilingual Pre-Trained Model (ICLR 2023)                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/THUDM/GLM-130B.svg?style=flat-square)                 |\n| [GPT-NeoX](https://github.com/EleutherAI/gpt-neox)                      | An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.                                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/EleutherAI/gpt-neox.svg?style=flat-square)            |\n| [Luotuo](https://github.com/LC1332/Luotuo-Chinese-LLM)                  | A Chinese LLM, Based on LLaMA and fine tune by Stanford Alpaca, Alpaca LoRA, Japanese-Alpaca-LoRA.                                                                                         | ![GitHub Badge](https://img.shields.io/github/stars/LC1332/Luotuo-Chinese-LLM.svg?style=flat-square)      |\n| [Mixtral-8x7B-v0.1](https://huggingface.co/mistralai/Mixtral-8x7B-v0.1) | The Mixtral-8x7B Large Language Model (LLM) is a pretrained generative Sparse Mixture of Experts.                                                                                          |                                                                                                           |\n| [StableLM](https://github.com/Stability-AI/StableLM)                    | StableLM: Stability AI Language Models                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/Stability-AI/StableLM.svg?style=flat-square)          |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n### CV Foundation Model\n\n| Project                                                                        | Details                                                                                                                                          | Repository                                                                                                   |\n| ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------ |\n| [disco-diffusion](https://github.com/alembics/disco-diffusion)                 | A frankensteinian amalgamation of notebooks, models and techniques for the generation of AI Art and Animations.                                  | ![GitHub Badge](https://img.shields.io/github/stars/alembics/disco-diffusion.svg?style=flat-square)          |\n| [midjourney](https://www.midjourney.com/home/)                                 | Midjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species.            |                                                                                                              |\n| [segment-anything (SAM)](https://github.com/facebookresearch/segment-anything) | produces high quality object masks from input prompts such as points or boxes, and it can be used to generate masks for all objects in an image. | ![GitHub Badge](https://img.shields.io/github/stars/facebookresearch/segment-anything.svg?style=flat-square) |\n| [stable-diffusion](https://github.com/CompVis/stable-diffusion)                | A latent text-to-image diffusion model                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/CompVis/stable-diffusion.svg?style=flat-square)          |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n### Audio Foundation Model\n\n| Project                                      | Details                                                                                                                                                                                                       | Repository                                                                                |\n| -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- |\n| [bark](https://github.com/suno-ai/bark)      | Bark is a transformer-based text-to-audio model created by Suno. Bark can generate highly realistic, multilingual speech as well as other audio - including music, background noise and simple sound effects. | ![GitHub Badge](https://img.shields.io/github/stars/suno-ai/bark.svg?style=flat-square)   |\n| [whisper](https://github.com/openai/whisper) | Robust Speech Recognition via Large-Scale Weak Supervision                                                                                                                                                    | ![GitHub Badge](https://img.shields.io/github/stars/openai/whisper.svg?style=flat-square) |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n### Robotics Foundation Model\n\n\u003e [!NOTE]\n\u003e **Emerging Architectures in VLA:**\n\u003e - **Continuous Diffusion Language Models:** Integrate diffusion heads or flow-matching to VLMs (e.g., DiVLA, OpenPI), enabling smooth, precise continuous action generation rather than discretized tokens.\n\u003e - **Recurrent Language Models:** Utilize State Space Models (SSMs) like Mamba or recurrent transformers (e.g., RoboMamba, RD-VLA) to reduce inference memory and handle temporal dependencies, allowing iterative reasoning for complex robotic decision-making.\n\n| Project                                                   | Details                                                                                                                                                                                                    | Repository                                                                                             |\n| --------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |\n| [DiVLA](https://github.com/hustvl/DiVLA)                  | A continuous diffusion-based Vision-Language-Action model that integrates diffusion policies into autoregressive VLMs for robust and precise continuous robotic control.                                   | ![GitHub Badge](https://img.shields.io/github/stars/hustvl/DiVLA.svg?style=flat-square)                |\n| [LeRobot](https://github.com/huggingface/lerobot)         | A central community library by Hugging Face for AI in robotics — end-to-end learning tools, data pipelines, and support for training/deploying VLA models.                                                 | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/lerobot.svg?style=flat-square)         |\n| [Octo](https://github.com/octo-models/octo)               | A transformer-based generalist robot policy pretrained on 800K+ robot trajectories from the Open X-Embodiment dataset. Supports language instructions, goal images, and fine-tuning to new embodiments.      | ![GitHub Badge](https://img.shields.io/github/stars/octo-models/octo.svg?style=flat-square)            |\n| [OpenPI](https://github.com/Physical-Intelligence/openpi) | Open-source VLA models from Physical Intelligence, including π₀ and π₀.5 — flow-based vision-language-action models pretrained on large-scale robot data with fine-tuning support.                         | ![GitHub Badge](https://img.shields.io/github/stars/Physical-Intelligence/openpi.svg?style=flat-square) |\n| [OpenVLA](https://github.com/openvla/openvla)             | A 7B-parameter open-source Vision-Language-Action model trained on 970K+ robot demonstrations from the Open X-Embodiment dataset for generalist robotic manipulation.                                      | ![GitHub Badge](https://img.shields.io/github/stars/openvla/openvla.svg?style=flat-square)             |\n| [RoboMamba](https://github.com/hustvl/RoboMamba)          | An efficient VLA model leveraging State Space Models (Mamba) instead of standard self-attention, offering linear inference complexity for efficient, recurrent robotic reasoning.                          | ![GitHub Badge](https://img.shields.io/github/stars/hustvl/RoboMamba.svg?style=flat-square)            |\n| [SmolVLA](https://huggingface.co/blog/smolvla)            | A compact ~450M parameter VLA by Hugging Face, designed to be computationally efficient and accessible, running on consumer GPUs or CPUs. Part of the LeRobot ecosystem.                                   |                                                                                                        |\n\n## Serving\n\n### Large Model Serving\n\n| Project                                                                               | Details                                                                                                         | Repository                                                                                                       |\n| ------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |\n| [Alpaca-LoRA-Serve](https://github.com/deep-diver/Alpaca-LoRA-Serve)                  | Alpaca-LoRA as Chatbot service                                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/deep-diver/Alpaca-LoRA-Serve.svg?style=flat-square)          |\n| [OneComp](https://github.com/FujitsuResearch/OneCompression)                               | Fujitsu Research's post-training quantization pipeline for LLMs (QEP, AutoBit, JointQ, rotation) with vLLM plugin (arXiv:2603.28845).                             | ![GitHub Badge](https://img.shields.io/github/stars/FujitsuResearch/OneCompression.svg?style=flat-square)        |\n| [CTranslate2](https://github.com/OpenNMT/CTranslate2)                                 | fast inference engine for Transformer models in C++                                                             | ![GitHub Badge](https://img.shields.io/github/stars/OpenNMT/CTranslate2.svg?style=flat-square)                   |\n| [Clip-as-a-service](https://github.com/jina-ai/clip-as-service)                       | serving the OpenAI CLIP model                                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/jina-ai/clip-as-service.svg?style=flat-square)               |\n| [DeepSpeed-MII](https://github.com/microsoft/DeepSpeed-MII)                           | MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.                             | ![GitHub Badge](https://img.shields.io/github/stars/microsoft/DeepSpeed-MII.svg?style=flat-square)               |\n| [Faster Whisper](https://github.com/guillaumekln/faster-whisper)                      | fast inference engine for whisper in C++ using CTranslate2.                                                     | ![GitHub Badge](https://img.shields.io/github/stars/guillaumekln/faster-whisper.svg?style=flat-square)           |\n| [FlexGen](https://github.com/FMInference/FlexGen)                                     | Running large language models on a single GPU for throughput-oriented scenarios. *(Archived)*                   | ![GitHub Badge](https://img.shields.io/github/stars/FMInference/FlexGen.svg?style=flat-square)                   |\n| [Flowise](https://github.com/FlowiseAI/Flowise)                                       | Drag \u0026 drop UI to build your customized LLM flow using LangchainJS.                                             | ![GitHub Badge](https://img.shields.io/github/stars/FlowiseAI/Flowise.svg?style=flat-square)                     |\n| [llama.cpp](https://github.com/ggerganov/llama.cpp)                                   | Port of Facebook's LLaMA model in C/C++                                                                         | ![GitHub Badge](https://img.shields.io/github/stars/ggerganov/llama.cpp.svg?style=flat-square)                   |\n| [LLMKube](https://github.com/defilantech/LLMKube)                                     | Kubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats. | ![GitHub Badge](https://img.shields.io/github/stars/defilantech/LLMKube.svg?style=flat-square)                   |\n| [Shimmy](https://github.com/Michael-A-Kuykendall/shimmy)                               | Python-free Rust inference server with OpenAI API compatibility and hot model swapping                        | ![GitHub Badge](https://img.shields.io/github/stars/Michael-A-Kuykendall/shimmy.svg?style=flat-square)        |\n| [Infinity](https://github.com/michaelfeil/infinity)                                   | Rest API server for serving text-embeddings                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/michaelfeil/infinity.svg?style=flat-square)                  |\n| [Modelz-LLM](https://github.com/tensorchord/modelz-llm)                               | OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)                          | ![GitHub Badge](https://img.shields.io/github/stars/tensorchord/modelz-llm.svg?style=flat-square)                |\n| [Off Grid](https://github.com/alichherawalla/off-grid-mobile-ai)                      | Open-source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline. | ![GitHub Badge](https://img.shields.io/github/stars/alichherawalla/off-grid-mobile-ai.svg?style=flat-square) |\n| [Ollama](https://github.com/jmorganca/ollama)                                         | Serve Llama 2 and other large language models locally from command line or through a browser interface.         | ![GitHub Badge](https://img.shields.io/github/stars/jmorganca/ollama.svg?style=flat-square)                      |\n| [Rapid-MLX](https://github.com/raullenchai/Rapid-MLX)                                 | OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching. | ![GitHub Badge](https://img.shields.io/github/stars/raullenchai/Rapid-MLX.svg?style=flat-square)                  |\n| [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM)                                | Inference engine for TensorRT on Nvidia GPUs                                                                    | ![GitHub Badge](https://img.shields.io/github/stars/NVIDIA/TensorRT-LLM.svg?style=flat-square)                   |\n| [text-generation-inference](https://github.com/huggingface/text-generation-inference) | Large Language Model Text Generation Inference                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/text-generation-inference.svg?style=flat-square) |\n| [text-embeddings-inference](https://github.com/huggingface/text-embeddings-inference) | Inference for text-embedding models                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/text-embeddings-inference.svg?style=flat-square) |\n| [tokenizers](https://github.com/huggingface/tokenizers)                               | 💥 Fast State-of-the-Art Tokenizers optimized for Research and Production                                       | ![GitHub Badge](https://img.shields.io/github/stars/huggingface/tokenizers.svg?style=flat-square)                |\n| [vllm](https://github.com/vllm-project/vllm)                                          | A high-throughput and memory-efficient inference and serving engine for LLMs.                                   | ![GitHub stars](https://img.shields.io/github/stars/vllm-project/vllm.svg?style=flat-square)                     |\n| [whisper-ctranslate2](https://github.com/Softcatala/whisper-ctranslate2)              |  is a 4x faster and low-memory usage drop-in cli replacement that supports word-level timestamps and VAD filter | ![GitHub Badge](https://img.shields.io/github/stars/Softcatala/whisper-cTranslate2?style=flat-square)                   |\n| [whisper.cpp](https://github.com/ggerganov/whisper.cpp)                               | Port of OpenAI's Whisper model in C/C++                                                                         | ![GitHub Badge](https://img.shields.io/github/stars/ggerganov/whisper.cpp.svg?style=flat-square)                 |\n| [x-stable-diffusion](https://github.com/stochasticai/x-stable-diffusion)              | Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. *(Archived)* | ![GitHub Badge](https://img.shields.io/github/stars/stochasticai/x-stable-diffusion.svg?style=flat-square)       |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n### Frameworks/Servers for Serving\n\n| Project                                                                    | Details                                                                                                                                                                                                                                                                                                                                            | Repository                                                                                                |\n| -------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |\n| [BentoML](https://github.com/bentoml/BentoML)                              | The Unified Model Serving Framework                                                                                                                                                                                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/bentoml/BentoML.svg?style=flat-square)                |\n| [Jina](https://github.com/jina-ai/jina)                                    | Build multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud Native                                                                                                                                                                                                                          | ![GitHub Badge](https://img.shields.io/github/stars/jina-ai/jina.svg?style=flat-square)                   |\n| [Mosec](https://github.com/mosecorg/mosec)                                 | A machine learning model serving framework with dynamic batching and pipelined stages, provides an easy-to-use Python interface.                                                                                                                                                                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/mosecorg/mosec?style=flat-square)                     |\n| [mcpproxy-go](https://github.com/smart-mcp-proxy/mcpproxy-go)              | Open-source MCP proxy with BM25 tool filtering, quarantine security, activity logging, and web UI. Routes multiple MCP servers through single endpoint, reducing context bloat by ~97%.                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/smart-mcp-proxy/mcpproxy-go.svg?style=flat-square)    |\n| [TFServing](https://github.com/tensorflow/serving)                         | A flexible, high-performance serving system for machine learning models.                                                                                                                                                                                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/tensorflow/serving.svg?style=flat-square)             |\n| [Torchserve](https://github.com/pytorch/serve)                             | Serve, optimize and scale PyTorch models in production *(Archived)*                                                                                                                                                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/pytorch/serve.svg?style=flat-square)                  |\n| [Triton Server (TRTIS)](https://github.com/triton-inference-server/server) | The Triton Inference Server provides an optimized cloud and edge inferencing solution.                                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/triton-inference-server/server.svg?style=flat-square) |\n| [langchain-serve](https://github.com/jina-ai/langchain-serve)              | Serverless LLM apps on Production with Jina AI Cloud *(Archived)*                                                                                                                                                                                                                                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/jina-ai/langchain-serve.svg?style=flat-square)        |\n| [lanarky](https://github.com/ajndkr/lanarky)                               | FastAPI framework to build production-grade LLM applications                                                                                                                                                                                                                                                                                       | ![GitHub Badge](https://img.shields.io/github/stars/ajndkr/lanarky.svg?style=flat-square)                 |\n| [ray-llm](https://github.com/ray-project/ray-llm)                          | LLMs on Ray - RayLLM *(Archived)*                                                                                                                                                                                                                                                                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/ray-project/ray-llm.svg?style=flat-square)            |\n| [Xinference](https://github.com/xorbitsai/inference)                       | Replace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop. | ![GitHub Badge](https://img.shields.io/github/stars/xorbitsai/inference.svg?style=flat-square)            |\n| [KubeAI](https://github.com/substratusai/kubeai)                       | Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text. | ![GitHub Badge](https://img.shields.io/github/stars/substratusai/kubeai.svg?style=flat-square)             |\n| [Kaito](https://github.com/kaito-project/kaito)                            | A Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.                                                                 | ![GitHub Badge](https://img.shields.io/github/stars/kaito-project/kaito.svg?style=flat-square)            |\n| [Open Responses](https://docs.julep.ai/open-responses) | Serverless open-source platform for building long-running LLM agents with tool use. | ![GitHub Badge](https://img.shields.io/github/stars/julep-ai/julep.svg?style=flat-square) |\n| [KubeStellar Console](https://github.com/kubestellar/console) | AI-powered multi-cluster Kubernetes dashboard for hybrid edge and cloud. GPU monitoring, LLM inference cluster management, benchmark streaming, and 20+ CNCF integrations. CNCF Sandbox (Apache 2.0). | ![GitHub Badge](https://img.shields.io/github/stars/kubestellar/console.svg?style=flat-square) |\n\n\n**[⬆ back to ToC](#table-of-contents)**\n\n## Security\n\n### Frameworks for LLM security\n\n| Project                                                 | Details                                                                                                                             | Repository                                                                                    |\n| ------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |\n| [Cordum](https://github.com/cordum-io/cordum) | Safety-first agent orchestration platform with pre-dispatch policy evaluation, output scanning (PII, secrets, injection), job scheduling, workflow engine, and full audit trail. | ![GitHub Badge](https://img.shields.io/github/stars/cordum-io/cordum.svg?style=flat-square) |\n| [brood-box](https://github.com/stacklok/brood-box) | CLI tool for running coding agents inside hardware-isolated microVMs with snapshot isolation, egress control, and MCP authorization. | ![GitHub Badge](https://img.shields.io/github/stars/stacklok/brood-box?style=flat-square) |\n| [dstack](https://github.com/Dstack-TEE/dstack)          | Open-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing. | ![GitHub Badge](https://img.shields.io/github/stars/Dstack-TEE/dstack?style=flat-square) |\n| [Plexiglass](https://github.com/kortex-labs/plexiglass) | A Python Machine Learning Pentesting Toolbox for Adversarial Attacks. Works with LLMs, DNNs, and other machine learning algorithms. | ![GitHub Badge](https://img.shields.io/github/stars/kortex-labs/plexiglass?style=flat-square) |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n### Observability\n\n| Project                                                                        | Details                                                                                                                                                                                                                        | Repository                                                                                                       |\n| ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- |\n| [Azure OpenAI Logger](https://github.com/aavetis/azure-openai-logger)          | \"Batteries included\" logging solution for your Azure OpenAI instance.                                                                                                                                                          | ![GitHub Badge](https://img.shields.io/github/stars/aavetis/azure-openai-logger?style=flat-square)               |\n| [ClevAgent](https://clevagent.io)                                              | Runtime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP API.                                                                                                    |                                                                                                                  |\n| [Deepchecks](https://github.com/deepchecks/deepchecks)                         | Tests for Continuous Validation of ML Models \u0026 Data. Deepchecks is a Python package for comprehensively validating your machine learning models and data with minimal effort.                                                  | ![GitHub Badge](https://img.shields.io/github/stars/deepchecks/deepchecks.svg?style=flat-square)                 |\n| [Evidently](https://github.com/evidentlyai/evidently)                          | An open-source framework to evaluate, test and monitor ML and LLM-powered systems.                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/evidentlyai/evidently.svg?style=flat-square)                 |\n| [EvalView](https://github.com/hidai25/eval-view)                              | Regression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.             | ![GitHub Badge](https://img.shields.io/github/stars/hidai25/eval-view.svg?style=flat-square)                     |\n| [Fiddler AI](https://github.com/fiddler-labs/fiddler-auditor)                  | Evaluate, monitor, analyze, and improve machine learning and generative models from pre-production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity. | ![GitHub Badge](https://img.shields.io/github/stars/fiddler-labs/fiddler-auditor.svg?style=flat-square)          |\n| [Giskard](https://github.com/Giskard-AI/giskard)                               | Testing framework dedicated to ML models, from tabular to LLMs. Detect risks of biases, performance issues and errors in 4 lines of code.                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/Giskard-AI/giskard.svg?style=flat-square)\n| [QWED](https://github.com/QWED-AI/qwed-verification) | Deterministic verification protocol for LLM outputs using 8 formal verification engines (SymPy, Z3, AST, SQLGlot). Prevents hallucinations through mathematical proofs rather than statistical methods. | ![GitHub Badge](https://img.shields.io/github/stars/QWED-AI/qwed-verification.svg?style=flat-square) |\n| [Great Expectations](https://github.com/great-expectations/great_expectations) | Always know what to expect from your data.                                                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/great-expectations/great_expectations.svg?style=flat-square) |\n| [Helicone](https://github.com/Helicone/helicone)                              | Open source LLM observability platform. One line of code to monitor, evaluate, and experiment with features like prompt management, agent tracing, and evaluations.                                                            | ![GitHub Badge](https://img.shields.io/github/stars/Helicone/helicone.svg?style=flat-square)                     |\n| [Traceloop OpenLLMetry](https://github.com/traceloop/openllmetry)                              | OpenTelemetry-based observability and monitoring for LLM and agents workflows.                                                           | ![GitHub Badge](https://img.shields.io/github/stars/traceloop/openllmetry.svg?style=flat-square)    \n| [Langfuse 🪢](https://langfuse.com) | Open-source LLM observability platform that helps teams collaboratively debug, analyze, and iterate on their LLM applications. | ![GitHub Badge](https://img.shields.io/github/stars/langfuse/langfuse.svg?style=flat-square)              |\n| [whylogs](https://github.com/whylabs/whylogs)                                  | The open standard for data logging                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/whylabs/whylogs.svg?style=flat-square)                       |\n| [Maxim AI](https://getmaxim.ai) | Platform for AI Agent Simulation, Evaluation \u0026 Observability |\n| [onWatch](https://github.com/onllm-dev/onwatch) | Lightweight Go CLI that tracks AI API quota usage across 7 providers (Anthropic, OpenAI, GitHub Copilot, MiniMax, and more). Background daemon, \u003c50MB RAM, zero telemetry, SQLite storage. | ![GitHub Badge](https://img.shields.io/github/stars/onllm-dev/onwatch.svg?style=flat-square) |\n| [RagTune](https://github.com/metawake/ragtune) | CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer. | ![GitHub Badge](https://img.shields.io/github/stars/metawake/ragtune.svg?style=flat-square) |\n| [traceAI](https://github.com/future-agi/traceAI)                                | Open-source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows.                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/traceAI?style=flat-square)                         |\n| [Future AGI](https://github.com/future-agi/futureagi-sdk)                    | Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.                                             | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/futureagi-sdk?style=flat-square)                   |\n| [semantic-coverage](https://github.com/aashirpersonal/semantic-coverage) | Visualizes RAG knowledge gaps and \"blind spots\" using 2D UMAP clustering and density detection.                                             | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/futureagi-sdk?style=flat-square)                   |\n| [Weco Observe](https://weco.ai) | Observability and debugging tool for AI research agents. Trace multi-step LLM agent runs, visualize decision trees, and identify failure modes in autonomous research workflows. Cloud hosted with open-source agent integration. |                                                                                                                    |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n## LLMOps\n\n| Project                                                            | Details                                                                                                                                                                                                                                                                                                                                                                                                 | Repository                                                                                                |\n| ------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |\n| [agenta](https://github.com/Agenta-AI/agenta)                      | The LLMOps platform to build robust LLM apps. Easily experiment and evaluate different prompts, models, and workflows to build robust apps.                                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/Agenta-AI/agenta.svg?style=flat-square)               |\n| [AgentMark](https://github.com/puzzlet-ai/agentmark)                      | Type-Safe Markdown-based Agents                                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/Puzzlet-ai/agentmark.svg?style=flat-square)               |\n| [AgentField](https://github.com/Agent-Field/agentfield)            | Open-source control plane for building and operating AI agents like APIs at scale, with routing, memory, observability, identity, auth, and policy controls.                                                                                                                        | ![GitHub Badge](https://img.shields.io/github/stars/Agent-Field/agentfield.svg?style=flat-square)             |\n| [AI studio](https://github.com/missingstudio/ai)                   | A Reliable Open Source AI studio to build core infrastructure stack for your LLM Applications. It allows you to gain visibility, make your application reliable, and prepare it for production with features such as caching, rate limiting, exponential retry, model fallback, and more.                                                                                                               | ![GitHub Badge](https://img.shields.io/github/stars/missingstudio/ai.svg?style=flat-square)               |\n| [Arize-Phoenix](https://github.com/Arize-ai/phoenix)               | ML observability for LLMs, vision, language, and tabular models.                                                                                                                                                                                                                                                                                                                                        | ![GitHub Badge](https://img.shields.io/github/stars/Arize-ai/phoenix.svg?style=flat-square)               |\n| [BudgetML](https://github.com/ebhy/budgetml)                       | Deploy a ML inference service on a budget in less than 10 lines of code.                                                                                                                                                                                                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/ebhy/budgetml.svg?style=flat-square)                  |\n| [Cheshire Cat AI](https://github.com/cheshire-cat-ai/core)         | Web framework to create vertical AI agents. FastAPI based, plugin system inspired to WordPress, admin panel, vector DB included                                                                                                                                                                                                                                                                         | ![GitHub Badge](https://img.shields.io/github/stars/cheshire-cat-ai/core.svg?style=flat-square)                  |\n| [Contexto](https://github.com/ekailabs/contexto) | Self-hosted context engine for AI agents with persistent conversation memory and recall. Works as a drop-in OpenAI-compatible proxy, OpenClaw plugin, or memory SDK — no code changes required. | ![GitHub Badge](https://img.shields.io/github/stars/ekailabs/contexto.svg?style=flat-square) |\n| [Dataoorts](https://dataoorts.com/ai)                              | Enjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.                                                                                                                                                                                                                                                                                        |                                                                                                           |\n| [deeplake](https://github.com/activeloopai/deeplake)               | Stream large multimodal datasets to achieve near 100% GPU utilization. Query, visualize, \u0026 version control data. Access data w/o the need to recompute the embeddings for the model finetuning.                                                                                                                                                                                                         | ![GitHub Badge](https://img.shields.io/github/stars/activeloopai/Hub.svg?style=flat-square)               |\n| [Dify](https://github.com/langgenius/dify)                         | Open-source framework aims to enable developers (and even non-developers) to quickly build useful applications based on large language models, ensuring they are visual, operable, and improvable.                                                                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/langgenius/dify.svg?style=flat-square)                |\n| [Dstack](https://github.com/dstackai/dstack)                       | Cost-effective LLM development in any cloud (AWS, GCP, Azure, Lambda, etc).                                                                                                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/dstackai/dstack.svg?style=flat-square)                |\n| [Embedchain](https://github.com/embedchain/embedchain)             | Framework to create ChatGPT like bots over your dataset.                                                                                                                                                                                                                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/embedchain/embedchain.svg?style=flat-square)          |\n| [Epsilla](https://epsilla.com)                                     | An all-in-one platform to create vertical AI agents powered by your private data and knowledge.                                                                                                                                                                                                      |               |\n| [Evidently](https://github.com/evidentlyai/evidently)              | An open-source framework to evaluate, test and monitor ML and LLM-powered systems.                                                                                                                                                                                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/evidentlyai/evidently.svg?style=flat-square)          |\n| [Fiddler AI](https://www.fiddler.ai/llmops)                        | Evaluate, monitor, analyze, and improve MLOps and LLMOps from pre-production to production.                                                                                                                                                                                                                                                                                                             |                                                                                                           |\n| [Glide](https://github.com/EinStack/glide)                         | Cloud-Native LLM Routing Engine. Improve LLM app resilience and speed.                                                                                                                                                                                                                                                                                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/einstack/glide.svg?style=flat-square)                 |\n| [gotoHuman](https://www.gotohuman.com)                             | Bring a **human into the loop** in your LLM-based and agentic workflows. Prompt users to approve actions, select next steps, or review and validate generated results.                                                                                                                                                                                                                                  |\n| [GPTCache](https://github.com/zilliztech/GPTCache)                 | Creating semantic cache to store responses from LLM queries.                                                                                                                                                                                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/zilliztech/GPTCache.svg?style=flat-square)            |\n| [GPUStack](https://github.com/gpustack/gpustack)                   | An open-source GPU cluster manager for running and managing LLMs                                                                                                                                                                                                                                                                                                                                        | ![GitHub Badge](https://img.shields.io/github/stars/gpustack/gpustack.svg?style=flat-square)              |\n| [Haystack](https://github.com/deepset-ai/haystack)                 | Quickly compose applications with LLM Agents, semantic search, question-answering and more.                                                                                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/deepset-ai/haystack.svg?style=flat-square)            |\n| [Hive](https://github.com/aden-hive/hive)                          | Open-source AI agent framework for building goal-driven, self-improving autonomous agents with auto-generated graphs, evolution loops, and MCP integration.                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/aden-hive/hive.svg?style=flat-square)                 |\n| [Helicone](https://github.com/Helicone/helicone)                   | Open-source LLM observability platform for logging, monitoring, and debugging AI applications. Simple 1-line integration to get started.                                                                                                                                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/helicone/helicone.svg?style=flat-square)              |\n| [Humanloop](https://humanloop.com)                                 | The LLM evals platform for enterprises, providing tools to develop, evaluate, and observe AI systems. |                                                                                                |\n| [Hypersigil](https://github.com/hypersigilhq/hypersigil)           | Open-source prompt lifecycle management and gateway with a Web UI.                         | ![GitHub Badge](https://img.shields.io/github/stars/hypersigilhq/hypersigil.svg?style=flat-square)        | \n| [Izlo](https://getizlo.com/)                                       | Prompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.                                                                                                                                                                                                                                                                                              |                                                                                                           |\n| [Keywords AI](https://keywordsai.co/)                              | A unified DevOps platform for AI software. Keywords AI makes it easy for developers to build LLM applications.                                                                                                                                                                                                                                                                                          |                                                                                                           |\n| [MLflow](https://github.com/mlflow/mlflow/tree/master)             | An open-source framework for the end-to-end machine learning lifecycle, helping developers track experiments, evaluate models/prompts, deploy models, and add observability with tracing. | ![GitHub Badge](https://img.shields.io/github/stars/mlflow/mlflow.svg?style=flat-square)  |\n| [Laminar](https://github.com/lmnr-ai/lmnr)                         | Open-source all-in-one platform for engineering AI products. Traces, Evals, Datasets, Labels.                                                                                                                                                                                                                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/lmnr-ai/lmnr.svg?style=flat-square)                   |\n| [langchain](https://github.com/hwchase17/langchain)                | Building applications with LLMs through composability                                                                                                                                                                                                                                                                                                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/hwchase17/langchain.svg?style=flat-square)            |\n| [LangFlow](https://github.com/logspace-ai/langflow)                | An effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface.                                                                                                                                                                                                                                                                                       | ![GitHub Badge](https://img.shields.io/github/stars/logspace-ai/langflow.svg?style=flat-square)           |\n| [Langfuse](https://github.com/langfuse/langfuse)                   | Open Source LLM Engineering Platform: Traces, evals, prompt management and metrics to debug and improve your LLM application.                                                                                                                                                                                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/langfuse/langfuse.svg?style=flat-square)              |\n| [LangKit](https://github.com/whylabs/langkit)                      | Out-of-the-box LLM telemetry collection library that extracts features and profiles prompts, responses and metadata about how your LLM is performing over time to find problems at scale.                                                                                                                                                                                                               | ![GitHub Badge](https://img.shields.io/github/stars/whylabs/langkit.svg?style=flat-square)                |\n| [LangWatch](https://github.com/langwatch/langwatch)                | LLM Ops platform with Analytics, Monitoring, Evaluations and an LLM Optimization Studio powered by DSPy | ![GitHub Badge](https://img.shields.io/github/stars/langwatch/langwatch.svg?style=flat-square) |\n| [LiteLLM 🚅](https://github.com/BerriAI/litellm/)                  | A simple \u0026 light 100 line package to **standardize LLM API calls** across OpenAI, Azure, Cohere, Anthropic, Replicate API Endpoints                                                                                                                                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/BerriAI/litellm.svg?style=flat-square)                |\n| [Literal AI](https://literalai.com/)                               | Multi-modal LLM observability and evaluation platform. Create prompt templates, deploy prompts versions, debug LLM runs, create datasets, run evaluations, monitor LLM metrics and collect human feedback.                                                                                                                                                                                              |                                                                                                           |\n| [LlamaIndex](https://github.com/jerryjliu/llama_index)             | Provides a central interface to connect your LLMs with external data.                                                                                                                                                                                                                                                                                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/jerryjliu/llama_index.svg?style=flat-square)          |\n| [LLMApp](https://github.com/pathwaycom/llm-app)                    | LLM App is a Python library that helps you build real-time LLM-enabled data pipelines with few lines of code.                                                                                                                                                                                                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/pathwaycom/llm-app.svg?style=flat-square)             |\n| [LLMFlows](https://github.com/stoyan-stoyanov/llmflows)            | LLMFlows is a framework for building simple, explicit, and transparent LLM applications such as chatbots, question-answering systems, and agents.                                                                                                                                                                                                                                                       | ![GitHub Badge](https://img.shields.io/github/stars/stoyan-stoyanov/llmflows.svg?style=flat-square)       |\n| [LRM](https://github.com/nickprotop/LocalizationManager)           | CLI/TUI tool for managing localization files (.resx, JSON, Android, iOS) with LLM-powered translation via Ollama, validation, and code scanning for unused/missing keys.                                                                                                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/nickprotop/LocalizationManager.svg?style=flat-square) |\n| [Lunary](https://github.com/lunary-ai/lunary)                      | Observability and prompt management for LLM chabots and agents. Debug agents with powerful tracing and logging. Usage analytics and dive deep into the history of your requests. Developer friendly modules with plug-and-play integration into LangChain.                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/lunary-ai/lunary.svg?style=flat-square)            |\n| [Mengram](https://github.com/alibaizhanov/mengram)                 | Open-source memory infrastructure for AI agents. Provides semantic (entities/facts), episodic (conversations), and procedural (learned behaviors) memory with auto-reflection. Python SDK, JS SDK, MCP server, and REST API.                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/alibaizhanov/mengram.svg?style=flat-square)           |\n| [magentic](https://github.com/jackmpcollins/magentic)              | Seamlessly integrate LLMs as Python functions. Use type annotations to specify structured output. Mix LLM queries and function calling with regular Python code to create complex LLM-powered functionality.                                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/jackmpcollins/magentic.svg?style=flat-square)         |\n| [Manag.ai](https://www.manag.ai)                                   | Your all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.                                                                                                                                                                                                                                                                                     |                                                                                                           |\n| [Mirascope](https://github.com/Mirascope/mirascope)                | Intuitive convenience tooling for lightning-fast, efficient development and ensuring quality in LLM-based applications                                                                                                                                                                                                                                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/Mirascope/mirascope.svg?style=flat-square)            |\n| [Neurolink](https://github.com/juspay/neurolink)                   | Multi-provider AI agent framework that unifies 12+ LLM providers (OpenAI, Google, Anthropic, AWS, Azure, Groq, etc.) with workflow orchestration. Production-grade platform for building LLM applications with streaming, tool calling, caching, and enterprise features. Battle-tested at 15M+ requests/month.                                                        | ![GitHub Badge](https://img.shields.io/github/stars/juspay/neurolink.svg?style=flat-square)               |\n| [OpenLIT](https://github.com/openlit/openlit)                      | OpenLIT is an OpenTelemetry-native GenAI and LLM Application Observability tool and provides OpenTelmetry Auto-instrumentation for monitoring LLMs, VectorDBs and Frameworks. It provides valuable insights into token \u0026 cost usage, user interaction, and performance related metrics.                                                                                                                 | ![GitHub Badge](https://img.shields.io/github/stars/dokulabs/doku.svg?style=flat-square)                  |\n| [Opik](https://github.com/comet-ml/opik)                           | Confidently evaluate, test, and ship LLM applications with a suite of observability tools to calibrate language model outputs across your dev and production lifecycle.                                                                                                                                                                                                                                 | ![GitHub Badge](https://img.shields.io/github/stars/comet-ml/opik.svg?style=flat-square)                  |\n| [Parea AI](https://www.parea.ai/)                                  | Platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground.                                                                                                                                                                                                                                                               | ![GitHub Badge](https://img.shields.io/github/stars/parea-ai/parea-sdk-py?style=flat-square)              |\n| [Pezzo 🕹️](https://github.com/pezzolabs/pezzo)                     | Pezzo is the open-source LLMOps platform built for developers and teams. In just two lines of code, you can seamlessly troubleshoot your AI operations, collaborate and manage your prompts in one place, and instantly deploy changes to any environment.                                                                                                                                              | ![GitHub Badge](https://img.shields.io/github/stars/pezzolabs/pezzo.svg?style=flat-square)                |\n| [PraisonAI](https://github.com/MervinPraison/PraisonAI)            | Production-ready Multi-AI Agents framework with self-reflection. Fastest agent instantiation (3.77μs), 100+ LLM support via LiteLLM, MCP integration, agentic workflows (route/parallel/loop/repeat), built-in memory, Python \u0026 JS SDKs.                                                                                                                                                               | ![GitHub Badge](https://img.shields.io/github/stars/MervinPraison/PraisonAI.svg?style=flat-square)        |\n| [PromptDX](https://github.com/puzzlet-ai/promptdx)                 | A declarative, extensible, and composable approach for developing LLM prompts using Markdown and JSX. | ![GitHub Badge](https://img.shields.io/github/stars/puzzlet-ai/promptdx.svg?style=flat-square) |\n| [PromptHub](https://www.prompthub.us)                              | Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, collaborate, deploy, and monitor, all from one place.                                                                                                                                                                                                                               |                                                                                                           |\n| [promptfoo](https://github.com/typpo/promptfoo)                    | Open-source tool for testing \u0026 evaluating prompt quality. Create test cases, automatically check output quality and catch regressions, and reduce evaluation cost.                                                                                                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/typpo/promptfoo.svg?style=flat-square)                |\n| [PromptFoundry](https://www.promptfoundry.ai)                      | The simple prompt engineering and evaluation tool designed for developers building AI applications.                                                                                                                                                                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/prompt-foundry/python-sdk.svg?style=flat-square)      |\n| [PromptLayer 🍰](https://www.promptlayer.com)                      | Prompt Engineering platform. Collaborate, test, evaluate, and monitor your LLM applications                                                                                                                                                                                                                                                                                                             | ![Github Badge](https://img.shields.io/github/stars/MagnivOrg/prompt-layer-library.svg?style=flat-square) |\n| [PromptMage](https://github.com/tsterbak/promptmage)               | Open-source tool to simplify the process of creating and managing LLM workflows and prompts as a self-hosted solution.                                                                                                                                                                                                                                                                                  | ![GitHub Badge](https://img.shields.io/github/stars/tsterbak/promptmage.svg?style=flat-square)            |\n| [PromptSite](https://github.com/dkuang1980/promptsite)               | A lightweight Python library for prompt lifecycle management that helps you version control, track, experiment and debug with your LLM prompts with ease. Minimal setup, no servers, databases, or API keys required - works directly with your local filesystem, ideal for data scientists and engineers to easily integrate into existing LLM workflows     |                   |\n| [Prompteams](https://www.prompteams.com)                           | Prompt management system. Version, test, collaborate, and retrieve prompts through real-time APIs. Have GitHub style with repos, branches, and commits (and commit history).                                                                                                                                                                                                                            |                                                                                                           |\n| [prompttools](https://github.com/hegelai/prompttools)              | Open-source tools for testing and experimenting with prompts. The core idea is to enable developers to evaluate prompts using familiar interfaces like code and notebooks. In just a few lines of codes, you can test your prompts and parameters across different models (whether you are using OpenAI, Anthropic, or LLaMA models). You can even evaluate the retrieval accuracy of vector databases. | ![GitHub Badge](https://img.shields.io/github/stars/hegelai/prompttools.svg?style=flat-square)            |\n| [Puzzlet AI](https://www.puzzlet.ai)                              | The Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM application - with version control, type-safety, and local development built-in.                                                                                                                                                                                                    |                                                                                                         |\n| [systemprompt.io](https://systemprompt.io)                         | Systemprompt.io is a Rest API with quality tooling to enable the creation, use and observability of prompts in any AI system. Control every detail of your prompt for a SOTA prompt management experience.                                                                                                                                                                                              |                                                                                                           |\n| [TeamoRouter](https://router.teamolab.com)                         | LLM routing gateway for OpenClaw. One API key to access Claude, GPT-4o, Gemini, DeepSeek, Kimi, MiniMax. Smart routing modes (teamo-best, teamo-balanced, teamo-eco) auto-pick the optimal model. Up to 50% off official prices. 2-second install via skill.md.                                                                                                                                        |                                                                                                           |\n| [TreeScale](https://treescale.com)                                 | All In One Dev Platform For LLM Apps. Deploy LLM-enhanced APIs seamlessly using tools for prompt optimization, semantic querying, version management, statistical evaluation, and performance tracking. As a part of the developer friendly API implementation TreeScale offers Elastic LLM product, which makes a unified API Endpoint for all major LLM providers and open source models.             |                                                                                                           |\n| [TrueFoundry](https://www.truefoundry.com/)                        | Deploy LLMOps tools like Vector DBs, Embedding server etc on your own Kubernetes (EKS,AKS,GKE,On-prem) Infra including deploying, Fine-tuning, tracking Prompts and serving Open Source LLM Models with full Data Security and Optimal GPU Management. Train and Launch your LLM Application at Production scale with best Software Engineering practices.                                              |                                                                                                           |\n| [ReliableGPT 💪](https://github.com/BerriAI/reliableGPT/)          | Handle OpenAI Errors (overloaded OpenAI servers, rotated keys, or context window errors) for your production LLM Applications.                                                                                                                                                                                                                                                                          | ![GitHub Badge](https://img.shields.io/github/stars/BerriAI/reliableGPT.svg?style=flat-square)            |\n| [Registry Broker](https://github.com/hashgraph-online/registry-broker) | Universal index and routing layer for AI agents. Aggregates agent metadata from multiple registries (NANDA, MCP, Virtuals, OpenRouter, A2A, X402 Bazaar) across web2 and web3, normalizes profiles, and provides protocol translation between agent ecosystems.                                                                                                                                        | ![GitHub Badge](https://img.shields.io/github/stars/hashgraph-online/registry-broker.svg?style=flat-square) |\n| [Rhesis](https://github.com/rhesis-ai/rhesis)                      | Open-source testing infrastructure for LLM and agentic applications. Collaborative platform enabling teams to define quality metrics, run evaluations, and ship confidently with version control and peer review workflows built for AI engineering.                                                                                                                                                | ![GitHub Badge](https://img.shields.io/github/stars/rhesis-ai/rhesis.svg?style=flat-square)               |\n| [Roundtable](https://github.com/askbudi/roundtable)                | Zero-configuration unified AI assistant management built on the FastMCP framework. Provides seamless integration with Claude, ChatGPT, and other AI assistants through a single MCP interface with session management, logging, and production-ready operations.                                                                                                                                        | ![GitHub Badge](https://img.shields.io/github/stars/askbudi/roundtable.svg?style=flat-square)             |\n| [Portkey](https://portkey.ai/)                                     | Control Panel with an observability suite \u0026 an AI gateway — to ship fast, reliable, and cost-efficient apps.                                                                                                                                                                                                                                                                                            |                                                                                                           |\n| [Semantic Cache Router](https://github.com/redjackfred/distributed-semantic-cache-and-stateful-routing-system) | Distributed semantic cache and stateful routing system that cuts LLM API costs by returning cached responses for semantically similar queries. Uses ANN vector search (cosine ≥ 0.8) and consistent hashing to pin requests to the same worker, achieving ~7× latency reduction on cache hits while scaling horizontally without cache thrash. | ![GitHub Badge](https://img.shields.io/github/stars/redjackfred/distributed-semantic-cache-and-stateful-routing-system.svg?style=flat-square) |\n| [Statewave](https://github.com/smaramwbc/statewave)                | Open-source memory runtime for AI agents. Compiles events into deterministic, provenance-tagged context bundles instead of query-time retrieval. Apache-2.0, self-hostable on Postgres + pgvector.                                                                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/smaramwbc/statewave.svg?style=flat-square)            |\n| [TensorZero](https://www.tensorzero.com/)                          | TensorZero is an open-source framework for building production-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluations, and experimentation.                                                                                                                                                                                                                        | ![GitHub Badge](https://img.shields.io/github/stars/tensorzero/tensorzero.svg?style=flat-square)          |\n| [Vellum](https://www.vellum.ai/)                                   | An AI product development platform to experiment with, evaluate, and deploy advanced LLM apps.                                                                                                                                                                                                                                                                                                          |                                                                                                           |\n| [Weights \u0026 Biases (Prompts)](https://docs.wandb.ai/guides/prompts) | A suite of LLMOps tools within the developer-first W\u0026B MLOps platform. Utilize W\u0026B Prompts for visualizing and inspecting LLM execution flow, tracking inputs and outputs, viewing intermediate results, securely managing prompts and LLM chain configurations.                                                                                                                                        |                                                                                                           |\n| [Wordware](https://www.wordware.ai)                                | A web-hosted IDE where non-technical domain experts work with AI Engineers to build task-specific AI agents. It approaches prompting as a new programming language rather than low/no-code blocks.                                                                                                                                                                                                      |                                                                                                           |\n| [xTuring](https://github.com/stochasticai/xturing)                 | Build and control your personal LLMs with fast and efficient fine-tuning.                                                                                                                                                                                                                                                                                                                               | ![GitHub Badge](https://img.shields.io/github/stars/stochasticai/xturing.svg?style=flat-square)           |\n| [ZenML](https://github.com/zenml-io/zenml)                         | Open-source framework for orchestrating, experimenting and deploying production-grade ML solutions, with built-in `langchain` \u0026 `llama_index` integrations.                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/zenml-io/zenml.svg?style=flat-square)                 |\n| [SwarmClaw](https://github.com/swarmclawai/swarmclaw) | Self-hosted multi-agent AI runtime with 23+ LLM providers, persistent memory, skills, schedules, sub-agent spawning, and MCP client + server support. Ships as desktop app, CLI, or Docker. | ![GitHub Badge](https://img.shields.io/github/stars/swarmclawai/swarmclaw.svg?style=flat-square) |\n| [ai-evaluation](https://github.com/future-agi/ai-evaluation) | Evaluation framework for automated, reproducible scoring of LLM, agent, and workflow performance. | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/ai-evaluation?style=flat-square) |\n| [future-agi](https://github.com/future-agi/future-agi) | Open-source self-hostable end-to-end agent engineering and optimization platform unifying tracing, evals, simulations, datasets, gateway, and guardrails for LLM and AI agent applications. | ![GitHub Badge](https://img.shields.io/github/stars/future-agi/future-agi?style=flat-square) |\n\n\n**[⬆ back to ToC](#table-of-contents)**\n\n## Search\n\n### Hybrid search\n| Project                                                   | Details                                                                                                                                                                                                                                                                                       | Repository                                                                                            |\n| --------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- |\n| [Airweave](https://github.com/airweave-ai/airweave)    | An easy way to turn any app into searchable data for LLMs.                                                                                                                                                                              | ![GitHub Badge](https://img.shields.io/github/stars/airweave-ai/airweave.svg?style=flat-square)    |\n\n\n### Vector search\n\n| Project                                                   | Details                                                                                                                                                                                                                                                                                       | Repository                                                                                            |\n| --------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- |\n| [AquilaDB](https://github.com/Aquila-Network/AquilaDB)    | An easy to use Neural Search Engine. Index latent vectors along with JSON metadata and do efficient k-NN search.                                                                                                                                                                              | ![GitHub Badge](https://img.shields.io/github/stars/Aquila-Network/AquilaDB.svg?style=flat-square)    |\n| [Awadb](https://github.com/awa-ai/awadb)                  | AI Native database for embedding vectors                                                                                                                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/awa-ai/awadb.svg?style=flat-square)               |\n| [Chroma](https://github.com/chroma-core/chroma)           | the open source embedding database                                                                                                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/chroma-core/chroma.svg?style=flat-square)         |\n| [Epsilla](https://github.com/epsilla-cloud/vectordb)      | A 10x faster, cheaper, and better vector database                                                                                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/epsilla-cloud/vectordb.svg?style=flat-square)         |\n| [Infinity](https://github.com/infiniflow/infinity)        | The AI-native database built for LLM applications, providing incredibly fast vector and full-text search                                                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/infiniflow/infinity.svg?style=flat-square)        |\n| [Lancedb](https://github.com/lancedb/lancedb)             | Developer-friendly, serverless vector database for AI applications. Easily add long-term memory to your LLM apps!                                                                                                                                                                             | ![GitHub Badge](https://img.shields.io/github/stars/lancedb/lancedb.svg?style=flat-square)            |\n| [Marqo](https://github.com/marqo-ai/marqo)                | Tensor search for humans.                                                                                                                                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/marqo-ai/marqo.svg?style=flat-square)             |\n| [Milvus](https://github.com/milvus-io/milvus)             | Vector database for scalable similarity search and AI applications.                                                                                                                                                                                                                           | ![GitHub Badge](https://img.shields.io/github/stars/milvus-io/milvus.svg?style=flat-square)           |\n| [Omnigraph](https://github.com/ModernRelay/omnigraph)     | Typed graph database where agents branch and merge like Git. S3-native, Rust, traversal + vector + BM25 in one runtime.                                                                                                                                                      | ![GitHub Badge](https://img.shields.io/github/stars/ModernRelay/omnigraph.svg?style=flat-square)      |\n| [ParadeDB](https://github.com/paradedb/paradedb)          | The transactional alternative to Elasticsearch, built on Postgres.                                                                                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/paradedb/paradedb.svg?style=flat-square)          |\n| [Pinecone](https://www.pinecone.io/)                      | The Pinecone vector database makes it easy to build high-performance vector search applications. Developer-friendly, fully managed, and easily scalable without infrastructure hassles.                                                                                                       |                                                                                                       |\n| [pgvector](https://github.com/pgvector/pgvector)          | Open-source vector similarity search for Postgres.                                                                                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/pgvector/pgvector.svg?style=flat-square)          |\n| [Rivestack](https://rivestack.io)                         | Managed PostgreSQL with pgvector for AI workloads. Built-in SQL editor lets you query your database with natural language (auto-converted to vector embeddings). Free tier includes 2GB storage.                                                                                                                                                                   |                                                                                                       |\n| [VectorChord](https://github.com/tensorchord/VectorChord) | Scalable, fast, and disk-friendly vector search in Postgres, the successor of `pgvecto.rs`.                                                                                                                                                                                                   | ![GitHub Badge](https://img.shields.io/github/stars/tensorchord/VectorChord.svg?style=flat-square)    |\n| [pgvecto.rs](https://github.com/tensorchord/pgvecto.rs)   | Vector database plugin for Postgres, written in Rust, specifically designed for LLM.                                                                                                                                                                                                          | ![GitHub Badge](https://img.shields.io/github/stars/tensorchord/pgvecto.rs.svg?style=flat-square)     |\n| [Qdrant](https://github.com/qdrant/qdrant)                | Vector Search Engine and Database for the next generation of AI applications. Also available in the cloud                                                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/qdrant/qdrant.svg?style=flat-square)              |\n| [txtai](https://github.com/neuml/txtai)                   | Build AI-powered semantic search applications                                                                                                                                                                                                                                                 | ![GitHub Badge](https://img.shields.io/github/stars/neuml/txtai.svg?style=flat-square)                |\n| [Vald](https://github.com/vdaas/vald)                     | A Highly Scalable Distributed Vector Search Engine                                                                                                                                                                                                                                            | ![GitHub Badge](https://img.shields.io/github/stars/vdaas/vald.svg?style=flat-square)                 |\n| [Vearch](https://github.com/vearch/vearch)                | A distributed system for embedding-based vector retrieval                                                                                                                                                                                                                                     | ![GitHub Badge](https://img.shields.io/github/stars/vearch/vearch.svg?style=flat-square)              |\n| [VectorDB](https://github.com/jina-ai/vectordb)           | A Python vector database you just need - no more, no less.                                                                                                                                                                                                                                    | ![GitHub Badge](https://img.shields.io/github/stars/jina-ai/vectordb.svg?style=flat-square)           |\n| [Vellum](https://www.vellum.ai/products/retrieval)        | A managed service for ingesting documents and performing hybrid semantic/keyword search across them. Comes with out-of-box support for OCR, text chunking, embedding model experimentation, metadata filtering, and production-grade APIs.                                                    |                                                                                                       |\n| [Weaviate](https://github.com/semi-technologies/weaviate) | Weaviate is an open source vector search engine that stores both objects and vectors, allowing for combining vector search with structured filtering with the fault-tolerance and scalability of a cloud-native database, all accessible through GraphQL, REST, and various language clients. | ![GitHub Badge](https://img.shields.io/github/stars/semi-technologies/weaviate.svg?style=flat-square) |\n\n**[⬆ back to ToC](#table-of-contents)**\n\n## Code AI\n\n| Project                                             | Details                                                                                                  | Repository                      ","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/tensorchord%2Fawesome-llmops/projects"}