Projects in Awesome Lists tagged with vllm
A curated list of projects in awesome lists tagged with vllm .
https://github.com/metainternal/llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
ai finetuning langchain llama llama2 llm machine-learning python pytorch vllm
Last synced: 13 Sep 2026
https://github.com/meta-llama/llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
ai finetuning langchain llama llama2 llm machine-learning python pytorch vllm
Last synced: 12 Feb 2026
https://github.com/xorbitsai/inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
artificial-intelligence chatglm deployment flan-t5 gemma ggml glm4 inference llama llama3 llamacpp llm machine-learning mistral openai-api pytorch qwen vllm whisper wizardlm
Last synced: 25 Apr 2026
https://github.com/openrlhf/openrlhf
An Easy-to-use, Scalable and High-performance RLHF Framework based on Ray (PPO & GRPO & REINFORCE++ & LoRA & vLLM & RFT)
large-language-models openai-o1 proximal-policy-optimization raylib reinforcement-learning reinforcement-learning-from-human-feedback transformers vllm
Last synced: 08 Mar 2026
https://github.com/OpenRLHF/OpenRLHF
An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT)
large-language-models openai-o1 proximal-policy-optimization raylib reinforcement-learning reinforcement-learning-from-human-feedback transformers vllm
Last synced: 04 Apr 2025
https://github.com/kserve/kserve
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
artificial-intelligence cncf genai hacktoberfest istio k8s knative kserve kubeflow kubernetes llm-inference machine-learning mlops model-interpretability model-serving pytorch service-mesh tensorflow vllm xgboost
Last synced: 15 Jul 2026
https://github.com/gpustack/gpustack
A GPU cluster manager that configures and orchestrates inference engines like vLLM and SGLang for high-performance AI model deployment.
ascend cuda deepseek distributed-inference genai high-performance-inference inference llama llm llm-inference llm-serving maas mindie openai qwen rocm sglang vllm
Last synced: 20 Apr 2026
https://github.com/katanaml/sparrow
Data processing and instruction calling with ML, LLM and Vision LLM
computer-vision gpt huggingface-transformers llm machinelearning nlp-machine-learning rag vllm
Last synced: 12 May 2025
https://github.com/Orchestra-Research/AI-Research-SKILLs
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
ai ai-research claude claude-code claude-skills codex gemini gpt-5 grpo huggingface machine-leanring megatron skills vllm
Last synced: 15 Feb 2026
https://github.com/Orchestra-Research/AI-research-SKILLs
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
ai ai-research claude claude-code claude-skills codex gemini gpt-5 grpo huggingface machine-leanring megatron skills vllm
Last synced: 12 Feb 2026
https://github.com/vllm-project/semantic-router
Intelligent Mixture-of-Models Router for Efficient LLM Inference
ai-gateway bert-classification envoy-ext-proc envoyproxy fine-tuning golang huggingface-candle huggingface-transformers kubernetes llm-tool-call mixture-of-models pii-detection prompt-engineering prompt-guard python rust semantic-router vllm
Last synced: 05 Jan 2026
https://github.com/containers/ramalama
The goal of RamaLama is to make working with AI boring.
ai containers inference-server llamacpp llm podman vllm
Last synced: 04 Feb 2026
https://github.com/bricks-cloud/bricksllm
🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.
ai anthropic api artificial-intelligence azure docker generative-ai golang gpt llm open-source openai postgresql privacy rest-api security self-hosted vllm ycombinator
Last synced: 14 Jan 2026
https://github.com/ruc-datalab/deepanalyze
DeepAnalyze is the first agentic LLM for autonomous data science.
agent agentic agentic-ai ai ai-scientist chatbot chatgpt data data-analysis data-engineering data-science data-visualization database gpt llama llm qwen science structured-data vllm
Last synced: 11 Nov 2025
https://github.com/prometheus-eval/prometheus-eval
Evaluate your LLM's response with Prometheus and GPT4 💯
evaluation gpt4 litellm llm llm-as-a-judge llm-as-evaluator llmops python vllm
Last synced: 06 Apr 2026
https://github.com/bricks-cloud/BricksLLM
🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.
ai anthropic api artificial-intelligence azure docker generative-ai golang gpt llm open-source openai postgresql privacy rest-api security self-hosted vllm ycombinator
Last synced: 09 Apr 2025
https://github.com/substratusai/kubeai
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
ai autoscaler faster-whisper inference-operator k8s kubernetes llm ollama ollama-operator openai-api vllm vllm-operator whisper
Last synced: 15 May 2025
https://github.com/harleyszhang/llm_note
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
cuda-programming kv-cache llm llm-inference transformer-models triton-kernels vllm
Last synced: 23 Aug 2025
https://github.com/vllm-project/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
ascend inference llm llm-serving llmops mlops model-serving transformer vllm
Last synced: 27 Feb 2026
https://github.com/ModelTC/LightCompress
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
awq benchmark deepseek-v3 deployment evaluation internlm2 large-language-models llm mixtral pruning quantization smoothquant token-merging token-pruning token-reduction tool vllm wan
Last synced: 12 Sep 2026
https://github.com/th1nhhdk/local_ai_ocr
An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR AI (running directly on your machine).
ai deepseek-ocr english llm local multilanguage multilingual ocr offline portable vietnamese vllm
Last synced: 25 Apr 2026
https://github.com/pegainfer-project/pegainfer
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
cuda cuda-kernels deepseek gpu inference inference-engine kimi kimi-k2 kv-cache llm llm-inference llm-serving model-serving moe openai-api paged-attention qwen qwen3 rust vllm
Last synced: 14 Sep 2026
https://github.com/verl-project/verl-omni
Multimodal RL training framework for diffusion & omni models
diffusion-models flow-matching grpo multimodal qwen reinforcement-learning rlhf vllm
Last synced: 02 Aug 2026
https://github.com/pgalko/bambooai
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
ai ai-agents anthropic data-analysis data-science docker gemini groq llm mistral ollama openai-api pandas pinecone python vector-database vllm
Last synced: 15 May 2025
https://github.com/apconw/sanic-web
一个轻量级、支持全链路且易于二次开发的大模型应用项目(Large Model Data Assistant) 支持DeepSeek/Qwen2.5等大模型 基于 Dify 、Ollama&Vllm、Sanic 和 Text2SQL 📊 等技术构建的一站式大模型应用开发项目,采用 Vue3、TypeScript 和 Vite 5 打造现代UI。它支持通过 ECharts 📈 实现基于大模型的数据图形化问答,具备处理 CSV 文件 📂 表格问答的能力。同时,能方便对接第三方开源 RAG 系统 检索系统 🌐等,以支持广泛的通用知识问答。
ai bigdata chat chatgpt deepseek-r1 dify echarts large-model-data-assistant llm ollama python qwen rag sanic text2sql vllm vue3
Last synced: 16 May 2025
https://github.com/jakobdylanc/llmcord
Make Discord your LLM frontend ● Supports any OpenAI compatible API (Ollama, LM Studio, vLLM, OpenRouter, xAI, Mistral, Groq and more)
bot chat chatbot discord frontend gpt gpt-4 grok groq llama llama3 llama4 llm mistral ollama oobabooga openai vllm xai
Last synced: 15 May 2025
https://github.com/mostlygeek/llama-swap
Model swapping for llama.cpp (or any local OpenAPI compatible server)
golang llama llamacpp localllama localllm openai openai-api vllm
Last synced: 16 May 2026
https://github.com/ome-projects/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
deepseek k8s kimi-k2 llama llm llm-inference model-as-a-service model-serving multi-node-kubernetes oracle-cloud pd-disaggregation qwen sglang vllm
Last synced: 09 Jun 2026
https://github.com/ModelTC/llmc
[EMNLP 2024 Industry Track] This is the official PyTorch implementation of "LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit".
awq benchmark deployment evaluation internlm2 large-language-models lightllm llama3 llm lvlm mixtral omniquant post-training-quantization pruning quantization quarot smoothquant spinquant tool vllm
Last synced: 23 Apr 2025
https://github.com/runpod-workers/worker-vllm
The Runpod worker template for serving our large language model endpoints. Powered by vLLM.
language-model llm runpod vllm
Last synced: 29 Jul 2026
https://github.com/huawei-csl/kvarn
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
agentic-ai kv-cache llm llm-inference long-context quantization vllm
Last synced: 01 Jul 2026
https://github.com/varunshenoy/super-json-mode
Low latency JSON generation using LLMs ⚡️
huggingface-transformers llm openai vllm
Last synced: 04 Apr 2025
https://github.com/sgl-project/ome
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
deepseek k8s kimi-k2 llama llm llm-inference model-as-a-service model-serving multi-node-kubernetes oracle-cloud pd-disaggregation qwen sglang vllm
Last synced: 17 Mar 2026
https://github.com/microsoft/vidur
A large-scale simulation framework for LLM inference
inference llm simulation transformer vllm
Last synced: 16 May 2025
https://github.com/micytao/vllm-playground
A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.
Last synced: 07 Feb 2026
https://github.com/chtmp223/topicGPT
TopicGPT: A Prompt-Based Framework for Topic Modeling (NAACL'24)
llm nlp openai python topic-modeling vllm
Last synced: 09 May 2025
https://github.com/jasonacox/tinyllm
Setup and run a local LLM and Chatbot using consumer grade hardware.
artificial-intelligence chatbot large-language-models llama-cpp-python llm openai rag retrieval-augmented-generation vllm
Last synced: 06 Sep 2025
https://github.com/pensarai/apex
AI-powered offensive security testing using autonomous agents, directly in your terminal.
agents ai ai-sdk anthropic cybersecurity offensive-security pentesting tui typescript vllm
Last synced: 11 May 2026
https://github.com/jaylfc/taOS
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).
agent-framework ai-agents ai-platform apple-silicon data-sovereignty distributed-computing kv-cache-quantization llm llm-inference local-first local-llm offline-first orange-pi privacy raspberry-pi rockchip-npu self-hosted turboquant vllm
Last synced: 26 Jun 2026
https://github.com/matrixhub-ai/matrixhub
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
artificial-intelligence huggingface kubernetes llm llm-inference mlops model-registry self-hosted sglang vllm
Last synced: 26 Jun 2026
https://github.com/shell-nlp/gpt_server
gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。
asr embedding fastchat function-calling gpt infinity llama llm lmdeploy openai prompt-injection rerank sglang text-moderation tts vllm
Last synced: 28 Feb 2026
https://github.com/MigoXLab/LMeterX
A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结
ai-gateway api-load-test benchmark http-load-testing llm llm-performance locust ocr performance-testing vllm vlm
Last synced: 25 Aug 2026
https://github.com/migoxlab/lmeterx
A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结
ai-gateway api-load-test benchmark http-load-testing llm llm-performance locust ocr performance-testing vllm vlm
Last synced: 15 Jun 2026
https://github.com/waybarrios/vllm-mlx
OpenAI-compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s.
apple-silicon audio-processing computer-vision image-understanding inference llm machine-learning macos mllm mlx multimodal-ai speech-to-text stt text-to-speech tts video-understanding vision-language-model vllm
Last synced: 23 Jan 2026
https://github.com/JackYFL/awesome-VLLMs
This repository collects papers on VLLM applications. We will update new papers irregularly.
application embodied llm mllm reasoning-agent survey vllm vlm vqa
Last synced: 06 Nov 2025
https://github.com/netease-media/grps
Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.
dynamic-batching serving tensorflow tensorrt tensorrt-llm torch triton-inference-server vllm
Last synced: 05 Apr 2025
https://github.com/NetEase-Media/grps
【深度学习模型部署框架】支持tf/torch/trt/trtllm/vllm以及更多nn框架,支持dynamic batching、streaming模式,支持python/c++双语言,可限制,可拓展,高性能。帮助用户快速地将模型部署到线上,并通过http/rpc接口方式提供服务。
dynamic-batching serving tensorflow tensorrt tensorrt-llm torch triton-inference-server vllm
Last synced: 04 Nov 2025
https://github.com/olegshulyakov/llama.ui
A minimal interface for AI Companion that runs entirely in your browser.
ai llamacpp llm llm-ui llm-webui ollama openai-api reactjs tailwindcss typescript vllm webui
Last synced: 07 Oct 2025
https://github.com/yoziru/nextjs-vllm-ui
Fully-featured, beautiful web interface for vLLM - built with NextJS.
ai llm-ui llm-webui nextjs openai-api self-hosted tailwindcss typescript ui vllm vllm-ui webui
Last synced: 13 Oct 2025
https://github.com/inftyai/llmaz
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
huggingface inference inference-platform kubernetes llamacpp llm modelscope ollama sglang text-generation-inference vllm
Last synced: 05 Apr 2025
https://github.com/azrtydxb/Fastllm-proxy
The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O on the request path, cache-affinity routing, RBAC, budgets and a 13-screen UI in the binary.
ai-gateway inference litellm llm llm-gateway load-balancing observability openai-api rbac reverse-proxy rust vllm
Last synced: 14 Sep 2026
https://github.com/Trainy-ai/llm-atc
Fine-tuning and serving LLMs on any cloud
Last synced: 06 May 2025
https://github.com/bike4mind/bike4mind
The open-core AI workbench — notebooks, agents, RAG, voice, and images across any model: OpenAI, Anthropic, Google, xAI, or local via Ollama/vLLM. BSL 1.1, auto-converting to Apache-2.0 on a two-year clock. Your AI keeps running when theirs doesn't.
agents ai ai-agents ai-workbench anthropic llm mcp mongodb multi-model nextjs ollama open-core openai rag self-hosted typescript vllm
Last synced: 27 Aug 2026
https://github.com/em-geeklab/llmone
Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play".
agent ai-server llm llm-inference llm-serving mindie ollama transformer vllm
Last synced: 07 Oct 2025
https://github.com/praetorian-inc/julius
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
ai-security attack-surface capability golang llm-evaluation llm-fingerprinting llm-pentesting llm-security llm-tools ollama reconnaissance security-scanner service-detection vllm
Last synced: 04 Jun 2026
https://github.com/llmariner/llmariner
Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.
ai autoscaling fine-tuning gpu inference k8s kubernetes llm ml multi-cluster openai operator vllm
Last synced: 11 Apr 2025
https://github.com/call518/logsentinelai
LLM-powered security log analyzer: detect threats & anomalies with zero regex — just declare a Pydantic schema. Real-time Telegram alerts, SIEM-ready with Elasticsearch/Kibana. Supports OpenAI, Ollama, vLLM.
ai ai-log-analysis anomaly-detection cybersecurity devsecops elasticsearch kibana llm monitoring ollama openai pydantic realtime security severity-analysis siem structured-output telegram threat-detection vllm
Last synced: 27 Aug 2026
https://github.com/intelligentnode/IntelliChat
Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.
anthropic azure-openai bard chatbot chatbot-ui chatgpt claude gemini gpt-4o llama2 llm mistral nextjs o3 openai opengpt replicate typescript vercel vllm
Last synced: 17 Apr 2025
https://github.com/intelligentnode/intellichat
Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.
anthropic azure-openai bard chatbot chatbot-ui chatgpt claude gemini gpt-4o llama2 llm mistral nextjs o3 openai opengpt replicate typescript vercel vllm
Last synced: 09 Apr 2025
https://github.com/defilantech/llmkube
Kubernetes operator for local LLM inference with llama.cpp, vLLM, and TGI - multi-GPU, autoscaling, air-gapped, production-ready
ai ai-infrastructure apple-silicon autoscaling edge-computing gguf gpu homelab inference kubernetes kubernetes-operator llama-cpp llm local-llm metal mlops multi-gpu nvidia self-hosted vllm
Last synced: 13 Jun 2026
https://github.com/automatika-robotics/embodied-agents
EmbodiedAgents is a fully-loaded ROS2 based framework for creating interactive physical agents that can understand, remember, and act upon contextual information from their environment.
deeplearning embodied-agent embodied-ai generative-ai godel-machine llm machine-learning multimodal ollama physical-ai roboml robotics ros2 vllm
Last synced: 18 Jan 2026
https://github.com/modal-labs/stopwatch
A tool for benchmarking LLMs on Modal
llms machine-learning sglang tensorrt-llm vllm
Last synced: 31 Jan 2026
https://github.com/niklasfrick/spark-dashboard
Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.
ai ai-monitoring dashboard dgx gpu gpu-monitoring llm llm-monitoring metrics nvidia observability spark vllm
Last synced: 30 May 2026
https://github.com/hadihonarvar/flock
Self-hosted LLM gateway. One Go binary turns your Macs and Linux boxes into a private inference cluster — multi-machine routing, sharding via llama.cpp-RPC, per-user keys + quotas + audit, OpenAI- and Anthropic-compatible APIs behind one endpoint. Point Cursor / Claude Code / Aider / SDKs at it.
ai-gateway aider anthropic claude-code cursor gguf golang inference llama-cpp llm local-llm mlx multi-tenant ollama openai-compatible opentelemetry prometheus self-hosted sharded-inference vllm
Last synced: 11 Jun 2026
https://github.com/generative-computing/granite-switch
Granite Switch — Build AI models like you build software
generative-computing llm-inference llms transformers vllm
Last synced: 28 May 2026
https://github.com/cpfiffer/self-expansion
Code for building self-expanding knowledge graphs with Outlines, vLLM, neo4j, and Modal.
ai graph-database knowledge-graph language-model modal vllm
Last synced: 08 Mar 2026
https://github.com/aws-samples/easy-model-deployer
A user-friendly Command-line/SDK tool that makes it quickly and easier to deploy open-source LLMs on AWS
comfyui-workflow deepseek deepseek-r1 ec2 ecs gemma3 huggingface inferentia-2 internlm2 langchain large-language-model ollama openai-compatible-api qwen2-5 qwq qwq-32b sagemaker vllm
Last synced: 13 Apr 2025
https://github.com/Netis/heron
Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required.
agentic-ai ai-agent-development ai-observability libpcap litellm llm-monitoring llm-observability llmops observability ollama packet-capture rust sglang vllm
Last synced: 11 Jul 2026
https://github.com/opencsgs/llm-inference
llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more.
deepspeed llama-cpp llm-inference ray transformer vllm
Last synced: 12 Apr 2025
https://github.com/argonne-lcf/llm-inference-bench
LLM-Inference-Bench
benchmark deepspeed inference llamacpp llm tensorrt-llm vllm
Last synced: 10 Apr 2025
https://github.com/phospho-app/fastassert
Dockerized LLM inference server with constrained output (JSON mode), built on top of vLLM and outlines. Faster, cheaper and without rate limits. Compare the quality and latency to your current LLM API provider.
docker llm llm-inference outlines vllm
Last synced: 30 Jul 2025
https://github.com/jagmarques/nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
attention compression e8-lattice inference kv-cache llama llm llm-inference long-context memory-efficient mistral pytorch quantization token-eviction transformers vector-quantization vllm
Last synced: 26 Jul 2026
https://github.com/outsourc-e/bench-loop
Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.
agent benchmark cli evaluation llm local-llm mlx ollama vllm
Last synced: 01 Jun 2026
https://github.com/theoddden/terradev
An imperative command-line-interface for AI workload orchestration
agentic-ai agentic-workflow cloud-gpu disaggregated-inference distributed-inference gpu-cluster gpu-provisioning huggingface kubernetes llm-inference mcp-server mixture-of-experts mlops multi-cloud ollama ray sglang vllm
Last synced: 28 Aug 2026
https://github.com/cloudpilot-ai/hermes
Policy-driven seamless lazy loading
hermes k8s kubernetes lazy-loading soci vllm
Last synced: 03 Jun 2026
https://github.com/VectorArc/avp-python
Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text
ai-agents inference kv-cache llm machine-learning multi-agent protocol python transformers vllm
Last synced: 04 Sep 2026
https://github.com/france-travail/happy_vllm
A REST API for vLLM, production ready
api-rest llm llm-serving production vllm
Last synced: 28 Apr 2025
https://github.com/vam876/localapi.ai
LocalAPI.AI is a local AI management tool for Ollama, offering Web UI management and compatibility with vLLM, LM Studio, llama.cpp, Mozilla-Llamafile, Jan Al, Cortex API, Local-LLM, LiteLLM, GPT4All, and more.
llamacpp lm-studio local-llm localai ollama ollama-api ollama-app ollama-gui ollama-ui ollama-webui ollama-webui-management vllm
Last synced: 01 Apr 2025
https://github.com/zrzrzrzrzrzrzr/lm-fly
大模型推理框架加速,让 LLM 飞起来
llm llm-inference mlx openvino tensorrt-llm tgi vllm
Last synced: 08 Aug 2025
https://github.com/xencon/aixcl
Agentic SDLC with local LLM's.
agentic-workflow ai-development development-platform idp-platform llamacpp llm local-first ollama opernsource self-hosted sldc vllm workflow-automation
Last synced: 03 May 2026
https://github.com/vectorarc/avp-python
Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text
ai-agents inference kv-cache llm machine-learning multi-agent protocol python transformers vllm
Last synced: 26 Apr 2026
https://github.com/aws-solutions-library-samples/guidance-for-scalable-model-inference-and-agentic-ai-on-amazon-eks
Comprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP
agentic-ai agentic-workflow huggingface inference inference-engine langfuse litellm-ai-gateway mcp-client mcp-server opensource-ai vllm
Last synced: 14 Oct 2025
https://github.com/paralleliq/piqc
Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.
ai-infrastructure cloud-native gpu introspection kubernetes llm llm-inference mlops model-discovery model-serving modelspec vllm
Last synced: 28 Aug 2026
https://github.com/hcd233/aris-ai-model-server
An OpenAI Compatible API which integrates LLM, Embedding and Reranker. 一个集成 LLM、Embedding 和 Reranker 的 OpenAI 兼容 API
ai awq embedding fastapi gptq llm mlx openai-compatible-api rag reranker sentence-transformers vllm
Last synced: 08 May 2025
https://github.com/slb350/open-agent-sdk-rust
Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and comprehensive examples.
llamacpp llm localllama localllm ollama openai rust rust-crate rust-library vllm
Last synced: 30 Apr 2026
https://github.com/ld-singh/ai-factory-ops-lab
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
ai-infrastructure dcgm devops gpu gpu-scheduling gpu-sharing grafana hami hpc inference k3s kai-scheduler kubernetes learning-by-doing mlops nvidia observability prometheus slurm vllm
Last synced: 01 Sep 2026
https://github.com/lamalab-org/macbench
Probing the limitations of multimodal language models for chemistry and materials research
benchmark chemistry llm materials mutlimodal vllm vlm
Last synced: 29 Jul 2025
https://github.com/0xnyn/periscope
LLM Performance Testing | K6 + Grafana + InfluxDB | A tiny toolkit for load testing and benchmarking OpenAI-like inference endpoints using K6 + Grafana + InfluxDB
grafana influxdb k6 llm-testing performance-testing vllm vllm-production-stack
Last synced: 27 Jul 2026