An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with vllm

A curated list of projects in awesome lists tagged with vllm .

https://github.com/metainternal/llama-cookbook

Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services

ai finetuning langchain llama llama2 llm machine-learning python pytorch vllm

Last synced: 13 Sep 2026

https://github.com/meta-llama/llama-cookbook

Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services

ai finetuning langchain llama llama2 llm machine-learning python pytorch vllm

Last synced: 12 Feb 2026

https://github.com/xorbitsai/inference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

artificial-intelligence chatglm deployment flan-t5 gemma ggml glm4 inference llama llama3 llamacpp llm machine-learning mistral openai-api pytorch qwen vllm whisper wizardlm

Last synced: 25 Apr 2026

https://github.com/lmcache/lmcache

Supercharge Your LLM with the Fastest KV Cache Layer

amd cuda fast inference kv-cache llm pytorch rocm speed vllm

Last synced: 03 Jul 2026

https://github.com/openrlhf/openrlhf

An Easy-to-use, Scalable and High-performance RLHF Framework based on Ray (PPO & GRPO & REINFORCE++ & LoRA & vLLM & RFT)

large-language-models openai-o1 proximal-policy-optimization raylib reinforcement-learning reinforcement-learning-from-human-feedback transformers vllm

Last synced: 08 Mar 2026

https://github.com/OpenRLHF/OpenRLHF

An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT)

large-language-models openai-o1 proximal-policy-optimization raylib reinforcement-learning reinforcement-learning-from-human-feedback transformers vllm

Last synced: 04 Apr 2025

https://github.com/kserve/kserve

Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

artificial-intelligence cncf genai hacktoberfest istio k8s knative kserve kubeflow kubernetes llm-inference machine-learning mlops model-interpretability model-serving pytorch service-mesh tensorflow vllm xgboost

Last synced: 15 Jul 2026

https://github.com/gpustack/gpustack

A GPU cluster manager that configures and orchestrates inference engines like vLLM and SGLang for high-performance AI model deployment.

ascend cuda deepseek distributed-inference genai high-performance-inference inference llama llm llm-inference llm-serving maas mindie openai qwen rocm sglang vllm

Last synced: 20 Apr 2026

https://github.com/katanaml/sparrow

Data processing and instruction calling with ML, LLM and Vision LLM

computer-vision gpt huggingface-transformers llm machinelearning nlp-machine-learning rag vllm

Last synced: 12 May 2025

https://github.com/Orchestra-Research/AI-Research-SKILLs

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.

ai ai-research claude claude-code claude-skills codex gemini gpt-5 grpo huggingface machine-leanring megatron skills vllm

Last synced: 15 Feb 2026

https://github.com/Orchestra-Research/AI-research-SKILLs

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.

ai ai-research claude claude-code claude-skills codex gemini gpt-5 grpo huggingface machine-leanring megatron skills vllm

Last synced: 12 Feb 2026

https://github.com/containers/ramalama

The goal of RamaLama is to make working with AI boring.

ai containers inference-server llamacpp llm podman vllm

Last synced: 04 Feb 2026

https://github.com/sybil-solutions/local-studio

Control panel for VLLM, Sglang, llama.cpp, exllamav3

ai exllama hosting llamacpp local local-ai self sglang vllm

Last synced: 01 Jul 2026

https://github.com/bricks-cloud/bricksllm

🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.

ai anthropic api artificial-intelligence azure docker generative-ai golang gpt llm open-source openai postgresql privacy rest-api security self-hosted vllm ycombinator

Last synced: 14 Jan 2026

https://github.com/sybil-solutions/vllm-studio

Control panel for VLLM, Sglang, llama.cpp, exllamav3

ai exllama hosting llamacpp local local-ai self sglang vllm

Last synced: 07 Jun 2026

https://github.com/prometheus-eval/prometheus-eval

Evaluate your LLM's response with Prometheus and GPT4 💯

evaluation gpt4 litellm llm llm-as-a-judge llm-as-evaluator llmops python vllm

Last synced: 06 Apr 2026

https://github.com/bricks-cloud/BricksLLM

🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.

ai anthropic api artificial-intelligence azure docker generative-ai golang gpt llm open-source openai postgresql privacy rest-api security self-hosted vllm ycombinator

Last synced: 09 Apr 2025

https://github.com/substratusai/kubeai

AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.

ai autoscaler faster-whisper inference-operator k8s kubernetes llm ollama ollama-operator openai-api vllm vllm-operator whisper

Last synced: 15 May 2025

https://github.com/harleyszhang/llm_note

LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.

cuda-programming kv-cache llm llm-inference transformer-models triton-kernels vllm

Last synced: 23 Aug 2025

https://github.com/vllm-project/vllm-ascend

Community maintained hardware plugin for vLLM on Ascend

ascend inference llm llm-serving llmops mlops model-serving transformer vllm

Last synced: 27 Feb 2026

https://github.com/ModelTC/LightCompress

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

awq benchmark deepseek-v3 deployment evaluation internlm2 large-language-models llm mixtral pruning quantization smoothquant token-merging token-pruning token-reduction tool vllm wan

Last synced: 12 Sep 2026

https://github.com/th1nhhdk/local_ai_ocr

An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR AI (running directly on your machine).

ai deepseek-ocr english llm local multilanguage multilingual ocr offline portable vietnamese vllm

Last synced: 25 Apr 2026

https://github.com/pegainfer-project/pegainfer

Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

cuda cuda-kernels deepseek gpu inference inference-engine kimi kimi-k2 kv-cache llm llm-inference llm-serving model-serving moe openai-api paged-attention qwen qwen3 rust vllm

Last synced: 14 Sep 2026

https://github.com/verl-project/verl-omni

Multimodal RL training framework for diffusion & omni models

diffusion-models flow-matching grpo multimodal qwen reinforcement-learning rlhf vllm

Last synced: 02 Aug 2026

https://github.com/pgalko/bambooai

A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.

ai ai-agents anthropic data-analysis data-science docker gemini groq llm mistral ollama openai-api pandas pinecone python vector-database vllm

Last synced: 15 May 2025

https://github.com/apconw/sanic-web

一个轻量级、支持全链路且易于二次开发的大模型应用项目(Large Model Data Assistant) 支持DeepSeek/Qwen2.5等大模型 基于 Dify 、Ollama&Vllm、Sanic 和 Text2SQL 📊 等技术构建的一站式大模型应用开发项目,采用 Vue3、TypeScript 和 Vite 5 打造现代UI。它支持通过 ECharts 📈 实现基于大模型的数据图形化问答,具备处理 CSV 文件 📂 表格问答的能力。同时,能方便对接第三方开源 RAG 系统 检索系统 🌐等,以支持广泛的通用知识问答。

ai bigdata chat chatgpt deepseek-r1 dify echarts large-model-data-assistant llm ollama python qwen rag sanic text2sql vllm vue3

Last synced: 16 May 2025

https://github.com/jakobdylanc/llmcord

Make Discord your LLM frontend ● Supports any OpenAI compatible API (Ollama, LM Studio, vLLM, OpenRouter, xAI, Mistral, Groq and more)

bot chat chatbot discord frontend gpt gpt-4 grok groq llama llama3 llama4 llm mistral ollama oobabooga openai vllm xai

Last synced: 15 May 2025

https://github.com/mostlygeek/llama-swap

Model swapping for llama.cpp (or any local OpenAPI compatible server)

golang llama llamacpp localllama localllm openai openai-api vllm

Last synced: 16 May 2026

https://github.com/ome-projects/ome

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

deepseek k8s kimi-k2 llama llm llm-inference model-as-a-service model-serving multi-node-kubernetes oracle-cloud pd-disaggregation qwen sglang vllm

Last synced: 09 Jun 2026

https://github.com/ModelTC/llmc

[EMNLP 2024 Industry Track] This is the official PyTorch implementation of "LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit".

awq benchmark deployment evaluation internlm2 large-language-models lightllm llama3 llm lvlm mixtral omniquant post-training-quantization pruning quantization quarot smoothquant spinquant tool vllm

Last synced: 23 Apr 2025

https://github.com/runpod-workers/worker-vllm

The Runpod worker template for serving our large language model endpoints. Powered by vLLM.

language-model llm runpod vllm

Last synced: 29 Jul 2026

https://github.com/huawei-csl/kvarn

KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

agentic-ai kv-cache llm llm-inference long-context quantization vllm

Last synced: 01 Jul 2026

https://github.com/varunshenoy/super-json-mode

Low latency JSON generation using LLMs ⚡️

huggingface-transformers llm openai vllm

Last synced: 04 Apr 2025

https://github.com/sgl-project/ome

Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

deepseek k8s kimi-k2 llama llm llm-inference model-as-a-service model-serving multi-node-kubernetes oracle-cloud pd-disaggregation qwen sglang vllm

Last synced: 17 Mar 2026

https://github.com/microsoft/vidur

A large-scale simulation framework for LLM inference

inference llm simulation transformer vllm

Last synced: 16 May 2025

https://github.com/micytao/vllm-playground

A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.

ai learning llms vllm

Last synced: 07 Feb 2026

https://github.com/chtmp223/topicGPT

TopicGPT: A Prompt-Based Framework for Topic Modeling (NAACL'24)

llm nlp openai python topic-modeling vllm

Last synced: 09 May 2025

https://github.com/jasonacox/tinyllm

Setup and run a local LLM and Chatbot using consumer grade hardware.

artificial-intelligence chatbot large-language-models llama-cpp-python llm openai rag retrieval-augmented-generation vllm

Last synced: 06 Sep 2025

https://github.com/pensarai/apex

AI-powered offensive security testing using autonomous agents, directly in your terminal.

agents ai ai-sdk anthropic cybersecurity offensive-security pentesting tui typescript vllm

Last synced: 11 May 2026

https://github.com/jaylfc/taOS

Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).

agent-framework ai-agents ai-platform apple-silicon data-sovereignty distributed-computing kv-cache-quantization llm llm-inference local-first local-llm offline-first orange-pi privacy raspberry-pi rockchip-npu self-hosted turboquant vllm

Last synced: 26 Jun 2026

https://github.com/matrixhub-ai/matrixhub

An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.

artificial-intelligence huggingface kubernetes llm llm-inference mlops model-registry self-hosted sglang vllm

Last synced: 26 Jun 2026

https://github.com/shell-nlp/gpt_server

gpt_server是一个用于生产级部署LLMs、Embedding、Reranker、ASR、TTS、文生图、图片编辑和文生视频的开源框架。

asr embedding fastchat function-calling gpt infinity llama llm lmdeploy openai prompt-injection rerank sglang text-moderation tts vllm

Last synced: 28 Feb 2026

https://github.com/guoqingbao/xinfer

Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.

agent llm qwen rust vllm

Last synced: 01 Jun 2026

https://github.com/MigoXLab/LMeterX

A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结

ai-gateway api-load-test benchmark http-load-testing llm llm-performance locust ocr performance-testing vllm vlm

Last synced: 25 Aug 2026

https://github.com/migoxlab/lmeterx

A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. 一站式通用 API 压测平台,支持大模型推理与业务 HTTP 接口,一键完成性能测试、结果对比与 AI 智能分析总结

ai-gateway api-load-test benchmark http-load-testing llm llm-performance locust ocr performance-testing vllm vlm

Last synced: 15 Jun 2026

https://github.com/waybarrios/vllm-mlx

OpenAI-compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s.

apple-silicon audio-processing computer-vision image-understanding inference llm machine-learning macos mllm mlx multimodal-ai speech-to-text stt text-to-speech tts video-understanding vision-language-model vllm

Last synced: 23 Jan 2026

https://github.com/JackYFL/awesome-VLLMs

This repository collects papers on VLLM applications. We will update new papers irregularly.

application embodied llm mllm reasoning-agent survey vllm vlm vqa

Last synced: 06 Nov 2025

https://github.com/netease-media/grps

Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.

dynamic-batching serving tensorflow tensorrt tensorrt-llm torch triton-inference-server vllm

Last synced: 05 Apr 2025

https://github.com/NetEase-Media/grps

【深度学习模型部署框架】支持tf/torch/trt/trtllm/vllm以及更多nn框架,支持dynamic batching、streaming模式,支持python/c++双语言,可限制,可拓展,高性能。帮助用户快速地将模型部署到线上,并通过http/rpc接口方式提供服务。

dynamic-batching serving tensorflow tensorrt tensorrt-llm torch triton-inference-server vllm

Last synced: 04 Nov 2025

https://github.com/mixa3607/ml-gfx906

ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60

amd-gpu comfyui llamacpp torch vllm

Last synced: 15 Apr 2026

https://github.com/olegshulyakov/llama.ui

A minimal interface for AI Companion that runs entirely in your browser.

ai llamacpp llm llm-ui llm-webui ollama openai-api reactjs tailwindcss typescript vllm webui

Last synced: 07 Oct 2025

https://github.com/gotzmann/booster

Booster - open accelerator for LLM models. Better inference and debugging for AI hackers

chatgpt exllama ggml gpt llama llama-cpp llamacpp llm ollama oobabooga openai vllm

Last synced: 11 Jun 2025

https://github.com/yoziru/nextjs-vllm-ui

Fully-featured, beautiful web interface for vLLM - built with NextJS.

ai llm-ui llm-webui nextjs openai-api self-hosted tailwindcss typescript ui vllm vllm-ui webui

Last synced: 13 Oct 2025

https://github.com/inftyai/llmaz

☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!

huggingface inference inference-platform kubernetes llamacpp llm modelscope ollama sglang text-generation-inference vllm

Last synced: 05 Apr 2025

https://github.com/azrtydxb/Fastllm-proxy

The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O on the request path, cache-affinity routing, RBAC, budgets and a 13-screen UI in the binary.

ai-gateway inference litellm llm llm-gateway load-balancing observability openai-api rbac reverse-proxy rust vllm

Last synced: 14 Sep 2026

https://github.com/Trainy-ai/llm-atc

Fine-tuning and serving LLMs on any cloud

finetuning llama2 llms vllm

Last synced: 06 May 2025

https://github.com/bike4mind/bike4mind

The open-core AI workbench — notebooks, agents, RAG, voice, and images across any model: OpenAI, Anthropic, Google, xAI, or local via Ollama/vLLM. BSL 1.1, auto-converting to Apache-2.0 on a two-year clock. Your AI keeps running when theirs doesn't.

agents ai ai-agents ai-workbench anthropic llm mcp mongodb multi-model nextjs ollama open-core openai rag self-hosted typescript vllm

Last synced: 27 Aug 2026

https://github.com/lightseekorg/torchspec

A PyTorch native library for training speculative decoding models

eagle3 fsdp lightseek llm mooncake pytorch sglang vllm

Last synced: 24 Apr 2026

https://github.com/em-geeklab/llmone

Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play".

agent ai-server llm llm-inference llm-serving mindie ollama transformer vllm

Last synced: 07 Oct 2025

https://github.com/praetorian-inc/julius

Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds

ai-security attack-surface capability golang llm-evaluation llm-fingerprinting llm-pentesting llm-security llm-tools ollama reconnaissance security-scanner service-detection vllm

Last synced: 04 Jun 2026

https://github.com/llmariner/llmariner

Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.

ai autoscaling fine-tuning gpu inference k8s kubernetes llm ml multi-cluster openai operator vllm

Last synced: 11 Apr 2025

https://github.com/call518/logsentinelai

LLM-powered security log analyzer: detect threats & anomalies with zero regex — just declare a Pydantic schema. Real-time Telegram alerts, SIEM-ready with Elasticsearch/Kibana. Supports OpenAI, Ollama, vLLM.

ai ai-log-analysis anomaly-detection cybersecurity devsecops elasticsearch kibana llm monitoring ollama openai pydantic realtime security severity-analysis siem structured-output telegram threat-detection vllm

Last synced: 27 Aug 2026

https://github.com/intelligentnode/IntelliChat

Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.

anthropic azure-openai bard chatbot chatbot-ui chatgpt claude gemini gpt-4o llama2 llm mistral nextjs o3 openai opengpt replicate typescript vercel vllm

Last synced: 17 Apr 2025

https://github.com/intelligentnode/intellichat

Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.

anthropic azure-openai bard chatbot chatbot-ui chatgpt claude gemini gpt-4o llama2 llm mistral nextjs o3 openai opengpt replicate typescript vercel vllm

Last synced: 09 Apr 2025

https://github.com/defilantech/llmkube

Kubernetes operator for local LLM inference with llama.cpp, vLLM, and TGI - multi-GPU, autoscaling, air-gapped, production-ready

ai ai-infrastructure apple-silicon autoscaling edge-computing gguf gpu homelab inference kubernetes kubernetes-operator llama-cpp llm local-llm metal mlops multi-gpu nvidia self-hosted vllm

Last synced: 13 Jun 2026

https://github.com/automatika-robotics/embodied-agents

EmbodiedAgents is a fully-loaded ROS2 based framework for creating interactive physical agents that can understand, remember, and act upon contextual information from their environment.

deeplearning embodied-agent embodied-ai generative-ai godel-machine llm machine-learning multimodal ollama physical-ai roboml robotics ros2 vllm

Last synced: 18 Jan 2026

https://github.com/modal-labs/stopwatch

A tool for benchmarking LLMs on Modal

llms machine-learning sglang tensorrt-llm vllm

Last synced: 31 Jan 2026

https://github.com/niklasfrick/spark-dashboard

Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.

ai ai-monitoring dashboard dgx gpu gpu-monitoring llm llm-monitoring metrics nvidia observability spark vllm

Last synced: 30 May 2026

https://github.com/hadihonarvar/flock

Self-hosted LLM gateway. One Go binary turns your Macs and Linux boxes into a private inference cluster — multi-machine routing, sharding via llama.cpp-RPC, per-user keys + quotas + audit, OpenAI- and Anthropic-compatible APIs behind one endpoint. Point Cursor / Claude Code / Aider / SDKs at it.

ai-gateway aider anthropic claude-code cursor gguf golang inference llama-cpp llm local-llm mlx multi-tenant ollama openai-compatible opentelemetry prometheus self-hosted sharded-inference vllm

Last synced: 11 Jun 2026

https://github.com/generative-computing/granite-switch

Granite Switch — Build AI models like you build software

generative-computing llm-inference llms transformers vllm

Last synced: 28 May 2026

https://github.com/cpfiffer/self-expansion

Code for building self-expanding knowledge graphs with Outlines, vLLM, neo4j, and Modal.

ai graph-database knowledge-graph language-model modal vllm

Last synced: 08 Mar 2026

https://github.com/aws-samples/easy-model-deployer

A user-friendly Command-line/SDK tool that makes it quickly and easier to deploy open-source LLMs on AWS

comfyui-workflow deepseek deepseek-r1 ec2 ecs gemma3 huggingface inferentia-2 internlm2 langchain large-language-model ollama openai-compatible-api qwen2-5 qwq qwq-32b sagemaker vllm

Last synced: 13 Apr 2025

https://github.com/Netis/heron

Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required.

agentic-ai ai-agent-development ai-observability libpcap litellm llm-monitoring llm-observability llmops observability ollama packet-capture rust sglang vllm

Last synced: 11 Jul 2026

https://github.com/opencsgs/llm-inference

llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more.

deepspeed llama-cpp llm-inference ray transformer vllm

Last synced: 12 Apr 2025

https://github.com/gameofdimension/vllm-cn

演示 vllm 对中文大语言模型的神奇效果

vllm

Last synced: 20 Jun 2025

https://github.com/phospho-app/fastassert

Dockerized LLM inference server with constrained output (JSON mode), built on top of vLLM and outlines. Faster, cheaper and without rate limits. Compare the quality and latency to your current LLM API provider.

docker llm llm-inference outlines vllm

Last synced: 30 Jul 2025

https://github.com/jagmarques/nexusquant

Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.

attention compression e8-lattice inference kv-cache llama llm llm-inference long-context memory-efficient mistral pytorch quantization token-eviction transformers vector-quantization vllm

Last synced: 26 Jul 2026

https://github.com/outsourc-e/bench-loop

Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.

agent benchmark cli evaluation llm local-llm mlx ollama vllm

Last synced: 01 Jun 2026

https://github.com/cloudpilot-ai/hermes

Policy-driven seamless lazy loading

hermes k8s kubernetes lazy-loading soci vllm

Last synced: 03 Jun 2026

https://github.com/VectorArc/avp-python

Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text

ai-agents inference kv-cache llm machine-learning multi-agent protocol python transformers vllm

Last synced: 04 Sep 2026

https://github.com/lework/llm-benchmark

LLM 并发性能测试工具,支持自动化压力测试和性能报告生成。

benchmark ollama vllm

Last synced: 13 Apr 2025

https://github.com/france-travail/happy_vllm

A REST API for vLLM, production ready

api-rest llm llm-serving production vllm

Last synced: 28 Apr 2025

https://github.com/vam876/localapi.ai

LocalAPI.AI is a local AI management tool for Ollama, offering Web UI management and compatibility with vLLM, LM Studio, llama.cpp, Mozilla-Llamafile, Jan Al, Cortex API, Local-LLM, LiteLLM, GPT4All, and more.

llamacpp lm-studio local-llm localai ollama ollama-api ollama-app ollama-gui ollama-ui ollama-webui ollama-webui-management vllm

Last synced: 01 Apr 2025

https://github.com/zrzrzrzrzrzrzr/lm-fly

大模型推理框架加速,让 LLM 飞起来

llm llm-inference mlx openvino tensorrt-llm tgi vllm

Last synced: 08 Aug 2025

https://github.com/vectorarc/avp-python

Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text

ai-agents inference kv-cache llm machine-learning multi-agent protocol python transformers vllm

Last synced: 26 Apr 2026

https://github.com/aws-solutions-library-samples/guidance-for-scalable-model-inference-and-agentic-ai-on-amazon-eks

Comprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP

agentic-ai agentic-workflow huggingface inference inference-engine langfuse litellm-ai-gateway mcp-client mcp-server opensource-ai vllm

Last synced: 14 Oct 2025

https://github.com/paralleliq/piqc

Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.

ai-infrastructure cloud-native gpu introspection kubernetes llm llm-inference mlops model-discovery model-serving modelspec vllm

Last synced: 28 Aug 2026

https://github.com/hcd233/aris-ai-model-server

An OpenAI Compatible API which integrates LLM, Embedding and Reranker. 一个集成 LLM、Embedding 和 Reranker 的 OpenAI 兼容 API

ai awq embedding fastapi gptq llm mlx openai-compatible-api rag reranker sentence-transformers vllm

Last synced: 08 May 2025

https://github.com/slb350/open-agent-sdk-rust

Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and comprehensive examples.

llamacpp llm localllama localllm ollama openai rust rust-crate rust-library vllm

Last synced: 30 Apr 2026

https://github.com/jeremyarancio/vlm-batch-deployment

Batch Deployment for Document Parsing with AWS Batch & Qwen-2.5-VL

aws batch llm vllm vlm

Last synced: 25 Oct 2025

https://github.com/ld-singh/ai-factory-ops-lab

Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.

ai-infrastructure dcgm devops gpu gpu-scheduling gpu-sharing grafana hami hpc inference k3s kai-scheduler kubernetes learning-by-doing mlops nvidia observability prometheus slurm vllm

Last synced: 01 Sep 2026

https://github.com/lamalab-org/macbench

Probing the limitations of multimodal language models for chemistry and materials research

benchmark chemistry llm materials mutlimodal vllm vlm

Last synced: 29 Jul 2025

https://github.com/0xnyn/periscope

LLM Performance Testing | K6 + Grafana + InfluxDB | A tiny toolkit for load testing and benchmarking OpenAI-like inference endpoints using K6 + Grafana + InfluxDB

grafana influxdb k6 llm-testing performance-testing vllm vllm-production-stack

Last synced: 27 Jul 2026

https://github.com/attach-dev/attach-gateway

Drop-in OIDC & Google A2A auth + Weaviate memory for Ollama, vLLM and any local LLM server.

a2a attach-dev auth0 gateway jwt llm memory ollama proxy vllm weaviate

Last synced: 10 Oct 2025

https://github.com/itelnov/skeernir

UI to deploy locally agents and customise interaction with them

agents ai fastapi htmx jinja2-templates langchain langgraph llamacpp localai ollama ui vllm

Last synced: 10 Oct 2025