{"id":34731322,"url":"https://github.com/sgl-project/mini-sglang","last_synced_at":"2026-04-04T02:09:59.183Z","repository":{"id":313869769,"uuid":"1048705942","full_name":"sgl-project/mini-sglang","owner":"sgl-project","description":"A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.","archived":false,"fork":false,"pushed_at":"2026-03-10T03:57:23.000Z","size":940,"stargazers_count":3657,"open_issues_count":29,"forks_count":486,"subscribers_count":15,"default_branch":"main","last_synced_at":"2026-03-10T10:37:19.209Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/sgl-project.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-09-01T22:31:45.000Z","updated_at":"2026-03-10T09:33:49.000Z","dependencies_parsed_at":"2025-09-09T09:27:17.338Z","dependency_job_id":"7fe0e81b-1f86-4b3f-b51b-ce64d5ea563b","html_url":"https://github.com/sgl-project/mini-sglang","commit_stats":null,"previous_names":["darksharpness/mini-sglang"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/sgl-project/mini-sglang","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sgl-project%2Fmini-sglang","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sgl-project%2Fmini-sglang/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sgl-project%2Fmini-sglang/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sgl-project%2Fmini-sglang/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/sgl-project","download_url":"https://codeload.github.com/sgl-project/mini-sglang/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sgl-project%2Fmini-sglang/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31384848,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-04T01:22:39.193Z","status":"online","status_checked_at":"2026-04-04T02:00:07.569Z","response_time":60,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-12-25T03:00:24.759Z","updated_at":"2026-04-04T02:09:59.173Z","avatar_url":"https://github.com/sgl-project.png","language":"Python","funding_links":[],"categories":["Repos","Python","Inference engines","Deployment and Serving","8. Inference Engines","Learning and Reference","3. Inference Engines \u0026 Serving","A01_文本生成_文本对话"],"sub_categories":["Server / Production","Tutorials and Books","大语言对话模型及数据"],"readme":"\u003cp align=\"center\"\u003e\n\u003cimg width=\"400\" src=\"/assets/logo.png\"\u003e\n\u003c/p\u003e\n\n# Mini-SGLang\n\nA **lightweight yet high-performance** inference framework for Large Language Models.\n\n---\n\nMini-SGLang is a compact implementation of [SGLang](https://github.com/sgl-project/sglang), designed to demystify the complexities of modern LLM serving systems. With a compact codebase of **~5,000 lines of Python**, it serves as both a capable inference engine and a transparent reference for researchers and developers.\n\n## ✨ Key Features\n\n- **High Performance**: Achieves state-of-the-art throughput and latency with advanced optimizations.\n- **Lightweight \u0026 Readable**: A clean, modular, and fully type-annotated codebase that is easy to understand and modify.\n- **Advanced Optimizations**:\n  - **Radix Cache**: Reuses KV cache for shared prefixes across requests.\n  - **Chunked Prefill**: Reduces peak memory usage for long-context serving.\n  - **Overlap Scheduling**: Hides CPU scheduling overhead with GPU computation.\n  - **Tensor Parallelism**: Scales inference across multiple GPUs.\n  - **Optimized Kernels**: Integrates **FlashAttention** and **FlashInfer** for maximum efficiency.\n  - ...\n\n## 🚀 Quick Start\n\n\u003e **⚠️ Platform Support**: Mini-SGLang currently supports **Linux only** (x86_64 and aarch64). Windows and macOS are not supported due to dependencies on Linux-specific CUDA kernels (`sgl-kernel`, `flashinfer`). We recommend using [WSL2](https://learn.microsoft.com/en-us/windows/wsl/install) on Windows or Docker for cross-platform compatibility.\n\n### 1. Environment Setup\n\nWe recommend using `uv` for a fast and reliable installation (note that `uv` does not conflict with `conda`).\n\n```bash\n# Create a virtual environment (Python 3.10+ recommended)\nuv venv --python=3.12\nsource .venv/bin/activate\n```\n\n**Prerequisites**: Mini-SGLang relies on CUDA kernels that are JIT-compiled. Ensure you have the **NVIDIA CUDA Toolkit** installed and that its version matches your driver's version. You can check your driver's CUDA capability with `nvidia-smi`.\n\n### 2. Installation\n\nInstall Mini-SGLang directly from the source:\n\n```bash\ngit clone https://github.com/sgl-project/mini-sglang.git\ncd mini-sglang \u0026\u0026 uv venv --python=3.12 \u0026\u0026 source .venv/bin/activate\nuv pip install -e .\n```\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e💡 Installing on Windows (WSL2)\u003c/b\u003e\u003c/summary\u003e\n\nSince Mini-SGLang requires Linux-specific dependencies, Windows users should use WSL2:\n\n1. **Install WSL2** (if not already installed):\n   ```powershell\n   # In PowerShell (as Administrator)\n   wsl --install\n   ```\n\n2. **Install CUDA on WSL2**:\n   - Follow [NVIDIA's WSL2 CUDA guide](https://docs.nvidia.com/cuda/wsl-user-guide/index.html)\n   - Ensure your Windows GPU drivers support WSL2\n\n3. **Install Mini-SGLang in WSL2**:\n   ```bash\n   # Inside WSL2 terminal\n   git clone https://github.com/sgl-project/mini-sglang.git\n   cd mini-sglang \u0026\u0026 uv venv --python=3.12 \u0026\u0026 source .venv/bin/activate\n   uv pip install -e .\n   ```\n\n4. **Access from Windows**: The server will be accessible at `http://localhost:8000` from Windows browsers and applications.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003e🐳 Running with Docker\u003c/b\u003e\u003c/summary\u003e\n\n**Prerequisites**:\n- [Docker](https://docs.docker.com/get-docker/)\n- [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html)\n\n1. **Build the Docker image**:\n   ```bash\n   docker build -t minisgl .\n   ```\n\n2. **Run the server**:\n   ```bash\n   docker run --gpus all -p 1919:1919 \\\n       minisgl --model Qwen/Qwen3-0.6B --host 0.0.0.0\n   ```\n\n3. **Run in interactive shell mode**:\n   ```bash\n   docker run -it --gpus all \\\n       minisgl --model Qwen/Qwen3-0.6B --shell\n   ```\n\n4. **Using Docker Volumes for persistent caches** (recommended for faster subsequent startups):\n   ```bash\n   docker run --gpus all -p 1919:1919 \\\n       -v huggingface_cache:/app/.cache/huggingface \\\n       -v tvm_cache:/app/.cache/tvm-ffi \\\n       -v flashinfer_cache:/app/.cache/flashinfer \\\n       minisgl --model Qwen/Qwen3-0.6B --host 0.0.0.0\n   ```\n\n\u003c/details\u003e\n\n### 3. Online Serving\n\nLaunch an OpenAI-compatible API server with a single command.\n\n```bash\n# Deploy Qwen/Qwen3-0.6B on a single GPU\npython -m minisgl --model \"Qwen/Qwen3-0.6B\"\n\n# Deploy meta-llama/Llama-3.1-70B-Instruct on 4 GPUs with Tensor Parallelism, on port 30000\npython -m minisgl --model \"meta-llama/Llama-3.1-70B-Instruct\" --tp 4 --port 30000\n```\n\nOnce the server is running, you can send requests using standard tools like `curl` or any OpenAI-compatible client.\n\n### 4. Interactive Shell\n\nChat with your model directly in the terminal by adding the `--shell` flag.\n\n```bash\npython -m minisgl --model \"Qwen/Qwen3-0.6B\" --shell\n```\n\n![shell-example](https://lmsys.org/images/blog/minisgl/shell.png)\n\nYou can also use `/reset` to clear the chat history.\n\n## Benchmark\n\n### Offline inference\n\nSee [bench.py](./benchmark/offline/bench.py) for more details. Set `MINISGL_DISABLE_OVERLAP_SCHEDULING=1` for ablation study on overlap scheduling.\n\nTest Configuration:\n\n- Hardware: 1xH200 GPU.\n- Model: Qwen3-0.6B, Qwen3-14B\n- Total Requests: 256 sequences\n- Input Length: Randomly sampled between 100-1024 tokens\n- Output Length: Randomly sampled between 100-1024 tokens\n\n![offline](https://lmsys.org/images/blog/minisgl/offline.png)\n\n### Online inference\n\nSee [benchmark_qwen.py](./benchmark/online/bench_qwen.py) for more details.\n\nTest Configuration:\n\n- Hardware: 4xH200 GPU, connected by NVLink.\n- Model: Qwen3-32B\n- Dataset: [Qwen trace](https://github.com/alibaba-edu/qwen-bailian-usagetraces-anon/blob/main/qwen_traceA_blksz_16.jsonl), replaying first 1000 requests.\n\nLaunch command:\n\n```bash\n# Mini-SGLang\npython -m minisgl --model \"Qwen/Qwen3-32B\" --tp 4 --cache naive\n\n# SGLang\npython3 -m sglang.launch_server --model \"Qwen/Qwen3-32B\" --tp 4 \\\n    --disable-radix --port 1919 --decode-attention flashinfer\n```\n\n\u003e **Note**: If you encounter network issues when downloading models from HuggingFace, try using `--model-source modelscope` to download from ModelScope instead:\n\u003e ```bash\n\u003e python -m minisgl --model \"Qwen/Qwen3-32B\" --tp 4 --model-source modelscope\n\u003e ```\n\n![online](https://lmsys.org/images/blog/minisgl/online.png)\n\n## 📚 Learn More\n\n- **[Detailed Features](./docs/features.md)**: Explore all available features and command-line arguments.\n- **[System Architecture](./docs/structures.md)**: Dive deep into the design and data flow of Mini-SGLang.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsgl-project%2Fmini-sglang","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsgl-project%2Fmini-sglang","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsgl-project%2Fmini-sglang/lists"}