{"id":51724615,"url":"https://github.com/xLLM-AI/xllm","last_synced_at":"2026-08-05T23:00:46.790Z","repository":{"id":310843959,"uuid":"1036704760","full_name":"xLLM-AI/xllm","owner":"xLLM-AI","description":"A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.","archived":false,"fork":false,"pushed_at":"2026-08-03T09:51:30.000Z","size":27030,"stargazers_count":1504,"open_issues_count":193,"forks_count":275,"subscribers_count":18,"default_branch":"main","last_synced_at":"2026-08-03T10:05:21.533Z","etag":null,"topics":["deepseek","glm","inference","inference-engine","large-language-models","llm-inference","qwen"],"latest_commit_sha":null,"homepage":"https://xllm-ai.com/","language":"C++","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/xLLM-AI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":".github/CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":".github/CODEOWNERS","security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","claude":"CLAUDE.md","gemini":null,"cursor":null,"copilot":null,"dco":null,"cla":null,"disclosure":null}},"created_at":"2025-08-12T13:16:07.000Z","updated_at":"2026-08-03T09:54:39.000Z","dependencies_parsed_at":"2025-12-29T07:06:24.233Z","dependency_job_id":"1f8b4016-5f86-4ce6-be34-769b2154f859","html_url":"https://github.com/xLLM-AI/xllm","commit_stats":null,"previous_names":["jd-opensource/xllm","xllm-ai/xllm"],"tags_count":10,"template":false,"template_full_name":null,"purl":"pkg:github/xLLM-AI/xllm","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xLLM-AI%2Fxllm","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xLLM-AI%2Fxllm/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xLLM-AI%2Fxllm/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xLLM-AI%2Fxllm/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/xLLM-AI","download_url":"https://codeload.github.com/xLLM-AI/xllm/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/xLLM-AI%2Fxllm/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36323912,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-08-05T02:00:06.619Z","response_time":104,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["deepseek","glm","inference","inference-engine","large-language-models","llm-inference","qwen"],"created_at":"2026-07-17T17:00:23.496Z","updated_at":"2026-08-05T23:00:46.772Z","avatar_url":"https://github.com/xLLM-AI.png","language":"C++","funding_links":[],"categories":["Model Serving \u0026 Inference"],"sub_categories":["Model Serving Frameworks"],"readme":"\u003c!-- Copyright 2022 JD Co.\n\nLicensed under the Apache License, Version 2.0 (the \"License\");\nyou may not use this project except in compliance with the License.\nYou may obtain a copy of the License at\n\n    http://www.apache.org/licenses/LICENSE-2.0\n\nUnless required by applicable law or agreed to in writing, software\ndistributed under the License is distributed on an \"AS IS\" BASIS,\nWITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\nSee the License for the specific language governing permissions and\nlimitations under the License. --\u003e\n\n[English](./README.md) | [中文](./README_zh.md)\n\n\u003cdiv align=\"center\"\u003e\n\u003cimg src=\"assets/logo_with_llm.png\" alt=\"xLLM\" style=\"width:50%; height:auto;\"\u003e\n    \n[![Document](https://img.shields.io/badge/Document-black?logo=html5\u0026labelColor=grey\u0026color=red)](https://docs.xllm-ai.com/) [![Docker](https://img.shields.io/badge/Docker-black?logo=docker\u0026labelColor=grey\u0026color=%231E90FF)](https://quay.io/repository/jd_xllm/xllm-ai?tab=tags) [![License](https://img.shields.io/badge/license-Apache%202.0-brightgreen?labelColor=grey)](https://opensource.org/licenses/Apache-2.0) [![report](https://img.shields.io/badge/Technical%20Report-red?logo=arxiv\u0026logoColor=%23B31B1B\u0026labelColor=%23F0EBEB\u0026color=%23D42626)](https://arxiv.org/abs/2510.14686) [![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/jd-opensource/xllm)\n    \n\u003c/div\u003e\n\n---------------------\n\n\n### 📢 News\n\u003c!-- only keep the latest 3 news, others should be folded --\u003e\n- 2026-07-06: 🎉 xLLM is officially donated to the OpenAtom Foundation!\n- 2026-06-13: 🎉 We day-0 support the [MiniMax-M3](https://huggingface.co/MiniMaxAI/MiniMax-M3) model, please refer to the [Deployment Document](https://github.com/jd-opensource/xllm/blob/preview/minimax-m3/testspace/run_minimax_m3.sh) for deployment.\n- 2026-04-24: 🎉 We day-0 support the [DeepSeek-V4](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash) model, please refer to the [Deployment Document](https://github.com/jd-opensource/xllm/blob/preview/deepseek-v4-mlu/testspace/run_deepseek_v4.sh) for deployment.\n\n\u003cdetails\u003e\n\u003csummary\u003eMore News\u003c/summary\u003e\n\n- 2026-02-12: 🎉 We day-0 support high-performance inference for the [GLM-5](https://github.com/zai-org/GLM-5) model, please refer to the [Deployment Document](https://github.com/zai-org/GLM-5/blob/main/example/ascend.md) for deployment.\n- 2025-12-21: 🎉 We day-0 support high-performance inference for the [GLM-4.7](https://github.com/zai-org) model.\n- 2025-12-08: 🎉 We day-0 support high-performance inference for the [GLM-4.6V](https://github.com/zai-org/GLM-V) model.\n- 2025-12-05: 🎉 We now support high-performance inference for the [GLM-4.5/GLM-4.6](https://github.com/zai-org/GLM-4.5/blob/main/README_zh.md) series models.\n- 2025-12-05: 🎉 We now support high-performance inference for the [VLM-R1](https://github.com/om-ai-lab/VLM-R1) model.\n- 2025-12-05: 🎉 We build hybrid KV cache management based on [Mooncake](https://github.com/kvcache-ai/Mooncake), supporting global KV cache management with intelligent offloading and prefetching.\n- 2025-10-16: 🎉 We recently have released our [xLLM Technical Report](https://arxiv.org/abs/2510.14686) on arXiv, providing comprehensive technical blueprints and implementation insights.\n\n\u003c/details\u003e\n\n## Overview\n\n**xLLM** is an **efficient LLM inference framework**, specifically optimized for **Chinese AI accelerators**, enabling enterprise-grade deployment with enhanced efficiency and reduced cost.\n\u003cdiv align=\"center\"\u003e\n\u003cimg src=\"assets/xllm_arch.png\" alt=\"xllm_arch\" style=\"width:90%; height:auto;\"\u003e\n\u003c/div\u003e\n\n## Highlights\n\n* **Top-tier Performance**: Delivers high-throughput, low-latency inference through many advanced features.\n* **Mainstream Hardware Support**: Purpose-built and deeply optimized for Chinese AI accelerators.\n* **Service-Engine Decoupled Architecture**: Service layer handles scheduling and availability; engine layer handles computation.\n* **Enterprise-grade Deployment**: Battle-tested at scale across JD.com's core retail business.\n\n## Hardware Support\n\n| Hardware           | Abbreviation | Example | Remark              |\n| ------------------ | ------------ | ------- | ------------------- |\n| Ascend NPU         | NPU          | A2, A3  | HDK Driver 25.2.0 + |\n| Cambricon MLU      | MLU          | MLU     |                     |\n| Moore Threads GPU  | MUSA         | S5000   |                     |\n| Hygon DCU          | DCU          | BW1000  |                     |\n| MetaX MACA         | MACA         | MXC500  |                     |\n| Iluvatar CoreX GPU | ILU          | BI150   |                     |\n\n\n## Getting Started\n\n* [Quick Start](https://docs.xllm-ai.com/en/getting_started/quick_start/)\n* [Launch xLLM](https://docs.xllm-ai.com/en/getting_started/launch_xllm/)\n* [Online Service](https://docs.xllm-ai.com/en/getting_started/online_service/)\n* [Offline Inference](https://docs.xllm-ai.com/en/getting_started/offline_service/)\n* [Supported Models](https://docs.xllm-ai.com/en/supported_models/)\n\n## Community \u0026 Support\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"assets/wechat_qrcode.png\" alt=\"qrcode3\" width=\"50%\" /\u003e\n\u003c/div\u003e\n\n## Acknowledgment\n\nThis project was made possible thanks to the following open-source projects:  \n- [ScaleLLM](https://github.com/vectorch-ai/ScaleLLM) - xLLM draws inspiration from ScaleLLM's graph construction method and references its runtime execution. \n- [Mooncake](https://github.com/kvcache-ai/Mooncake) - Build xLLM hybrid KV cache management based on Mooncake.\n- [brpc](https://github.com/apache/brpc) - Build high-performance http service based on brpc.\n- [tokenizers-cpp](https://github.com/mlc-ai/tokenizers-cpp) - Build C++ tokenizer based on tokenizers-cpp.\n- [safetensors](https://github.com/huggingface/safetensors) - xLLM relies on the C binding safetensors capability.\n- [Partial JSON Parser](https://github.com/promplate/partial-json-parser) - Implement xLLM's C++ JSON parser with insights from Python and Go implementations.\n- [concurrentqueue](https://github.com/cameron314/concurrentqueue) - A fast multi-producer, multi-consumer lock-free concurrent queue for C++11.\n\n\nThanks to the following collaborating university laboratories:\n\n- [THU-MIG](https://ise.thss.tsinghua.edu.cn/mig/projects.html) (School of Software, BNRist, Tsinghua University)\n- USTC-Cloudlab (Cloud Computing Lab, University of Science and Technology of China)\n- [Beihang-HiPO](https://github.com/buaa-hipo) (Beihang HiPO research group)\n- PKU-DS-LAB (Data Structure Laboratory, Peking University)\n- PKU-NetSys-LAB (NetSys Lab, Peking University)\n- [TJU-TANKLab](https://flashserve.org/) (TANK Lab, Tianjin University)\n\nThanks to all the following [developers](https://github.com/jd-opensource/xllm/graphs/contributors) who have contributed to xLLM.\n\n\u003ca href=\"https://github.com/jd-opensource/xllm/graphs/contributors\"\u003e\n  \u003cimg src=\"https://contrib.rocks/image?repo=jd-opensource/xllm\" /\u003e\n\u003c/a\u003e\n\n\n## Citation\n\nIf you think this repository is helpful to you, welcome to cite us:\n```\n@article{liu2025xllm,\n  title={xLLM Technical Report},\n  author={Liu, Tongxuan and Peng, Tao and Yang, Peijun and Zhao, Xiaoyang and Lu, Xiusheng and Huang, Weizhe and Liu, Zirui and Chen, Xiaoyu and Liang, Zhiwei and Xiong, Jun and others},\n  journal={arXiv preprint arXiv:2510.14686},\n  year={2025}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FxLLM-AI%2Fxllm","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FxLLM-AI%2Fxllm","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FxLLM-AI%2Fxllm/lists"}