{"id":26896958,"url":"https://github.com/wdlctc/headinfer","last_synced_at":"2025-04-01T04:02:31.667Z","repository":{"id":276908860,"uuid":"930700333","full_name":"wdlctc/headinfer","owner":"wdlctc","description":null,"archived":false,"fork":false,"pushed_at":"2025-03-26T03:47:16.000Z","size":12,"stargazers_count":39,"open_issues_count":0,"forks_count":5,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-03-26T04:34:55.014Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/wdlctc.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2025-02-11T04:08:17.000Z","updated_at":"2025-03-23T21:03:05.000Z","dependencies_parsed_at":"2025-02-11T05:22:32.759Z","dependency_job_id":"925c0ed9-15f3-46f3-aca5-6ac7f796b3a7","html_url":"https://github.com/wdlctc/headinfer","commit_stats":null,"previous_names":["wdlctc/headinfer"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wdlctc%2Fheadinfer","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wdlctc%2Fheadinfer/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wdlctc%2Fheadinfer/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wdlctc%2Fheadinfer/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/wdlctc","download_url":"https://codeload.github.com/wdlctc/headinfer/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":246580468,"owners_count":20800111,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-04-01T04:02:25.327Z","updated_at":"2025-04-01T04:02:31.662Z","avatar_url":"https://github.com/wdlctc.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"# HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading  \n\n![License](https://img.shields.io/badge/license-MIT-blue.svg)  \n![Python](https://img.shields.io/badge/python-3.8%2B-blue)  \n![PyTorch](https://img.shields.io/badge/PyTorch-1.12%2B-orange)  \n\n[[paper](https://arxiv.org/abs/2502.12574)]\n\n## Overview  \n\n**HeadInfer** is a memory-efficient inference framework for large language models (LLMs) that significantly reduces GPU memory consumption by leveraging a **head-wise offloading** strategy. Unlike traditional layer-wise KV cache offloading, **HeadInfer** dynamically manages attention heads, maintaining only a subset of the KV cache on the GPU while offloading the rest to CPU memory.  \n\nWith **HeadInfer**, an **8B model can process up to 4 million tokens on a single consumer-grade GPU** (e.g., RTX 4090 with 24GB VRAM), **reducing GPU KV cache memory from 128GB to just 1GB** without approximation.  \n\n## Features  \n\n- ✅ **Head-wise KV cache offloading**: Fine-grained memory optimization for long-context inference.  \n- ✅ **Supports million-token inference**: Achieves up to **4M context length** on consumer GPUs.  \n- ✅ **Asynchronous data transfer**: Overlaps computation with offloading to minimize bottlenecks.  \n- ✅ **Compatible with major LLMs**: Works with LLaMA, Mistral, Qwen, and more.  \n- ✅ **Minimal changes to existing inference frameworks**: Easy integration with Hugging Face models.  \n\n## Installation  \n\n#### Training and Evaluation Environment\n\n```bash\nconda create -yn duo python=3.10\nconda activate duo\n\nconda install -y git\nconda install -y nvidia/label/cuda-12.4.0::cuda-toolkit\nconda install -y nvidia::cuda-cudart-dev\nconda install -y pytorch torchvision torchaudio pytorch-cuda=12.4 -c pytorch -c nvidia\n\npip install transformers==4.45.2 accelerate \npip install flash-attn --no-build-isolation\n```\n\n\n## Usage  \n\nOne-click-run with HeadInfer\n\n```bash\npython main.py\n```\n\nRunning Inference with HeadInfer\n```python\n\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\n\n\nfrom headinfer.cache import OffloadedCache\nfrom headinfer.mp import mp_headinfer, mp_simulate_decode\n\nmodel_name = \"meta-llama/Meta-Llama-3-8B\"\ntokenizer = AutoTokenizer.from_pretrained(model_name)\nmodel = AutoModelForCausalLM.from_pretrained(model_name)\n\n# Wrap the model with HeadInfer\nheadinfer_model = HeadInferModel(model)\n\n# Generate text with long context\ninput_text = \"Once upon a time in a galaxy far, far away...\"\ninput_ids = tokenizer(input_text, return_tensors=\"pt\").input_ids\n\n\nwith torch.inference_mode():\n\n    # patch the model\n    mp_headinfer(model)\n    past_key_values = OffloadedCache()\n\n    model(input_ids=input_ids, past_key_values=past_key_values, use_cache=True, num_logits_to_keep=1)\n\n```\n\n## Citation\nIf you find HeadInfer useful for your research, please cite:\n\n```bibtex\n@article{luo2025headinfer,\n  title={HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading},\n  author={Luo, Cheng and Cai, Zefan and Sun, Hanshi and Xiao, Jinqi and Yuan, Bo and Xiao, Wen and Hu, Junjie and Zhao, Jiawei and Chen, Beidi and Anandkumar, Anima},\n  journal={arXiv preprint arXiv:2502.12574},\n  year={2025}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwdlctc%2Fheadinfer","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fwdlctc%2Fheadinfer","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwdlctc%2Fheadinfer/lists"}