{"id":52773,"url":"https://github.com/horseee/Awesome-Efficient-LLM","name":"Awesome-Efficient-LLM","description":"A curated list for Efficient Large Language Models","projects_count":217,"last_synced_at":"2026-09-13T22:00:27.843Z","repository":{"id":168186275,"uuid":"643793540","full_name":"horseee/Awesome-Efficient-LLM","owner":"horseee","description":"A curated list for Efficient Large Language Models","archived":false,"fork":false,"pushed_at":"2025-06-17T02:35:25.000Z","size":67840,"stargazers_count":2035,"open_issues_count":11,"forks_count":168,"subscribers_count":43,"default_branch":"main","last_synced_at":"2026-08-25T03:11:32.984Z","etag":null,"topics":["compression","efficient-llm","knowledge-distillation","language-model","llm","llm-compression","model-quantization","pruning-algorithms"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/horseee.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2023-05-22T07:07:49.000Z","updated_at":"2026-08-24T07:44:09.000Z","dependencies_parsed_at":"2024-11-08T08:01:09.489Z","dependency_job_id":"f82a28ea-4512-42da-a63f-9829c101589f","html_url":"https://github.com/horseee/Awesome-Efficient-LLM","commit_stats":{"total_commits":455,"total_committers":17,"mean_commits":"26.764705882352942","dds":"0.12967032967032965","last_synced_commit":"fdefd207ab2490d827beefce81dbf490e6e1eec3"},"previous_names":["horseee/awesome-efficient-llm"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/horseee/Awesome-Efficient-LLM","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/horseee%2FAwesome-Efficient-LLM","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/horseee%2FAwesome-Efficient-LLM/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/horseee%2FAwesome-Efficient-LLM/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/horseee%2FAwesome-Efficient-LLM/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/horseee","download_url":"https://codeload.github.com/horseee/Awesome-Efficient-LLM/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/horseee%2FAwesome-Efficient-LLM/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":341189360,"owners_count":37294215,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-13T02:00:22.117Z","response_time":122,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-01-14T13:39:35.121Z","updated_at":"2026-09-13T22:00:27.843Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Survey","Quantization","Network Pruning","Paper from Sep 30, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Knowledge Distillation","Text Compression","Inference Acceleration","Low-Rank Decomposition","Tuning","Hardware/System","Leaderboard","Full List","Efficient MOE","Hardware","Efficient Architecture of LLM","KV Cache Compression","Paper from Sep 2, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from May 26, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from June 6, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from 05/26/2024 - Now (see Full List from 05/22/2023 [here](#full-list))","Paper from June 2, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from June 21, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from June 13, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from July 13, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from July 4, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from August 17, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))","Paper from August 24, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))"],"sub_categories":["KV Cache Compression","Knowledge Distillation","Please check out all the papers by selecting the sub-area you're interested in. On this main page, we're showing papers released in the past 90 days.","Hardware/System/Serving","Please check out all the papers by selecting the sub-area you're interested in. On this page, we're showing papers released in the past 30 days.","Network Pruning / Sparsity","Quantization","Inference Acceleration","Hardware/System","Tuning","Text Compression","Please check out all the papers by selecting the sub-area you're interested in. On this page, we're showing papers released in the past 60 days.","Low-Rank Decomposition","Survey","Efficient MOE","Efficient Architecture of LLM","Survey (or Benchmark)","Efficient Fine-tuning","Efficient Training","Please check out all the papers by selecting the sub-area you're interested in. On this main page, only papers released in the past 90 days are shown."],"readme":"# Awesome-Efficient-LLM\nA curated list for **Efficient Large Language Models**\n\n## Full List\n  - [Network Pruning / Sparsity](pruning.md)\n  - [Knowledge Distillation](knowledge_distillation.md)\n  - [Quantization](quantization.md)\n  - [Inference Acceleration](inference_acceleration.md)\n  - [Efficient MOE](efficient_moe.md)\n  - [Efficient Architecture of LLM](efficient_architecture_llm.md)\n  - [KV Cache Compression](kv_cache_compression.md)\n  - [Text Compression](text_compression.md)\n  - [Low-Rank Decomposition](low_rank_decomposition.md)\n  - [Hardware / System / Serving](hardware.md)\n  - [Efficient Fine-tuning](tuning.md)\n  - [Efficient Training](efficient_training.md)\n  - [Survey or Benchmark](survey.md)\n  - [Reasoning Model](https://github.com/fscdc/Awesome-Efficient-Reasoning-Models)\n\n### Please check out all the papers by selecting the sub-area you're interested in. On this main page, only papers released in the past 90 days are shown.\n\n#### 🚀 Updates\n* April 15, 2025: We have a new [curated list](https://github.com/fscdc/Awesome-Efficient-Reasoning-Models) for **efficient reasoning model**!\n* May 29, 2024: We've had this awesome list for a year now :smiling_face_with_three_hearts:! \n* Sep 6, 2023: Add a new subdirectory [project/](project/) to organize efficient LLM projects.\n* July 11, 2023: A new subdirectory [efficient_plm/](efficient_plm/) is created to house papers that are applicable to PLMs. \n\n#### 💮 Contributing\n\nIf you'd like to include your paper, or need to update any details such as conference information or code URLs, please feel free to submit a pull request. You can generate the required markdown format for each paper by filling in the information in `generate_item.py` and execute `python generate_item.py`. We warmly appreciate your contributions to this list. Alternatively, you can email me with the links to your paper and code, and I would add your paper to the list at my earliest convenience. \n\n#### :star: Recommended Paper\n\nFor each topic, we have curated a list of recommended papers that have garnered a lot of GitHub stars or citations.\n\n\n## Paper from Sep 30, 2024 - Now (see Full List from May 22, 2023 [here](#full-list))\n\n### Quick Link \n  - [Network Pruning / Sparsity](#network-pruning--sparsity)\n  - [Knowledge Distillation](#knowledge-distillation)\n  - [Quantization](#quantization)\n  - [Inference Acceleration](#inference-acceleration)\n  - [Efficient MOE](#efficient_moe)\n  - [Efficient Architecture of LLM](#efficient-architecture-of-llm)\n  - [KV Cache Compression](#kv-cache-compression)\n  - [Text Compression](#text-compression)\n  - [Low-Rank Decomposition](#low-rank-decomposition)\n  - [Hardware / System / Serving](#hardwaresystemserving)\n  - [Efficient Fine-tuning](#efficient-fine-tuning)\n  - [Efficient Training](#efficient-training)\n  - [Survey](#survey-or-benchmark)\n\n### Network Pruning / Sparsity\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n| [![Star](https://img.shields.io/github/stars/IST-DASLab/sparsegpt.svg?style=social\u0026label=Star)](https://github.com/IST-DASLab/sparsegpt) [![Publish](https://img.shields.io/badge/Conference-ICML'23-blue)]() [![Type](https://img.shields.io/badge/Unstructured-C2A4A6)]() \u003cbr\u003e :star: [SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot](https://github.com/IST-DASLab/sparsegpt) \u003cbr\u003e Elias Frantar, Dan Alistarh| \u003cimg width=\"522\" alt=\"image\" src=\"figures/sparsegpt.png\"\u003e |[Github](https://github.com/IST-DASLab/sparsegpt) [paper](https://arxiv.org/abs/2301.00774) | [//]: #Recommend\n| [![Star](https://img.shields.io/github/stars/horseee/LLM-Pruner.svg?style=social\u0026label=Star)](https://github.com/horseee/LLM-Pruner) [![Publish](https://img.shields.io/badge/Conference-NeurIPS'23-blue)]() [![Type](https://img.shields.io/badge/Structural-C2A4A6)]() \u003cbr\u003e :star: [LLM-Pruner: On the Structural Pruning of Large Language Models](https://arxiv.org/abs/2305.11627) \u003cbr\u003e Xinyin Ma, Gongfan Fang, Xinchao Wang | \u003cimg width=\"561\" alt=\"image\" src=\"figures/llm_pruner.png\"\u003e| [Github](https://github.com/horseee/LLM-Pruner) [paper](https://arxiv.org/abs/2305.11627)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/locuslab/wanda.svg?style=social\u0026label=Star)](https://github.com/locuslab/wanda) [![Publish](https://img.shields.io/badge/Conference-ICLR'24-blue)]() [![Type](https://img.shields.io/badge/Unstructured-C2A4A6)]()  \u003cbr\u003e :star: [A Simple and Effective Pruning Approach for Large Language Models](https://arxiv.org/abs/2306.11695) \u003cbr\u003e Mingjie Sun, Zhuang Liu, Anna Bair, J. Zico Kolter |\u003cimg width=\"1002\" alt=\"image\" src=\"https://user-images.githubusercontent.com/20168304/245999360-f951de47-269d-491d-826a-8e6d85627849.png\"\u003e |[Github](https://github.com/locuslab/wanda) \u003cbr\u003e [Paper](https://arxiv.org/abs/2306.11695)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/princeton-nlp/LLM-Shearing.svg?style=social\u0026label=Star)](https://github.com/princeton-nlp/LLM-Shearing) [![Publish](https://img.shields.io/badge/Conference-ICLR'24-blue)]() [![Type](https://img.shields.io/badge/Structural-C2A4A6)]() \u003cbr\u003e :star: [Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning](https://arxiv.org/abs/2310.06694) \u003cbr\u003e Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/LLM-shearing.png\"\u003e |[Github](https://github.com/princeton-nlp/LLM-Shearing) \u003cbr\u003e [Paper](https://arxiv.org/abs/2310.06694)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/NVlabs/MaskLLM.svg?style=social\u0026label=Star)](https://github.com/NVlabs/MaskLLM) [![Publish](https://img.shields.io/badge/Conference-NeurIPS'24-blue)]() [![Type](https://img.shields.io/badge/Semi_Structured-C2A4A6)]() \u003cbr\u003e :star: [MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models](https://arxiv.org/abs/2409.17481) \u003cbr\u003e Gongfan Fang, Hongxu Yin, Saurav Muralidharan, Greg Heinrich, Jeff Pool, Jan Kautz, Pavlo Molchanov, Xinchao Wang |\u003cimg width=\"302\" alt=\"image\" src=\"https://github.com/NVlabs/MaskLLM/blob/main/assets/animation-LQ.gif\"\u003e |[Github](https://github.com/NVlabs/MaskLLM) \u003cbr\u003e [Paper](https://arxiv.org/abs/2409.17481)|[//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/IntelLabs/Hardware-Aware-Automated-Machine-Learning.svg?style=social\u0026label=Star)](https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning/tree/main/Mamba-Shedder) [![Publish](https://img.shields.io/badge/Conference-NAACL'25-blue)]() [![Type](https://img.shields.io/badge/Structural-C2A4A6)]() \u003cbr\u003e[Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models](https://arxiv.org/abs/2501.17088) \u003cbr\u003e Juan Pablo Munoz, Jinjie Yuan, Nilesh Jain |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/Mamba-Shedder.png\"\u003e |[Github](https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning/tree/main/Mamba-Shedder) \u003cbr\u003e [Paper](https://arxiv.org/abs/2501.17088)|[//]: #01/28\n|[![Star](https://img.shields.io/github/stars/IntelLabs/Hardware-Aware-Automated-Machine-Learning.svg?style=social\u0026label=Star)](https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning/tree/main/MultiPruner) [![Type](https://img.shields.io/badge/Structural-C2A4A6)]() \u003cbr\u003e[MultiPruner: Balanced Structure Removal in Foundation Models](https://arxiv.org/abs/2501.09949) \u003cbr\u003e Juan Pablo Munoz, Jinjie Yuan, Nilesh Jain |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/MultiPruner.png\"\u003e |[Github](https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning/tree/main/MultiPruner) \u003cbr\u003e [Paper](https://arxiv.org/abs/2501.09949)|[//]: #01/17\n|[HashAttention: Semantic Sparsity for Faster Inference](https://arxiv.org/abs/2412.14468) \u003cbr\u003e Aditya Desai, Shuo Yang, Alejandro Cuadron, Ana Klimovic, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.14468v1/extracted/6081011/images/sparseatt.png\"\u003e |[Paper](https://arxiv.org/abs/2412.14468)|[//]: #12/30\n|[Adaptive Pruning for Large Language Models with Structural Importance Awareness](https://arxiv.org/abs/2412.15127) \u003cbr\u003e Haotian Zheng, Jinke Ren, Yushan Sun, Ruichen Zhang, Wenbo Zhang, Zhen Li, Dusit Niyato, Shuguang Cui, Yatong Han |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.15127v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2412.15127)|[//]: #12/30\n|[SlimGPT: Layer-wise Structured Pruning for Large Language Models](https://arxiv.org/abs/2412.18110) \u003cbr\u003e Gui Ling, Ziyang Wang, Yuliang Yan, Qingwen Liu |\u003cimg width=\"302\" alt=\"image\" src=\"https://arxiv.org/html/2412.18110v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.18110)|[//]: #12/30\n|[Less is More: Towards Green Code Large Language Models via Unified Structural Pruning](https://arxiv.org/abs/2412.15921) \u003cbr\u003e Guang Yang, Yu Zhou, Xiangyu Zhang, Wei Cheng, Ke Liu, Xiang Chen, Terry Yue Zhuo, Taolue Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/Flab-Pruner.png\"\u003e |[Paper](https://arxiv.org/abs/2412.15921)|[//]: #12/30\n|[Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking](https://arxiv.org/abs/2412.01380) \u003cbr\u003e Marco Federici, Davide Belli, Mart van Baalen, Amir Jalalirad, Andrii Skliar, Bence Major, Markus Nagel, Paul Whatmough |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.01380v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.01380)|[//]: #12/09\n|[Puzzle: Distillation-Based NAS for Inference-Optimized LLMs](https://arxiv.org/abs/2411.19146) \u003cbr\u003e Akhiad Bercovich, Tomer Ronen, Talor Abramovich, Nir Ailon, Nave Assaf, Mohammad Dabbah et al |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.19146v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.19146)|[//]: #12/09\n|[![Star](https://img.shields.io/github/stars/yaolu-zjut/Navigation-LLM-layer-pruning.svg?style=social\u0026label=Star)](https://github.com/yaolu-zjut/Navigation-LLM-layer-pruning)\u003cbr\u003e[Reassessing Layer Pruning in LLMs: New Insights and Methods](https://arxiv.org/abs/2411.15558) \u003cbr\u003e Yao Lu, Hao Cheng, Yujie Fang, Zeyu Wang, Jiaheng Wei, Dongwei Xu, Qi Xuan, Xiaoniu Yang, Zhaowei Zhu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/yaolu-zjut/Navigation-LLM-layer-pruning/raw/main/framework.JPG\"\u003e |[Github](https://github.com/yaolu-zjut/Navigation-LLM-layer-pruning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.15558)|[//]: #12/03\n|[Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity](https://arxiv.org/abs/2411.10069) \u003cbr\u003e Zichen Song, Sitan Huang, Yuxin Wu, Zhongfeng Kang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.10069v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.10069)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/GATECH-EIC/AmoebaLLM.svg?style=social\u0026label=Star)](https://github.com/GATECH-EIC/AmoebaLLM)[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24-blue)]()\u003cbr\u003e[AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment](https://arxiv.org/abs/2411.10606) \u003cbr\u003e Yonggan Fu, Zhongzhi Yu, Junwei Li, Jiayi Qian, Yongan Zhang, Xiangchi Yuan, Dachuan Shi, Roman Yakunin, Yingyan Celine Lin |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.10606v1/x2.png\"\u003e |[Github](https://github.com/GATECH-EIC/AmoebaLLM) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.10606)|[//]: #11/24\n|[Scaling Law for Post-training after Model Pruning](https://arxiv.org/abs/2411.10272) \u003cbr\u003e Xiaodong Chen, Yuxuan Hu, Jing Zhang, Xiaokang Zhang, Cuiping Li, Hong Chen | |[Paper](https://arxiv.org/abs/2411.10272)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/hexuandeng/DRPruning.svg?style=social\u0026label=Star)](https://github.com/hexuandeng/DRPruning)\u003cbr\u003e[DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization](https://arxiv.org/abs/2411.14055) \u003cbr\u003e Hexuan Deng, Wenxiang Jiao, Xuebo Liu, Min Zhang, Zhaopeng Tu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/hexuandeng/DRPruning/raw/main/pic/main.png\"\u003e |[Github](https://github.com/hexuandeng/DRPruning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.14055)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/thunlp/SparsingLaw.svg?style=social\u0026label=Star)](https://github.com/thunlp/SparsingLaw)\u003cbr\u003e[Sparsing Law: Towards Large Language Models with Greater Activation Sparsity](https://arxiv.org/abs/2411.02335) \u003cbr\u003e Yuqi Luo, Chenyang Song, Xu Han, Yingfa Chen, Chaojun Xiao, Zhiyuan Liu, Maosong Sun |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/thunlp/SparsingLaw/raw/master/figs/sample.jpg\"\u003e |[Github](https://github.com/thunlp/SparsingLaw) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.02335)|[//]: #11/18\n|[AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis](https://arxiv.org/abs/2411.02117) \u003cbr\u003e Zichen Song, Yuxin Wu, Sitan Huang, Zhongfeng Kang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.02117v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.02117)|[//]: #11/18\n|[Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts](https://arxiv.org/abs/2410.19185) \u003cbr\u003e Danyal Aftab, Steven Davy |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.19185v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.19185)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/AboveParadise/LLMCBench.svg?style=social\u0026label=Star)](https://github.com/AboveParadise/LLMCBench)\u003cbr\u003e[LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment](https://arxiv.org/abs/2410.21352) \u003cbr\u003e Ge Yang, Changyi He, Jinyang Guo, Jianyu Wu, Yifu Ding, Aishan Liu, Haotong Qin, Pengliang Ji, Xianglong Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/AboveParadise/LLMCBench/raw/main/figs/f1.png\"\u003e |[Github](https://github.com/AboveParadise/LLMCBench) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.21352)|[//]: #11/17\n|[Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs](https://arxiv.org/abs/2410.16135) \u003cbr\u003e Kang Zhao, Tao Yuan, Han Bao, Zhenfeng Su, Chang Gao, Zhaofeng Sun, Zichen Liang, Liping Jing, Jianfei Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.16135v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.16135)|[//]: #10/30\n|[![Star](https://img.shields.io/github/stars/IST-DASLab/EvoPress.svg?style=social\u0026label=Star)](https://github.com/IST-DASLab/EvoPress)\u003cbr\u003e[EvoPress: Towards Optimal Dynamic Model Compression via Evolutionary Search](https://arxiv.org/abs/2410.14649) \u003cbr\u003e Oliver Sieberling, Denis Kuznedelev, Eldar Kurtic, Dan Alistarh |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/evopress.png\"\u003e |[Github](https://github.com/IST-DASLab/EvoPress) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.14649)|[//]: #10/30\n|[FedSpaLLM: Federated Pruning of Large Language Models](https://arxiv.org/abs/2410.14852) \u003cbr\u003e Guangji Bai, Yijiang Li, Zilinghan Li, Liang Zhao, Kibaek Kim |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.14852v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.14852)|[//]: #10/30\n|[![Star](https://img.shields.io/github/stars/piuzha/APT.svg?style=social\u0026label=Star)](https://github.com/piuzha/APT)\u003cbr\u003e[Pruning Foundation Models for High Accuracy without Retraining](https://arxiv.org/abs/2410.15567) \u003cbr\u003e Pu Zhao, Fei Sun, Xuan Shen, Pinrui Yu, Zhenglun Kong, Yanzhi Wang, Xue Lin | |[Github](https://github.com/piuzha/APT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.15567)|[//]: #10/30\n|[Self-calibration for Language Model Quantization and Pruning](https://arxiv.org/abs/2410.17170) \u003cbr\u003e Miles Williams, George Chrysostomou, Nikolaos Aletras |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.17170v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.17170)|[//]: #10/29\n|[Beware of Calibration Data for Pruning Large Language Models](https://arxiv.org/abs/2410.17711) \u003cbr\u003e Yixin Ji, Yang Xiang, Juntao Li, Qingrong Xia, Ping Li, Xinyu Duan, Zhefeng Wang, Min Zhang | |[Paper](https://arxiv.org/abs/2410.17711)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/haiquanlu/AlphaPruning.svg?style=social\u0026label=Star)](https://github.com/haiquanlu/AlphaPruning)[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24-blue)]()\u003cbr\u003e[AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models](https://arxiv.org/abs/2410.10912) \u003cbr\u003e Haiquan Lu, Yefan Zhou, Shiwei Liu, Zhangyang Wang, Michael W. Mahoney, Yaoqing Yang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.10912v1/x1.png\"\u003e |[Github](https://github.com/haiquanlu/AlphaPruning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.10912)|[//]: #10/21\n|[Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix](https://arxiv.org/abs/2410.11261) \u003cbr\u003e Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, Yufa Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.11261v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.11261)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/ZhengaoLi/DISP-LLM-Dimension-Independent-Structural-Pruning.svg?style=social\u0026label=Star)](https://github.com/ZhengaoLi/DISP-LLM-Dimension-Independent-Structural-Pruning)[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24-blue)]()\u003cbr\u003e[DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models](https://arxiv.org/abs/2410.11988) \u003cbr\u003e Shangqian Gao, Chi-Heng Lin, Ting Hua, Tang Zheng, Yilin Shen, Hongxia Jin, Yen-Chang Hsu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.11988v1/x1.png\"\u003e |[Github](https://github.com/ZhengaoLi/DISP-LLM-Dimension-Independent-Structural-Pruning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.11988)|[//]: #10/21\n|[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24%20Workshop-blue)]()\u003cbr\u003e[Self-Data Distillation for Recovering Quality in Pruned Large Language Models](https://arxiv.org/abs/2410.09982) \u003cbr\u003e Vithursan Thangarasa, Ganesh Venkatesh, Nish Sinnadurai, Sean Lie |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.09982v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.09982)|[//]: #10/21\n|[LLM-Rank: A Graph Theoretical Approach to Pruning Large Language Models](https://arxiv.org/abs/2410.13299) \u003cbr\u003e David Hoffmann, Kailash Budhathoki, Matthaeus Kleindessner |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.13299v1/extracted/5931028/img/llm_to_mlp.png\"\u003e |[Paper](https://arxiv.org/abs/2410.13299)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/abx393/llm-pruning-calibration-data.svg?style=social\u0026label=Star)](https://github.com/abx393/llm-pruning-calibration-data)[![Publish](https://img.shields.io/badge/Conference-EMNLP'24-blue)]()\u003cbr\u003e[Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning](https://arxiv.org/abs/2410.07461) \u003cbr\u003e Abhinav Bandari, Lu Yin, Cheng-Yu Hsieh, Ajay Kumar Jaiswal, Tianlong Chen, Li Shen, Ranjay Krishna, Shiwei Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.07461v1/x1.png\"\u003e |[Github](https://github.com/abx393/llm-pruning-calibration-data) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.07461)|[//]: #10/13\n|[Mitigating Copy Bias in In-Context Learning through Neuron Pruning](https://arxiv.org/abs/2410.01288) \u003cbr\u003e Ameen Ali, Lior Wolf, Ivan Titov |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/copy_icl.png\"\u003e |[Paper](https://arxiv.org/abs/2410.01288)|[//]: #10/04\n|[![Star](https://img.shields.io/github/stars/IntelLabs/Hardware-Aware-Automated-Machine-Learning.svg?style=social\u0026label=Star)](https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning/tree/main/SQFT)[![Publish](https://img.shields.io/badge/Conference-EMNLP'24%20Findings-blue)]() [![Type](https://img.shields.io/badge/Unstructured-C2A4A6)]() [![Type](https://img.shields.io/badge/w/Quantization-39B0A9)]() \u003cbr\u003e[SQFT: Low-cost Model Adaptation in Low-precision Sparse Foundation Models](https://arxiv.org/abs/2410.03750) \u003cbr\u003e Juan Pablo Munoz, Jinjie Yuan, Nilesh Jain |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/SQFT.png\"\u003e |[Github](https://github.com/IntelLabs/Hardware-Aware-Automated-Machine-Learning/tree/main/SQFT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.03750)|[//]: #10/01\n|[![Star](https://img.shields.io/github/stars/PiotrNawrot/sparse-frontier.svg?style=social\u0026label=Star)](https://github.com/PiotrNawrot/sparse-frontier)\u003cbr\u003e[The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs](https://arxiv.org/abs/2504.17768) \u003cbr\u003e Piotr Nawrot, Robert Li, Renjie Huang, Sebastian Ruder, Kelly Marchisio, Edoardo M. Ponti |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/SparseFrontier.png\"\u003e |[Github](https://github.com/PiotrNawrot/sparse-frontier) \u003cbr\u003e [Paper](https://arxiv.org/abs/2504.17768)|[//]: #05/05\n|[![Star](https://img.shields.io/github/stars/woominsong/Simba.svg?style=social\u0026label=Star)](https://github.com/woominsong/Simba)[![Publish](https://img.shields.io/badge/Journal-TMLR_2025-blue)]()\u003cbr\u003e[Sparsified State-Space Models are Efficient Highway Networks](https://arxiv.org/abs/2505.20698) \u003cbr\u003e Woomin Song, Jihoon Tack, Sangwoo Mo, Seunghyuk Oh, Jinwoo Shin |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/simba.png\"\u003e |[Github](https://github.com/woominsong/Simba) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.20698)|[//]: #06/03\n\n\n\n\n### Knowledge Distillation\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|:star: [Knowledge Distillation of Large Language Models](https://arxiv.org/abs/2306.08543) \u003cbr\u003e Yuxian Gu, Li Dong, Furu Wei, Minlie Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/microsoft/LMOps/blob/main/minillm/figures/method.png\"\u003e |[Github](https://github.com/microsoft/LMOps/tree/main/minillm) \u003cbr\u003e [Paper](https://arxiv.org/abs/2306.08543)| [//]: #Recommend\n|[![Publish](https://img.shields.io/badge/Conference-COLING'25-blue)]()\u003cbr\u003e[Self-Evolution Knowledge Distillation for LLM-based Machine Translation](https://arxiv.org/abs/2412.15303) \u003cbr\u003e Yuncheng Song, Liang Ding, Changtong Zan, Shujian Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.15303v1/extracted/6081708/model_two.png\"\u003e |[Paper](https://arxiv.org/abs/2412.15303)|[//]: #12/30\n|[Large Language Models Compression via Low-Rank Feature Distillation](https://arxiv.org/abs/2412.16719) \u003cbr\u003e Yaya Sy, Christophe Cerisara, Irina Illina |\u003cimg width=\"302\" alt=\"image\" src=\"https://arxiv.org/html/2412.16719v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.16719)|[//]: #12/30\n|[![Star](https://img.shields.io/github/stars/HITSZ-HLT/FSA-Distillation.svg?style=social\u0026label=Star)](https://github.com/HITSZ-HLT/FSA-Distillation)\u003cbr\u003e[Distilling Fine-grained Sentiment Understanding from Large Language Models](https://arxiv.org/abs/2412.18552) \u003cbr\u003e Yice Zhang, Guangyu Xie, Hongling Xu, Kaiheng Hou, Jianzhu Bao, Qianlong Wang, Shiwei Chen, Ruifeng Xu |\u003cimg width=\"302\" alt=\"image\" src=\"https://arxiv.org/html/2412.18552v1/x1.png\"\u003e |[Github](https://github.com/HITSZ-HLT/FSA-Distillation) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.18552)|[//]: #12/30\n|[![Star](https://img.shields.io/github/stars/alonso130r/knowledge-distillation.svg?style=social\u0026label=Star)](https://github.com/alonso130r/knowledge-distillation)\u003cbr\u003e[Enhancing Knowledge Distillation for LLMs with Response-Priming Prompting](https://arxiv.org/abs/2412.17846) \u003cbr\u003e Vijay Goyal, Mustafa Khan, Aprameya Tirupati, Harveer Saini, Michael Lam, Kevin Zhu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.17846v1/extracted/6080471/prompt-example.png\"\u003e |[Github](https://github.com/alonso130r/knowledge-distillation) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.17846)|[//]: #12/30\n|[Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation](https://arxiv.org/abs/2411.14698) \u003cbr\u003e Xunyu Zhu, Jian Li, Can Ma, Weiping Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.14698v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.14698)|[//]: #12/03\n|[![Star](https://img.shields.io/github/stars/kaistai/GenPI.svg?style=social\u0026label=Star)](https://github.com/kaistai/GenPI)\u003cbr\u003e[Generative Prompt Internalization](https://arxiv.org/abs/2411.15927) \u003cbr\u003e Haebin Shin, Lei Ji, Yeyun Gong, Sungdong Kim, Eunbi Choi, Minjoon Seo |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/GCD.png\"\u003e |[Github](https://github.com/kaistai/GenPI) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.15927)|[//]: #12/02\n|[SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models](https://arxiv.org/abs/2410.19503) \u003cbr\u003e Jahyun Koo, Yerin Hwang, Yongil Kim, Taegwan Kang, Hyunkyung Bae, Kyomin Jung |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/switch.png\"\u003e |[Paper](https://arxiv.org/abs/2410.19503)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/jdeschena/sdtt.svg?style=social\u0026label=Star)](https://github.com/jdeschena/sdtt)\u003cbr\u003e[Beyond Autoregression: Fast LLMs via Self-Distillation Through Time](https://arxiv.org/abs/2410.21035) \u003cbr\u003e Justin Deschenaux, Caglar Gulcehre |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.21035v1/x3.png\"\u003e |[Github](https://github.com/jdeschena/sdtt) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.21035)|[//]: #11/17\n|[Pre-training Distillation for Large Language Models: A Design Space Exploration](https://arxiv.org/abs/2410.16215) \u003cbr\u003e Hao Peng, Xin Lv, Yushi Bai, Zijun Yao, Jiajie Zhang, Lei Hou, Juanzi Li | |[Paper](https://arxiv.org/abs/2410.16215)|[//]: #10/30\n|[![Star](https://img.shields.io/github/stars/thu-coai/MiniPLM.svg?style=social\u0026label=Star)](https://github.com/thu-coai/MiniPLM)\u003cbr\u003e[MiniPLM: Knowledge Distillation for Pre-Training Language Models](https://arxiv.org/abs/2410.17215) \u003cbr\u003e Yuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou, Minlie Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/thu-coai/MiniPLM/raw/main/figures/method.png\"\u003e |[Github](https://github.com/thu-coai/MiniPLM) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.17215)|[//]: #10/29\n|[Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling](https://arxiv.org/abs/2410.11325) \u003cbr\u003e Wenda Xu, Rujun Han, Zifeng Wang, Long T. Le, Dhruv Madeka, Lei Li, William Yang Wang, Rishabh Agarwal, Chen-Yu Lee, Tomas Pfister |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.11325v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.11325)|[//]: #10/21\n|[Evolutionary Contrastive Distillation for Language Model Alignment](https://arxiv.org/abs/2410.07513) \u003cbr\u003e Julian Katz-Samuels, Zheng Li, Hyokun Yun, Priyanka Nigam, Yi Xu, Vaclav Petricek, Bing Yin, Trishul Chilimbi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.07513v1/extracted/5913898/figures/main_alg_v3.png\"\u003e |[Paper](https://arxiv.org/abs/2410.07513)|[//]: #10/13\n\n\n\n### Quantization\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/IST-DASLab/gptq.svg?style=social\u0026label=Star)](https://github.com/IST-DASLab/gptq)[![Publish](https://img.shields.io/badge/Conference-ICLR'22-blue)]()\u003cbr\u003e :star: [GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers](https://arxiv.org/abs/2210.17323) \u003cbr\u003e Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh |\u003cimg width=\"202\" alt=\"image\" src=\"figures/GPTQ.png\"\u003e |[Github](https://github.com/IST-DASLab/gptq) \u003cbr\u003e [Paper](https://arxiv.org/abs/2210.17323)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/mit-han-lab/smoothquant.svg?style=social\u0026label=Star)](https://github.com/mit-han-lab/smoothquant)[![Publish](https://img.shields.io/badge/Conference-ICML'23-blue)]() \u003cbr\u003e :star: [SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models](https://arxiv.org/abs/2211.10438) \u003cbr\u003e Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/mit-han-lab/smoothquant/blob/main/figures/intuition.png\"\u003e |[Github](https://github.com/mit-han-lab/smoothquant) \u003cbr\u003e [Paper](https://arxiv.org/abs/2211.10438)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/mit-han-lab/llm-awq.svg?style=social\u0026label=Star)](https://github.com/mit-han-lab/llm-awq) \u003cbr\u003e :star: [AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration](https://arxiv.org/abs/2306.00978) \u003cbr\u003e Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, Song Han |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/mit-han-lab/llm-awq/blob/main/figures/overview.png\"\u003e |[Github](https://github.com/mit-han-lab/llm-awq) \u003cbr\u003e [Paper](https://arxiv.org/abs/2306.00978)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/OpenGVLab/OmniQuant.svg?style=social\u0026label=Star)](https://github.com/OpenGVLab/OmniQuant)[![Publish](https://img.shields.io/badge/Conference-ICLR'24-blue)]()\u003cbr\u003e :star: [OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models](https://arxiv.org/abs/2308.13137) \u003cbr\u003e Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, Ping Luo |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/omniquant.png\"\u003e |[Github](https://github.com/OpenGVLab/OmniQuant) \u003cbr\u003e [Paper](https://arxiv.org/abs/2308.13137)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/utkarsh-dmx/project-resq.svg?style=social\u0026label=Star)](https://github.com/utkarsh-dmx/project-resq)\u003cbr\u003e[ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals](https://arxiv.org/abs/2412.14363) \u003cbr\u003e Utkarsh Saxena, Sayeh Sharify, Kaushik Roy, Xin Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/ResQ.png\"\u003e |[Github](https://github.com/utkarsh-dmx/project-resq) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.14363)|[//]: #12/30\n|[MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design](https://arxiv.org/abs/2412.14590) \u003cbr\u003e Zhen Zheng, Xiaonan Song, Chuanjie Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.14590v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.14590)|[//]: #12/30\n|[GQSA: Group Quantization and Sparsity for Accelerating Large Language Model Inference](https://arxiv.org/abs/2412.17560) \u003cbr\u003e Chao Zeng, Songwei Liu, Shu Yang, Fangmin Chen, Xing Mei, Lean Fu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.17560v1/extracted/6090667/GQS_block.png\"\u003e |[Paper](https://arxiv.org/abs/2412.17560)|[//]: #12/30\n|[LSAQ: Layer-Specific Adaptive Quantization for Large Language Model Deployment](https://arxiv.org/abs/2412.18135) \u003cbr\u003e Binrui Zeng, Bin Ji, Xiaodong Liu, Jie Yu, Shasha Li, Jun Ma, Xiaopeng Li, Shangwen Wang, Xinran Hong |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.18135v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.18135)|[//]: #12/30\n|[SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization](https://arxiv.org/abs/2412.04180) \u003cbr\u003e Runsheng Bai, Qiang Liu, Bo Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.04180v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2412.04180)|[//]: #12/09\n|[CPTQuant -- A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models](https://arxiv.org/abs/2412.03599) \u003cbr\u003e Amitash Nanda, Sree Bhargavi Balija, Debashis Sahoo |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.03599v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2412.03599)|[//]: #12/09\n|[![Publish](https://img.shields.io/badge/Conference-HPCA'25-blue)]()\u003cbr\u003e[Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format](https://arxiv.org/abs/2411.15982) \u003cbr\u003e Chao Fang, Man Shi, Robin Geens, Arne Symons, Zhongfeng Wang, Marian Verhelst |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.15982v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.15982)|[//]: #12/03\n|[MixPE: Quantization and Hardware Co-design for Efficient LLM Inference](https://arxiv.org/abs/2411.16158) \u003cbr\u003e Yu Zhang, Mingzi Wang, Lancheng Zou, Wulong Liu, Hui-Ling Zhen, Mingxuan Yuan, Bei Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.16158v1/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2411.16158)|[//]: #12/03\n|[![Star](https://img.shields.io/github/stars/abdelfattah-lab/BitMoD-HPCA-25.svg?style=social\u0026label=Star)](https://github.com/abdelfattah-lab/BitMoD-HPCA-25)[![Publish](https://img.shields.io/badge/Conference-HPCA'25-blue)]()\u003cbr\u003e[BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration](https://arxiv.org/abs/2411.11745) \u003cbr\u003e Yuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang, Marta Andronic, George A. Constantinides, Mohamed S. Abdelfattah |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.11745v1/x5.png\"\u003e |[Github](https://github.com/abdelfattah-lab/BitMoD-HPCA-25) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.11745)|[//]: #11/24\n|[AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference](https://arxiv.org/abs/2411.09909) \u003cbr\u003e Janghwan Lee, Jiwoong Park, Jinseok Kim, Yongjik Kim, Jungju Oh, Jinwook Oh, Jungwook Choi |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/AMXFP4.png\"\u003e |[Paper](https://arxiv.org/abs/2411.09909)|[//]: #11/24\n|[Bi-Mamba: Towards Accurate 1-Bit State Space Models](https://arxiv.org/abs/2411.11843) \u003cbr\u003e Shengkun Tang, Liqun Ma, Haonan Li, Mingjie Sun, Zhiqiang Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.11843v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2411.11843)|[//]: #11/24\n|[\"Give Me BF16 or Give Me Death\"? Accuracy-Performance Trade-Offs in LLM Quantization](https://arxiv.org/abs/2411.02355) \u003cbr\u003e Eldar Kurtic, Alexandre Marques, Shubhra Pandit, Mark Kurtz, Dan Alistarh | |[Paper](https://arxiv.org/abs/2411.02355)|[//]: #11/18\n|[GWQ: Gradient-Aware Weight Quantization for Large Language Models](https://arxiv.org/abs/2411.00850) \u003cbr\u003e Yihua Shao, Siyu Liang, Xiaolin Lin, Zijian Ling, Zixian Zhu et al  |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.00850v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2411.00850)|[//]: #11/18\n|[A Comprehensive Study on Quantization Techniques for Large Language Models](https://arxiv.org/abs/2411.02530) \u003cbr\u003e Jiedong Lang, Zhehao Guo, Shuyu Huang | |[Paper](https://arxiv.org/abs/2411.02530)|[//]: #11/18\n|[BitNet a4.8: 4-bit Activations for 1-bit LLMs](https://arxiv.org/abs/2411.04965) \u003cbr\u003e Hongyu Wang, Shuming Ma, Furu Wei |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.04965v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.04965)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/Intelligent-Computing-Lab-Yale/TesseraQ.svg?style=social\u0026label=Star)](https://github.com/Intelligent-Computing-Lab-Yale/TesseraQ)\u003cbr\u003e[TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction](https://arxiv.org/abs/2410.19103) \u003cbr\u003e Yuhang Li, Priyadarshini Panda |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/Intelligent-Computing-Lab-Yale/TesseraQ/raw/main/imgs/tesseraq.png\"\u003e |[Github](https://github.com/Intelligent-Computing-Lab-Yale/TesseraQ) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.19103)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/xinghaow99/BitStack.svg?style=social\u0026label=Star)](https://github.com/xinghaow99/BitStack)\u003cbr\u003e[BitStack: Fine-Grained Size Control for Compressed Large Language Models in Variable Memory Environments](https://arxiv.org/abs/2410.23918) \u003cbr\u003e Xinghao Wang, Pengyu Wang, Bo Wang, Dong Zhang, Yunhua Zhou, Xipeng Qiu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/xinghaow99/BitStack/raw/main/assets/bitstack.png\"\u003e |[Github](https://github.com/xinghaow99/BitStack) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.23918)|[//]: #11/17\n|[The Impact of Inference Acceleration Strategies on Bias of LLMs](https://arxiv.org/abs/2410.22118) \u003cbr\u003e Elisabeth Kirsten, Ivan Habernal, Vedant Nanda, Muhammad Bilal Zafar | |[Paper](https://arxiv.org/abs/2410.22118)|[//]: #11/17\n|[Understanding the difficulty of low-precision post-training quantization of large language models](https://arxiv.org/abs/2410.14570) \u003cbr\u003e Zifei Xu, Sayeh Sharify, Wanzin Yazar, Tristan Webb, Xin Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.14570v1/extracted/5935973/figures/fig1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.14570)|[//]: #10/30\n|[![Star](https://img.shields.io/github/stars/microsoft/BitNet.svg?style=social\u0026label=Star)](https://github.com/microsoft/BitNet)\u003cbr\u003e[1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs](https://arxiv.org/abs/2410.16144) \u003cbr\u003e Jinheng Wang, Hansong Zhou, Ting Song, Shaoguang Mao, Shuming Ma, Hongyu Wang, Yan Xia, Furu Wei |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.16144v2/x1.png\"\u003e |[Github](https://github.com/microsoft/BitNet) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.16144)|[//]: #10/30\n|[QuAILoRA: Quantization-Aware Initialization for LoRA](https://arxiv.org/abs/2410.14713) \u003cbr\u003e Neal Lawton, Aishwarya Padmakumar, Judith Gaspers, Jack FitzGerald, Anoop Kumar, Greg Ver Steeg, Aram Galstyan | |[Paper](https://arxiv.org/abs/2410.14713)|[//]: #10/30\n|[Evaluating Quantized Large Language Models for Code Generation on Low-Resource Language Benchmarks](https://arxiv.org/abs/2410.14766) \u003cbr\u003e Enkhbold Nyamsuren | |[Paper](https://arxiv.org/abs/2410.14766)|[//]: #10/30\n| [![Star](https://img.shields.io/github/stars/SqueezeAILab/SqueezeLLM.svg?style=social\u0026label=Star)](https://github.com/SqueezeAILab/SqueezeLLM) \u003cbr\u003e :star: [SqueezeLLM: Dense-and-Sparse Quantization](https://arxiv.org/pdf/2306.07629.pdf) \u003cbr\u003eSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, Kurt Keutzer | \u003cimg width=\"1102\" alt=\"image\" src=\"figures/SqueezeLLM.png\"\u003e |[Github](https://github.com/SqueezeAILab/SqueezeLLM) \u003cbr\u003e [Paper](https://arxiv.org/pdf/2306.07629.pdf)| [//]: #Recommend\n|[Pyramid Vector Quantization for LLMs](https://arxiv.org/abs/2410.16926) \u003cbr\u003e Tycho F. A. van der Ouderaa, Maximilian L. Croci, Agrin Hilmkil, James Hensman |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.16926v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.16926)|[//]: #10/29\n|[SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators](https://arxiv.org/abs/2410.10714) \u003cbr\u003e Rasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker, Houman Bedayat, Sachin Mehta, Mohammad Rastegari, Mahyar Najibi, Saman Naderiparizi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.10714v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.10714)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/ruikangliu/FlatQuant.svg?style=social\u0026label=Star)](https://github.com/ruikangliu/FlatQuant)\u003cbr\u003e[FlatQuant: Flatness Matters for LLM Quantization](https://arxiv.org/abs/2410.09426) \u003cbr\u003e Yuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao, Kang Zhao, Yuening Li, Jiaxin Hu, Xianzhi Yu, Lu Hou, Chun Yuan, Xin Jiang, Wulong Liu, Jun Yao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.09426v1/x11.png\"\u003e |[Github](https://github.com/ruikangliu/FlatQuant) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.09426)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/Mohammad-Mozaffari/slim.svg?style=social\u0026label=Star)](https://github.com/Mohammad-Mozaffari/slim)\u003cbr\u003e[SLiM: One-shot Quantized Sparse Plus Low-rank Approximation of LLMs](https://arxiv.org/abs/2410.09615) \u003cbr\u003e Mohammad Mozaffari, Maryam Mehri Dehnavi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.09615v1/x1.png\"\u003e |[Github](https://github.com/Mohammad-Mozaffari/slim) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.09615)|[//]: #10/21\n|[Scaling laws for post-training quantized large language models](https://arxiv.org/abs/2410.12119) \u003cbr\u003e Zifei Xu, Alexander Lan, Wanzin Yazar, Tristan Webb, Sayeh Sharify, Xin Wang |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.12119v1/extracted/5929616/figures/fig_12.png\"\u003e |[Paper](https://arxiv.org/abs/2410.12119)|[//]: #10/21\n|[Continuous Approximations for Improving Quantization Aware Training of LLMs](https://arxiv.org/abs/2410.10849) \u003cbr\u003e He Li, Jianhang Hong, Yuanzhuo Wu, Snehal Adbol, Zonglin Li | |[Paper](https://arxiv.org/abs/2410.10849)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/LuoYingSong/DAQ.svg?style=social\u0026label=Star)](https://github.com/LuoYingSong/DAQ)\u003cbr\u003e[DAQ: Density-Aware Post-Training Weight-Only Quantization For LLMs](https://arxiv.org/abs/2410.12187) \u003cbr\u003e Yingsong Luo, Ling Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.12187v2/x1.png\"\u003e |[Github](https://github.com/LuoYingSong/DAQ) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.12187)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/enyac-group/Quamba.svg?style=social\u0026label=Star)](https://github.com/enyac-group/Quamba)\u003cbr\u003e[Quamba: A Post-Training Quantization Recipe for Selective State Space Models](https://arxiv.org/abs/2410.13229) \u003cbr\u003e Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Diana Marculescu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.13229v1/extracted/5933363/figures/outliers.png\"\u003e |[Github](https://github.com/enyac-group/Quamba) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.13229)|[//]: #10/21\n|[AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations](https://arxiv.org/abs/2410.13212) \u003cbr\u003e Qian Tao, Wenyuan Yu, Jingren Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.13212v1/extracted/5933292/figures/kvmix.png\"\u003e |[Paper](https://arxiv.org/abs/2410.13212)|[//]: #10/21\n|[Channel-Wise Mixed-Precision Quantization for Large Language Models](https://arxiv.org/abs/2410.13056) \u003cbr\u003e Zihan Chen, Bike Xie, Jundong Li, Cong Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.13056v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.13056)|[//]: #10/21\n|[Progressive Mixed-Precision Decoding for Efficient LLM Inference](https://arxiv.org/abs/2410.13461) \u003cbr\u003e Hao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee, Hongxiang Fan, Stylianos I. Venieris |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.13461v1/x4.png\"\u003e |[Paper](https://arxiv.org/abs/2410.13461)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/Anonymous1252022/EXAQ.svg?style=social\u0026label=Star)](https://github.com/Anonymous1252022/EXAQ)\u003cbr\u003e[EXAQ: Exponent Aware Quantization For LLMs Acceleration](https://arxiv.org/abs/2410.03185) \u003cbr\u003e Moran Shkolnik, Maxim Fishman, Brian Chmiel, Hilla Ben-Yaacov, Ron Banner, Kfir Yehuda Levy |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/EXAQ.png\"\u003e |[Github](https://github.com/Anonymous1252022/EXAQ) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.03185)|[//]: #10/14\n|[![Star](https://img.shields.io/github/stars/ChenMnZ/PrefixQuant.svg?style=social\u0026label=Star)](https://github.com/ChenMnZ/PrefixQuant)\u003cbr\u003e[PrefixQuant: Static Quantization Beats Dynamic through Prefixed Outliers in LLMs](https://arxiv.org/abs/2410.05265) \u003cbr\u003e Mengzhao Chen, Yi Liu, Jiahao Wang, Yi Bin, Wenqi Shao, Ping Luo |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.05265v1/x1.png\"\u003e |[Github](https://github.com/ChenMnZ/PrefixQuant) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.05265)|[//]: #10/14\n|[![Star](https://img.shields.io/github/stars/vahe1994/AQLM.svg?style=social\u0026label=Star)](https://github.com/vahe1994/AQLM)\u003cbr\u003e :star: [Extreme Compression of Large Language Models via Additive Quantization](https://arxiv.org/abs/2401.06118) \u003cbr\u003e Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/MCQ.png\"\u003e |[Github](https://github.com/vahe1994/AQLM) \u003cbr\u003e [Paper](https://arxiv.org/abs/2401.06118)| [//]: #Recommend\n|[Scaling Laws for Mixed quantization in Large Language Models](https://arxiv.org/abs/2410.06722) \u003cbr\u003e Zeyu Cao, Cheng Zhang, Pedro Gimenes, Jianqiao Lu, Jianyi Cheng, Yiren Zhao |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/LLM-MPQ.png\"\u003e |[Paper](https://arxiv.org/abs/2410.06722)|[//]: #10/14\n|[PalmBench: A Comprehensive Benchmark of Compressed Large Language Models on Mobile Platforms](https://arxiv.org/abs/2410.05315) \u003cbr\u003e Yilong Li, Jingyu Liu, Hao Zhang, M Badri Narayanan, Utkarsh Sharma, Shuai Zhang, Pan Hu, Yijing Zeng, Jayaram Raghuram, Suman Banerjee |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/PalmBench.png\"\u003e |[Paper](https://arxiv.org/abs/2410.05315)|[//]: #10/14\n|[CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression](https://arxiv.org/abs/2410.07505) \u003cbr\u003e Wenyuan Liu, Xindian Ma, Peng Zhang, Yan Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.07505v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.07505)|[//]: #10/13\n|[SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration](https://arxiv.org/abs/2410.02367) \u003cbr\u003e Jintao Zhang, Jia wei, Pengle Zhang, Jun Zhu, Jianfei Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.02367v1/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2410.02367)|[//]: #10/04\n|[Addition is All You Need for Energy-efficient Language Models](https://arxiv.org/abs/2410.00907) \u003cbr\u003e Hongyin Luo, Wei Sun |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.00907v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.00907)|[//]: #10/02\n|[![Star](https://img.shields.io/github/stars/snu-mllab/GuidedQuant.svg?style=social\u0026label=Star)](https://github.com/snu-mllab/GuidedQuant)[![Publish](https://img.shields.io/badge/Conference-ICML'25-blue)]()\u003cbr\u003e[GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance](https://arxiv.org/abs/2505.07004) \u003cbr\u003e Jinuk Kim, Marwa El Halabi, Wonpyo Park, Clemens JS Schaefer, Deokjae Lee, Yeonhong Park, Jae W. Lee, Hyun Oh Song |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/GuidedQuant.png\"\u003e |[Github](https://github.com/snu-mllab/GuidedQuant) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.07004)|[//]: #06/15\n\n\n### Inference Acceleration\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/FMInference/DejaVu.svg?style=social\u0026label=Star)](https://github.com/FMInference/DejaVu)[![Publish](https://img.shields.io/badge/Conference-ICML'23%20Oral-blue)]()\u003cbr\u003e :star: [Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time](https://openreview.net/forum?id=wIPIhHd00i) \u003cbr\u003e Zichang Liu, Jue WANG, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, Beidi Chen |\u003cimg width=\"202\" alt=\"image\" src=\"figures/DajeVu.png\"\u003e |[Github](https://github.com/FMInference/DejaVu) \u003cbr\u003e [Paper](https://openreview.net/forum?id=wIPIhHd00i)| [//]: #Recommend\n| [![Star](https://img.shields.io/github/stars/flexflow/FlexFlow.svg?style=social\u0026label=Star)](https://github.com/flexflow/FlexFlow/tree/inference) \u003cbr\u003e :star: [SpecInfer: Accelerating Generative LLM Serving with Speculative Inference and Token Tree Verification](https://arxiv.org/abs/2305.09781) \u003cbr\u003e Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, Zhihao Jia| \u003cimg width=\"600\" alt=\"image\" src=\"https://github.com/flexflow/FlexFlow/blob/inference/img/overview.png\"\u003e| [Github](https://github.com/flexflow/FlexFlow/tree/inference) \u003cbr\u003e [paper](https://arxiv.org/abs/2305.09781) | [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/mit-han-lab/streaming-llm.svg?style=social\u0026label=Star)](https://github.com/mit-han-lab/streaming-llm)\u003cbr\u003e :star: [Efficient Streaming Language Models with Attention Sinks](https://arxiv.org/abs/2309.17453) \u003cbr\u003e Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/mit-han-lab/streaming-llm/blob/main/figures/schemes.png\"\u003e |[Github](https://github.com/mit-han-lab/streaming-llm) \u003cbr\u003e [Paper](https://arxiv.org/abs/2309.17453)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/SafeAILab/EAGLE.svg?style=social\u0026label=Star)](https://github.com/SafeAILab/EAGLE)\u003cbr\u003e:star: [EAGLE: Lossless Acceleration of LLM Decoding by Feature Extrapolation](https://sites.google.com/view/eagle-llm) \u003cbr\u003e Yuhui Li, Chao Zhang, and Hongyang Zhang |\u003cimg width=\"302\" alt=\"image\" src=\"https://github.com/SafeAILab/EAGLE/blob/main/figs/fig1.png\"\u003e |[Github](https://github.com/SafeAILab/EAGLE) \u003cbr\u003e [Blog](https://sites.google.com/view/eagle-llm)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/FasterDecoding/Medusa.svg?style=social\u0026label=Star)](https://github.com/FasterDecoding/Medusa)\u003cbr\u003e :star: [Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads](https://arxiv.org/abs/2401.10774) \u003cbr\u003e Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2401.10774v1/x1.png\"\u003e |[Github](https://github.com/FasterDecoding/Medusa) \u003cbr\u003e [Paper](https://arxiv.org/abs/2401.10774)| [//]: #Recommend\n|[Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration](https://arxiv.org/abs/2412.00061) \u003cbr\u003e Zhuofan Wen, Shangtong Gui, Yang Feng |\u003cimg width=\"302\" alt=\"image\" src=\"https://arxiv.org/html/2412.00061v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.00061)|[//]: #12/09\n|[PLD+: Accelerating LLM inference by leveraging Language Model Artifacts](https://arxiv.org/abs/2412.01447) \u003cbr\u003e Shwetha Somasundaram, Anirudh Phukan, Apoorv Saxena |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.01447v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.01447)|[//]: #12/09\n|[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24%20ENLSP-blue)]()\u003cbr\u003e[FastDraft: How to Train Your Draft](https://arxiv.org/abs/2411.11055) \u003cbr\u003e Ofir Zafrir, Igor Margulis, Dorin Shteyman, Guy Boudoukh | |[Paper](https://arxiv.org/abs/2411.11055)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/David-Li0406/SMoA.svg?style=social\u0026label=Star)](https://github.com/David-Li0406/SMoA)\u003cbr\u003e[SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents](https://arxiv.org/abs/2411.03284) \u003cbr\u003e Dawei Li, Zhen Tan, Peijia Qian, Yifan Li, Kumar Satvik Chaudhary, Lijie Hu, Jiayi Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/SMoA.png\"\u003e |[Github](https://github.com/David-Li0406/SMoA) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.03284)|[//]: #11/18\n|[The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation](https://arxiv.org/abs/2411.03786) \u003cbr\u003e Lawrence Stewart, Matthew Trager, Sujan Kumar Gonugondla, Stefano Soatto | |[Paper](https://arxiv.org/abs/2411.03786)|[//]: #11/18\n|[Accelerated AI Inference via Dynamic Execution Methods](https://arxiv.org/abs/2411.00853) \u003cbr\u003e Haim Barad, Jascha Achterberg, Tien Pei Chou, Jean Yu | |[Paper](https://arxiv.org/abs/2411.00853)|[//]: #11/18\n|[SuffixDecoding: A Model-Free Approach to Speeding Up Large Language Model Inference](https://arxiv.org/abs/2411.04975) \u003cbr\u003e Gabriele Oliaro, Zhihao Jia, Daniel Campos, Aurick Qiao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.04975v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.04975)|[//]: #11/18\n|[Dynamic Strategy Planning for Efficient Question Answering with Large Language Models](https://arxiv.org/abs/2410.23511) \u003cbr\u003e Tanmay Parekh, Pradyot Prakash, Alexander Radovic, Akshay Shekher, Denis Savenkov |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.23511v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.23511)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/Infini-AI-Lab/MagicPIG.svg?style=social\u0026label=Star)](https://github.com/Infini-AI-Lab/MagicPIG)\u003cbr\u003e[MagicPIG: LSH Sampling for Efficient LLM Generation](https://arxiv.org/abs/2410.16179) \u003cbr\u003e Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye, Yang Zhou, Jianyu Zhang, Niklas Nolte, Yuandong Tian, Matthijs Douze, Leon Bottou, Zhihao Jia, Beidi Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.16179v2/x15.png\"\u003e |[Github](https://github.com/Infini-AI-Lab/MagicPIG) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.16179)|[//]: #10/30\n|[Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition](https://arxiv.org/abs/2410.17765) \u003cbr\u003e Artem Basharin, Andrei Chertkov, Ivan Oseledets |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/canonical_tensor_decomposition.png\"\u003e |[Paper](https://arxiv.org/abs/2410.17765)|[//]: #10/29\n|[Efficient Inference for Augmented Large Language Models](https://arxiv.org/abs/2410.18248) \u003cbr\u003e Rana Shahout, Cong Liang, Shiji Xin, Qianru Lao, Yong Cui, Minlan Yu, Michael Mitzenmacher |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.18248v1/extracted/5949546/figures/illustrations/api_example_png.png\"\u003e |[Paper](https://arxiv.org/abs/2410.18248)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/MatteoNulli/Vocabulary_pruning.svg?style=social\u0026label=Star)](https://github.com/MatteoNulli/Vocabulary_pruning)\u003cbr\u003e[Dynamic Vocabulary Pruning in Early-Exit LLMs](https://arxiv.org/abs/2410.18952) \u003cbr\u003e Jort Vincenti, Karim Abdel Sadek, Joan Velja, Matteo Nulli, Metod Jazbec |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/MatteoNulli/Vocabulary_pruning/raw/main/src/images/final_nips.svg\"\u003e |[Github](https://github.com/MatteoNulli/Vocabulary_pruning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.18952)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/wangqinsi1/CoreInfer.svg?style=social\u0026label=Star)](https://github.com/wangqinsi1/CoreInfer)\u003cbr\u003e[CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation](https://arxiv.org/abs/2410.18311#) \u003cbr\u003e Qinsi Wang, Saeed Vahidian, Hancheng Ye, Jianyang Gu, Jianyi Zhang, Yiran Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://wangqinsi1.github.io/coreinfer_page/static/images/overview.png\"\u003e |[Github](https://github.com/wangqinsi1/CoreInfer) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.18311#)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/mit-han-lab/duo-attention.svg?style=social\u0026label=Star)](https://github.com/mit-han-lab/duo-attention)\u003cbr\u003e[DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads](https://arxiv.org/abs/2410.10819) \u003cbr\u003e Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, Song Han |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/mit-han-lab/duo-attention/raw/main/figures/method1.jpg\"\u003e |[Github](https://github.com/mit-han-lab/duo-attention) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.10819)|[//]: #10/21\n|[DySpec: Faster Speculative Decoding with Dynamic Token Tree Structure](https://arxiv.org/abs/2410.11744) \u003cbr\u003e Yunfan Xiong, Ruoyu Zhang, Yanzeng Li, Tianhao Wu, Lei Zou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.11744v1/extracted/5913908/figures/tree_bold.png\"\u003e |[Paper](https://arxiv.org/abs/2410.11744)|[//]: #10/21\n|[QSpec: Speculative Decoding with Complementary Quantization Schemes](https://arxiv.org/abs/2410.11305) \u003cbr\u003e Juntao Zhao, Wenhao Lu, Sheng Wang, Lingpeng Kong, Chuan Wu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.11305v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.11305)|[//]: #10/21\n|[TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention](https://arxiv.org/abs/2410.05076) \u003cbr\u003e Lijie Yang, Zhihao Zhang, Zhuofu Chen, Zikun Li, Zhihao Jia |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.05076v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.05076)|[//]: #10/14\n|[ParallelSpec: Parallel Drafter for Efficient Speculative Decoding](https://arxiv.org/abs/2410.05589) \u003cbr\u003e Zilin Xiao, Hongming Zhang, Tao Ge, Siru Ouyang, Vicente Ordonez, Dong Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.05589v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.05589)|[//]: #10/14\n|[![Star](https://img.shields.io/github/stars/hemingkx/SWIFT.svg?style=social\u0026label=Star)](https://github.com/hemingkx/SWIFT)\u003cbr\u003e[SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration](https://arxiv.org/abs/2410.06916) \u003cbr\u003e Heming Xia, Yongqi Li, Jun Zhang, Cunxiao Du, Wenjie Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/hemingkx/SWIFT/raw/main/assets/swift.png\"\u003e |[Github](https://github.com/hemingkx/SWIFT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.06916)|[//]: #10/14\n|[![Star](https://img.shields.io/github/stars/MooreThreads/TurboRAG.svg?style=social\u0026label=Star)](https://github.com/MooreThreads/TurboRAG)\u003cbr\u003e[TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text](https://arxiv.org/abs/2410.07590) \u003cbr\u003e Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, Yaohua Tang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/MooreThreads/TurboRAG/raw/main/assets/image/TurboRAG.png\"\u003e |[Github](https://github.com/MooreThreads/TurboRAG) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.07590)|[//]: #10/13\n|[A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts](https://arxiv.org/abs/2410.01485) \u003cbr\u003e Suyu Ge, Xihui Lin, Yunan Zhang, Jiawei Han, Hao Peng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.01485v1/extracted/5895696/figures/model_architecture.png\"\u003e |[Paper](https://arxiv.org/abs/2410.01485)|[//]: #10/04\n|[![Publish](https://img.shields.io/badge/Conference-SIGMOD'25-blue)]()\u003cbr\u003e[Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation](https://www.arxiv.org/pdf/2502.15734) \u003cbr\u003e Shubham Agarwal, Sai Sundaresan, Subrata Mitra, Debabrata Mahapatra, Archit Gupta, Rounak Sharma, Nirmal Joshua Kapu, Tong Yu, Shiv Saini |\u003cimg width=\"600\" alt=\"image\" src=\"figures/cachecraft.png\"\u003e | \u003cbr\u003e [Paper](https://www.arxiv.org/pdf/2502.15734)|[//]: #02/05\n|[Mamba Drafters for Speculative Decoding](https://arxiv.org/abs/2506.01206) \u003cbr\u003e Daewon Choi, Seunghyuk Oh, Saket Dingliwal, Jihoon Tack, Kyuyoung Kim, Woomin Song, Seojin Kim, Insu Han, Jinwoo Shin, Aram Galstyan, Shubham Katiyar, Sravan Babu Bodapati |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/mamba_drafters.png\"\u003e |[Paper](https://arxiv.org/abs/2506.01206)|[//]: #06/03\n|[Accelerated Test-Time Scaling with Model-Free Speculative Sampling](https://arxiv.org/abs/2506.04708) \u003cbr\u003e Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/stand.png\"\u003e |[Paper](https://arxiv.org/abs/2506.04708)|[//]: #06/05\n\n### Efficient MOE\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/dvmazur/mixtral-offloading.svg?style=social\u0026label=Star)](https://github.com/dvmazur/mixtral-offloading)\u003cbr\u003e:star: [Fast Inference of Mixture-of-Experts Language Models with Offloading](https://arxiv.org/abs/2312.17238) \u003cbr\u003e Artyom Eliseev, Denis Mazur |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/mixtral_offloading.png\"\u003e |[Github](https://github.com/dvmazur/mixtral-offloading) \u003cbr\u003e [Paper](https://arxiv.org/abs/2312.17238)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/duterscmy/CD-MoE.svg?style=social\u0026label=Star)](https://github.com/duterscmy/CD-MoE)\u003cbr\u003e[Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning](https://arxiv.org/abs/2412.00069) \u003cbr\u003e Mingyu Cao, Gen Li, Jie Ji, Jiaqi Zhang, Xiaolong Ma, Shiwei Liu, Lu Yin |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.00069v1/x2.png\"\u003e |[Github](https://github.com/duterscmy/CD-MoE) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.00069)|[//]: #12/09\n|[Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference](https://arxiv.org/abs/2412.00099) \u003cbr\u003e Andrii Skliar, Ties van Rozendaal, Romain Lepert, Todor Boinovski, Mart van Baalen, Markus Nagel, Paul Whatmough, Babak Ehteshami Bejnordi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.00099v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.00099)|[//]: #12/09\n|[![Star](https://img.shields.io/github/stars/EnflameTechnology/DeepSpeed.svg?style=social\u0026label=Star)](https://github.com/EnflameTechnology/DeepSpeed)\u003cbr\u003e[MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization](https://arxiv.org/abs/2411.00662) \u003cbr\u003e Jingming Guo, Yan Liu, Yu Meng, Zhiwei Tao, Banglan Liu, Gang Chen, Xiang Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.00662v1/x1.png\"\u003e |[Github](https://github.com/EnflameTechnology/DeepSpeed) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.00662)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/xiaochengsky/MoEI-2.svg?style=social\u0026label=Star)](https://github.com/xiaochengsky/MoEI-2)\u003cbr\u003e[MoE-I2: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition](https://arxiv.org/abs/2411.01016) \u003cbr\u003e Cheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang, Yu Gong, Yuanlin Duan, Wenqi Jia, Miao Yin, Yu Cheng, Bo Yuan |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.01016v1/x1.png\"\u003e |[Github](https://github.com/xiaochengsky/MoEI-2) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.01016)|[//]: #11/18\n|[HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference](https://arxiv.org/abs/2411.01433) \u003cbr\u003e Peng Tang, Jiacheng Liu, Xiaofeng Hou, Yifei Pu, Jing Wang, Pheng-Ann Heng, Chao Li, Minyi Guo |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.01433v2/extracted/5980843/figures/overview5.png\"\u003e |[Paper](https://arxiv.org/abs/2411.01433)|[//]: #11/18\n|[ProMoE: Fast MoE-based LLM Serving using Proactive Caching](https://arxiv.org/abs/2410.22134) \u003cbr\u003e Xiaoniu Song, Zihang Zhong, Rong Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.22134v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.22134)|[//]: #11/17\n|[ExpertFlow: Optimized Expert Activation and Token Allocation for Efficient Mixture-of-Experts Inference](https://arxiv.org/abs/2410.17954) \u003cbr\u003e Xin He, Shunkang Zhang, Yuxin Wang, Haiyan Yin, Zihao Zeng, Shaohuai Shi, Zhenheng Tang, Xiaowen Chu, Ivor Tsang, Ong Yew Soon |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.17954v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.17954)|[//]: #10/29\n|[EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference](https://arxiv.org/abs/2410.12247) \u003cbr\u003e Yulei Qian, Fengcun Li, Xiangyang Ji, Xiaoyu Zhao, Jianchao Tan, Kefeng Zhang, Xunliang Cai | |[Paper](https://arxiv.org/abs/2410.12247)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/Aaronhuang-778/MC-MoE.svg?style=social\u0026label=Star)](https://github.com/Aaronhuang-778/MC-MoE)\u003cbr\u003e[MC-MoE: Mixture Compressor for Mixture-of-Experts LLMs Gains More](https://arxiv.org/abs/2410.06270) \u003cbr\u003e Wei Huang, Yue Liao, Jianhui Liu, Ruifei He, Haoru Tan, Shiming Zhang, Hongsheng Li, Si Liu, Xiaojuan Qi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/Aaronhuang-778/MC-MoE/raw/main/imgs/WX20241009-191322@2x.png\"\u003e |[Github](https://github.com/Aaronhuang-778/MC-MoE) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.06270)|[//]: #10/14\n\n\n\n### Efficient Architecture of LLM\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/NVlabs/hymba.svg?style=social\u0026label=Star)](https://github.com/NVlabs/hymba) ![Publish](https://img.shields.io/badge/Conference-ICLR'25-blue) \u003cbr\u003e[Hymba: A Hybrid-head Architecture for Small Language Models](https://www.arxiv.org/abs/2411.13676) \u003cbr\u003e Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Sunil Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, Yingyan Lin, Jan Kautz, Pavlo Molchanov |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/hymba.png\"\u003e |[Paper](https://www.arxiv.org/pdf/2411.13676)|\n|[![Star](https://img.shields.io/github/stars/mbzuai-oryx/MobiLlama.svg?style=social\u0026label=Star)](https://github.com/mbzuai-oryx/MobiLlama)\u003cbr\u003e:star: [MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT](https://arxiv.org/abs/2402.16840) \u003cbr\u003e Omkar Thawakar, Ashmal Vayani, Salman Khan, Hisham Cholakal, Rao M. Anwer, Michael Felsberg, Tim Baldwin, Eric P. Xing, Fahad Shahbaz Khan |\u003cimg width=\"402\" alt=\"image\" src=\"https://github.com/mbzuai-oryx/MobiLlama/raw/main/images/mobillama_generation.gif\"\u003e |[Github](https://github.com/mbzuai-oryx/MobiLlama) \u003cbr\u003e [Paper](https://arxiv.org/abs/2402.16840) \u003cbr\u003e[Model](https://huggingface.co/MBZUAI/MobiLlama-05B) | [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/XuezheMax/megalodon.svg?style=social\u0026label=Star)](https://github.com/XuezheMax/megalodon)\u003cbr\u003e:star: [Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length](https://arxiv.org/abs/2404.08801) \u003cbr\u003e Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, Chunting Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/megalodon.png\"\u003e |[Github](https://github.com/XuezheMax/megalodon) \u003cbr\u003e [Paper](https://arxiv.org/abs/2404.08801)| [//]: #Recommend\n|[Taipan: Efficient and Expressive State Space Language Models with Selective Attention](https://arxiv.org/abs/2410.18572) \u003cbr\u003e Chien Van Nguyen, Huy Huu Nguyen, Thang M. Pham, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Ryan A. Rossi, Trung Bui, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.18572v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.18572)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/microsoft/SeerAttention.svg?style=social\u0026label=Star)](https://github.com/microsoft/SeerAttention)\u003cbr\u003e[SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs](https://arxiv.org/abs/2410.13276) \u003cbr\u003e Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Hayden Kwok-Hay So, Ting Cao, Fan Yang, Mao Yang |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.13276v1/x4.png\"\u003e |[Github](https://github.com/microsoft/SeerAttention) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.13276)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/TUDa-HWAI/Basis_Sharing.svg?style=social\u0026label=Star)](https://github.com/TUDa-HWAI/Basis_Sharing)\u003cbr\u003e[Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model Compression](https://arxiv.org/abs/2410.03765) \u003cbr\u003e Jingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li, Grace Li Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.03765v1/x1.png\"\u003e |[Github](https://github.com/TUDa-HWAI/Basis_Sharing) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.03765)|[//]: #10/14\n|[Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions](https://arxiv.org/abs/2410.06577) \u003cbr\u003e Zhihao He, Hang Yu, Zi Gong, Shizhan Liu, Jianguo Li, Weiyao Lin |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.06577v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2410.06577)|[//]: #10/14\n|[Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers](https://arxiv.org/abs/2506.01215) \u003cbr\u003e Woomin Song, Sai Muralidhar Jayanthi, Srikanth Ronanki, Kanthashree Mysore Sathyendra, Jinwoo Shin, Aram Galstyan, Shubham Katiyar, Sravan Babu Bodapati |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/reform.png\"\u003e |[Paper](https://arxiv.org/abs/2506.01215)|[//]: #06/03\n\n\n### KV Cache Compression\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|:star: [Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs](https://arxiv.org/abs/2310.01801) \u003cbr\u003e Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, Jianfeng Gao |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/FastGen.png\"\u003e |[Paper](https://arxiv.org/abs/2310.01801)| [//]: #Recommend\n|[ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression](https://arxiv.org/abs/2412.03213) \u003cbr\u003e Guangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang, Minyi Guo |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.03213v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.03213)|[//]: #12/09\n|[Unifying KV Cache Compression for Large Language Models with LeanKV](https://arxiv.org/abs/2412.03131) \u003cbr\u003e Yanqi Zhang, Yuwei Hu, Runyuan Zhao, John C.S. Lui, Haibo Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.03131v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2412.03131)|[//]: #12/09\n|[Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity](https://arxiv.org/abs/2412.02252) \u003cbr\u003e Da Ma, Lu Chen, Situo Zhang, Yuxun Miao, Su Zhu, Zhi Chen, Hongshen Xu, Hanqi Li, Shuai Fan, Lei Pan, Kai Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.02252v1/extracted/6041612/figs/intro.png\"\u003e |[Paper](https://arxiv.org/abs/2412.02252)|[//]: #12/09\n|[MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache](https://arxiv.org/abs/2411.18077) \u003cbr\u003e Akshat Sharma, Hangliang Ding, Jianping Li, Neel Dani, Minjia Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.18077v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.18077)|[//]: #12/07\n|[TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection](https://arxiv.org/abs/2411.02886) \u003cbr\u003e Wei Wu, Zhuoshi Pan, Chao Wang, Liyi Chen, Yunchu Bai, Kun Fu, Zheng Wang, Hui Xiong |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.02886v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.02886)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/FYYFU/HeadKV.svg?style=social\u0026label=Star)](https://github.com/FYYFU/HeadKV)\u003cbr\u003e[Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning](https://arxiv.org/abs/2410.19258) \u003cbr\u003e Yu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong, Yue Dong, Wen Xiao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/FYYFU/HeadKV/raw/main/main.png\"\u003e |[Github](https://github.com/FYYFU/HeadKV) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.19258)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/JunqiZhao888/buzz-llm.svg?style=social\u0026label=Star)](https://github.com/JunqiZhao888/buzz-llm)\u003cbr\u003e[BUZZ: Beehive-structured Sparse KV Cache with Segmented Heavy Hitters for Efficient LLM Inference](https://arxiv.org/abs/2410.23079) \u003cbr\u003e Junqi Zhao, Zhijin Fang, Shu Li, Shaohui Yang, Shichao He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.23079v1/x1.png\"\u003e |[Github](https://github.com/JunqiZhao888/buzz-llm) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.23079)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/whyNLP/LCKV.svg?style=social\u0026label=Star)](https://github.com/whyNLP/LCKV)\u003cbr\u003e[A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference](https://arxiv.org/abs/2410.14442) \u003cbr\u003e You Wu, Haoyi Wu, Kewei Tu |\u003cimg width=\"202\" alt=\"image\" src=\"figures/cross-layer-kv.png\"\u003e |[Github](https://github.com/whyNLP/LCKV) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.14442)|[//]: #10/30\n|[Lossless KV Cache Compression to 2%](https://arxiv.org/abs/2410.15252) \u003cbr\u003e Zhen Yang, J.N.Han, Kan Wu, Ruobing Xie, An Wang, Xingwu Sun, Zhanhui Kang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.15252v1/extracted/5937225/images/CLLA_Overview.png\"\u003e |[Paper](https://arxiv.org/abs/2410.15252)|[//]: #10/30\n|[MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection](https://arxiv.org/abs/2410.14731) \u003cbr\u003e Bokai Lin, Zihao Zeng, Zipeng Xiao, Siqi Kou, Tianqi Hou, Xiaofeng Gao, Hao Zhang, Zhijie Deng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.14731v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.14731)|[//]: #10/30\n|[![Star](https://img.shields.io/github/stars/iankur/vqllm.svg?style=social\u0026label=Star)](https://github.com/iankur/vqllm)\u003cbr\u003e[Residual vector quantization for KV cache compression in large language model](https://arxiv.org/abs/2410.15704) \u003cbr\u003e Ankur Kumar | |[Github](https://github.com/iankur/vqllm) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.15704)|[//]: #10/30\n|[![Star](https://img.shields.io/github/stars/yangyifei729/KVSharer.svg?style=social\u0026label=Star)](https://github.com/yangyifei729/KVSharer)\u003cbr\u003e[KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing](https://arxiv.org/abs/2410.18517) \u003cbr\u003e Yifei Yang, Zouying Cao, Qiguang Chen, Libo Qin, Dongjie Yang, Hai Zhao, Zhi Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/yangyifei729/KVSharer/raw/main/img/main_fig.jpg\"\u003e |[Github](https://github.com/yangyifei729/KVSharer) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.18517)|[//]: #10/29\n|[LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy](https://arxiv.org/abs/2410.03111) \u003cbr\u003e Rongzhi Zhang, Kuang Wang, Liyuan Liu, Shuohang Wang, Hao Cheng, Chao Zhang, Yelong Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/LoRC.png\"\u003e |[Paper](https://arxiv.org/abs/2410.03111)|[//]: #10/14\n|[SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation](https://arxiv.org/abs/2410.03960) \u003cbr\u003e Aurick Qiao, Zhewei Yao, Samyam Rajbhandari, Yuxiong He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.03960v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.03960)|[//]: #10/14\n|[![Publish](https://img.shields.io/badge/Conference-ICML'24-blue)]()\u003cbr\u003e[Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference](https://arxiv.org/abs/2403.09636) \u003cbr\u003e Piotr Nawrot, Adrian Łańcucki, Marcin Chochowski, David Tarjan, Edoardo M. Ponti |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/DMC.png\"\u003e |[Paper](https://arxiv.org/abs/2403.09636)|[//]: #10/02\n|[KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head](https://arxiv.org/abs/2410.00161) \u003cbr\u003e Isaac Rehg |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.00161v1/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2410.00161)|[//]: #10/02\n|[![Star](https://img.shields.io/github/stars/FFY0/AdaKV.svg?style=social\u0026label=Star)](https://github.com/FFY0/AdaKV)\u003cbr\u003e[Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference](https://arxiv.org/abs/2407.11550) \u003cbr\u003e Yuan Feng, Junlin Lv, Yukun Cao, Xike Xie, S. Kevin Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/adakv.png\"\u003e |[Github](https://github.com/FFY0/AdaKV) \u003cbr\u003e [Paper](https://arxiv.org/abs/2407.11550)|[//]: #10/13\n|[![Star](https://img.shields.io/github/stars/snu-mllab/Context-Memory.svg?style=social\u0026label=Star)](https://github.com/snu-mllab/Context-Memory) [![Publish](https://img.shields.io/badge/Conference-ICLR'24-blue)]() \u003cbr\u003e[Compressed Context Memory for Online Language Model Interaction](https://arxiv.org/abs/2312.03414) \u003cbr\u003e Jang-Hyun Kim, Junyoung Yeom, Sangdoo Yun, Hyun Oh Song |\u003cimg width=\"902\" alt=\"image\" src=\"figures/CCM.png\"\u003e |[Github](https://github.com/snu-mllab/Context-Memory) \u003cbr\u003e [Paper](https://arxiv.org/abs/2312.03414)|[//]: #10/13\n|[![Star](https://img.shields.io/github/stars/snu-mllab/KVzip.svg?style=social\u0026label=Star)](https://github.com/snu-mllab/KVzip)\u003cbr\u003e[KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction](https://arxiv.org/abs/2505.23416) \u003cbr\u003e Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/kvzip.png\"\u003e |[Github](https://github.com/snu-mllab/KVzip) [Paper](https://arxiv.org/abs/2505.23416)|[//]: #05/29\n\n\n### Text Compression\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/microsoft/LLMLingua.svg?style=social\u0026label=Star)](https://github.com/microsoft/LLMLingua)[![Publish](https://img.shields.io/badge/Conference-EMNLP'23-blue)]()\u003cbr\u003e:star: [LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models](https://arxiv.org/abs/2310.05736) \u003cbr\u003e Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/microsoft/LLMLingua/blob/main/images/LLMLingua_framework.png\"\u003e |[Github](https://github.com/microsoft/LLMLingua) \u003cbr\u003e [Paper](https://arxiv.org/abs/2310.05736)| [//]: #Recommend\n|[![Star](https://img.shields.io/github/stars/alipay/L3TC-leveraging-rwkv-for-learned-lossless-low-complexity-text-compression.svg?style=social\u0026label=Star)](https://github.com/alipay/L3TC-leveraging-rwkv-for-learned-lossless-low-complexity-text-compression)\u003cbr\u003e[L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression](https://arxiv.org/abs/2412.16642) \u003cbr\u003e Junxuan Zhang, Zhengxue Cheng, Yan Zhao, Shihao Wang, Dajiang Zhou, Guo Lu, Li Song |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.16642v2/x2.png\"\u003e |[Github](https://github.com/alipay/L3TC-leveraging-rwkv-for-learned-lossless-low-complexity-text-compression) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.16642)|[//]: #12/30\n|[![Star](https://img.shields.io/github/stars/NL2G/promptoptme.svg?style=social\u0026label=Star)](https://github.com/NL2G/promptoptme)\u003cbr\u003e[PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics](https://arxiv.org/abs/2412.16120) \u003cbr\u003e Daniil Larionov, Steffen Eger |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.16120v1/x1.png\"\u003e |[Github](https://github.com/NL2G/promptoptme) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.16120)|[//]: #12/30\n|[![Star](https://img.shields.io/github/stars/microsoft/LLMLingua.svg?style=social\u0026label=Star)](https://github.com/microsoft/LLMLingua)\u003cbr\u003e:star: [LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression](https://arxiv.org/abs/2310.06839) \u003cbr\u003e Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/longllmlingua.png\"\u003e |[Github](https://github.com/microsoft/LLMLingua) \u003cbr\u003e [Paper](https://arxiv.org/abs/2310.06839)| [//]: #Recommend\n|[A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression](https://arxiv.org/abs/2412.17483) \u003cbr\u003e Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu, Zhicheng Dou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.17483v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.17483)|[//]: #12/30\n|[JPPO: Joint Power and Prompt Optimization for Accelerated Large Language Model Services](https://arxiv.org/abs/2411.18010) \u003cbr\u003e Feiran You, Hongyang Du, Kaibin Huang, Abbas Jamalipour |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.18010v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.18010)|[//]: #12/07\n|[![Star](https://img.shields.io/github/stars/kaistai/GenPI.svg?style=social\u0026label=Star)](https://github.com/kaistai/GenPI)\u003cbr\u003e[Generative Prompt Internalization](https://arxiv.org/abs/2411.15927) \u003cbr\u003e Haebin Shin, Lei Ji, Yeyun Gong, Sungdong Kim, Eunbi Choi, Minjoon Seo |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/GCD.png\"\u003e |[Github](https://github.com/kaistai/GenPI) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.15927)|[//]: #12/02\n|[![Star](https://img.shields.io/github/stars/noelkelias/multitok.svg?style=social\u0026label=Star)](https://github.com/noelkelias/multitok)\u003cbr\u003e[MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression](https://arxiv.org/abs/2410.21548) \u003cbr\u003e Noel Elias, Homa Esfahanizadeh, Kaan Kale, Sriram Vishwanath, Muriel Medard |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.21548v1/extracted/5960495/Figures/MultiTok.png\"\u003e |[Github](https://github.com/noelkelias/multitok) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.21548)|[//]: #11/17\n|[![Publish](https://img.shields.io/badge/Conference-EMNLP'24%20Findings-blue)]()\u003cbr\u003e[Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability](https://arxiv.org/abs/2410.11786) \u003cbr\u003e Tsz Ting Chung, Leyang Cui, Lemao Liu, Xinting Huang, Shuming Shi, Dit-Yan Yeung |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.11786v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.11786)|[//]: #10/21\n|[![Publish](https://img.shields.io/badge/Conference-EMNLP'24%20Findings-blue)]()\u003cbr\u003e[From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression](https://arxiv.org/abs/2410.04139) \u003cbr\u003e Eunseong Choi, Sunkyung Lee, Minjin Choi, June Park, Jongwuk Lee |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.04139v1/extracted/5902409/Figures/fig_R2C_framework_2col_v4.png\"\u003e |[Paper](https://arxiv.org/abs/2410.04139)|[//]: #10/14\n|[Perception Compressor:A training-free prompt compression method in long context scenarios](https://arxiv.org/abs/2409.19272) \u003cbr\u003e Jiwei Tang, Jin Xu, Tingwei Lu, Hai Lin, Yiming Zhao, Hai-Tao Zheng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2409.19272v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2409.19272)|[//]: #10/02\n| [![Star](https://img.shields.io/github/stars/Workday/cpc.svg?style=social\u0026label=Star)](https://github.com/Workday/cpc)![Publish](https://img.shields.io/badge/Conference-AAAI'25-blue)\u003cbr\u003e[Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference](https://arxiv.org/abs/2409.01227) \u003cbr\u003e Barys Liskavets, Maxim Ushakov, Shuvendu Roy, Mark Klibanov, Ali Etemad, Shane Luke |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2409.01227v3/x1.png\"\u003e |[Github](https://github.com/Workday/cpc) \u003cbr\u003e [Paper](https://arxiv.org/abs/2409.01227)|[//]: #12/30\n| [Task-agnostic Prompt Compression with Context-aware Sentence Embedding and Reward-guided Task Descriptor](https://arxiv.org/abs/2502.13374v1) \u003cbr\u003e Barys Liskavets, Shuvendu Roy, Maxim Ushakov, Mark Klibanov, Ali Etemad, Shane Luke |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.13374v1/x2.png\"\u003e | [Paper](https://arxiv.org/abs/2502.13374v1)|[//]: #12/30\n\n### Low-Rank Decomposition\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24-blue)]()\u003cbr\u003e[ESPACE: Dimensionality Reduction of Activations for Model Compression](https://arxiv.org/abs/2410.05437) \u003cbr\u003e Charbel Sakr, Brucek Khailany |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/ESPACE.png\"\u003e |[Paper](https://arxiv.org/abs/2410.05437)|[//]: #10/14\n|[![Star](https://img.shields.io/github/stars/selfsupervised-ai/Natural-GaLore.svg?style=social\u0026label=Star)](https://github.com/selfsupervised-ai/Natural-GaLore)\u003cbr\u003e[Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning](https://arxiv.org/abs/2410.16029) \u003cbr\u003e Arijit Das | |[Github](https://github.com/selfsupervised-ai/Natural-GaLore) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.16029)|[//]: #10/30\n|[CompAct: Compressed Activations for Memory-Efficient LLM Training](https://arxiv.org/abs/2410.15352) \u003cbr\u003e Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.15352v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.15352)|[//]: #10/30\n\n### Hardware/System/Serving\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[KunServe: Elastic and Efficient Large Language Model Serving with Parameter-centric Memory Management](https://arxiv.org/abs/2412.18169) \u003cbr\u003e Rongxin Cheng, Yifan Peng, Yuxin Lai, Xingda Wei, Rong Chen, Haibo Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.18169v2/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2412.18169)|[//]: #12/30\n|[FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving](https://arxiv.org/abs/2411.18424) \u003cbr\u003e Ao Shen, Zhiyao Li, Mingyu Gao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.18424v1/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2411.18424)|[//]: #12/07\n|[CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration](https://arxiv.org/abs/2411.02829) \u003cbr\u003e Hongpeng Jin, Yanzhao Wu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.02829v1/extracted/5978301/images/method_overview_sm.png\"\u003e |[Paper](https://arxiv.org/abs/2411.02829)|[//]: #11/18\n|[Ripple: Accelerating LLM Inference on Smartphones with Correlation-Aware Neuron Management](https://arxiv.org/abs/2410.19274) \u003cbr\u003e Tuowei Wang, Ruwen Fan, Minxing Huang, Zixu Hao, Kun Li, Ting Cao, Youyou Lu, Yaoxue Zhang, Ju Ren |\u003cimg width=\"302\" alt=\"image\" src=\"https://arxiv.org/html/2410.19274v2/x7.png\"\u003e |[Paper](https://arxiv.org/abs/2410.19274)|[//]: #11/17\n|[![Publish](https://img.shields.io/badge/Conference-ICCAD'24-blue)]()\u003cbr\u003e[ALISE: Accelerating Large Language Model Serving with Speculative Scheduling](https://arxiv.org/abs/2410.23537) \u003cbr\u003e Youpeng Zhao, Jun Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.23537v1/extracted/5967257/imgs/b1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.23537)|[//]: #11/17\n|[EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models](https://arxiv.org/abs/2410.15332) \u003cbr\u003e Junhao Hu, Wenrui Huang, Haoyi Wang, Weidong Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, Tao Xie |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.15332v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2410.15332)|[//]: #10/30\n|[![Publish](https://img.shields.io/badge/Conference-NeurIPS'24-blue)]()\u003cbr\u003e[SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training](https://arxiv.org/abs/2410.15526) \u003cbr\u003e Jinda Jia, Cong Xie, Hanlin Lu, Daoce Wang, Hao Feng, Chengming Zhang, Baixi Sun, Haibin Lin, Zhi Zhang, Xin Liu, Dingwen Tao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.15526v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.15526)|[//]: #10/30\n|[FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs](https://arxiv.org/abs/2410.16663) \u003cbr\u003e Haoran Lin, Xianzhi Yu, Kang Zhao, Lu Hou, Zongyuan Zhan et al |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.16663v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2410.16663)|[//]: #10/29\n|[POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference](https://arxiv.org/abs/2410.18038) \u003cbr\u003e Aditya K Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachandran Ramjee, Ashish Panwar |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.18038v1/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2410.18038)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/Lizonghang/TPI-LLM.svg?style=social\u0026label=Star)](https://github.com/Lizonghang/TPI-LLM)\u003cbr\u003e[TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices](https://arxiv.org/abs/2410.00531) \u003cbr\u003e Zonghang Li, Wenjiao Feng, Mohsen Guizani, Hongfang Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.00531v1/x4.png\"\u003e |[Github](https://github.com/Lizonghang/TPI-LLM) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.00531)|[//]: #10/02\n\n\n\n### Efficient Fine-tuning\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Publish](https://img.shields.io/badge/Conference-ACL'25%20Findings-blue)]()\u003cbr\u003e[LoRMA: Low-Rank Multiplicative Adaptation for LLMs](https://arxiv.org/abs/2506.07621) \u003cbr\u003e Harsh Bihany, Shubham Patel, Ashutosh Modi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.07621v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2506.07621)|[//]: #06/16\n|[HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization](https://arxiv.org/abs/2411.10696) \u003cbr\u003e Huaqin Zhao, Jiaxi Li, Yi Pan, Shizhe Liang, Xiaofeng Yang, Wei Liu, Xiang Li, Fei Dou, Tianming Liu, Jin Lu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.10696v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.10696)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/LCS2-IIITD/MonteCLoRA.svg?style=social\u0026label=Star)](https://github.com/LCS2-IIITD/MonteCLoRA)\u003cbr\u003e[Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation](https://arxiv.org/abs/2411.04358) \u003cbr\u003e Ayan Sengupta, Vaibhav Seth, Arinjay Pathak, Natraj Raman, Sriram Gopalakrishnan, Tanmoy Chakraborty |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.04358v2/x3.png\"\u003e |[Github](https://github.com/LCS2-IIITD/MonteCLoRA) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.04358)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/selfsupervised-ai/Natural-GaLore.svg?style=social\u0026label=Star)](https://github.com/selfsupervised-ai/Natural-GaLore)\u003cbr\u003e[Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning](https://arxiv.org/abs/2410.16029) \u003cbr\u003e Arijit Das | |[Github](https://github.com/selfsupervised-ai/Natural-GaLore) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.16029)|[//]: #10/30\n|[Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs](https://arxiv.org/abs/2410.19694) \u003cbr\u003e Yifei Zhang, Hao Zhu, Aiwei Liu, Han Yu, Piotr Koniusz, Irwin King |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.19694v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2410.19694)|[//]: #11/18\n|[![Publish](https://img.shields.io/badge/Conference-EMNLP'24%20Findings-blue)]()\u003cbr\u003e[MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning](https://arxiv.org/abs/2410.18035) \u003cbr\u003e Jingfan Zhang, Yi Zhao, Dan Chen, Xing Tian, Huanran Zheng, Wei Zhu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.18035v1/extracted/5949512/em_lora_framework.png\"\u003e |[Paper](https://arxiv.org/abs/2410.18035)|[//]: #10/29\n|[![Star](https://img.shields.io/github/stars/Kowsher/RoCoFT.svg?style=social\u0026label=Star)](https://github.com/Kowsher/RoCoFT)\u003cbr\u003e[RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates](https://arxiv.org/abs/2410.10075) \u003cbr\u003e Md Kowsher, Tara Esmaeilbeig, Chun-Nam Yu, Mojtaba Soltanalian, Niloofar Yousefi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/Kowsher/RoCoFT/blob/main/figures/rocoft.png\"\u003e |[Github](https://github.com/Kowsher/RoCoFT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.10075)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/Kaiseem/IST.svg?style=social\u0026label=Star)](https://github.com/Kaiseem/IST)[![Publish](https://img.shields.io/badge/Conference-EMNLP'24-blue)]()\u003cbr\u003e[Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models](https://arxiv.org/abs/2410.11772) \u003cbr\u003e Kai Yao, Penlei Gao, Lichun Li, Yuan Zhao, Xiaofeng Wang, Wei Wang, Jianke Zhu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.11772v1/x3.png\"\u003e |[Github](https://github.com/Kaiseem/IST) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.11772)|[//]: #10/21\n|[![Publish](https://img.shields.io/badge/Conference-Nature%20Scientific%20Reports-blue)]()\u003cbr\u003e[Parameter-Efficient Fine-Tuning of Large Language Models using Semantic Knowledge Tuning](https://arxiv.org/abs/2410.08598) \u003cbr\u003e Nusrat Jahan Prottasha, Asif Mahmud, Md. Shohanur Islam Sobuj, Prakash Bhat, Md Kowsher, Niloofar Yousefi, Ozlem Ozmen Garibay |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.08598v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.08598)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/xvyaward/qeft.svg?style=social\u0026label=Star)](https://github.com/xvyaward/qeft)[![Publish](https://img.shields.io/badge/Conference-EMNLP'24%20Findings-blue)]()\u003cbr\u003e[QEFT: Quantization for Efficient Fine-Tuning of LLMs](https://arxiv.org/abs/2410.08661) \u003cbr\u003e Changhun Lee, Jun-gyu Jin, Younghyun Cho, Eunhyeok Park |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.08661v1/x2.png\"\u003e |[Github](https://github.com/xvyaward/qeft) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.08661)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/Aofei-Chang/BIPEFT.svg?style=social\u0026label=Star)](https://github.com/Aofei-Chang/BIPEFT)[![Publish](https://img.shields.io/badge/Conference-EMNLP'24%20Findings-blue)]()\u003cbr\u003e[BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models](https://arxiv.org/abs/2410.09079) \u003cbr\u003e Aofei Chang, Jiaqi Wang, Han Liu, Parminder Bhatia, Cao Xiao, Ting Wang, Fenglong Ma |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.09079v1/x1.png\"\u003e |[Github](https://github.com/Aofei-Chang/BIPEFT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.09079)|[//]: #10/21\n|[![Star](https://img.shields.io/github/stars/sayankotor/sparse_grads.svg?style=social\u0026label=Star)](https://github.com/sayankotor/sparse_grads)\u003cbr\u003e[SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers](https://arxiv.org/abs/2410.07383) \u003cbr\u003e Viktoriia Chekalina, Anna Rudenko, Gleb Mezentsev, Alexander Mikhalev, Alexander Panchenko, Ivan Oseledets |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.07383v1/x1.png\"\u003e |[Github](https://github.com/sayankotor/sparse_grads) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.07383)|[//]: #10/13\n|[SpaLLM: Unified Compressive Adaptation of Large Language Models with Sketching](https://arxiv.org/abs/2410.06364) \u003cbr\u003e Tianyi Zhang, Junda Su, Oscar Wu, Zhaozhuo Xu, Anshumali Shrivastava |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.06364v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.06364)|[//]: #10/13\n\n\n\n### Efficient Training\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/neiterman21/LDB.svg?style=social\u0026label=Star)](https://github.com/neiterman21/LDB)\u003cbr\u003e[LayerDropBack: A Universally Applicable Approach for Accelerating Training of Deep Networks](https://arxiv.org/abs/2412.18027) \u003cbr\u003e Evgeny Hershkovitch Neiterman, Gil Ben-Artzi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.18027v1/x1.png\"\u003e |[Github](https://github.com/neiterman21/LDB) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.18027)|[//]: #12/30\n|[AutoMixQ: Self-Adjusting Quantization for High Performance Memory-Efficient Fine-Tuning](https://arxiv.org/abs/2411.13814) \u003cbr\u003e Changhai Zhou, Shiyang Zhang, Yuhua Zhou, Zekai Liu, Shichao Weng |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/AutoMixQ.png\"\u003e |[Paper](https://arxiv.org/abs/2411.13814)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/TsinghuaC3I/LPA.svg?style=social\u0026label=Star)](https://github.com/TsinghuaC3I/LPA)[![Publish](https://img.shields.io/badge/Conference-EMNLP'24-blue)]()\u003cbr\u003e[Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention](https://arxiv.org/abs/2411.02063) \u003cbr\u003e Xingtai Lv, Ning Ding, Kaiyan Zhang, Ermo Hua, Ganqu Cui, Bowen Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.02063v1/x1.png\"\u003e |[Github](https://github.com/TsinghuaC3I/LPA) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.02063)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/NVlabs/COAT.svg?style=social\u0026label=Star)](https://github.com/NVlabs/COAT)\u003cbr\u003e[COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training](https://arxiv.org/abs/2410.19313) \u003cbr\u003e Haocheng Xi, Han Cai, Ligeng Zhu, Yao Lu, Kurt Keutzer, Jianfei Chen, Song Han |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/NVlabs/COAT/blob/main/docs/figs/FP8PrecisionFlow.png\"\u003e |[Github](https://github.com/NVlabs/COAT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.19313)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/wuhouming/BitPipe.svg?style=social\u0026label=Star)](https://github.com/wuhouming/BitPipe)\u003cbr\u003e[BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training](https://arxiv.org/abs/2410.19367) \u003cbr\u003e Houming Wu, Ling Chen, Wenjie Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/wuhouming/BitPipe/raw/main/docs/BitPipe_images/BitPipe-v.svg\"\u003e |[Github](https://github.com/wuhouming/BitPipe) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.19367)|[//]: #11/17\n|[![Star](https://img.shields.io/github/stars/selfsupervised-ai/Natural-GaLore.svg?style=social\u0026label=Star)](https://github.com/selfsupervised-ai/Natural-GaLore)\u003cbr\u003e[Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning](https://arxiv.org/abs/2410.16029) \u003cbr\u003e Arijit Das | |[Github](https://github.com/selfsupervised-ai/Natural-GaLore) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.16029)|[//]: #10/30\n|[CompAct: Compressed Activations for Memory-Efficient LLM Training](https://arxiv.org/abs/2410.15352) \u003cbr\u003e Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster |\u003cimg width=\"202\" alt=\"image\" src=\"https://arxiv.org/html/2410.15352v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.15352)|[//]: #10/30\n\n\n\n### Survey (or Benchmark)\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding](https://arxiv.org/abs/2411.13157) \u003cbr\u003e Hyun Ryu, Eric Kim |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.13157v1/extracted/6012092/figure2.png\"\u003e |[Paper](https://arxiv.org/abs/2411.13157)|[//]: #11/24\n|[![Star](https://img.shields.io/github/stars/argonne-lcf/LLM-Inference-Bench.svg?style=social\u0026label=Star)](https://github.com/argonne-lcf/LLM-Inference-Bench)\u003cbr\u003e[LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators](https://arxiv.org/abs/2411.00136) \u003cbr\u003e Krishna Teja Chitty-Venkata, Siddhisanket Raskar, Bharat Kale, Farah Ferdaus et al | |[Github](https://github.com/argonne-lcf/LLM-Inference-Bench) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.00136)|[//]: #11/18\n|[![Star](https://img.shields.io/github/stars/ZongqianLi/Prompt-Compression-Survey.svg?style=social\u0026label=Star)](https://github.com/ZongqianLi/Prompt-Compression-Survey)\u003cbr\u003e[Prompt Compression for Large Language Models: A Survey](https://arxiv.org/abs/2410.12388) \u003cbr\u003e Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.12388v2/extracted/5933385/Figures/tree_overview.png\"\u003e |[Github](https://github.com/ZongqianLi/Prompt-Compression-Survey) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.12388)|[//]: #10/21\n|[Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective](https://arxiv.org/abs/2410.04466) \u003cbr\u003e Jinhao Li, Jiaming Xu, Shan Huang, Yonghua Chen, Wen Li, Jun Liu, Yaoxiu Lian, Jiayi Pan, Li Ding, Hao Zhou, Guohao Dai |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.04466v1/x4.png\"\u003e |[Paper](https://arxiv.org/abs/2410.04466)|[//]: #10/14\n\n\n\n\n\n\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/horseee%2Fawesome-efficient-llm/projects"}