Awesome-LLM-Inference
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
https://github.com/xlite-dev/Awesome-LLM-Inference
Last synced: 1 day ago
JSON representation
-
🎉Awesome LLM Inference Papers with Codes
-
📙Awesome LLM Inference Papers with Codes
-
📖KV Cache Scheduling/Quantize/Dropping ([©️back👆🏻](#paperlist))
-
📖LLM Algorithmic/Eval Survey ([©️back👆🏻](#paperlist))
-
📖LLM Train/Inference Framework ([©️back👆🏻](#paperlist))
- **DeepSpeed-FastGen 2x vLLM?** - fastgen]](https://github.com/microsoft/DeepSpeed)  |⭐️⭐️ |
- inferflow
-
📖Long Context Attention/KV Cache Optimization ([©️back👆🏻](#paperlist))
-
📖Non Transformer Architecture ([©️back👆🏻](#paperlist))
- **Mamba** - spaces/mamba) |⭐️⭐️ |
-
📖Parallel Decoding/Sampling ([©️back👆🏻](#paperlist))
-
📖Structured Prune/KD/Weight Sparse ([©️back👆🏻](#paperlist))
- **Admm Pruning** - pruning]](https://github.com/fmfi-compbio/admm-pruning) |⭐️ |
- FFSplit
-
📖Weight/Activation Quantize/Compress ([©️back👆🏻](#paperlist))
-
-
📖Contents
-
📖Continuous/In-flight Batching ([©️back👆🏻](#paperlist))
- **DeepSpeed-FastGen 2x vLLM?** - fastgen]](https://github.com/microsoft/DeepSpeed)  |⭐️⭐️ |
- LightSeq
- **Continuous Batching**
- **In-flight Batching** - LLM]](https://github.com/NVIDIA/TensorRT-LLM)  |⭐️⭐️ |
- Splitwise
- SpotServe
- **vTensor** - machine-learning/glake/tree/master/GLakeServe) |⭐️⭐️ |
- LightSeq
- Splitwise
- SpotServe
- Automatic Inference Engine Tuning
- **SJF Scheduling**
- **BatchLLM**
- **DeepSpeed-FastGen 2x vLLM?** - fastgen]](https://github.com/microsoft/DeepSpeed)  |⭐️⭐️ |
-
📖CPU/Single GPU/FPGA/Mobile Inference ([©️back👆🏻](#paperlist))
- FlexGen
- OpenVINO
- LLM CPU Inference - extension-for-transformers]](https://github.com/intel/intel-extension-for-transformers)  |⭐️ |
- LinguaLinked
- FlightLLM
-
📖CPU/Single GPU/FPGA/NPU/Mobile Inference ([©️back👆🏻](#paperlist))
- Transformer-Lite
- **xFasterTransformer**
- Summary
- FlexGen
- OpenVINO
- LLM CPU Inference - extension-for-transformers]](https://github.com/intel/intel-extension-for-transformers)  |⭐️ |
- LinguaLinked
- FlightLLM
- **FastAttention**
- **NITRO** - lab/nitro) |⭐️ |
- **Off Grid** - grid-mobile]](https://github.com/alichherawalla/off-grid-mobile) |⭐️ |
- **llama-cpp-power8** - cpp-power8]](https://github.com/Scottcjn/llama-cpp-power8) |⭐️ |
- **RAM Coffers** - coffers]](https://github.com/Scottcjn/ram-coffers) |⭐️ |
-
📖DeepSeek/Multi-head Latent Attention(MLA) ([©️back👆🏻](#paperlist))
- **DeepSeek-R1** - R1]](https://github.com/deepseek-ai/DeepSeek-R1)  | ⭐️⭐️ |
- **TransMLA**
- **DeepSeek-NSA**
- **FlashMLA** - ai/FlashMLA.svg?style=social) |⭐️⭐️ |
- **DualPipe** - ai/DualPipe.svg?style=social) |⭐️⭐️ |
- **DeepEP** - ai/DeepEP.svg?style=social) |⭐️⭐️ |
- **DeepGEMM** - ai/DeepGEMM.svg?style=social) |⭐️⭐️ |
- **EPLB** - ai/EPLB.svg?style=social) |⭐️⭐️ |
- **3FS** - ai/3FS.svg?style=social) |⭐️⭐️ |
- **推理系统**
- **MHA2MLA** - Ushio/MHA2MLA)  |⭐️⭐️ |
- **X-EcoMLA**
-
📖Disaggregating Prefill and Decoding ([©️back👆🏻](#paperlist))
-
📖Early-Exit/Intermediate Layer Decoding ([©️back👆🏻](#paperlist))
- DeeBERT
- BERxiT
- **LITE**
- **EE-LLM** - LLM]](https://github.com/pan-x-c/EE-LLM)  |⭐️⭐️ |
- **FREE**
- DeeBERT
- **LITE**
- **EE-LLM** - LLM]](https://github.com/pan-x-c/EE-LLM)  |⭐️⭐️ |
- **FREE**
- Skip Attention
- **KOALA**
- FastBERT
- **SkipDecode**
- **EE-Tuning** - Tuning]](https://github.com/pan-x-c/EE-LLM)  |⭐️⭐️ |
-
📖GEMM/Tensor Cores/MMA/Parallel ([©️back👆🏻](#paperlist))
- Microbenchmark
- Tensor Parallel
- **flute**
- Tensor Core
- Intra-SM Parallelism
- FP8
- Tensor Cores
- **LUT TENSOR CORE**
- **MARLIN** - DASLab/marlin) |⭐️⭐️ |
- **SpMM**
- **TEE**
- **HiFloat8**
- **Tensor Cores**
- **cutlass/cute**
- **HADACORE** - labs/applied-ai/tree/main/kernels/cuda/inference/hadamard_transform) |⭐️ |
- **FLASH-ATTENTION RNG**
- **TRITONBENCH**
- **Triton-distributed** - distributed]](https://github.com/ByteDance-Seed/Triton-distributed) |⭐️⭐️ |
- QUICK
-
📖GEMM/Tensor Cores/WMMA/Parallel ([©️back👆🏻](#paperlist))
-
📖IO/FLOPs-Aware/Sparse Attention ([©️back👆🏻](#paperlist))
- Online Softmax
- Hash Attention
- **FlashAttention** - attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️ |
- Online Softmax
- FlashAttention
- FLOP, I/O
- **FlashAttention-2** - attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️ |
- **Flash-Decoding** - attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️ |
- Flash-Decoding++
- SparseGPT - DASLab/sparsegpt)  |⭐️ |
- **GLA**
- SCCA
- **FlashLLM**
- CHAI
- Flash Tree Attention
- DeFT
- MoA - nics/MoA)  | ⭐️ |
- CHAI
- **FlashAttention-3** - attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️ |
- Shared Attention
- Online Softmax
- Hash Attention
- **FlashAttention** - attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️ |
- Online Softmax
- **FlashAttention-2** - attention]](https://github.com/Dao-AILab/flash-attention) |⭐️⭐️ |
- Flash-Decoding++
- SparseGPT - DASLab/sparsegpt)  |⭐️ |
- **GLA**
- SCCA
- **FlashLLM**
- **CHESS**
- INT-FLASHATTENTION - FlashAttention]](https://github.com/INT-FlashAttention2024/INT-FlashAttention)  | ⭐️ |
- **SageAttention** - ml/SageAttention)  | ⭐️⭐️ |
- **SageAttention-2** - ml/SageAttention)  | ⭐️⭐️ |
- **Squeezed Attention**
- **TurboAttention**
- **FFPA** - attn-mma]](https://github.com/DefTruth/ffpa-attn-mma) |⭐️⭐️ |
- **SpargeAttention** - ml/SpargeAttn)  | ⭐️⭐️ |
- **FFPA** - attn-mma]](https://github.com/xlite-dev/ffpa-attn-mma) |⭐️⭐️ |
- **SeerAttention**
- **Slim attention** - ai/transformer-tricks)  | ⭐️⭐️⭐️ |
- **MMInference**
- **Sparse Frontier** - frontier)  | ⭐️⭐️ |
- **Flex Attention** - gym]](https://github.com/pytorch-labs/attention-gym)  | ⭐️⭐️ |
- **FFPA** - attn]](https://github.com/xlite-dev/ffpa-attn) |⭐️⭐️ |
- **SageAttention-3** - ml/SageAttention)  | ⭐️⭐️ |
- **Parallel Encoding** - AI-Lab/APE)  | ⭐️⭐️ |
- **Parallel Encoding** - attention]](https://github.com/TemporaryLoRA/Block-attention)  | ⭐️⭐️ |
-
📖KV Cache Scheduling/Quantize/Dropping ([©️back👆🏻](#paperlist))
- **PagedAttention** - project/vllm) |⭐️⭐️ |
- KV Cache FP8 + WINT4
- MQA
- **GQA**
- QK-Sparse/Dropping Attention - sparse-flash-attention]](https://github.com/epfml/dynamic-sparse-flash-attention) |⭐️ |
- LTP
- KV Cache Compress
- H2O
- **TensorRT-LLM KV Cache FP8** - LLM]](https://github.com/NVIDIA/TensorRT-LLM)  |⭐️⭐️ |
- **Adaptive KV Cache Compress**
- CacheGen
- Prompt Caching
- QAQ - KVCacheQuantization]](https://github.com/ClubieDong/QAQ-KVCacheQuantization)  |⭐️⭐️ |
- DMC
- Keyformer - matrix-ai/keyformer-llm) |⭐️⭐️ |
- Less
- MiKV
- KV Cache Compress with LoRA - Context-Memory]](https://github.com/snu-mllab/Context-Memory)  |⭐️⭐️ |
- FASTDECODE
- Sparsity-Aware KV Caching
- SqueezeAttention
- **Shared Prefixes**
- Chunked Prefills
- **RadixAttention** - project/sglang)  |⭐️⭐️ |
- **ChunkAttention** - attention]](https://github.com/microsoft/chunk-attention)  |⭐️⭐️ |
- GEAR - project/GEAR) |⭐️ |
- SnapKV
- vAttention
- KVCache-1Bit
- KV-Runahead
- ZipCache
- MiniCache
- CacheBlend
- CompressKV
- **DistKV-LLM**
- Prompt Caching
- MemServe
- QAQ - KVCacheQuantization]](https://github.com/ClubieDong/QAQ-KVCacheQuantization)  |⭐️⭐️ |
- DMC
- MLKV - mlkv]](https://github.com/zaydzuhri/pythia-mlkv) |⭐️ |
- **PagedAttention** - project/vllm) |⭐️⭐️ |
- MQA
- **GQA**
- QK-Sparse/Dropping Attention - sparse-flash-attention]](https://github.com/epfml/dynamic-sparse-flash-attention) |⭐️ |
- LTP
- KV Cache Compress
- **Adaptive KV Cache Compress**
- CacheGen
- KV Cache Compress with LoRA - Context-Memory]](https://github.com/snu-mllab/Context-Memory)  |⭐️⭐️ |
- Less
- MiKV
- FASTDECODE
- Sparsity-Aware KV Caching
-
Categories
Sub Categories
📖KV Cache Scheduling/Quantize/Dropping ([©️back👆🏻](#paperlist))
73
📖Weight/Activation Quantize/Compress ([©️back👆🏻](#paperlist))
52
📖IO/FLOPs-Aware/Sparse Attention ([©️back👆🏻](#paperlist))
48
📖Long Context Attention/KV Cache Optimization ([©️back👆🏻](#paperlist))
40
📖Parallel Decoding/Sampling ([©️back👆🏻](#paperlist))
34
📖LLM Algorithmic/Eval Survey ([©️back👆🏻](#paperlist))
23
📖LLM Train/Inference Framework/Design ([©️back👆🏻](#paperlist))
21
📖GEMM/Tensor Cores/MMA/Parallel ([©️back👆🏻](#paperlist))
19
📖Continuous/In-flight Batching ([©️back👆🏻](#paperlist))
14
📖Early-Exit/Intermediate Layer Decoding ([©️back👆🏻](#paperlist))
14
📖CPU/Single GPU/FPGA/NPU/Mobile Inference ([©️back👆🏻](#paperlist))
14
📖Prompt/Context/KV Compression ([©️back👆🏻](#paperlist))
13
📖Structured Prune/KD/Weight Sparse ([©️back👆🏻](#paperlist))
12
📖DeepSeek/Multi-head Latent Attention(MLA) ([©️back👆🏻](#paperlist))
12
📖Multi-GPUs/Multi-Nodes Parallelism ([©️back👆🏻](#paperlist))
10
📖Mixture-of-Experts(MoE) LLM Inference ([©️back👆🏻](#paperlist))
9
📖Non Transformer Architecture ([©️back👆🏻](#paperlist))
7
📖LLM Train/Inference Framework ([©️back👆🏻](#paperlist))
6
📖GEMM/Tensor Cores/WMMA/Parallel ([©️back👆🏻](#paperlist))
6
📖CPU/Single GPU/FPGA/Mobile Inference ([©️back👆🏻](#paperlist))
5
📖VLM/Position Embed/Others ([©️back👆🏻](#paperlist))
5
📖Trending LLM/VLM Topics ([©️back👆🏻](#paperlist))
5
📖Prompt/Context Compression ([©️back👆🏻](#paperlist))
4
📖Disaggregating Prefill and Decoding ([©️back👆🏻](#paperlist))
4
📖Position Embed/Others ([©️back👆🏻](#paperlist))
2
Part of the Elyan Labs Ecosystem
1
Keywords
mlsys
3
sdpa
3
tensor-cores
3
flash-attention
3
cuda
3
attention
3
deepseek
2
deepseek-r1
2
deepseek-v3
2
flash-mla
2
fused-mla
2
mla
2
openai-triton
1
nlp
1
model-serving
1
llm
1
llama
1
gpt
1
deep-learning
1
large-language-models
1
machine-learning-systems
1
natural-language-processing
1
acceleration
1
cogvideox
1
diffusion
1
dit
1
flux
1
transformers
1
wan
1