Projects in Awesome Lists tagged with flashattention
A curated list of projects in awesome lists tagged with flashattention .
https://github.com/egaoharu-kensei/flash-attention-triton
Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode
ampere attention attention-mechanism blackwell deep-learning flash-attention flash-attention-2 flashattention gpu hopper implementation implementation-from-scratch large-language-models llm machine-learning optimization pytorch transfromers triton turing
Last synced: 15 Jan 2026
https://github.com/manishklach/mlx-metal-kernels
Experimental MLX custom Metal kernels for Apple Silicon — fast attention, decode, KV-cache, and future Mac GPU inference primitives.
apple-gpu apple-silicon attention custom-kernels deep-learning flashattention gpu-kernels inference kv-cache llm llm-inference machine-learning macos metal metal-kernels mlx mps python transformers
Last synced: 01 Jul 2026
https://github.com/naidezhujimo/triton-flashattention
This repository contains multiple implementations of Flash Attention optimized with Triton kernels, showcasing progressive performance improvements through hardware-aware optimizations. The implementations range from basic block-wise processing to advanced techniques like FP8 quantization and prefetching
attention flashattention triton
Last synced: 09 Apr 2025
https://github.com/manishklach/ghostkv-lab
Research harness for evaluating query-time bounded elimination of reconstructable KV-cache witnesses in long-context transformer inference workloads. Related provisional filing: IN 202641062451.
ai-infrastructure attention-optimization cxl flashattention gpu-memory kv-cache llm-inference long-context long-context-inference memory-systems systems-research transformer transformer-memory transformer-optimization
Last synced: 09 Jun 2026