An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with flashattention

A curated list of projects in awesome lists tagged with flashattention .

https://github.com/manishklach/mlx-metal-kernels

Experimental MLX custom Metal kernels for Apple Silicon — fast attention, decode, KV-cache, and future Mac GPU inference primitives.

apple-gpu apple-silicon attention custom-kernels deep-learning flashattention gpu-kernels inference kv-cache llm llm-inference machine-learning macos metal metal-kernels mlx mps python transformers

Last synced: 01 Jul 2026

https://github.com/naidezhujimo/triton-flashattention

This repository contains multiple implementations of Flash Attention optimized with Triton kernels, showcasing progressive performance improvements through hardware-aware optimizations. The implementations range from basic block-wise processing to advanced techniques like FP8 quantization and prefetching

attention flashattention triton

Last synced: 09 Apr 2025

https://github.com/manishklach/ghostkv-lab

Research harness for evaluating query-time bounded elimination of reconstructable KV-cache witnesses in long-context transformer inference workloads. Related provisional filing: IN 202641062451.

ai-infrastructure attention-optimization cxl flashattention gpu-memory kv-cache llm-inference long-context long-context-inference memory-systems systems-research transformer transformer-memory transformer-optimization

Last synced: 09 Jun 2026