An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with gqa

A curated list of projects in awesome lists tagged with gqa .

https://github.com/bruce-lee-ly/decoding_attention

Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.

cuda cuda-core decoding-attention flash-attention flashinfer flashmla gpu gqa inference large-language-model llm mha mla mqa multi-head-attention nvidia

Last synced: 19 Aug 2025

https://github.com/the-swarm-corporation/hyena-y

A PyTorch implementation of the Hyena-Y model, a convolution-based multi-hybrid architecture optimized for edge devices.

agents ai attention gqa hyena loguru ml model pytorch ssms tensorflow transformers

Last synced: 14 Sep 2025

https://github.com/sahasourav17/intellianswer

A RAG-based question-answering system that processes user queries using local documents. It extracts relevant information to answer questions, falling back to a large language model when local sources are insufficient, ensuring accurate and contextual responses.

chromadb generative-qa gqa llm local-rag ollama openai rag

Last synced: 07 Mar 2026

https://github.com/bob798/ohmygpt

从0到1手写的中文小型大语言模型 — minimind 风格教学项目:RMSNorm·RoPE·GQA·SwiGLU,完整 tokenizer→pretrain→SFT 管线,单张 RTX 3060 可训。A from-scratch Chinese LLM for learning.

chinese-nlp education from-scratch gqa llm minimind pretraining pytorch rope sft transformer

Last synced: 04 Jul 2026

https://github.com/blackwell-systems/merge-barriers

Merge Barriers in BPE Tokenization: How Tokenizer Design Causally Determines Attention Head Specialization. Paper, 23 eval scripts, 86 result files, tokenizer definitions, 4 model checkpoints.

architecture-independence attention-heads attention-mechanism bpe causal-ablation deep-learning delimiter-specialization gpt-neox gqa llama llm mechanistic-interpretability merge-barriers nlp reproducible-research sentencepiece structured-data subword-tokenization tokenization transformer

Last synced: 08 Jul 2026

https://github.com/pathcosmos/frankenstallm

Korean 3B LLM (pure Transformer) pretrained from scratch on 8× NVIDIA B200 GPUs with SFT + ORPO alignment

flash-attention fp8 gguf gqa korean-llm nvidia-b200 orpo pretraining sft transformer

Last synced: 29 May 2026

https://github.com/joe0731/modelsig

Compare LLM architectures without downloading weights — structural fingerprint & proxy-test advisor for vLLM, TensorRT-LLM, SGLang, ONNX Runtime

architecture fingerprint gqa huggingface inference llama llm mistral model-analysis moe onnx onnxruntime proxy-testing qwen safetensors tensorrt-llm vllm

Last synced: 19 Apr 2026