{"id":102397,"url":"https://github.com/kcxain/Awesome-LLM4Kernel","name":"Awesome-LLM4Kernel","description":"LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development","projects_count":76,"last_synced_at":"2026-07-30T08:00:23.983Z","repository":{"id":328793902,"uuid":"1112269024","full_name":"kcxain/Awesome-LLM4Kernel","owner":"kcxain","description":"LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development","archived":false,"fork":false,"pushed_at":"2026-03-31T03:03:40.000Z","size":80,"stargazers_count":77,"open_issues_count":0,"forks_count":3,"subscribers_count":3,"default_branch":"master","last_synced_at":"2026-07-11T08:03:44.775Z","etag":null,"topics":["agent","awesome","cuda","llm","triton"],"latest_commit_sha":null,"homepage":"https://kechang.xin/Awesome-LLM4Kernel/","language":"HTML","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kcxain.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-12-08T11:45:05.000Z","updated_at":"2026-06-30T09:41:28.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/kcxain/Awesome-LLM4Kernel","commit_stats":null,"previous_names":["kcxain/awesome-llm4kernel"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/kcxain/Awesome-LLM4Kernel","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kcxain%2FAwesome-LLM4Kernel","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kcxain%2FAwesome-LLM4Kernel/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kcxain%2FAwesome-LLM4Kernel/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kcxain%2FAwesome-LLM4Kernel/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kcxain","download_url":"https://codeload.github.com/kcxain/Awesome-LLM4Kernel/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kcxain%2FAwesome-LLM4Kernel/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36067204,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-30T02:00:05.956Z","response_time":106,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-01-02T00:00:35.741Z","updated_at":"2026-07-30T08:00:23.983Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["📖 Benchmarks","🔧 Method"],"sub_categories":["Search-based piplines","Agent-based pipelines","Domain specific Models","Agentic RL","Domain-specific Models"],"readme":"# \u003cdiv align=\"center\"\u003eAwesome-LLM4Kernel\u003c/div\u003e\n\n\u003cdiv align=\"center\"\u003e\n\n[![Awesome](https://awesome.re/badge.svg)](https://awesome.re)\n[![Paper](https://img.shields.io/badge/Paper-75-green.svg)](https://github.com/kcxain/Awesome-LLM4Kernel)\n[![Last Commit](https://img.shields.io/github/last-commit/kcxain/Awesome-LLM4Kernel)](https://github.com/kcxain/Awesome-LLM4Kernel)\n[![Website](https://img.shields.io/badge/Website-Live-orange)](https://kechang.xin/Awesome-LLM4Kernel/)\n[![Contribution Welcome](https://img.shields.io/badge/Contributions-welcome-blue)]()\n\n\u003c/div\u003e\n\n\nGPU kernels are central to modern compute stacks and directly determine training and inference efficiency. Kernel development is difficult because it requires hardware expertise and iterative refinement with multi step tool feedback. Since Stanford released KernelBench in February 2025, the LLM4Kernel field has grown rapidly, with increasing interest in using large language models to support or automate kernel generation, optimization, and verification.\n\nThis project provides a continuous and comprehensive survey of the field, covering both benchmarks and methods. On the methodological side, we categorize existing work into four major directions:\n\n - Search-based piplines\n - Agent-based pipelines\n - Domain-specific Models\n - Agentic RL\n\nWe include all relevant top conference papers, arXiv preprints, open source projects, technical reports, and blogs, aiming to build the most complete resource hub for LLM4Kernel research.\n\nOnline page: https://kechang.xin/Awesome-LLM4Kernel/\n\n## 📖 Benchmarks\n\n- **SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits** [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.19173) [![Code](https://img.shields.io/github/stars/NVIDIA/SOL-ExecBench)](https://github.com/NVIDIA/SOL-ExecBench)  \n\t- Edward Lin, Sahil Modi, Siva Kumar Sastry Hari, Qijing Huang, Zhifan Ye, Nestor Qin, Fengzhe Zhou, Yuan Zhang, Jingquan Wang, Sana Damani, Dheeraj Peri, Ouye Xie, Aditya Kane, Moshe Maor, Michael Behar, Triston Cao, Rishabh Mehta, Vartika Singh, Vikram Sharma Mailthody, Terry Chen, Zihao Ye, Hanfeng Chen, Tianqi Chen, Vinod Grover, Wei Chen, Wei Liu, Eric Chung, Luis Ceze, Roger Bringmann, Cyril Zeller, Michael Lightstone, Christos Kozyrakis, Humphrey Shi\n\t- **Institution:** NVIDIA\n\t- **Task:** CUDA Kernel Optimization Benchmarking\n\n- **FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems** [![Paper](https://img.shields.io/badge/arXiv-26.01-red)](https://arxiv.org/abs/2601.00227v1) [![Code](https://img.shields.io/github/stars/flashinfer-ai/flashinfer-bench)](https://github.com/flashinfer-ai/flashinfer-bench)  \n\t- Shanli Xing, Yiyan Zhai, Alexander Jiang, Yixin Dong, Yong Wu, Zihao Ye, Charlie Ruan, Yingyi Huang, Yineng Zhang, Liangsheng Yin, Aksara Bayyapu, Luis Ceze, Tianqi Chen\n\t- **Institution:** University of Washington, Carnegie Mellon University, NVIDIA\n\t- **Task:** CUDA/Triton Optimization\n\n- **KernelBench: Can LLMs Write Efficient GPU Kernels?** [![Paper](https://img.shields.io/badge/ICML-25-green)](https://arxiv.org/pdf/2502.10517) [![Code](https://img.shields.io/github/stars/ScalingIntelligence/KernelBench)](https://github.com/ScalingIntelligence/KernelBench)  \n\t- Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, Azalia Mirhoseini  \n\t- **Institution:** Stanford University  \n\t- **Task:** Torch -\u003e CUDA  \n\n- **TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators** [![Paper](https://img.shields.io/badge/ACL_findings-25-green)](https://aclanthology.org/2025.findings-acl.1183.pdf) [![Code](https://img.shields.io/github/stars/thunlp/TritonBench)](https://github.com/thunlp/TritonBench)  \n\t- Jianling Li, ShangZhan Li, Zhenye Gao, Qi Shi, Yuxuan Li, Zefan Wang, Jiacheng Huang, WangHaojie WangHaojie, Jianrong Wang, Xu Han, Zhiyuan Liu, Maosong Sun  \n\t- **Institution:** Tianjin University, Tsinghua University  \n\t- **Task:** Torch | NL -\u003e Triton  \n\n- **ComputeEval: Evaluating Large Language Models for CUDA Code Generation** [![Code](https://img.shields.io/github/stars/NVIDIA/compute-eval)](https://github.com/NVIDIA/compute-eval)  \n\t- **Institution:** NVIDIA  \n\t- **Task:** NL -\u003e CUDA  \n\n- **BackendBench: An Evaluation Suite for Testing How Well LLMs and Humans Can Write PyTorch Backends** [![Blog](https://img.shields.io/badge/Blog-Meta-blue)](https://github.com/meta-pytorch/BackendBench/blob/main/docs/correctness.md) [![Code](https://img.shields.io/github/stars/meta-pytorch/BackendBench)](https://github.com/meta-pytorch/BackendBench)  \n\t- **Institution:** Meta  \n\t- **Task:** Torch -\u003e CUDA | Triton  \n\n- **MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation** [![Paper](https://img.shields.io/badge/arXiv-25.07-red)](https://arxiv.org/pdf/2507.17773) [![Code](https://img.shields.io/github/stars/wzzll123/MultiKernelBench)](https://github.com/wzzll123/MultiKernelBench)  \n\t- Zhongzhen Wen, Yinghui Zhang, Zhong Li, Zhongxin Liu, Linna Xie, Tian Zhang  \n\t- **Institution:** Nanjing University  \n\t- **Task:** Torch -\u003e CUDA | Pallas | AscendC  \n\n- **robust-kbench: Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.09-red)](https://arxiv.org/pdf/2509.14279) [![Code](https://img.shields.io/github/stars/SakanaAI/robust-kbench)](https://github.com/SakanaAI/robust-kbench)  \n\t- Robert Tjarko Lange, Qi Sun, Aaditya Prasad, Maxence Faldor, Yujin Tang, David Ha  \n\t- **Institution:** Sakana AI  \n\t- **Task:** Torch -\u003e CUDA\n\n- **gpuFLOPBench: Counting Without Running: Evaluating LLMs’ Reasoning About Code Complexity** [![Paper](https://img.shields.io/badge/arXiv-25.12-red)](https://arxiv.org/abs/2512.04355) [![Code](https://img.shields.io/github/stars/Scientific-Computing-Lab/gpuFLOPBench)](https://github.com/Scientific-Computing-Lab/gpuFLOPBench)  \n\t- Gregory Bolet, Giorgis Georgakoudis, Konstantinos Parasyris, Harshitha Menon, Niranjan Hasabnis, Kirk W. Cameron, Gal Oren\n\t- **Institution:** Stanford University  \n\t- **Task:** CUDA -\u003e FLOPs  \n\n- **RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis** [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.11506)  \n\t- Zhen Bi, Qian Fan, Renjie Liu, Xing Di, Zihao Zhu, Borui Wang, Yiru Chen, Xiaoyi Dong, Rui Liu, Cheng Tan, Nian Liu, Xuhui Fan, Mark Shirman, Gal Oren, Anton Fonin, Konstantinos Parasyris, Yusong Gao, Song Han\n\t- **Institution:** Huzhou University, Banbu AI Foundation, Chinese Academy of Sciences, Carnegie Mellon University, University of Edinburgh\n\t- **Task:** On-device LLM -\u003e Roofline\n\n- **Can Large Language Models Predict Parallel Code Performance** [![Paper](https://img.shields.io/badge/Conference-25-green)](https://dl.acm.org/doi/abs/10.1145/3731545.3743645) [![Code](https://img.shields.io/github/stars/Scientific-Computing-Lab/ParallelCodeEstimation)](https://github.com/Scientific-Computing-Lab/ParallelCodeEstimation)  \n\t- Gregory Bolet, Giorgis Georgakoudis, Harshitha Menon, Konstantinos Parasyris, Niranjan Hasabnis, Hayden Estes, Kirk W. Cameron, Gal Oren\n\t- **Institution:** Virginia Tech, LLNL, Code Metal, Stanford University, Technion\n\t- **Task:** CUDA | OpenMP -\u003e Roofline Class\n\n- **NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers** [![Paper](https://img.shields.io/badge/arXiv-25.07-red)](https://arxiv.org/abs/2507.14403) [![Code](https://img.shields.io/github/stars/AMDResearch/NPUEval)](https://github.com/AMDResearch/NPUEval)  \n\t- Sarunas Kalade, Graham Schelle\n\t- **Institution:** Advanced Micro Devices\n\t- **Task:** NL -\u003e NPU Kernel\n\n- **CUDABench: Benchmarking LLMs for Text-to-CUDA Generation** [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.02236) [![Code](https://img.shields.io/github/stars/CUDA-Bench/CUDABench)](https://github.com/CUDA-Bench/CUDABench)  \n\t- Jiace Zhu, Wentao Chen, Qi Fan, Zhixing Ren, Junying Wu, Xing Zhe Chai, Chotiwit Rungrueangwutthinon, Yehan Ma, An Zou\n\t- **Institution:** Shanghai Jiao Tong University\n\t- **Task:** NL -\u003e CUDA\n\n- **KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware** [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.08721)  \n\t- Jiayi Nie, Haoran Wu, Yao Lai, Zeyu Cao, Cheng Zhang, Binglei Lou, Erwei Wang, Jianyi Cheng, Timothy M. Jones, Robert Mullins, Rika Antonova, Yiren Zhao\n\t- **Task:** NL -\u003e Accelerator Kernel\n\n- **KernelBook: PyTorch to Triton Code Translation Dataset** [![Dataset](https://img.shields.io/badge/Dataset-HuggingFace-yellow)](https://huggingface.co/datasets/GPUMODE/KernelBook)  \n\t- Sahan Paliskara, Mark Saroufim\n\t- **Institution:** GPUMODE\n\t- **Task:** Torch -\u003e Triton\n\n## 🔧 Method\n\n### Search-based piplines\n\n- **KernelFoundry: Hardware-aware evolutionary GPU kernel optimization**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.12440)  \n\t- Nina Wiedemann, Quentin Leboutet, Michael Paulitsch, Diana Wofk, Benjamin Ummenhofer  \n\t- **Institution:** University of Freiburg, Infineon Technologies AG  \n\t- **Task:** CUDA | SYCL Kernel Optimization\n\n- **OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.12305)  \n\t- Arijit Bhattacharjee, Heng Ping, Son Vu Le, Paul Bogdan, Nesreen K. Ahmed, Ali Jannesari  \n\t- **Institution:** Iowa State University  \n\t- **Task:** NL -\u003e CUDA + MCTS Optimization\n\n- **K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.19128) [![Code](https://img.shields.io/github/stars/caoshiyi/K-Search)](https://github.com/caoshiyi/K-Search)\n\t- Shiyi Cao, Ziming Mao, Joseph E. Gonzalez, Ion Stoica  \n\t- **Institution:** UC Berkeley  \n\t- **Task:** CUDA Optimization\n\n- **KernelBand: Boosting LLM-based Kernel Optimization with a Hierarchical and Hardware-aware Multi-armed Bandit** [![Paper](https://img.shields.io/badge/aiXiv-25.11-red)](https://arxiv.org/pdf/2511.18868)\n\t- Dezhi Ran, Shuxiao Xie, Mingfang Ji, Ziyue Hua, Mengzhou Wu, Yuan Cao, Yuzhe Guo, Yu Hao, Linyi Li, Yitao Hu, Tao Xie\n\t- **Institution:** Peking University\n\t- **Task:** Torch -\u003e Triton\n\n- **Automating GPU Kernel Generation with DeepSeek-R1 and Inference Time Scaling** [![Blog](https://img.shields.io/badge/Blog-NVIDIA-blue)](https://developer.nvidia.com/blog/automating-gpu-kernel-generation-with-deepseek-r1-and-inference-time-scaling/)  \n\t- Terry Chen, Bing Xu, Kirthi Devleker\n\t- **Institution:** NVIDIA\n\t- **Task:** NL -\u003e CUDA Attention Kernel\n\n- **Tutoring LLM into a Better CUDA Optimizer** [![Paper](https://img.shields.io/badge/Euro--Par-25-green)](https://link.springer.com/chapter/10.1007/978-3-031-99857-7_18) [![Code](https://img.shields.io/github/stars/matyas-brabec/2025-europar-llm)](https://github.com/matyas-brabec/2025-europar-llm)  \n\t- Matyas Brabec, Jiri Klepl, Michal Topfer, Martin Krulis\n\t- **Institution:** Charles University\n\t- **Task:** NL -\u003e CUDA\n\n- **GPU Performance Portability needs Autotuning** [![Paper](https://img.shields.io/badge/arXiv-25.05-red)](https://arxiv.org/abs/2505.03780)  \n\t- Burkhard Ringlein, Thomas Parnell, Radu Stoica\n\t- **Institution:** IBM Research Europe\n\t- **Task:** Triton Attention -\u003e Cross-GPU Optimization\n\n- **TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.12-red)](https://arxiv.org/abs/2512.09196) [![Code](https://img.shields.io/github/stars/RLsys-Foundation/TritonForge)](https://github.com/RLsys-Foundation/TritonForge)  \n\t- Haonan Li, Keyu Man, Partha Kanuparthy, Hanning Chen, Wei Sun, Sreen Tallam, Chenguang Zhu, Kevin Zhu, Zhiyun Qian\n\t- **Task:** Triton Optimization\n\n- **EVOENGINEER: Mastering Automated CUDA Kernel Code Evolution with Large Language Models** [![Paper](https://img.shields.io/badge/arXiv-25.10-red)](https://arxiv.org/abs/2510.03760)  \n\t- Ping Guo, Chenyu Zhu, Siyuan Chen, Fei Liu, Xi Lin, Zhichao Lu, Qingfu Zhang\n\t- **Institution:** City University of Hong Kong\n\t- **Task:** CUDA Optimization\n\n- **From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph** [![Paper](https://img.shields.io/badge/arXiv-25.10-red)](https://arxiv.org/abs/2510.19873) [![Code](https://img.shields.io/github/stars/blacknickwield/ReGraphT)](https://github.com/blacknickwield/ReGraphT)  \n\t- Junfeng Gong, Zhiyi Wei, Junying Chen, Cheng Liu, Huawei Li\n\t- **Institution:** Institute of Computing Technology, Chinese Academy of Sciences, University of Chinese Academy of Sciences, South China University of Technology\n\t- **Task:** Sequential Code -\u003e CUDA\n\n- **MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization** [![Paper](https://img.shields.io/badge/arXiv-26.01-red)](https://arxiv.org/abs/2601.05475)  \n\t- Jiefu Ou, Sapana Chaudhary, Kaj Bostrom, Nathaniel Weir, Shuai Zhang, Huzefa Rangwala, George Karypis\n\t- **Institution:** Johns Hopkins University, Amazon Web Services\n\t- **Task:** CUDA | C++ Optimization\n\n- **AVO: Agentic Variation Operators for Autonomous Evolutionary Search** [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.24517)  \n\t- Terry Chen, Zhifan Ye, Bing Xu, Zihao Ye, Timmy Liu, Ali Hassani, Tianqi Chen, Andrew Kerr, Haicheng Wu, Yang Xu, Yu-Jung Chen, Hanfeng Chen, Aditya Kane, Ronny Krashinsky, Ming-Yu Liu, Vinod Grover, Luis Ceze, Roger Bringmann, John Tran, Wei Liu, Fung Xie, Michael Lightstone, Humphrey Shi  \n\t- **Institution:** NVIDIA  \n\t- **Task:** Attention CUDA Kernel Optimization\n\n### Agent-based pipelines\n\n- **StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning** [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.02637)  \n\t- Shiyang Li, Zijian Zhang, Winson Chen, Yuebo Luo, Mingyi Hong, Caiwen Ding\n\t- **Institution:** University of Minnesota, Twin Cities\n\t- **Task:** Torch -\u003e End-to-End CUDA\n\n- **KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.10085) [![Code](https://img.shields.io/github/stars/0satan0/KernelMem)](https://github.com/0satan0/KernelMem)  \n\t- Qitong Sun, Jun Han, Tianlin Li, Zhe Tang, Sheng Chen, Fei Yang, Aishan Liu, Xianglong Liu, Yang Liu  \n\t- **Institution:** Beihang University, Beijing Academy of Artificial Intelligence  \n\t- **Task:** Torch -\u003e CUDA\n\n- **Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.07169)  \n\t- Yuxuan Han, Meng-Hao Guo, Zhengning Liu, Wenguang Chen, Shi-Min Hu  \n\t- **Institution:** Tsinghua University  \n\t- **Task:** CUDA Optimization\n\n- **CUCo: An Agentic Framework for Compute and Communication Co-design**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.02376)  \n\t- Bodun Hu, Yoga Sri Varshan V, Saurabh Agarwal, Aditya Akella  \n\t- **Institution:** University of Wisconsin-Madison  \n\t- **Task:** CUDA Compute \u003c-\u003e Communication Co-design\n\n- **AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.01-red)](https://arxiv.org/abs/2601.22760)  \n\t- Zhongzhen Wen, Shudi Shao, Zhong Li, Yu Ge, Tongtong Xu, Yuanyi Lin, Tian Zhang  \n\t- **Institution:** Nanjing University  \n\t- **Task:** Torch | NL -\u003e AscendC\n\n- **KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.14293)  \n\t- Kris Shengjun Dong, Sahil Modi, Dima Nikiforov, Sana Damani, Edward Lin, Siva Kumar Sastry Hari, Christos Kozyrakis  \n\t- **Institution:** NVIDIA, UC Berkeley  \n\t- **Task:** CUDA Optimization\n\n- **AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis** [![Paper](https://img.shields.io/badge/aiXiv-25.12-red)](https://arxiv.org/pdf/2512.23424v1) [![Code](https://img.shields.io/github/stars/mindspore-ai/akg)](https://github.com/mindspore-ai/akg/blob/master/aikg/README_CN.md)  \n\t- Jinye Du, Quan Yuan, Zuyao Zhang, Yanzhi Yi, Jiahui Hu, Wangyi Chen, Yiyang Zhu, Qishui Zheng, Wenxiang Zou, Xiangyu Chang, Zuohe Zheng, Zichun Ye, Chao Liu, Shanni Li, Renwei Zhang, Yiping Deng, Xinwei Hu, Xuefeng Jin, Jie Zhao\n\t- **Institution:** Huawei\n\t- **Task:** Torch -\u003e CUDA | Triton | Tilelang | AscendC\n\n- **TritorX: Agentic Operator Generation for ML ASICs** [![Paper](https://img.shields.io/badge/aiXiv-25.12-red)](https://www.arxiv.org/abs/2512.10977)\n\t- Alec M. Hammond, Aram Markosyan, Aman Dontula, Simon Mahns, Zacharias Fisches, Dmitrii Pedchenko, Keyur Muzumdar, Natacha Supper, Mark Saroufim, Joe Isaacson, Laura Wang, Warren Hunt, Kaustubh Gondkar, Roman Levenstein, Gabriel Synnaeve, Richard Li, Jacob Kahn, Ajit Mathews\n\t- **Institution:** Meta\n\t- **Task:** torch ATen Docstring -\u003e Triton\n\n- **KernelFalcon: Autonomous GPU Kernel Generation via Deep Agents** [![Paper](https://img.shields.io/badge/blog-25.11-blue)](https://pytorch.org/blog/kernelfalcon-autonomous-gpu-kernel-generation-via-deep-agents/) [![Code](https://img.shields.io/github/stars/meta-pytorch/KernelAgent)](https://github.com/meta-pytorch/KernelAgent)  \n\t- Laura Wang\n\t- **Institution:** PyTorch Team at Meta\n\t- **Task:** Torch -\u003e Triton\n\n- **STARK: Strategic Team of Agents for Refining Kernels** [![Paper](https://img.shields.io/badge/arXiv-25.10-red)](https://arxiv.org/pdf/2510.16996)\n\t- Juncheng Dong, Yang Yang, Tao Liu, Yang Wang, Feng Qi, Vahid Tarokh, Kaushik Rangadurai, Shuang Yang\n\t- **Institution:** Meta Ranking AI Research\n\t- **Task:** Torch -\u003e CUDA\n\n- **QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach** [![Paper](https://img.shields.io/badge/OSDI-25-green)](https://arxiv.org/abs/2505.02146) [![Code](https://img.shields.io/github/stars/QiMeng-IPRC/QiMeng-Xpiler)](https://github.com/QiMeng-IPRC/QiMeng-Xpiler)  \n\t- Shouyang Dong, Yuanbo Wen, Jun Bi, Di Huang, Jiaming Guo, Jianxing Xu, Ruibai Xu, Xinkai Song, Yifan Hao, Xuehai Zhou, Tianshi Chen, Qi Guo, Yunji Chen\n\t- **Institution:** University of Science and Technology of China, Cambricon Technologies, Institute of Computing Technology, Institute of Software\n\t- **Task:** CUDA \u003c-\u003e BangC \u003c-\u003e Hip \u003c-\u003e VNNI  \n\n- **QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm** [![Paper](https://img.shields.io/badge/ACL-25-green)](https://arxiv.org/abs/2506.12355) [![Code](https://img.shields.io/github/stars/chris-chow/QiMeng-Attention)](https://github.com/chris-chow/QiMeng-Attention)  \n\t- Qirui Zhou, Shaohui Peng, Weiqiang Xiong, Haixin Chen, Yuanbo Wen, Haochen Li, Ling Li, Qi Guo, Yongwei Zhao, Ke Gao, Ruizhi Chen, Yanjun Wu, Chen Zhao, Yunji Chen\n\t- **Institution:** Institute of Software, Institute of Computing Technology  \n\t- **Task:** NL -\u003e CUDA (Attention)  \n\n- **QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives** [![Paper](https://img.shields.io/badge/IJCAI-25-green)](https://arxiv.org/pdf/2505.06302) [![Code](https://img.shields.io/github/stars/zhangxuzhi/QiMeng-TensorOp)](https://github.com/zhangxuzhi/QiMeng-TensorOp)  \n\t- Xuzhi Zhang, Shaohui Peng, Qirui Zhou, Yuanbo Wen, Qi Guo, Ruizhi Chen, Xinguo Zhu, Weiqiang Xiong, Haixin Chen, Congying Ma, Ke Gao, Chen Zhao, Yanjun Wu, Yunji Chen, Ling Li  \n\t- **Institution:** Institute of Computing Technology, Institute of Software \n\t- **Task:** NL -\u003e Hardware-specific Tensor Operators (RISC-V, ARM, GPU)\n\n- **QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models** [![Paper](https://img.shields.io/badge/AAAI-25-green)](https://ojs.aaai.org/index.php/AAAI/article/view/34461) [![Code](https://img.shields.io/github/stars/chris-chow/QiMeng-GEMM)](https://github.com/chris-chow/QiMeng-GEMM)  \n\t- Qirui Zhou, Yuanbo Wen, Ruizhi Chen, Ke Gao, Weiqiang Xiong, Ling Li, Qi Guo, Yanjun Wu, Yunji Chen \n\t- **Institution:** Institute of Computing Technology, Institute of Software\n\t- **Task:** NL -\u003e CUDA (GEMM)\n\n- **GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.06-red)](https://arxiv.org/abs/2506.20807)  \n\t- Martin Andrews, Sam Witteveen\n\t- **Task:** CUDA Optimization\n\n- **Geak: Introducing Triton Kernel AI Agent \u0026 Evaluation Benchmarks** [![Paper](https://img.shields.io/badge/arXiv-25.07-red)](https://arxiv.org/abs/2507.23194) [![Code](https://img.shields.io/github/stars/AMD-AIG-AIMA/GEAK-agent)](https://github.com/AMD-AIG-AIMA/GEAK-agent) [![Eval](https://img.shields.io/github/stars/AMD-AIG-AIMA/GEAK-eval)](https://github.com/AMD-AIG-AIMA/GEAK-eval)  \n\t- Jianghui Wang, Vinay Joshi, Saptarshi Majumder, Xu Chao, Bin Ding, Ziqiong Liu, Pratik Prabhanjan Brahma, Dong Li, Zicheng Liu, Emad Barsoum\n\t- **Task:** NL -\u003e Triton\n\n- **How Many Agents Does it Take to Beat PyTorch? (surprisingly not that much)** [![Blog](https://img.shields.io/badge/Blog-Lossfunk-blue)](https://letters.lossfunk.com/p/how-many-agents-does-it-take-to-beat)  \n\t- Shikhar Mishra, Ayush Nangia\n\t- **Task:** Torch -\u003e CUDA\n\n- **Astra: A Multi-Agent System for GPU Kernel Performance Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.09-red)](https://arxiv.org/abs/2509.07506) [![Code](https://img.shields.io/github/stars/Anjiang-Wei/Astra)](https://github.com/Anjiang-Wei/Astra)  \n\t- Anjiang Wei, Tianran Sun, Yogesh Seenichamy, Hang Song, Anne Ouyang, Azalia Mirhoseini, Ke Wang, Alex Aiken\n\t- **Task:** CUDA Optimization\n\n- **CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.11-red)](https://arxiv.org/abs/2511.01884) [![Code](https://img.shields.io/github/stars/OptimAI-Lab/CudaForge)](https://github.com/OptimAI-Lab/CudaForge)  \n\t- Zijian Zhang, Rong Wang, Shiyang Li, Yuebo Luo, Mingyi Hong, Caiwen Ding\n\t- **Institution:** University of Minnesota, Twin Cities\n\t- **Task:** Torch -\u003e CUDA\n\n- **KForge: Program Synthesis for Diverse AI Hardware Accelerators** [![Paper](https://img.shields.io/badge/arXiv-25.11-red)](https://arxiv.org/abs/2511.13274)  \n\t- **Task:** NL -\u003e Accelerator Kernel\n\n- **The AI CUDA engineer: Agentic CUDA kernel discovery, optimization and composition** [![Report](https://img.shields.io/badge/Report-Sakana%20AI-blue)](https://pub.sakana.ai/static/paper.pdf)  \n\t- **Task:** Torch -\u003e CUDA\n\n- **Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems** [![Paper](https://img.shields.io/badge/arXiv-25.11-red)](https://arxiv.org/abs/2511.16964)  \n\t- **Task:** PyTorch Inference -\u003e CUDA\n\n- **PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.11-red)](https://arxiv.org/abs/2511.06345)  \n\t- **Task:** CUDA Optimization\n\n- **cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution** [![Paper](https://img.shields.io/badge/arXiv-25.12-red)](https://arxiv.org/abs/2512.16465) [![Code](https://img.shields.io/github/stars/champloo2878/cuPilot-Kernels)](https://github.com/champloo2878/cuPilot-Kernels)  \n\t- **Task:** CUDA Optimization\n\n- **AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.11-red)](https://arxiv.org/abs/2511.15915) [![Code](https://img.shields.io/github/stars/zhang677/AccelOpt)](https://github.com/zhang677/AccelOpt)  \n\t- Genghan Zhang, Shaowei Zhu, Anjiang Wei, Zhenyu Song, Allen Nie, Zhen Jia, Nandita Vijaykumar, Yida Wang, Kunle Olukotun\n\t- **Institution:** Stanford University, Amazon Web Services, University of Toronto\n\t- **Task:** NKI -\u003e Trainium Kernel Optimization\n\n- **Adaptive Self-improvement LLM Agentic System for ML Library Development** [![Paper](https://img.shields.io/badge/ICML-25-green)](https://arxiv.org/abs/2502.02534) [![Code](https://img.shields.io/github/stars/zhang677/PCL-liteLLM)](https://github.com/zhang677/PCL-liteLLM)  \n\t- Genghan Zhang, Weixin Liang, Olivia Hsu, Kunle Olukotun\n\t- **Institution:** Stanford University\n\t- **Task:** NL -\u003e ASPL ML Library\n\n### Domain-specific Models\n\n- **InCoder-32B: Code Foundation Model for Industrial Scenarios** [![Paper](https://img.shields.io/badge/arXiv-26.03-red)](https://arxiv.org/abs/2603.16790)  \n\t- Jian Yang, Wei Zhang, Jiajun Wu, Junhang Cheng, Shawn Guo, Haowen Wang, Weicheng Gu, Yaxin Du, Joseph Li, Fanglin Xu, Yizhi Li, Lin Jing, Yuanbo Wang, Yuhan Gao, Ruihao Gong, Chuan Hao, Ran Tao, Aishan Liu, Tuney Zheng, Ganqu Cui, Zhoujun Li, Mingjie Tang, Chenghua Lin, Wayne Xin Zhao, Xianglong Liu, Ming Zhou, Bryan Dai, Weifeng Lv\n\t- **Institution:** Beihang University, iQuest Research, Shanghai Jiao Tong University, ELLIS, University of Manchester, Shanghai Artificial Intelligence Laboratory, Sichuan University, Renmin University of China, Langboat\n\t- **Task:** Code -\u003e GPU Kernel Optimization\n\n- **DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.11715)  \n\t- Haolei Bai, Lingcheng Kong, Xueyi Chen, Jianmian Wang, Zhiqiang Tao, Huan Wang  \n\t- **Institution:** Westlake University  \n\t- **Task:** Torch -\u003e CUDA\n\n- **AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs** [![Paper](https://img.shields.io/badge/arXiv-25.07-red)](https://arxiv.org/abs/2507.05687) [![Code](https://img.shields.io/github/stars/AI9Stars/AutoTriton)](https://github.com/AI9Stars/AutoTriton)  \n\t- Shangzhan Li, Zefan Wang, Ye He, Yuxuan Li, Qi Shi, Jianling Li, Yonggang Hu, Wanxiang Che, Xu Han, Zhiyuan Liu, Maosong Sun\n\t- **Institution:** Tsinghua University\n\t- **Task:** Torch -\u003e Triton\n\n- **QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation** [![Paper](https://img.shields.io/badge/NeurIPS-25-green)](https://arxiv.org/pdf/2506.11153) [![Code](https://img.shields.io/github/stars/QiMeng-IPRC/QiMeng-MuPa)](https://github.com/QiMeng-IPRC/QiMeng-MuPa)  \n\t- Changxin Ke, Rui Zhang, Shuo Wang, Li Ding, Guangli Li, Yuanbo Wen, Shuoming Zhang, Ruiyuan Xu, Jin Qin, Jiaming Guo, Chenxi Wang, Ling Li, Qi Guo, Yunji Chen\n\t- **Institution:** Institute of Computing Technology\n\t- **Task:** C -\u003e CUDA\n\n - **AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units** [![Paper](https://img.shields.io/badge/arXiv-26.01-red)](https://arxiv.org/abs/2601.07160)\n \t- Xinzi Cao, Jianyang Zhai, Pengfei Li, Zhiheng Hu, Cen Yan, Bingxu Mu, Guanghuan Fang, Bin She, Jiayu Li, Yihan Su, Dongyang Tao, Xiansong Huang, Fan Xu, Feidiao Yang, Yao Lu, Chang-Dong Wang, Yutong Lu, Weicheng Xue, Bin Zhou, Yonghong Tian\n\t- **Institution:** Pengcheng Laboratory, HUAWEI, Sun Yat-sen University\n\t- **Task:** Torch -\u003e AscendC\n\n- **CUDA-LLM: LLMs Can Write Efficient CUDA Kernels** [![Paper](https://img.shields.io/badge/arXiv-25.06-red)](https://arxiv.org/abs/2506.09092)  \n\t- Wentao Chen, Jiace Zhu, Qi Fan, Yehan Ma, An Zou\n\t- **Institution:** Shanghai Jiao Tong University\n\t- **Task:** NL -\u003e CUDA\n\n- **KernelLLM** [![Model](https://img.shields.io/badge/Model-HuggingFace-yellow)](https://huggingface.co/facebook/KernelLLM)  \n\t- **Institution:** Meta\n\t- **Task:** Torch -\u003e Triton\n\n- **Scaling LLM Test-Time Compute with Mobile NPU on Smartphones** [![Paper](https://img.shields.io/badge/arXiv-25.09-red)](https://arxiv.org/abs/2509.23324) [![Code](https://img.shields.io/github/stars/haozixu/llama.cpp-npu)](https://github.com/haozixu/llama.cpp-npu) [![Library](https://img.shields.io/github/stars/haozixu/htp-ops-lib)](https://github.com/haozixu/htp-ops-lib)  \n\t- **Task:** LLM Inference -\u003e Mobile NPU\n\n- **CudaLLM: Training Language Models to Generate High-Performance CUDA Kernels** [![Code](https://img.shields.io/github/stars/ByteDance-Seed/cudaLLM)](https://github.com/ByteDance-Seed/cudaLLM) [![Model](https://img.shields.io/badge/Model-HuggingFace-yellow)](https://huggingface.co/ByteDance-Seed/cudaLLM-8B)  \n\t- **Institution:** ByteDance Seed\n\t- **Task:** NL -\u003e CUDA\n\n- **Omniwise: Predicting GPU Kernels Performance with LLMs** [![Paper](https://img.shields.io/badge/arXiv-25.06-red)](https://arxiv.org/abs/2506.20886)  \n\t- Zixian Wang, Cole Ramos, Muhammad A. Awad, Keith Lowery\n\t- **Institution:** University of Illinois Urbana-Champaign, AMD\n\t- **Task:** CUDA -\u003e Performance Metrics\n\n- **ConCuR: Conciseness Makes State-of-the-Art Kernel Generation** [![Paper](https://img.shields.io/badge/arXiv-25.10-red)](https://arxiv.org/abs/2510.07356) [![Model](https://img.shields.io/badge/Model-HuggingFace-yellow)](https://huggingface.co/lkongam/KernelCoder)  \n\t- **Task:** Torch -\u003e CUDA\n\n### Agentic RL\n\n- **CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.24286) [![Code](https://img.shields.io/github/stars/BytedTsinghua-SIA/CUDA-Agent)](https://github.com/BytedTsinghua-SIA/CUDA-Agent)  \n\t- Weinan Dai, Hanlin Wu, Qiying Yu, Huan-ang Gao, Jiahao Li, Chengquan Jiang, Weiqiang Lou, Yufan Song, Hongli Yu, Jiaze Chen, Wei-Ying Ma, Ya-Qin Zhang, Jingjing Liu, Mingxuan Wang, Xin Liu, Hao Zhou\n\t- **Institution:** ByteDance Seed, Tsinghua AIR\n\t- **Task:** Torch -\u003e CUDA\n\n- **Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.05885) [![Code](https://img.shields.io/github/stars/hkust-nlp/KernelGYM)](https://github.com/hkust-nlp/KernelGYM)  \n\t- Wei Liu, Jiawei Xu, Yingru Li, Longtao Zheng, Tianjian Li, Qian Liu, Junxian He  \n\t- **Institution:** HKUST, TikTok  \n\t- **Task:** Torch -\u003e Triton\n\n- **Fine-Tuning GPT-5 for GPU Kernel Generation**  \n  [![Paper](https://img.shields.io/badge/arXiv-26.02-red)](https://arxiv.org/abs/2602.11000)  \n\t- Ali Tehrani, Yahya Emara, Essam Wissam, Wojciech Paluch, Waleed Atallah, Łukasz Dudziak, Mohamed S. Abdelfattah  \n\t- **Institution:** Makora  \n\t- **Task:** Torch -\u003e CUDA\n\n- **QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation** [![Paper](https://img.shields.io/badge/AAAI-26-green)](https://arxiv.org/abs/2511.20100) [![Code](https://img.shields.io/github/stars/QiMeng-IPRC/QiMeng-Kernel)](https://github.com/QiMeng-IPRC/QiMeng-Kernel)  \n\t- Xinguo Zhu, Shaohui Peng, Jiaming Guo, Yunji Chen, Qi Guo, Yuanbo Wen, Hang Qin, Ruizhi Chen, Qirui Zhou, Ke Gao, Yanjun Wu, Chen Zhao, Ling Li\n\t- **Institution:** Institute of Software, Institute of Computing Technology\n\t- **Task:** Torch -\u003e Triton\n\n- **CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning** [![Paper](https://img.shields.io/badge/arXiv-25.07-red)](https://arxiv.org/abs/2507.14111) [![Code](https://img.shields.io/github/stars/deepreinforce-ai/CUDA-L1)](https://github.com/deepreinforce-ai/CUDA-L1) [![Project](https://img.shields.io/badge/Project-Page-blue)](https://deepreinforce-ai.github.io/cudal1_blog/)  \n\t- Xiaoya Li, Xiaofei Sun, Albert Wang, Jiwei Li, Chris Shum\n\t- **Task:** CUDA Optimization\n\n- **TRITONRL: Training LLMs to Think and Code Triton Without Cheating** [![Paper](https://img.shields.io/badge/arXiv-25.10-red)](https://arxiv.org/abs/2510.17891)  \n\t- Jiin Woo, Shaowei Zhu, Allen Nie, Zhen Jia, Yida Wang, Youngsuk Park\n\t- **Task:** Torch -\u003e Triton\n\n- **CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning** [![Paper](https://img.shields.io/badge/CGO-25-green)](https://dl.acm.org/doi/abs/10.1145/3696443.3708943)  \n\t- Guoliang He, Eiko Yoneki\n\t- **Institution:** University of Cambridge\n\t- **Task:** SASS Scheduling Optimization\n\n- **Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement Learning** [![Paper](https://img.shields.io/badge/OpenReview-25-green)](https://openreview.net/forum?id=VdLEaGPYWT)  \n\t- Yaoyu Wang, Hankun Dai, Zhidong Yang, Junmin Xiao, Guangming Tan\n\t- **Task:** Sparse Matrix -\u003e CUDA\n\n- **SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.08-red)](https://arxiv.org/abs/2508.20258)  \n\t- Arya Tschand, Muhammad Awad, Ryan Swann, Kesavan Ramakrishnan, Jeffrey Ma, Keith Lowery, Ganesh Dasika, Vijay Janapa Reddi\n\t- **Task:** CUDA Swizzling Optimization\n\n- **Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization** [![Paper](https://img.shields.io/badge/arXiv-25.10-red)](https://arxiv.org/abs/2510.17158)  \n\t- Daniel Nichols, Konstantinos Parasyris, Charles Jekel, Abhinav Bhatele, Harshitha Menon\n\t- **Task:** CUDA Optimization\n\n- **CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning** [![Paper](https://img.shields.io/badge/arXiv-25.12-red)](https://arxiv.org/abs/2512.02551) [![Code](https://img.shields.io/github/stars/deepreinforce-ai/CUDA-L2)](https://github.com/deepreinforce-ai/CUDA-L2)  \n\t- Songqiao Su, Xiaofei Sun, Xiaoya Li, Albert Wang, Jiwei Li, Chris Shum\n\t- **Institution:** DeepReinforce Team\n\t- **Task:** HGEMM Optimization\n\n- **Kevin: Multi-Turn RL for Generating CUDA Kernels** [![Paper](https://img.shields.io/badge/arXiv-25.07-red)](https://arxiv.org/abs/2507.11948)\n\t- Carlo Baronio, Pietro Marsella, Ben Pan, Simon Guo, Silas Alberti\n\t- **Institution:** Stanford University\n\t- **Task:** Torch -\u003e CUDA\n\n## Contribution\n\nFeel free to open an [issue](https://github.com/kcxain/Awesome-LLM4Kernel/issues/new) or submit a [pull request](https://github.com/kcxain/Awesome-LLM4Kernel/fork) to correct errors or add work that has not yet been included in this project. You can also email us at kcxain@gmail.com for any form of discussion and collaboration.\n\n\n## Citation\n\nIf you find this work useful, welcome to cite us.\n\n```bib\n@article{llm4kernel,\n  title={LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development},\n  author={Changxin Ke},\n  year={2025}\n  url={https://github.com/kcxain/Awesome-LLM4Kernel}\n}\n```\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/kcxain%2Fawesome-llm4kernel/projects"}