{"id":72716,"url":"https://github.com/coderonion/awesome-llm-and-aigc","name":"awesome-llm-and-aigc","description":"🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.","projects_count":1402,"last_synced_at":"2026-10-08T04:00:28.908Z","repository":{"id":129483790,"uuid":"602069450","full_name":"coderonion/awesome-llm-and-aigc","owner":"coderonion","description":"🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.","archived":false,"fork":false,"pushed_at":"2025-08-01T16:22:21.000Z","size":276,"stargazers_count":815,"open_issues_count":13,"forks_count":80,"subscribers_count":15,"default_branch":"main","last_synced_at":"2026-10-02T21:55:28.912Z","etag":null,"topics":["ai4s","ai4science","aigc","awesome-list","cuda","datasets","deepseek","gpt","langchain","llama","llm","mllm","qwen","qwen3","r1","reinforcement-learning","triton","vla","vlm","yolo"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/coderonion.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2023-02-15T12:40:08.000Z","updated_at":"2026-09-25T12:03:02.000Z","dependencies_parsed_at":"2023-11-14T15:30:17.487Z","dependency_job_id":"bac7a583-bd62-48b0-bea4-c73ab4a37e47","html_url":"https://github.com/coderonion/awesome-llm-and-aigc","commit_stats":{"total_commits":138,"total_committers":4,"mean_commits":34.5,"dds":0.3405797101449275,"last_synced_commit":"2fbde9dc388dce14810effb312b5822cd132ad12"},"previous_names":["codingonion/awesome-llm-and-aigc","sjinzh/awesome-llm-and-aigc","coderonion/awesome-llm-and-aigc"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/coderonion/awesome-llm-and-aigc","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coderonion%2Fawesome-llm-and-aigc","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coderonion%2Fawesome-llm-and-aigc/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coderonion%2Fawesome-llm-and-aigc/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coderonion%2Fawesome-llm-and-aigc/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/coderonion","download_url":"https://codeload.github.com/coderonion/awesome-llm-and-aigc/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coderonion%2Fawesome-llm-and-aigc/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":343072799,"owners_count":37963027,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-10-03T02:00:07.567Z","response_time":59,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-10-05T16:05:19.846Z","updated_at":"2026-10-08T04:00:28.908Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Blogs","Interview","Summary","Open API","Datasets","Applications","Prompts","Videos","Jobs and Interview"],"sub_categories":["数据集","提示语（魔法）"],"readme":"# Awesome-llm-and-aigc\r\n[![Awesome](https://cdn.rawgit.com/sindresorhus/awesome/d7305f38d29fed78fa85652e3a63e154dd8e8829/media/badge.svg)](https://github.com/sindresorhus/awesome)\r\n\r\n🚀🚀🚀 This repository lists some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.\r\n\r\n## Contents\r\n- [Awesome-llm-and-aigc](#awesome-llm-and-aigc)\r\n  - [Summary](#summary)\r\n    - [Frameworks](#frameworks)\r\n      - [Official Version](#official-version)\r\n        - [Neural Network Architecture](#neural-network-architecture)\r\n        - [Large Language Model](#large-language-model)\r\n        - [Large Vision Language Model](#large-vision-language-model)\r\n        - [Vision Language Action](#vision-language-action)\r\n        - [AI Generated Content](#ai-generated-content)\r\n      - [Performance Analysis and Visualization](#performance-analysis-and-visualization)\r\n      - [Training and Fine-Tuning Framework](#training-and-fine-tuning-framework)\r\n      - [Reinforcement Learning Framework](#reinforcement-learning-framework)\r\n      - [LLM Inference Framework](#llm-inference-framework)\r\n        - [LLM Inference and Serving Engine](#llm-inference-and-serving-engine)\r\n        - [High Performance Kernel Library](#high-performance-kernel-library)\r\n        - [C and CPP Implementation](#c-and-cpp-implementation)\r\n        - [Triton Implementation](#triton-implementation)\r\n        - [Python Implementation](#python-implementation)\r\n        - [Mojo Implementation](#mojo-implementation)\r\n        - [Rust Implementation](#rust-implementation)\r\n        - [zig Implementation](#zig-implementation)\r\n        - [Go Implementation](#go-implementation)\r\n      - [LLM Quantization Framework](#llm-quantization-framework)\r\n      - [Application Development Platform](#application-development-platform)\r\n      - [RAG Framework](#rag-framework)\r\n      - [Vector Database](#vector-database)\r\n      - [Memory Management](#memory-management)\r\n    - [Awesome List](#awesome-list)\r\n    - [Paper Overview](#paper-overview)\r\n    - [Learning Resources](#learning-resources)\r\n    - [Community](#community)\r\n  - [Prompts](#prompts)\r\n  - [Open API](#open-api)\r\n    - [Python API](#python-api)\r\n    - [Rust API](#rust-api)\r\n    - [Csharp API](#csharp-api)\r\n    - [Node.js API](#node.js-api)\r\n  - [Applications](#applications)\r\n    - [IDE](#ide)\r\n    - [Chatbot](#chatbot)\r\n    - [Object Detection Field](#object-detection-field)\r\n    - [Autonomous Driving Field](#autonomous-driving-field)\r\n    - [Robotics and Embodied AI](#robotics-and-embodied-ai)\r\n    - [Code Assistant](#code-assistant)\r\n    - [Translator](#translator)\r\n    - [Local knowledge Base](#local-knowledge-base)\r\n    - [Long-Term Memory](#long-term-memory)\r\n    - [Question Answering System](#question-answering-system)\r\n    - [Academic Field](#academic-field)\r\n    - [Medical Field](#medical-field)\r\n    - [Mental Health Field](#mental-health-field)\r\n    - [Legal Field](#legal-field)\r\n    - [Financial Field](#Financial-field)\r\n    - [Math Field](#math-field)\r\n    - [Music Field](#music-field)\r\n    - [Speech and audio Field](#speech-and-audio-field)\r\n    - [Humor Generation](#humor-generation)\r\n    - [Animation Field](#animation-field)\r\n    - [Food Field](#food-field)\r\n    - [PPT Field](#ppt-field)\r\n    - [Tool Learning](#tool-learning)\r\n    - [Adversarial Attack Field](#adversarial-attack-field)\r\n    - [Multi-Agent Collaboration](#multi-agent-collaboration)\r\n    - [AI Avatar and Digital Human](#ai-avatar-and-digital-human)\r\n    - [GUI](#gui)\r\n  - [Datasets](#datasets)\r\n    - [Awesome Datasets List](#awesome-datasets-list)\r\n    - [Open Datasets Platform](#open-datasets-platform)\r\n    - [Humanoid Robotics Datasets](#humanoid-robotics-datasets)\r\n    - [Text Datasets](#text-datasets)\r\n    - [Multimodal Datasets](#multimodal-datasets)\r\n    - [SFT Datasets](#sft-datasets)\r\n    - [Datasets Tools](#datasets-tools)\r\n        - [Data Annotation](#data-annotation)\r\n  - [Blogs](#blogs)\r\n  - [Interview](#interview)\r\n\r\n\r\n## Summary\r\n\r\n  - ### Frameworks\r\n\r\n    - #### Official Version\r\n\r\n\r\n      - ##### Neural Network Architecture\r\n        ###### 神经网络架构\r\n\r\n        - [Transformer](https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/models/transformer.py) \u003cimg src=\"https://img.shields.io/github/stars/tensorflow/tensor2tensor?style=social\"/\u003e : \"Attention is All You Need\". (**[arXiv 2017](https://arxiv.org/abs/1706.03762)**).\r\n\r\n        - [KAN](https://github.com/KindXiaoming/pykan) \u003cimg src=\"https://img.shields.io/github/stars/KindXiaoming/pykan?style=social\"/\u003e : \"KAN: Kolmogorov-Arnold Networks\". (**[arXiv 2024](https://arxiv.org/abs/2404.19756)**).\r\n\r\n        - [FlashAttention](https://github.com/Dao-AILab/flash-attention) \u003cimg src=\"https://img.shields.io/github/stars/Dao-AILab/flash-attention?style=social\"/\u003e : Fast and memory-efficient exact attention. \"FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness\". (**[arXiv 2022](https://arxiv.org/abs/2205.14135)**).\r\n\r\n\r\n\r\n      - ##### Large Language Model\r\n        ###### 大语言模型（LLM）\r\n\r\n        - GPT-1 : \"Improving Language Understanding by Generative Pre-Training\". (**[cs.ubc.ca, 2018](https://www.cs.ubc.ca/~amuham01/LING530/papers/radford2018improving.pdf)**).\r\n\r\n        - [GPT-2](https://github.com/openai/gpt-2) \u003cimg src=\"https://img.shields.io/github/stars/openai/gpt-2?style=social\"/\u003e : \"Language Models are Unsupervised Multitask Learners\". (**[OpenAI blog, 2019](https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf)**). [Better language models and their implications](https://openai.com/research/better-language-models).\r\n\r\n        - [GPT-3](https://github.com/openai/gpt-3) \u003cimg src=\"https://img.shields.io/github/stars/openai/gpt-3?style=social\"/\u003e : \"GPT-3: Language Models are Few-Shot Learners\". (**[arXiv 2020](https://arxiv.org/abs/2005.14165)**).\r\n\r\n        - InstructGPT : \"Training language models to follow instructions with human feedback\". (**[arXiv 2022](https://arxiv.org/abs/2203.02155)**). \"Aligning language models to follow instructions\". (**[OpenAI blog, 2022](https://openai.com/research/instruction-following)**).\r\n\r\n        - [ChatGPT](https://chat.openai.com/): [Optimizing Language Models for Dialogue](https://openai.com/blog/chatgpt).\r\n\r\n        - [GPT-4](https://openai.com/product/gpt-4): GPT-4 is OpenAI’s most advanced system, producing safer and more useful responses. \"Sparks of Artificial General Intelligence: Early experiments with GPT-4\". (**[arXiv 2023](https://arxiv.org/abs/2303.12712)**). \"GPT-4 Architecture, Infrastructure, Training Dataset, Costs, Vision, MoE\". (**[SemianAlysis, 2023](https://www.semianalysis.com/p/gpt-4-architecture-infrastructure)**).\r\n\r\n        - [Llama 2](https://github.com/facebookresearch/llama) \u003cimg src=\"https://img.shields.io/github/stars/facebookresearch/llama?style=social\"/\u003e : Inference code for LLaMA models. \"LLaMA: Open and Efficient Foundation Language Models\". (**[arXiv 2023](https://arxiv.org/abs/2302.13971)**). \"Llama 2: Open Foundation and Fine-Tuned Chat Models\". (**[ai.meta.com, 2023-07-18](https://ai.meta.com/research/publications/llama-2-open-foundation-and-fine-tuned-chat-models/)**). (**[2023-07-18, Llama 2 is here - get it on Hugging Face](https://huggingface.co/blog/llama2)**).\r\n\r\n        - [Llama 3](https://github.com/meta-llama/llama3) \u003cimg src=\"https://img.shields.io/github/stars/meta-llama/llama3?style=social\"/\u003e : The official Meta Llama 3 GitHub site.\r\n\r\n        - [Qwen（通义千问）](https://github.com/QwenLM/Qwen) \u003cimg src=\"https://img.shields.io/github/stars/QwenLM/Qwen?style=social\"/\u003e : The official repo of Qwen (通义千问) chat \u0026 pretrained large language model proposed by Alibaba Cloud.\r\n\r\n        - [Qwen3](https://github.com/QwenLM/Qwen3) \u003cimg src=\"https://img.shields.io/github/stars/QwenLM/Qwen3?style=social\"/\u003e : Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud. \"Qwen3: Think Deeper, Act Faster\". (**[Qwen Blog](https://qwenlm.github.io/blog/qwen3/)**). \"Qwen2.5 Technical Report\". (**[arXiv 2024](https://arxiv.org/abs/2412.15115)**). \"Qwen2 Technical Report\". (**[arXiv 2024](https://arxiv.org/abs/2407.10671)**).\r\n\r\n        - [DeepSeek-V3](https://github.com/deepseek-ai/DeepSeek-V3) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/DeepSeek-V3?style=social\"/\u003e : \"DeepSeek-V3 Technical Report\". (**[arXiv 2024](https://arxiv.org/abs/2412.19437)**).\r\n\r\n        - [DeepSeek-R1](https://github.com/deepseek-ai/DeepSeek-R1) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/DeepSeek-R1?style=social\"/\u003e : \"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning\". (**[arXiv 2025](https://arxiv.org/abs/2501.12948)**).\r\n\r\n        - [Open R1](https://github.com/huggingface/open-r1) \u003cimg src=\"https://img.shields.io/github/stars/huggingface/open-r1?style=social\"/\u003e : Fully open reproduction of [DeepSeek-R1](https://github.com/deepseek-ai/DeepSeek-R1).\r\n\r\n        - [TinyZero](https://github.com/Jiayi-Pan/TinyZero) \u003cimg src=\"https://img.shields.io/github/stars/Jiayi-Pan/TinyZero?style=social\"/\u003e : Clean, minimal, accessible reproduction of DeepSeek R1-Zero. TinyZero is a reproduction of [DeepSeek R1 Zero](https://github.com/deepseek-ai/DeepSeek-R1) in countdown and multiplication tasks. We built upon [veRL](https://github.com/volcengine/verl).\r\n\r\n        - [GRPO-Zero](https://github.com/policy-gradient/GRPO-Zero) \u003cimg src=\"https://img.shields.io/github/stars/policy-gradient/GRPO-Zero?style=social\"/\u003e : GRPO training with minimal dependencies. We implement almost everything from scratch and only depend on tokenizers for tokenization and pytorch for training.\r\n\r\n        - [Search-R1](https://github.com/PeterGriffinJin/Search-R1) \u003cimg src=\"https://img.shields.io/github/stars/PeterGriffinJin/Search-R1?style=social\"/\u003e : Search-R1: An Efficient, Scalable RL Training Framework for Reasoning \u0026 Search Engine Calling interleaved LLM based on veRL. \"Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning\". (**[arXiv 2025](https://arxiv.org/abs/2503.09516)**).\r\n\r\n        - [Logic-RL](https://github.com/Unakar/Logic-RL) \u003cimg src=\"https://img.shields.io/github/stars/Unakar/Logic-RL?style=social\"/\u003e : Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning. \"Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning\". (**[arXiv 2025](https://arxiv.org/abs/2502.14768)**).\r\n\r\n        - [X-R1](https://github.com/dhcode-cpp/X-R1) \u003cimg src=\"https://img.shields.io/github/stars/dhcode-cpp/X-R1?style=social\"/\u003e : X-R1 aims to build an easy-to-use, low-cost training framework based on end-to-end reinforcement learning to accelerate the development of Scaling Post-Training. Inspired by [DeepSeek-R1](https://github.com/deepseek-ai/DeepSeek-R1) and [open-r1](https://github.com/huggingface/open-r1) , we produce minimal-cost for training 0.5B R1-Zero \"Aha Moment\"💡 from base model\r\n\r\n        - [DeepScaleR](https://github.com/agentica-project/deepscaler) \u003cimg src=\"https://img.shields.io/github/stars/agentica-project/deepscaler?style=social\"/\u003e : Democratizing Reinforcement Learning for LLMs. [www.agentica-project.com](https://www.agentica-project.com/). [\"DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL\"](https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2)\r\n\r\n        - [OpenSeek](https://github.com/FlagAI-Open/OpenSeek) \u003cimg src=\"https://img.shields.io/github/stars/FlagAI-Open/OpenSeek?style=social\"/\u003e : OpenSeek aims to unite the global open source community to drive collaborative innovation in algorithms, data and systems to develop next-generation models that surpass DeepSeek.\r\n\r\n\r\n\r\n\r\n\r\n        - [Gemma](https://github.com/google/gemma_pytorch) \u003cimg src=\"https://img.shields.io/github/stars/google/gemma_pytorch?style=social\"/\u003e : The official PyTorch implementation of Google's Gemma models. [ai.google.dev/gemma](https://ai.google.dev/gemma)\r\n\r\n        - [Grok-1](https://github.com/xai-org/grok-1) \u003cimg src=\"https://img.shields.io/github/stars/xai-org/grok-1?style=social\"/\u003e : This repository contains JAX example code for loading and running the Grok-1 open-weights model.\r\n\r\n        - [Claude](https://www.anthropic.com/product) : Claude is a next-generation AI assistant based on Anthropic’s research into training helpful, honest, and harmless AI systems.\r\n\r\n        - [Whisper](https://github.com/openai/whisper) \u003cimg src=\"https://img.shields.io/github/stars/openai/whisper?style=social\"/\u003e : Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification. \"Robust Speech Recognition via Large-Scale Weak Supervision\". (**[arXiv 2022](https://arxiv.org/abs/2212.04356)**).\r\n\r\n        - [OpenChat](https://github.com/imoneoi/openchat) \u003cimg src=\"https://img.shields.io/github/stars/imoneoi/openchat?style=social\"/\u003e : OpenChat: Advancing Open-source Language Models with Imperfect Data. [huggingface.co/openchat/openchat](https://huggingface.co/openchat/openchat)\r\n\r\n        - [GPT-Engineer](https://github.com/AntonOsika/gpt-engineer) \u003cimg src=\"https://img.shields.io/github/stars/AntonOsika/gpt-engineer?style=social\"/\u003e : Specify what you want it to build, the AI asks for clarification, and then builds it. GPT Engineer is made to be easy to adapt, extend, and make your agent learn how you want your code to look. It generates an entire codebase based on a prompt.\r\n\r\n        - [StableLM](https://github.com/Stability-AI/StableLM) \u003cimg src=\"https://img.shields.io/github/stars/Stability-AI/StableLM?style=social\"/\u003e : StableLM: Stability AI Language Models.\r\n\r\n        - [JARVIS](https://github.com/microsoft/JARVIS) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/JARVIS?style=social\"/\u003e : JARVIS, a system to connect LLMs with ML community. \"HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace\". (**[arXiv 2023](https://arxiv.org/abs/2303.17580)**).\r\n\r\n        - [MiniGPT-4](https://github.com/Vision-CAIR/MiniGPT-4) \u003cimg src=\"https://img.shields.io/github/stars/Vision-CAIR/MiniGPT-4?style=social\"/\u003e : MiniGPT-4: Enhancing Vision-language Understanding with Advanced Large Language Models. [minigpt-4.github.io](https://minigpt-4.github.io/)\r\n\r\n        - [minGPT](https://github.com/karpathy/minGPT) \u003cimg src=\"https://img.shields.io/github/stars/karpathy/minGPT?style=social\"/\u003e : A minimal PyTorch re-implementation of the OpenAI GPT (Generative Pretrained Transformer) training.\r\n\r\n        - [nanoGPT](https://github.com/karpathy/nanoGPT) \u003cimg src=\"https://img.shields.io/github/stars/karpathy/nanoGPT?style=social\"/\u003e : The simplest, fastest repository for training/finetuning medium-sized GPTs.\r\n\r\n        - [MicroGPT](https://github.com/muellerberndt/micro-gpt) \u003cimg src=\"https://img.shields.io/github/stars/muellerberndt/micro-gpt?style=social\"/\u003e : A simple and effective autonomous agent compatible with GPT-3.5-Turbo and GPT-4. MicroGPT aims to be as compact and reliable as possible.\r\n\r\n        - [Dolly](https://github.com/databrickslabs/dolly) \u003cimg src=\"https://img.shields.io/github/stars/databrickslabs/dolly?style=social\"/\u003e : Databricks’ Dolly, a large language model trained on the Databricks Machine Learning Platform. [Hello Dolly: Democratizing the magic of ChatGPT with open models](https://www.databricks.com/blog/2023/03/24/hello-dolly-democratizing-magic-chatgpt-open-models.html)\r\n\r\n        - [LMFlow](https://github.com/OptimalScale/LMFlow) \u003cimg src=\"https://img.shields.io/github/stars/OptimalScale/LMFlow?style=social\"/\u003e : An extensible, convenient, and efficient toolbox for finetuning large machine learning models, designed to be user-friendly, speedy and reliable, and accessible to the entire community. Large Language Model for All. [optimalscale.github.io/LMFlow/](https://optimalscale.github.io/LMFlow/)\r\n\r\n        - [Colossal-AI](https://github.com/hpcaitech/ColossalAI) \u003cimg src=\"https://img.shields.io/github/stars/hpcaitech/ColossalAI?style=social\"/\u003e : Making big AI models cheaper, easier, and scalable. [www.colossalai.org](www.colossalai.org). \"Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training\". (**[arXiv 2021](https://arxiv.org/abs/2110.14883)**).\r\n\r\n        - [Lit-LLaMA](https://github.com/Lightning-AI/lit-llama) \u003cimg src=\"https://img.shields.io/github/stars/Lightning-AI/lit-llama?style=social\"/\u003e : ⚡ Lit-LLaMA. Implementation of the LLaMA language model based on nanoGPT. Supports flash attention, Int8 and GPTQ 4bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed.\r\n\r\n        - [GPT-4-LLM](https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM) \u003cimg src=\"https://img.shields.io/github/stars/Instruction-Tuning-with-GPT-4/GPT-4-LLM?style=social\"/\u003e : \"Instruction Tuning with GPT-4\". (**[arXiv 2023](https://arxiv.org/abs/2304.03277)**). [instruction-tuning-with-gpt-4.github.io/](https://instruction-tuning-with-gpt-4.github.io/)\r\n\r\n        - [Stanford Alpaca](https://github.com/tatsu-lab/stanford_alpaca) \u003cimg src=\"https://img.shields.io/github/stars/tatsu-lab/stanford_alpaca?style=social\"/\u003e : Stanford Alpaca: An Instruction-following LLaMA Model.\r\n\r\n        - [Liger-Kernel](https://github.com/linkedin/Liger-Kernel) \u003cimg src=\"https://img.shields.io/github/stars/linkedin/Liger-Kernel?style=social\"/\u003e : Efficient Triton Kernels for LLM Training. [arxiv.org/pdf/2410.10989](https://arxiv.org/pdf/2410.10989)\r\n\r\n        - [FlagGems](https://github.com/FlagOpen/FlagGems) \u003cimg src=\"https://img.shields.io/github/stars/FlagOpen/FlagGems?style=social\"/\u003e : FlagGems is a high-performance general operator library implemented in [OpenAI Triton](https://github.com/openai/triton). It aims to provide a suite of kernel functions to accelerate LLM training and inference.\r\n\r\n        - [feizc/Visual-LLaMA](https://github.com/feizc/Visual-LLaMA) \u003cimg src=\"https://img.shields.io/github/stars/feizc/Visual-LLaMA?style=social\"/\u003e : Open LLaMA Eyes to See the World. This project aims to optimize LLaMA model for visual information understanding like GPT-4 and further explore the potentional of large language model.\r\n\r\n        - [Lightning-AI/lightning-colossalai](https://github.com/Lightning-AI/lightning-colossalai) \u003cimg src=\"https://img.shields.io/github/stars/Lightning-AI/lightning-colossalai?style=social\"/\u003e : Efficient Large-Scale Distributed Training with [Colossal-AI](https://colossalai.org/) and [Lightning AI](https://lightning.ai/).\r\n\r\n        - [GPT4All](https://github.com/nomic-ai/gpt4all) \u003cimg src=\"https://img.shields.io/github/stars/nomic-ai/gpt4all?style=social\"/\u003e : GPT4All: An ecosystem of open-source on-edge large language models. GTP4All is an ecosystem to train and deploy powerful and customized large language models that run locally on consumer grade CPUs.\r\n\r\n        - [ChatALL](https://github.com/sunner/ChatALL) \u003cimg src=\"https://img.shields.io/github/stars/sunner/ChatALL?style=social\"/\u003e :  Concurrently chat with ChatGPT, Bing Chat, bard, Alpaca, Vincuna, Claude, ChatGLM, MOSS, iFlytek Spark, ERNIE and more, discover the best answers. [chatall.ai](http://chatall.ai/)\r\n\r\n        - [1595901624/gpt-aggregated-edition](https://github.com/1595901624/gpt-aggregated-edition) \u003cimg src=\"https://img.shields.io/github/stars/1595901624/gpt-aggregated-edition?style=social\"/\u003e : 聚合ChatGPT官方版、ChatGPT免费版、文心一言、Poe、chatchat等多平台，支持自定义导入平台。\r\n\r\n        - [FreedomIntelligence/LLMZoo](https://github.com/FreedomIntelligence/LLMZoo) \u003cimg src=\"https://img.shields.io/github/stars/FreedomIntelligence/LLMZoo?style=social\"/\u003e : ⚡LLM Zoo is a project that provides data, models, and evaluation benchmark for large language models.⚡ [Tech Report](https://github.com/FreedomIntelligence/LLMZoo/blob/main/assets/llmzoo.pdf)\r\n\r\n        - [shm007g/LLaMA-Cult-and-More](https://github.com/shm007g/LLaMA-Cult-and-More) \u003cimg src=\"https://img.shields.io/github/stars/shm007g/LLaMA-Cult-and-More?style=social\"/\u003e : News about 🦙 Cult and other AIGC models.\r\n\r\n        - [X-PLUG/mPLUG-Owl](https://github.com/X-PLUG/mPLUG-Owl) \u003cimg src=\"https://img.shields.io/github/stars/X-PLUG/mPLUG-Owl?style=social\"/\u003e : mPLUG-Owl🦉: Modularization Empowers Large Language Models with Multimodality.\r\n\r\n        - [i-Code](https://github.com/microsoft/i-Code) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/i-Code?style=social\"/\u003e : The ambition of the i-Code project is to build integrative and composable multimodal Artificial Intelligence. The \"i\" stands for integrative multimodal learning. \"CoDi: Any-to-Any Generation via Composable Diffusion\". (**[arXiv 2023](https://arxiv.org/abs/2305.11846)**).\r\n\r\n        - [WorkGPT](https://github.com/h2oai/h2ogpt) \u003cimg src=\"https://img.shields.io/github/stars/h2oai/h2ogpt?style=social\"/\u003e : WorkGPT is an agent framework in a similar fashion to AutoGPT or LangChain.\r\n\r\n        - [h2oGPT](https://github.com/team-openpm/workgpt) \u003cimg src=\"https://img.shields.io/github/stars/team-openpm/workgpt?style=social\"/\u003e : h2oGPT is a large language model (LLM) fine-tuning framework and chatbot UI with document(s) question-answer capabilities. \"h2oGPT: Democratizing Large Language Models\". (**[arXiv 2023](https://arxiv.org/abs/2306.08161)**).\r\n\r\n        - [LongLLaMA ](https://github.com/CStanKonrad/long_llama) \u003cimg src=\"https://img.shields.io/github/stars/CStanKonrad/long_llama?style=social\"/\u003e : LongLLaMA is a large language model capable of handling long contexts. It is based on OpenLLaMA and fine-tuned with the Focused Transformer (FoT) method.\r\n\r\n        - [LLaMA-Adapter](https://github.com/OpenGVLab/LLaMA-Adapter) \u003cimg src=\"https://img.shields.io/github/stars/OpenGVLab/LLaMA-Adapter?style=social\"/\u003e : Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters. LLaMA-Adapter: Efficient Fine-tuning of LLaMA 🚀\r\n\r\n        - [DemoGPT](https://github.com/melih-unsal/DemoGPT) \u003cimg src=\"https://img.shields.io/github/stars/melih-unsal/DemoGPT?style=social\"/\u003e : Create 🦜️🔗 LangChain apps by just using prompts with the power of Llama 2 🌟 Star to support our work! | 只需使用句子即可创建 LangChain 应用程序。 给个star支持我们的工作吧！DemoGPT: Auto Gen-AI App Generator with the Power of Llama 2. ⚡ With just a prompt, you can create interactive Streamlit apps via 🦜️🔗 LangChain's transformative capabilities \u0026 Llama 2.⚡ [demogpt.io](https://www.demogpt.io/)\r\n\r\n        - [Lamini](https://github.com/lamini-ai/lamini) \u003cimg src=\"https://img.shields.io/github/stars/lamini-ai/lamini?style=social\"/\u003e : Lamini: The LLM engine for rapidly customizing models 🦙\r\n\r\n        - [xorbitsai/inference](https://github.com/xorbitsai/inference) \u003cimg src=\"https://img.shields.io/github/stars/xorbitsai/inference?style=social\"/\u003e : Xorbits Inference (Xinference) is a powerful and versatile library designed to serve LLMs, speech recognition models, and multimodal models, even on your laptop. It supports a variety of models compatible with GGML, such as llama, chatglm, baichuan, whisper, vicuna, orac, and many others.\r\n\r\n        - [epfLLM/Megatron-LLM](https://github.com/epfLLM/Megatron-LLM) \u003cimg src=\"https://img.shields.io/github/stars/epfLLM/Megatron-LLM?style=social\"/\u003e : distributed trainer for LLMs.\r\n\r\n        - [AmineDiro/cria](https://github.com/AmineDiro/cria) \u003cimg src=\"https://img.shields.io/github/stars/AmineDiro/cria?style=social\"/\u003e : OpenAI compatible API for serving LLAMA-2 model.\r\n\r\n        - [Llama-2-Onnx](https://github.com/microsoft/Llama-2-Onnx) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/Llama-2-Onnx?style=social\"/\u003e : Llama 2 Powered By ONNX.\r\n\r\n        - [gpt-llm-trainer](https://github.com/mshumer/gpt-llm-trainer) \u003cimg src=\"https://img.shields.io/github/stars/mshumer/gpt-llm-trainer?style=social\"/\u003e : The goal of this project is to explore an experimental new pipeline to train a high-performing task-specific model. We try to abstract away all the complexity, so it's as easy as possible to go from idea -\u003e performant fully-trained model.\r\n\r\n\r\n\r\n\r\n\r\n\r\n\r\n        - [ChatGLM-6B](https://github.com/THUDM/ChatGLM-6B) \u003cimg src=\"https://img.shields.io/github/stars/THUDM/ChatGLM-6B?style=social\"/\u003e : ChatGLM-6B: An Open Bilingual Dialogue Language Model | 开源双语对话语言模型。 ChatGLM-6B 是一个开源的、支持中英双语的对话语言模型，基于 [General Language Model (GLM)](https://github.com/THUDM/GLM) 架构，具有 62 亿参数。 \"GLM: General Language Model Pretraining with Autoregressive Blank Infilling\". (**[ACL 2022](https://aclanthology.org/2022.acl-long.26/)**).  \"GLM-130B: An Open Bilingual Pre-trained Model\". (**[ICLR 2023](https://openreview.net/forum?id=-Aw0rrrPUF)**).\r\n\r\n        - [ChatGLM2-6B](https://github.com/THUDM/ChatGLM2-6B) \u003cimg src=\"https://img.shields.io/github/stars/THUDM/ChatGLM2-6B?style=social\"/\u003e : ChatGLM2-6B: An Open Bilingual Chat LLM | 开源双语对话语言模型。ChatGLM2-6B 是开源中英双语对话模型 ChatGLM-6B 的第二代版本，在保留了初代模型对话流畅、部署门槛较低等众多优秀特性的基础之上，ChatGLM2-6B 引入了更强大的性能、更强大的性能、更高效的推理、更开放的协议。\r\n\r\n        - [ChatGLM3](https://github.com/THUDM/ChatGLM3) \u003cimg src=\"https://img.shields.io/github/stars/THUDM/ChatGLM3?style=social\"/\u003e : ChatGLM3 series: Open Bilingual Chat LLMs | 开源双语对话语言模型。\r\n\r\n        - [InternLM（书生·浦语）](https://github.com/InternLM/InternLM) \u003cimg src=\"https://img.shields.io/github/stars/InternLM/InternLM?style=social\"/\u003e : Official release of InternLM2 7B and 20B base and chat models. 200K context support. [internlm.intern-ai.org.cn/](https://internlm.intern-ai.org.cn/)\r\n\r\n        - [Baichuan-7B（百川-7B）](https://github.com/baichuan-inc/Baichuan-7B) \u003cimg src=\"https://img.shields.io/github/stars/baichuan-inc/Baichuan-7B?style=social\"/\u003e : A large-scale 7B pretraining language model developed by BaiChuan-Inc. Baichuan-7B 是由百川智能开发的一个开源可商用的大规模预训练语言模型。基于 Transformer 结构，在大约 1.2 万亿 tokens 上训练的 70 亿参数模型，支持中英双语，上下文窗口长度为 4096。在标准的中文和英文 benchmark（C-Eval/MMLU）上均取得同尺寸最好的效果。[huggingface.co/baichuan-inc/baichuan-7B](https://huggingface.co/baichuan-inc/Baichuan-7B)\r\n\r\n        - [Baichuan-13B（百川-13B）](https://github.com/baichuan-inc/Baichuan-13B) \u003cimg src=\"https://img.shields.io/github/stars/baichuan-inc/Baichuan-13B?style=social\"/\u003e : A 13B large language model developed by Baichuan Intelligent Technology. Baichuan-13B 是由百川智能继 Baichuan-7B 之后开发的包含 130 亿参数的开源可商用的大规模语言模型，在权威的中文和英文 benchmark 上均取得同尺寸最好的效果。本次发布包含有预训练 (Baichuan-13B-Base) 和对齐 (Baichuan-13B-Chat) 两个版本。[huggingface.co/baichuan-inc/Baichuan-13B-Chat](https://huggingface.co/baichuan-inc/Baichuan-13B-Chat)\r\n\r\n        - [Baichuan2](https://github.com/baichuan-inc/Baichuan2) \u003cimg src=\"https://img.shields.io/github/stars/baichuan-inc/Baichuan2?style=social\"/\u003e : A series of large language models developed by Baichuan Intelligent Technology. Baichuan 2 是百川智能推出的新一代开源大语言模型，采用 2.6 万亿 Tokens 的高质量语料训练。Baichuan 2 在多个权威的中文、英文和多语言的通用、领域 benchmark 上取得同尺寸最佳的效果。本次发布包含有 7B、13B 的 Base 和 Chat 版本，并提供了 Chat 版本的 4bits 量化。[huggingface.co/baichuan-inc](https://huggingface.co/baichuan-inc). \"Baichuan 2: Open Large-scale Language Models\". (**[arXiv 2023](https://arxiv.org/abs/2309.10305)**).\r\n\r\n        - [MOSS](https://github.com/OpenLMLab/MOSS) \u003cimg src=\"https://img.shields.io/github/stars/OpenLMLab/MOSS?style=social\"/\u003e : An open-source tool-augmented conversational language model from Fudan University. MOSS是一个支持中英双语和多种插件的开源对话语言模型，moss-moon系列模型具有160亿参数，在FP16精度下可在单张A100/A800或两张3090显卡运行，在INT4/8精度下可在单张3090显卡运行。MOSS基座语言模型在约七千亿中英文以及代码单词上预训练得到，后续经过对话指令微调、插件增强学习和人类偏好训练具备多轮对话能力及使用多种插件的能力。[txsun1997.github.io/blogs/moss.html](https://txsun1997.github.io/blogs/moss.html)\r\n\r\n        - [BayLing（百聆）](https://github.com/ictnlp/BayLing) \u003cimg src=\"https://img.shields.io/github/stars/OpenLMLab/MOSS?style=social\"/\u003e : “百聆”是一个具有增强的语言对齐的英语/中文大语言模型，具有优越的英语/中文能力，在多项测试中取得ChatGPT 90%的性能。BayLing is an English/Chinese LLM equipped with advanced language alignment, showing superior capability in English/Chinese generation, instruction following and multi-turn interaction. [nlp.ict.ac.cn/bayling](http://nlp.ict.ac.cn/bayling). \"BayLing: Bridging Cross-lingual Alignment and Instruction Following through Interactive Translation for Large Language Models\". (**[arXiv 2023](https://arxiv.org/abs/2306.10968)**).\r\n\r\n        - [FlagAI（悟道·天鹰（Aquila））](https://github.com/FlagAI-Open/FlagAI) \u003cimg src=\"https://img.shields.io/github/stars/FlagAI-Open/FlagAI?style=social\"/\u003e : FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model. Our goal is to support training, fine-tuning, and deployment of large-scale models on various downstream tasks with multi-modality.\r\n\r\n        - [YuLan-Chat（玉兰）](https://github.com/RUC-GSAI/YuLan-Chat/) \u003cimg src=\"https://img.shields.io/github/stars/RUC-GSAI/YuLan-Chat?style=social\"/\u003e : YuLan-Chat models are chat-based large language models, which are developed by the researchers in GSAI, Renmin University of China (YuLan, which represents Yulan Magnolia, is the campus flower of Renmin University of China). The newest version is developed by continually-pretraining and instruction-tuning [LLaMA-2](https://github.com/facebookresearch/llama) with high-quality English and Chinese data. YuLan-Chat系列模型是中国人民大学高瓴人工智能学院师生共同开发的支持聊天的大语言模型（名字\"玉兰\"取自中国人民大学校花）。 最新版本基于LLaMA-2进行了中英文双语的继续预训练和指令微调。\r\n\r\n        - [Yi-1.5](https://github.com/01-ai/Yi-1.5) \u003cimg src=\"https://img.shields.io/github/stars/01-ai/Yi-1.5?style=social\"/\u003e : Yi-1.5 is an upgraded version of Yi, delivering stronger performance in coding, math, reasoning, and instruction-following capability.\r\n\r\n        - [智海-录问](https://github.com/zhihaiLLM/wisdomInterrogatory) \u003cimg src=\"https://img.shields.io/github/stars/zhihaiLLM/wisdomInterrogatory?style=social\"/\u003e : 智海-录问(wisdomInterrogatory)是由浙江大学、阿里巴巴达摩院以及华院计算三家单位共同设计研发的法律大模型。核心思想：以“普法共享和司法效能提升”为目标，从推动法律智能化体系入司法实践、数字化案例建设、虚拟法律咨询服务赋能等方面提供支持，形成数字化和智能化的司法基座能力。\r\n\r\n        - [活字](https://github.com/HIT-SCIR/huozi) \u003cimg src=\"https://img.shields.io/github/stars/HIT-SCIR/huozi?style=social\"/\u003e : 活字是由哈工大自然语言处理研究所多位老师和学生参与开发的一个开源可商用的大规模预训练语言模型。 该模型基于 Bloom 结构的70 亿参数模型，支持中英双语，上下文窗口长度为 2048。 在标准的中文和英文基准以及主观评测上均取得同尺寸中优异的结果。\r\n\r\n\r\n\r\n        - [MiLM-6B](https://github.com/XiaoMi/MiLM-6B) \u003cimg src=\"https://img.shields.io/github/stars/XiaoMi/MiLM-6B?style=social\"/\u003e : MiLM-6B 是由小米开发的一个大规模预训练语言模型，参数规模为64亿。在 C-Eval 和 CMMLU 上均取得同尺寸最好的效果。\r\n\r\n        - [Chinese LLaMA and Alpaca](https://github.com/ymcui/Chinese-LLaMA-Alpaca) \u003cimg src=\"https://img.shields.io/github/stars/ymcui/Chinese-LLaMA-Alpaca?style=social\"/\u003e : 中文LLaMA\u0026Alpaca大语言模型+本地CPU/GPU训练部署 (Chinese LLaMA \u0026 Alpaca LLMs)。\"Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca\". (**[arXiv 2023](https://arxiv.org/abs/2304.08177)**).\r\n\r\n        - [Chinese-LLaMA-Alpaca-2](https://github.com/ymcui/Chinese-LLaMA-Alpaca-2) \u003cimg src=\"https://img.shields.io/github/stars/ymcui/Chinese-LLaMA-Alpaca-2?style=social\"/\u003e : 中文 LLaMA-2 \u0026 Alpaca-2 大模型二期项目 (Chinese LLaMA-2 \u0026 Alpaca-2 LLMs).\r\n\r\n        - [FlagAlpha/Llama2-Chinese](https://github.com/FlagAlpha/Llama2-Chinese) \u003cimg src=\"https://img.shields.io/github/stars/FlagAlpha/Llama2-Chinese?style=social\"/\u003e : Llama中文社区，最好的中文Llama大模型，完全开源可商用。\r\n\r\n        - [michael-wzhu/Chinese-LlaMA2](https://github.com/michael-wzhu/Chinese-LlaMA2) \u003cimg src=\"https://img.shields.io/github/stars/michael-wzhu/Chinese-LlaMA2?style=social\"/\u003e : Repo for adapting Meta LlaMA2 in Chinese! META最新发布的LlaMA2的汉化版！ （完全开源可商用）\r\n\r\n        - [CPM-Bee](https://github.com/OpenBMB/CPM-Bee) \u003cimg src=\"https://img.shields.io/github/stars/OpenBMB/CPM-Bee?style=social\"/\u003e : CPM-Bee是一个完全开源、允许商用的百亿参数中英文基座模型，也是[CPM-Live](https://live.openbmb.org/)训练的第二个里程碑。\r\n\r\n        - [PandaLM](https://github.com/WeOpenML/PandaLM) \u003cimg src=\"https://img.shields.io/github/stars/WeOpenML/PandaLM?style=social\"/\u003e : PandaLM: Reproducible and Automated Language Model Assessment.\r\n\r\n        - [SpeechGPT](https://github.com/0nutation/SpeechGPT) \u003cimg src=\"https://img.shields.io/github/stars/0nutation/SpeechGPT?style=social\"/\u003e : \"SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities\". (**[arXiv 2023](https://arxiv.org/abs/2305.11000)**).\r\n\r\n        - [GPT2-Chinese](https://github.com/Morizeyao/GPT2-Chinese) \u003cimg src=\"https://img.shields.io/github/stars/Morizeyao/GPT2-Chinese?style=social\"/\u003e : Chinese version of GPT2 training code, using BERT tokenizer.\r\n\r\n        - [Chinese-Tiny-LLM](https://github.com/Chinese-Tiny-LLM/Chinese-Tiny-LLM) \u003cimg src=\"https://img.shields.io/github/stars/Chinese-Tiny-LLM/Chinese-Tiny-LLM?style=social\"/\u003e : \"Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model\". (**[arXiv 2024](https://arxiv.org/abs/2404.04167)**).\r\n\r\n        - [潘多拉 (Pandora)](https://github.com/pengzhile/pandora) \u003cimg src=\"https://img.shields.io/github/stars/pengzhile/pandora?style=social\"/\u003e : 潘多拉，一个让你呼吸顺畅的ChatGPT。Pandora, a ChatGPT that helps you breathe smoothly.\r\n\r\n        - [百度-文心大模型](https://wenxin.baidu.com/) : 百度全新一代知识增强大语言模型，文心大模型家族的新成员，能够与人对话互动，回答问题，协助创作，高效便捷地帮助人们获取信息、知识和灵感。\r\n\r\n        - [百度智能云-千帆大模型](https://cloud.baidu.com/product/wenxinworkshop) : 百度智能云千帆大模型平台一站式企业级大模型平台，提供先进的生成式AI生产及应用全流程开发工具链。\r\n\r\n        - [华为云-盘古大模型](https://www.huaweicloud.com/product/pangu.html) : 盘古大模型致力于深耕行业，打造金融、政务、制造、矿山、气象、铁路等领域行业大模型和能力集，将行业知识know-how与大模型能力相结合，重塑千行百业，成为各组织、企业、个人的专家助手。\"Accurate medium-range global weather forecasting with 3D neural networks\". (**[Nature 2023](https://www.nature.com/articles/s41586-023-06185-3)**).\r\n\r\n        - [商汤科技-日日新SenseNova](https://techday.sensetime.com/list) : 日日新（SenseNova），是商汤科技宣布推出的大模型体系，包括自然语言处理模型“商量”（SenseChat）、文生图模型“秒画”和数字人视频生成平台“如影”（SenseAvatar）等。\r\n\r\n        - [科大讯飞-星火认知大模型](https://xinghuo.xfyun.cn/) : 新一代认知智能大模型，拥有跨领域知识和语言理解能力，能够基于自然对话方式理解与执行任务。\r\n\r\n        - [字节跳动-豆包](https://www.doubao.com/) : 豆包。\r\n\r\n        - [CrazyBoyM/llama3-Chinese-chat](https://github.com/CrazyBoyM/llama3-Chinese-chat) \u003cimg src=\"https://img.shields.io/github/stars/CrazyBoyM/llama3-Chinese-chat?style=social\"/\u003e : Llama3 中文版。\r\n\r\n\r\n\r\n\r\n\r\n      - ##### Large Vision Language Model\r\n        ###### 视觉语言大模型（LVLM）\r\n\r\n        - [Qwen2.5-VL](https://github.com/QwenLM/Qwen2.5-VL) \u003cimg src=\"https://img.shields.io/github/stars/QwenLM/Qwen2-VL?style=social\"/\u003e : Qwen2-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud. \"Qwen2.5-VL Technical Report\". (**[arXiv 2025](https://arxiv.org/abs/2502.13923)**). [2025-01-26，Qwen2.5 VL! Qwen2.5 VL! Qwen2.5 VL!](https://qwenlm.github.io/blog/qwen2.5-vl/). \"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution\". (**[arXiv 2024](https://arxiv.org/abs/2409.12191)**). \"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond\". (**[arXiv 2023](https://arxiv.org/abs/2308.12966)**).\r\n\r\n        - [Kimi-VL](https://github.com/MoonshotAI/Kimi-VL) \u003cimg src=\"https://img.shields.io/github/stars/MoonshotAI/Kimi-VL?style=social\"/\u003e : Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities. \"Kimi-VL Technical Report\". (**[arXiv 2025](https://arxiv.org/abs/2504.07491)**).\r\n\r\n        - [Visual-RFT](https://github.com/Liuziyu77/Visual-RFT) \u003cimg src=\"https://img.shields.io/github/stars/Liuziyu77/Visual-RFT?style=social\"/\u003e : 🌈We introduce Visual Reinforcement Fine-tuning (Visual-RFT), the first comprehensive adaptation of Deepseek-R1's RL strategy to the multimodal field. We use the Qwen2-VL-2/7B model as our base model and design a rule-based verifiable reward, which is integrated into a GRPO-based reinforcement fine-tuning framework to enhance the performance of LVLMs across various visual perception tasks. ViRFT extends R1's reasoning capabilities to multiple visual perception tasks, including various detection tasks like Open Vocabulary Detection, Few-shot Detection, Reasoning Grounding, and Fine-grained Image Classification. \"Visual-RFT: Visual Reinforcement Fine-Tuning\". (**[arXiv 2025](https://arxiv.org/abs/2503.01785)**).\r\n\r\n        - [VLM-R1](https://github.com/om-ai-lab/VLM-R1) \u003cimg src=\"https://img.shields.io/github/stars/om-ai-lab/VLM-R1?style=social\"/\u003e : VLM-R1: A stable and generalizable R1-style Large Vision-Language Model. Solve Visual Understanding with Reinforced VLMs. [2025-03-20，Improving Object Detection through Reinforcement Learning with VLM-R1](https://om-ai-lab.github.io/2025_03_20.html).\r\n\r\n        - [Video-R1](https://github.com/tulerfeng/Video-R1) \u003cimg src=\"https://img.shields.io/github/stars/tulerfeng/Video-R1?style=social\"/\u003e : \"Video-R1: Reinforcing Video Reasoning in MLLMs\". (**[arXiv 2025](https://arxiv.org/abs/2503.21776)**).\r\n\r\n        - [MAYE](https://github.com/GAIR-NLP/MAYE) \u003cimg src=\"https://img.shields.io/github/stars/GAIR-NLP/MAYE?style=social\"/\u003e : This project presents MAYE, a transparent and reproducible framework and a comprehensive evaluation scheme for applying reinforcement learning (RL) to vision-language models (VLMs). The codebase is built entirely from scratch without relying on existing RL toolkits. \"Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme\". (**[arXiv 2025](https://arxiv.org/abs/2504.02587)**).\r\n\r\n        - [Osilly/Vision-R1](https://github.com/Osilly/Vision-R1) \u003cimg src=\"https://img.shields.io/github/stars/Osilly/Vision-R1?style=social\"/\u003e : \"Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models\". (**[arXiv 2025](https://arxiv.org/abs/2503.06749)**).\r\n\r\n        - [Griffon/Vision-R1](https://github.com/jefferyZhan/Griffon/tree/master/Vision-R1) \u003cimg src=\"https://img.shields.io/github/stars/jefferyZhan/Griffon?style=social\"/\u003e : \"Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning\". (**[arXiv 2025](https://arxiv.org/abs/2503.18013)**).\r\n\r\n        - [Janus](https://github.com/deepseek-ai/Janus) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/Janus?style=social\"/\u003e : 🚀 Janus-Series: Unified Multimodal Understanding and Generation Models. \"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling\". (**[arXiv 2025](https://arxiv.org/abs/2501.17811)**). \"Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation\". (**[arXiv 2024](https://arxiv.org/abs/2410.13848)**). \"JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation\". (**[arXiv 2024](https://arxiv.org/abs/2411.07975)**).\r\n\r\n        - [VisualThinker-R1-Zero](https://github.com/turningpoint-ai/VisualThinker-R1-Zero) \u003cimg src=\"https://img.shields.io/github/stars/turningpoint-ai/VisualThinker-R1-Zero?style=social\"/\u003e : VisualThinker-R1-Zero: First ever R1-Zero's Aha Moment on just a 2B non-SFT Model. VisualThinker-R1-Zero is a replication of [DeepSeek-R1-Zero](https://arxiv.org/abs/2501.12948) in visual reasoning. We are the first to successfully observe the emergent “aha moment” and increased response length in visual reasoning on just a 2B non-SFT models. For more details, please refer to the notion [report](https://turningpointai.notion.site/the-multimodal-aha-moment-on-2b-model).\r\n\r\n        - [R1-V](https://github.com/Deep-Agent/R1-V) \u003cimg src=\"https://img.shields.io/github/stars/Deep-Agent/R1-V?style=social\"/\u003e : R1-V: Reinforcing Super Generalization Ability in Vision Language Models with Less Than $3.\r\n\r\n        - [LLaVA](https://github.com/haotian-liu/LLaVA) \u003cimg src=\"https://img.shields.io/github/stars/haotian-liu/LLaVA?style=social\"/\u003e : 🌋 LLaVA: Large Language and Vision Assistant. Visual instruction tuning towards large language and vision models with GPT-4 level capabilities. [llava.hliu.cc](https://llava.hliu.cc/). \"Visual Instruction Tuning\". (**[arXiv 2023](https://arxiv.org/abs/2304.08485)**).\r\n\r\n        - [NVILA](https://github.com/NVlabs/VILA) \u003cimg src=\"https://img.shields.io/github/stars/NVlabs/VILA?style=social\"/\u003e : VILA - a multi-image visual language model with training, inference and evaluation recipe, deployable from cloud to edge (Jetson Orin and laptops). \"NVILA: Efficient Frontier Visual Language Models\". (**[arXiv 2024](https://arxiv.org/abs/2412.04468)**).\r\n\r\n        - [Visual ChatGPT](https://github.com/microsoft/visual-chatgpt) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/visual-chatgpt?style=social\"/\u003e : Visual ChatGPT connects ChatGPT and a series of Visual Foundation Models to enable sending and receiving images during chatting. \"Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models\". (**[arXiv 2023](https://arxiv.org/abs/2303.04671)**).\r\n\r\n        - [CLIP](https://github.com/openai/CLIP) \u003cimg src=\"https://img.shields.io/github/stars/openai/CLIP?style=social\"/\u003e : CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image. \"Learning Transferable Visual Models From Natural Language Supervision\". (**[arXiv 2021](https://arxiv.org/abs/2103.00020)**).\r\n\r\n        - [OpenCLIP](https://github.com/mlfoundations/open_clip) \u003cimg src=\"https://img.shields.io/github/stars/mlfoundations/open_clip?style=social\"/\u003e : Welcome to an open source implementation of OpenAI's [CLIP](https://arxiv.org/abs/2103.00020) (Contrastive Language-Image Pre-training). \"Reproducible scaling laws for contrastive language-image learning\". (**[arXiv 2022](https://arxiv.org/abs/2212.07143)**).\r\n\r\n        - [GLIP](https://github.com/microsoft/GLIP) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/GLIP?style=social\"/\u003e : \"Grounded Language-Image Pre-training\". (**[CVPR 2022](https://arxiv.org/abs/2112.03857)**).\r\n\r\n        - [GLIPv2](https://github.com/microsoft/GLIP) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/GLIP?style=social\"/\u003e : \"GLIPv2: Unifying Localization and Vision-Language Understanding\". (**[arXiv 2022](https://arxiv.org/abs/2206.05836)**).\r\n\r\n        - [InternImage](https://github.com/OpenGVLab/InternImage) \u003cimg src=\"https://img.shields.io/github/stars/OpenGVLab/InternImage?style=social\"/\u003e : \"InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions\". (**[CVPR 2023](https://arxiv.org/abs/2211.05778)**).\r\n\r\n        - [SAM](https://github.com/facebookresearch/segment-anything) \u003cimg src=\"https://img.shields.io/github/stars/facebookresearch/segment-anything?style=social\"/\u003e : The repository provides code for running inference with the Segment Anything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model. \"Segment Anything\". (**[arXiv 2023](https://arxiv.org/abs/2304.02643)**).\r\n\r\n        - [Grounded-SAM](https://github.com/IDEA-Research/Grounded-Segment-Anything) \u003cimg src=\"https://img.shields.io/github/stars/IDEA-Research/Grounded-Segment-Anything?style=social\"/\u003e : Marrying Grounding DINO with Segment Anything \u0026 Stable Diffusion \u0026 Tag2Text \u0026 BLIP \u0026 Whisper \u0026 ChatBot - Automatically Detect , Segment and Generate Anything with Image, Text, and Audio Inputs. We plan to create a very interesting demo by combining [Grounding DINO](https://github.com/IDEA-Research/GroundingDINO) and [Segment Anything](https://github.com/facebookresearch/segment-anything) which aims to detect and segment Anything with text inputs!\r\n\r\n        - [SEEM](https://github.com/UX-Decoder/Segment-Everything-Everywhere-All-At-Once) \u003cimg src=\"https://img.shields.io/github/stars/UX-Decoder/Segment-Everything-Everywhere-All-At-Once?style=social\"/\u003e : We introduce SEEM that can Segment Everything Everywhere with Multi-modal prompts all at once. SEEM allows users to easily segment an image using prompts of different types including visual prompts (points, marks, boxes, scribbles and image segments) and language prompts (text and audio), etc. It can also work with any combinations of prompts or generalize to custom prompts! \"Segment Everything Everywhere All at Once\". (**[arXiv 2023](https://arxiv.org/abs/2304.06718)**).\r\n\r\n        - [SAM3D](https://github.com/DYZhang09/SAM3D) \u003cimg src=\"https://img.shields.io/github/stars/DYZhang09/SAM3D?style=social\"/\u003e : \"SAM3D: Zero-Shot 3D Object Detection via [Segment Anything](https://github.com/facebookresearch/segment-anything) Model\". (**[arXiv 2023](https://arxiv.org/abs/2306.02245)**).\r\n\r\n        - [ImageBind](https://github.com/facebookresearch/ImageBind) \u003cimg src=\"https://img.shields.io/github/stars/facebookresearch/ImageBind?style=social\"/\u003e : \"ImageBind: One Embedding Space To Bind Them All\". (**[CVPR 2023](https://arxiv.org/abs/2305.05665)**).\r\n\r\n        - [Track-Anything](https://github.com/gaomingqi/Track-Anything) \u003cimg src=\"https://img.shields.io/github/stars/gaomingqi/Track-Anything?style=social\"/\u003e : Track-Anything is a flexible and interactive tool for video object tracking and segmentation, based on Segment Anything, XMem, and E2FGVI. \"Track Anything: Segment Anything Meets Videos\". (**[arXiv 2023](https://arxiv.org/abs/2304.11968)**).\r\n\r\n        - [qianqianwang68/omnimotion](https://github.com/qianqianwang68/omnimotion) \u003cimg src=\"https://img.shields.io/github/stars/qianqianwang68/omnimotion?style=social\"/\u003e : \"Tracking Everything Everywhere All at Once\". (**[arXiv 2023](https://arxiv.org/abs/2306.05422)**).\r\n\r\n        - [M3I-Pretraining](https://github.com/OpenGVLab/M3I-Pretraining) \u003cimg src=\"https://img.shields.io/github/stars/OpenGVLab/M3I-Pretraining?style=social\"/\u003e : \"Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information\". (**[arXiv 2022](https://arxiv.org/abs/2211.09807)**).\r\n\r\n        - [BEVFormer](https://github.com/fundamentalvision/BEVFormer) \u003cimg src=\"https://img.shields.io/github/stars/fundamentalvision/BEVFormer?style=social\"/\u003e : BEVFormer: a Cutting-edge Baseline for Camera-based Detection. \"BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers\". (**[arXiv 2022](https://arxiv.org/abs/2203.17270)**).\r\n\r\n        - [Uni-Perceiver](https://github.com/fundamentalvision/Uni-Perceiver) \u003cimg src=\"https://img.shields.io/github/stars/fundamentalvision/Uni-Perceiver?style=social\"/\u003e : \"Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks\". (**[CVPR 2022](https://openaccess.thecvf.com/content/CVPR2022/html/Zhu_Uni-Perceiver_Pre-Training_Unified_Architecture_for_Generic_Perception_for_Zero-Shot_and_CVPR_2022_paper.html)**).\r\n\r\n        - [AnyLabeling](https://github.com/vietanhdev/anylabeling) \u003cimg src=\"https://img.shields.io/github/stars/vietanhdev/anylabeling?style=social\"/\u003e : 🌟 AnyLabeling 🌟. Effortless data labeling with AI support from YOLO and Segment Anything! Effortless data labeling with AI support from YOLO and Segment Anything!\r\n\r\n        - [X-AnyLabeling](https://github.com/CVHub520/X-AnyLabeling) \u003cimg src=\"https://img.shields.io/github/stars/CVHub520/X-AnyLabeling?style=social\"/\u003e : 💫 X-AnyLabeling 💫. Effortless data labeling with AI support from Segment Anything and other awesome models!\r\n\r\n        - [Label Anything](https://github.com/open-mmlab/playground/tree/main/label_anything) \u003cimg src=\"https://img.shields.io/github/stars/open-mmlab/playground?style=social\"/\u003e : OpenMMLab PlayGround: Semi-Automated Annotation with Label-Studio and SAM.\r\n\r\n        - [RevCol](https://github.com/megvii-research/RevCol) \u003cimg src=\"https://img.shields.io/github/stars/megvii-research/RevCol?style=social\"/\u003e : \"Reversible Column Networks\". (**[arXiv 2023](https://arxiv.org/abs/2212.11696)**).\r\n\r\n        - [Macaw-LLM](https://github.com/lyuchenyang/Macaw-LLM) \u003cimg src=\"https://img.shields.io/github/stars/lyuchenyang/Macaw-LLM?style=social\"/\u003e : Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.\r\n\r\n        - [SAM-PT](https://github.com/SysCV/sam-pt) \u003cimg src=\"https://img.shields.io/github/stars/SysCV/sam-pt?style=social\"/\u003e : SAM-PT: Extending SAM to zero-shot video segmentation with point-based tracking. \"Segment Anything Meets Point Tracking\". (**[arXiv 2023](https://arxiv.org/abs/2307.01197)**).\r\n\r\n        - [Video-LLaMA](https://github.com/DAMO-NLP-SG/Video-LLaMA) \u003cimg src=\"https://img.shields.io/github/stars/DAMO-NLP-SG/Video-LLaMA?style=social\"/\u003e : \"Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding\". (**[arXiv 2023](https://arxiv.org/abs/2306.02858)**).\r\n\r\n        - [Video-LLaVA](https://github.com/PKU-YuanGroup/Video-LLaVA) \u003cimg src=\"https://img.shields.io/github/stars/PKU-YuanGroup/Video-LLaVA?style=social\"/\u003e : \"Video-LLaVA: Learning United Visual Representation by Alignment Before Projection\". (**[EMNLP 2024](https://arxiv.org/pdf/2311.10122.pdf)**).\r\n\r\n        - [MobileSAM](https://github.com/ChaoningZhang/MobileSAM) \u003cimg src=\"https://img.shields.io/github/stars/ChaoningZhang/MobileSAM?style=social\"/\u003e : \"Faster Segment Anything: Towards Lightweight SAM for Mobile Applications\". (**[arXiv 2023](https://arxiv.org/abs/2306.14289)**).\r\n\r\n        - [BuboGPT](https://github.com/magic-research/bubogpt) \u003cimg src=\"https://img.shields.io/github/stars/magic-research/bubogpt?style=social\"/\u003e : \"BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs\". (**[arXiv 2023](https://arxiv.org/abs/2307.08581)**).\r\n\r\n\r\n\r\n\r\n      - ##### Vision Language Action\r\n        ###### 视觉语言动作大模型（VLA）\r\n\r\n        - [Embodied-R](https://github.com/EmbodiedCity/Embodied-R.code) \u003cimg src=\"https://img.shields.io/github/stars/EmbodiedCity/Embodied-R.code?style=social\"/\u003e : \"Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning\". (**[arXiv 2025](https://arxiv.org/abs/2504.12680)**).\r\n\r\n\r\n\r\n\r\n      - ##### AI Generated Content\r\n        ###### 人工智能生成内容（AIGC）\r\n\r\n        - [Wan2.1](https://github.com/Wan-Video/Wan2.1) \u003cimg src=\"https://img.shields.io/github/stars/Wan-Video/Wan2.1?style=social\"/\u003e : Wan: Open and Advanced Large-Scale Video Generative Models.\r\n\r\n        - [Sora](https://openai.com/sora) : Sora is an AI model that can create realistic and imaginative scenes from text instructions.\r\n\r\n        - [Open Sora Plan](https://github.com/PKU-YuanGroup/Open-Sora-Plan) \u003cimg src=\"https://img.shields.io/github/stars/PKU-YuanGroup/Open-Sora-Plan?style=social\"/\u003e : This project aim to reproducing [Sora](https://openai.com/sora) (Open AI T2V model), but we only have limited resource. We deeply wish the all open source community can contribute to this project. 本项目希望通过开源社区的力量复现Sora，由北大-兔展AIGC联合实验室共同发起，当前我们资源有限仅搭建了基础架构，无法进行完整训练，希望通过开源社区逐步增加模块并筹集资源进行训练，当前版本离目标差距巨大，仍需持续完善和快速迭代，欢迎Pull request！！！[Project Page](https://pku-yuangroup.github.io/Open-Sora-Plan/) [中文主页](https://pku-yuangroup.github.io/Open-Sora-Plan/blog_cn.html)\r\n\r\n        - [Mini Sora](https://github.com/mini-sora/minisora) \u003cimg src=\"https://img.shields.io/github/stars/mini-sora/minisora?style=social\"/\u003e : The Mini Sora project aims to explore the implementation path and future development direction of Sora.\r\n\r\n        - [EMO](https://github.com/HumanAIGC/EMO) \u003cimg src=\"https://img.shields.io/github/stars/HumanAIGC/EMO?style=social\"/\u003e : \"EMO: Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions\". (**[arXiv 2024](https://arxiv.org/abs/2402.17485)**).\r\n\r\n        - [Stable Diffusion](https://github.com/CompVis/stable-diffusion) \u003cimg src=\"https://img.shields.io/github/stars/CompVis/stable-diffusion?style=social\"/\u003e : Stable Diffusion is a latent text-to-image diffusion model. Stable Diffusion was made possible thanks to a collaboration with [Stability AI](https://stability.ai/) and [Runway](https://runwayml.com/) and builds upon our previous work \"High-Resolution Image Synthesis with Latent Diffusion Models\". (**[CVPR 2022](https://openaccess.thecvf.com/content/CVPR2022/html/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper.html)**).\r\n\r\n        - [Stable Diffusion Version 2](https://github.com/Stability-AI/stablediffusion) \u003cimg src=\"https://img.shields.io/github/stars/Stability-AI/stablediffusion?style=social\"/\u003e : This repository contains [Stable Diffusion](https://github.com/CompVis/stable-diffusion) models trained from scratch and will be continuously updated with new checkpoints. \"High-Resolution Image Synthesis with Latent Diffusion Models\". (**[CVPR 2022](https://openaccess.thecvf.com/content/CVPR2022/html/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper.html)**).\r\n\r\n        - [StableStudio](https://github.com/Stability-AI/StableStudio) \u003cimg src=\"https://img.shields.io/github/stars/Stability-AI/StableStudio?style=social\"/\u003e : StableStudio by [Stability AI](https://stability.ai/). 👋 Welcome to the community repository for StableStudio, the open-source version of [DreamStudio](https://dreamstudio.ai/).\r\n\r\n        - [AudioCraft](https://github.com/facebookresearch/audiocraft) \u003cimg src=\"https://img.shields.io/github/stars/facebookresearch/audiocraft?style=social\"/\u003e : Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.\r\n\r\n        - [InvokeAI](https://github.com/invoke-ai/InvokeAI) \u003cimg src=\"https://img.shields.io/github/stars/invoke-ai/InvokeAI?style=social\"/\u003e : Invoke AI - Generative AI for Professional Creatives. Professional Creative Tools for Stable Diffusion, Custom-Trained Models, and more. [invoke-ai.github.io/InvokeAI/](https://invoke-ai.github.io/InvokeAI/)\r\n\r\n        - [DragGAN](https://github.com/XingangPan/DragGAN) \u003cimg src=\"https://img.shields.io/github/stars/XingangPan/DragGAN?style=social\"/\u003e : \"Stable Diffusion Training with MosaicML. This repo contains code used to train your own Stable Diffusion model on your own data\". (**[SIGGRAPH 2023](https://vcai.mpi-inf.mpg.de/projects/DragGAN/)**).\r\n\r\n        - [AudioGPT](https://github.com/AIGC-Audio/AudioGPT) \u003cimg src=\"https://img.shields.io/github/stars/AIGC-Audio/AudioGPT?style=social\"/\u003e : AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head.\r\n\r\n        - [PandasAI](https://github.com/gventuri/pandas-ai) \u003cimg src=\"https://img.shields.io/github/stars/gventuri/pandas-ai?style=social\"/\u003e : Pandas AI is a Python library that adds generative artificial intelligence capabilities to Pandas, the popular data analysis and manipulation tool. It is designed to be used in conjunction with Pandas, and is not a replacement for it.\r\n\r\n        - [mosaicml/diffusion](https://github.com/mosaicml/diffusion) \u003cimg src=\"https://img.shields.io/github/stars/mosaicml/diffusion?style=social\"/\u003e : Stable Diffusion Training with MosaicML. This repo contains code used to train your own Stable Diffusion model on your own data.\r\n\r\n        - [VisorGPT](https://github.com/Sierkinhane/VisorGPT) \u003cimg src=\"https://img.shields.io/github/stars/Sierkinhane/VisorGPT?style=social\"/\u003e : Customize spatial layouts for conditional image synthesis models, e.g., ControlNet, using GPT. \"VisorGPT: Learning Visual Prior via Generative Pre-Training\". (**[arXiv 2023](https://arxiv.org/abs/2305.13777)**).\r\n\r\n        - [ControlNet](https://github.com/lllyasviel/ControlNet) \u003cimg src=\"https://img.shields.io/github/stars/lllyasviel/ControlNet?style=social\"/\u003e : Let us control diffusion models! \"Adding Conditional Control to Text-to-Image Diffusion Models\". (**[arXiv 2023](https://arxiv.org/abs/2302.05543)**).\r\n\r\n        - [Fooocus](https://github.com/lllyasviel/Fooocus) \u003cimg src=\"https://img.shields.io/github/stars/lllyasviel/Fooocus?style=social\"/\u003e : Fooocus is an image generating software. Fooocus is a rethinking of Stable Diffusion and Midjourney’s designs. \"微信公众号「GitHubStore」《[Fooocus : 集Stable Diffusion 和 Midjourney 优点于一身的开源AI绘图软件](https://mp.weixin.qq.com/s/adyXek6xcz5aOPAGqZBrvg)》\"。\r\n\r\n        - [MindDiffuser](https://github.com/ReedOnePeck/MindDiffuser) \u003cimg src=\"https://img.shields.io/github/stars/ReedOnePeck/MindDiffuser?style=social\"/\u003e : \"MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion\". (**[arXiv 2023](https://arxiv.org/abs/2308.04249)**).\r\n\r\n\r\n\r\n        - [World Labs](https://www.worldlabs.ai/) : We are a spatial intelligence company building Large World Models to perceive, generate, and interact with the 3D world.\r\n\r\n        - [Genie 2](https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/) : Genie 2: A large-scale foundation world model.\r\n\r\n        - [Midjourney](https://www.midjourney.com/) : Midjourney is an independent research lab exploring new mediums of thought and expanding the imaginative powers of the human species.\r\n\r\n        - [DreamStudio](https://dreamstudio.ai/) : Effortless image generation for creators with big dreams.\r\n\r\n        - [Firefly](https://www.adobe.com/sensei/generative-ai/firefly.html) : Adobe Firefly: Experiment, imagine, and make an infinite range of creations with Firefly, a family of creative generative AI models coming to Adobe products.\r\n\r\n        - [Jasper](https://www.jasper.ai/) : Meet Jasper. On-brand AI content wherever you create.\r\n\r\n        - [Copy.ai](https://www.copy.ai/) : Whatever you want to ask, our chat has the answers.\r\n\r\n        - [Peppertype.ai](https://www.peppercontent.io/peppertype-ai/) : Leverage the AI-powered platform to ideate, create, distribute, and measure your content and prove your content marketing ROI.\r\n\r\n        - [ChatPPT](https://chat-ppt.com/) : ChatPPT来袭命令式一键生成PPT。\r\n\r\n\r\n\r\n\r\n\r\n\r\n    - #### Performance Analysis and Visualization\r\n      ##### 性能分析及可视化\r\n\r\n        - [FlagPerf](https://github.com/FlagOpen/FlagPerf) \u003cimg src=\"https://img.shields.io/github/stars/FlagOpen/FlagPerf?style=social\"/\u003e : FlagPerf is an open-source software platform for benchmarking AI chips. FlagPerf是智源研究院联合AI硬件厂商共建的一体化AI硬件评测引擎，旨在建立以产业实践为导向的指标体系，评测AI硬件在软件栈组合（模型+框架+编译器）下的实际能力。\r\n\r\n        - [hahnyuan/LLM-Viewer](https://github.com/hahnyuan/LLM-Viewer) \u003cimg src=\"https://img.shields.io/github/stars/hahnyuan/LLM-Viewer?style=social\"/\u003e : Analyze the inference of Large Language Models (LLMs). Analyze aspects like computation, storage, transmission, and hardware roofline model in a user-friendly interface.\r\n\r\n        - [harleyszhang/llm_counts](https://github.com/harleyszhang/llm_counts) \u003cimg src=\"https://img.shields.io/github/stars/harleyszhang/llm_counts?style=social\"/\u003e : llm theoretical performance analysis tools and support params, flops, memory and latency analysis.\r\n\r\n\r\n\r\n    - #### Training and Fine-Tuning Framework\r\n      ##### 训练和微调框架\r\n\r\n        - [DeepSpeed](https://github.com/deepspeedai/DeepSpeed) \u003cimg src=\"https://img.shields.io/github/stars/deepspeedai/DeepSpeed?style=social\"/\u003e : DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective. [www.deepspeed.ai/](https://www.deepspeed.ai/)\r\n\r\n        - [unsloth](https://github.com/unslothai/unsloth) \u003cimg src=\"https://img.shields.io/github/stars/unslothai/unsloth?style=social\"/\u003e : Finetune Llama 3.3, DeepSeek-R1 \u0026 Reasoning LLMs 2x faster with 70% less memory. [unsloth.ai](https://unsloth.ai/)\r\n\r\n        - [LLaMA-Factory](https://github.com/hiyouga/LLaMA-Factory) \u003cimg src=\"https://img.shields.io/github/stars/hiyouga/LLaMA-Factory?style=social\"/\u003e : Unified Efficient Fine-Tuning of 100+ LLMs \u0026 VLMs (ACL 2024). \"LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models\". (**[arXiv 2024](https://arxiv.org/abs/2403.13372)**).\r\n\r\n\r\n    - #### Reinforcement Learning Framework\r\n      ##### 强化学习框架\r\n\r\n        - [TTRL](https://github.com/PRIME-RL/TTRL) \u003cimg src=\"https://img.shields.io/github/stars/PRIME-RL/TTRL?style=social\"/\u003e : \"TTRL: Test-Time Reinforcement Learning\". (**[arXiv 2025](https://arxiv.org/abs/2504.16084)**).\r\n\r\n\r\n\r\n\r\n\r\n    - #### LLM Inference Framework\r\n      ##### 大语言模型推理框架\r\n\r\n\r\n        - ##### LLM Inference and Serving Engine\r\n\r\n            - [TensorRT](https://github.com/NVIDIA/TensorRT) \u003cimg src=\"https://img.shields.io/github/stars/NVIDIA/TensorRT?style=social\"/\u003e : NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT. [developer.nvidia.com/tensorrt](https://developer.nvidia.com/tensorrt)\r\n\r\n            - [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) \u003cimg src=\"https://img.shields.io/github/stars/NVIDIA/TensorRT-LLM?style=social\"/\u003e : TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines. [nvidia.github.io/TensorRT-LLM](https://nvidia.github.io/TensorRT-LLM)\r\n\r\n            - [NVIDIA/TensorRT-Model-Optimizer](https://github.com/NVIDIA/TensorRT-Model-Optimizer) \u003cimg src=\"https://img.shields.io/github/stars/NVIDIA/TensorRT-Model-Optimizer?style=social\"/\u003e : TensorRT Model Optimizer is a unified library of state-of-the-art model optimization techniques such as quantization, pruning, distillation, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM or TensorRT to optimize inference speed on NVIDIA GPUs. [nvidia.github.io/TensorRT-Model-Optimizer](https://nvidia.github.io/TensorRT-Model-Optimizer/)\r\n\r\n            - [Ollama](https://github.com/ollama/ollama) \u003cimg src=\"https://img.shields.io/github/stars/ollama/ollama?style=social\"/\u003e : Get up and running with Llama 3.3, DeepSeek-R1, Phi-4, Gemma 2, and other large language models. [ollama.com](https://ollama.com/)\r\n\r\n            - [vLLM](https://github.com/vllm-project/vllm) \u003cimg src=\"https://img.shields.io/github/stars/vllm-project/vllm?style=social\"/\u003e : A high-throughput and memory-efficient inference and serving engine for LLMs. [docs.vllm.ai](https://docs.vllm.ai/)\r\n\r\n            - [Nano-vLLM](https://github.com/GeeeekExplorer/nano-vllm) \u003cimg src=\"https://img.shields.io/github/stars/GeeeekExplorer/nano-vllm?style=social\"/\u003e : A lightweight vLLM implementation built from scratch.\r\n\r\n            - [SGLang](https://github.com/sgl-project/sglang) \u003cimg src=\"https://img.shields.io/github/stars/sgl-project/sglang?style=social\"/\u003e : SGLang is a fast serving framework for large language models and vision language models. [docs.sglang.ai/](https://docs.sglang.ai/)\r\n\r\n            - [llama.cpp](https://github.com/ggerganov/llama.cpp) \u003cimg src=\"https://img.shields.io/github/stars/ggerganov/llama.cpp?style=social\"/\u003e : LLM inference in C/C++.\r\n\r\n            - [FlashInfer](https://github.com/flashinfer-ai/flashinfer) \u003cimg src=\"https://img.shields.io/github/stars/flashinfer-ai/flashinfer?style=social\"/\u003e : FlashInfer: Kernel Library for LLM Serving . [flashinfer.ai](flashinfer.ai)\r\n\r\n            - [Chitu（赤兔）](https://github.com/thu-pacman/chitu) \u003cimg src=\"https://img.shields.io/github/stars/thu-pacman/chitu?style=social\"/\u003e : High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability. Chitu (赤兔) 是一个专注于效率、灵活性和可用性的高性能大语言模型推理框架。\r\n\r\n            - [ztxz16/fastllm](https://github.com/ztxz16/fastllm) \u003cimg src=\"https://img.shields.io/github/stars/ztxz16/fastllm?style=social\"/\u003e : fastllm是c++实现自有算子替代Pytorch的高性能全功能大模型推理库，可以推理Qwen, Llama, Phi等稠密模型，以及DeepSeek, Qwen-moe等moe模型。fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型，任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型，单并发20tps；INT4量化模型单并发30tps，多并发可达60+。\r\n\r\n            - [MLC LLM](https://github.com/mlc-ai/mlc-llm) \u003cimg src=\"https://img.shields.io/github/stars/mlc-ai/mlc-llm?style=social\"/\u003e : Universal LLM Deployment Engine with ML Compilation. [llm.mlc.ai/](https://llm.mlc.ai/)\r\n\r\n            - [KTransformers](https://github.com/kvcache-ai/ktransformers) \u003cimg src=\"https://img.shields.io/github/stars/kvcache-ai/ktransformers?style=social\"/\u003e : A Flexible Framework for Experiencing Cutting-edge LLM Inference Optimizations. [kvcache-ai.github.io/ktransformers/](https://kvcache-ai.github.io/ktransformers/)\r\n\r\n            - [Aphrodite](https://github.com/aphrodite-engine/aphrodite-engine) \u003cimg src=\"https://img.shields.io/github/stars/aphrodite-engine/aphrodite-engine?style=social\"/\u003e : Large-scale LLM inference engine. [aphrodite.pygmalion.chat](https://aphrodite.pygmalion.chat/)\r\n\r\n            - [GPUStack](https://github.com/gpustack/gpustack) \u003cimg src=\"https://img.shields.io/github/stars/gpustack/gpustack?style=social\"/\u003e : GPUStack is an open-source GPU cluster manager for running AI models. Manage GPU clusters for running AI models. [gpustack.ai](https://gpustack.ai/)\r\n\r\n            - [Lamini](https://github.com/lamini-ai/lamini) \u003cimg src=\"https://img.shields.io/github/stars/lamini-ai/lamini?style=social\"/\u003e : The Official Python Client for Lamini's API. [lamini.ai/](https://lamini.ai/)\r\n\r\n            - [triton-inference-server/tensorrtllm_backend](https://github.com/triton-inference-server/tensorrtllm_backend) \u003cimg src=\"https://img.shields.io/github/stars/triton-inference-server/tensorrtllm_backend?style=social\"/\u003e : The Triton TensorRT-LLM Backend.\r\n\r\n            - [datawhalechina/self-llm](https://github.com/datawhalechina/self-llm) \u003cimg src=\"https://img.shields.io/github/stars/datawhalechina/self-llm?style=social\"/\u003e :  《开源大模型食用指南》基于Linux环境快速部署开源大模型，更适合中国宝宝的部署教程。\r\n\r\n            - [ninehills/llm-inference-benchmark](https://github.com/ninehills/llm-inference-benchmark) \u003cimg src=\"https://img.shields.io/github/stars/ninehills/llm-inference-benchmark?style=social\"/\u003e : LLM Inference benchmark.\r\n\r\n            - [csbench/csbench](https://github.com/csbench/csbench) \u003cimg src=\"https://img.shields.io/github/stars/csbench/csbench?style=social\"/\u003e : \"CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery\". (**[arXiv 2024](https://arxiv.org/abs/2406.08587)**).\r\n\r\n            - [MooreThreads/vllm_musa](https://github.com/MooreThreads/vllm_musa) \u003cimg src=\"https://img.shields.io/github/stars/MooreThreads/vllm_musa?style=social\"/\u003e : A high-throughput and memory-efficient inference and serving engine for LLMs. [docs.vllm.ai](https://docs.vllm.ai/)\r\n\r\n\r\n\r\n\r\n        - ##### High Performance Kernel Library\r\n\r\n            - [Liger-Kernel](https://github.com/linkedin/Liger-Kernel) \u003cimg src=\"https://img.shields.io/github/stars/linkedin/Liger-Kernel?style=social\"/\u003e : Efficient Triton Kernels for LLM Training. [arxiv.org/pdf/2410.10989](https://arxiv.org/pdf/2410.10989)\r\n\r\n            - [FlashInfer](https://github.com/flashinfer-ai/flashinfer) \u003cimg src=\"https://img.shields.io/github/stars/flashinfer-ai/flashinfer?style=social\"/\u003e : FlashInfer: Kernel Library for LLM Serving . [flashinfer.ai](flashinfer.ai)\r\n\r\n            - [DeepGEMM](https://github.com/deepseek-ai/DeepGEMM) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/DeepGEMM?style=social\"/\u003e : DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling.\r\n\r\n            - [FlashMLA](https://github.com/deepseek-ai/FlashMLA) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/FlashMLA?style=social\"/\u003e : FlashMLA: Efficient MLA Decoding Kernel for Hopper GPUs.\r\n\r\n            - [DeepEP](https://github.com/deepseek-ai/DeepEP) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/DeepEP?style=social\"/\u003e : DeepEP: an efficient expert-parallel communication library.\r\n\r\n            - [FlagGems](https://github.com/FlagOpen/FlagGems) \u003cimg src=\"https://img.shields.io/github/stars/FlagOpen/FlagGems?style=social\"/\u003e : FlagGems is a high-performance general operator library implemented in [OpenAI Triton](https://github.com/openai/triton). It aims to provide a suite of kernel functions to accelerate LLM training and inference.\r\n\r\n\r\n\r\n\r\n        - ##### C and CPP Implementation\r\n\r\n            - [llama.cpp](https://github.com/ggerganov/llama.cpp) \u003cimg src=\"https://img.shields.io/github/stars/ggerganov/llama.cpp?style=social\"/\u003e : LLM inference in C/C++.\r\n\r\n            - [llm.c](https://github.com/karpathy/llm.c) \u003cimg src=\"https://img.shields.io/github/stars/karpathy/llm.c?style=social\"/\u003e : LLM training in simple, pure C/CUDA. There is no need for 245MB of PyTorch or 107MB of cPython. For example, training GPT-2 (CPU, fp32) is ~1,000 lines of clean code in a single file. It compiles and runs instantly, and exactly matches the PyTorch reference implementation.\r\n\r\n            - [llama2.c](https://github.com/karpathy/llama2.c) \u003cimg src=\"https://img.shields.io/github/stars/karpathy/llama2.c?style=social\"/\u003e : Inference Llama 2 in one file of pure C. Train the Llama 2 LLM architecture in PyTorch then inference it with one simple 700-line C file (run.c).\r\n\r\n            - [TinyChatEngine](https://github.com/mit-han-lab/TinyChatEngine) \u003cimg src=\"https://img.shields.io/github/stars/mit-han-lab/TinyChatEngine?style=social\"/\u003e : TinyChatEngine: On-Device LLM Inference Library. Running large language models (LLMs) and visual language models (VLMs) on the edge is useful: copilot services (coding, office, smart reply) on laptops, cars, robots, and more. Users can get instant responses with better privacy, as the data is local. This is enabled by LLM model compression technique: [SmoothQuant](https://github.com/mit-han-lab/smoothquant) and [AWQ (Activation-aware Weight Quantization)](https://github.com/mit-han-lab/llm-awq), co-designed with TinyChatEngine that implements the compressed low-precision model. Feel free to check out our [slides](https://github.com/mit-han-lab/TinyChatEngine/blob/main/assets/slides.pdf) for more details!\r\n\r\n            - [gemma.cpp](https://github.com/google/gemma.cpp) \u003cimg src=\"https://img.shields.io/github/stars/google/gemma.cpp?style=social\"/\u003e :  gemma.cpp is a lightweight, standalone C++ inference engine for the Gemma foundation models from Google.\r\n\r\n            - [whisper.cpp](https://github.com/ggerganov/whisper.cpp) \u003cimg src=\"https://img.shields.io/github/stars/ggerganov/whisper.cpp?style=social\"/\u003e : High-performance inference of [OpenAI's Whisper](https://github.com/openai/whisper) automatic speech recognition (ASR) model.\r\n\r\n            - [ztxz16/fastllm](https://github.com/ztxz16/fastllm) \u003cimg src=\"https://img.shields.io/github/stars/ztxz16/fastllm?style=social\"/\u003e : fastllm是c++实现自有算子替代Pytorch的高性能全功能大模型推理库，可以推理Qwen, Llama, Phi等稠密模型，以及DeepSeek, Qwen-moe等moe模型。fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型，任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型，单并发20tps；INT4量化模型单并发30tps，多并发可达60+。\r\n\r\n            - [ChatGLM.cpp](https://github.com/li-plus/chatglm.cpp) \u003cimg src=\"https://img.shields.io/github/stars/li-plus/chatglm.cpp?style=social\"/\u003e : C++ implementation of [ChatGLM-6B](https://github.com/THUDM/ChatGLM-6B) and [ChatGLM2-6B](https://github.com/THUDM/ChatGLM2-6B).\r\n\r\n            - [MegEngine/InferLLM](https://github.com/MegEngine/InferLLM) \u003cimg src=\"https://img.shields.io/github/stars/MegEngine/InferLLM?style=social\"/\u003e : InferLLM is a lightweight LLM model inference framework that mainly references and borrows from the llama.cpp project.\r\n\r\n            - [DeployAI/nndeploy](https://github.com/DeployAI/nndeploy) \u003cimg src=\"https://img.shields.io/github/stars/DeployAI/nndeploy?style=social\"/\u003e : nndeploy是一款模型端到端部署框架。以多端推理以及基于有向无环图模型部署为内核，致力为用户提供跨平台、简单易用、高性能的模型部署体验。[nndeploy-zh.readthedocs.io/zh/latest/](https://nndeploy-zh.readthedocs.io/zh/latest/)\r\n\r\n            - [zjhellofss/KuiperInfer (自制深度学习推理框架)](https://github.com/zjhellofss/KuiperInfer) \u003cimg src=\"https://img.shields.io/github/stars/zjhellofss/KuiperInfer?style=social\"/\u003e :  带你从零实现一个高性能的深度学习推理库，支持llama 、Unet、Yolov5、Resnet等模型的推理。Implement a high-performance deep learning inference library step by step.\r\n\r\n            - [skeskinen/llama-lite](https://github.com/skeskinen/llama-lite) \u003cimg src=\"https://img.shields.io/github/stars/skeskinen/llama-lite?style=social\"/\u003e : Embeddings focused small version of Llama NLP model.\r\n\r\n            - [Const-me/Whisper](https://github.com/Const-me/Whisper) \u003cimg src=\"https://img.shields.io/github/stars/Const-me/Whisper?style=social\"/\u003e : High-performance GPGPU inference of OpenAI's Whisper automatic speech recognition (ASR) model.\r\n\r\n            - [wangzhaode/ChatGLM-MNN](https://github.com/wangzhaode/ChatGLM-MNN) \u003cimg src=\"https://img.shields.io/github/stars/wangzhaode/ChatGLM-MNN?style=social\"/\u003e : Pure C++, Easy Deploy ChatGLM-6B.\r\n\r\n            - [davidar/eigenGPT](https://github.com/davidar/eigenGPT) \u003cimg src=\"https://img.shields.io/github/stars/davidar/eigenGPT?style=social\"/\u003e : Minimal C++ implementation of GPT2.\r\n\r\n            - [Tlntin/Qwen-TensorRT-LLM](https://github.com/Tlntin/Qwen-TensorRT-LLM) \u003cimg src=\"https://img.shields.io/github/stars/Tlntin/Qwen-TensorRT-LLM?style=social\"/\u003e : 使用TRT-LLM完成对Qwen-7B-Chat实现推理加速。\r\n\r\n            - [FeiGeChuanShu/trt2023](https://github.com/FeiGeChuanShu/trt2023) \u003cimg src=\"https://img.shields.io/github/stars/FeiGeChuanShu/trt2023?style=social\"/\u003e : NVIDIA TensorRT Hackathon 2023复赛选题：通义千问Qwen-7B用TensorRT-LLM模型搭建及优化。\r\n\r\n            - [TRT2022/trtllm-llama](https://github.com/TRT2022/trtllm-llama) \u003cimg src=\"https://img.shields.io/github/stars/TRT2022/trtllm-llama?style=social\"/\u003e : ☢️ TensorRT 2023复赛——基于TensorRT-LLM的Llama模型推断加速优化。\r\n\r\n            - [AmeyaWagh/llama2.cpp](https://github.com/AmeyaWagh/llama2.cpp) \u003cimg src=\"https://img.shields.io/github/stars/AmeyaWagh/llama2.cpp?style=social\"/\u003e : Inference Llama 2 in C++.\r\n\r\n\r\n\r\n        - ##### Triton Implementation\r\n\r\n            - [Liger-Kernel](https://github.com/linkedin/Liger-Kernel) \u003cimg src=\"https://img.shields.io/github/stars/linkedin/Liger-Kernel?style=social\"/\u003e : Efficient Triton Kernels for LLM Training. [arxiv.org/pdf/2410.10989](https://arxiv.org/pdf/2410.10989)\r\n\r\n            - [FlagGems](https://github.com/FlagOpen/FlagGems) \u003cimg src=\"https://img.shields.io/github/stars/FlagOpen/FlagGems?style=social\"/\u003e : FlagGems is a high-performance general operator library implemented in [OpenAI Triton](https://github.com/openai/triton). It aims to provide a suite of kernel functions to accelerate LLM training and inference.\r\n\r\n            - [harleyszhang/lite_llama](https://github.com/harleyszhang/lite_llama) \u003cimg src=\"https://img.shields.io/github/stars/harleyszhang/lite_llama?style=social\"/\u003e : A light llama-like llm inference framework based on the triton kernel.\r\n\r\n            - [linxihui/dkernel](https://github.com/linxihui/dkernel) \u003cimg src=\"https://img.shields.io/github/stars/linxihui/dkernel?style=social\"/\u003e : This repo contains customized CUDA kernels written in OpenAI Triton. As of now, it contains the sparse attention kernel used in [phi-3-small models](https://huggingface.co/microsoft/Phi-3-small-8k-instruct). The sparse attention is also supported in vLLM for efficient inference.\r\n\r\n\r\n\r\n\r\n\r\n\r\n        - ##### Python Implementation\r\n\r\n            - [llama-cpp-python](https://github.com/abetlen/llama-cpp-python) \u003cimg src=\"https://img.shields.io/github/stars/abetlen/llama-cpp-python?style=social\"/\u003e : Python bindings for llama.cpp. [llama-cpp-python.readthedocs.io](https://llama-cpp-python.readthedocs.io/)\r\n\r\n            - [ggml-python](https://github.com/abetlen/ggml-python) \u003cimg src=\"https://img.shields.io/github/stars/abetlen/ggml-python?style=social\"/\u003e : Python bindings for ggml. [ggml-python.readthedocs.io](https://ggml-python.readthedocs.io/)\r\n\r\n\r\n\r\n\r\n        - ##### Mojo Implementation\r\n\r\n            - [llama2.mojo](https://github.com/tairov/llama2.mojo) \u003cimg src=\"https://img.shields.io/github/stars/tairov/llama2.mojo?style=social\"/\u003e : Inference Llama 2 in one file of pure 🔥\r\n\r\n            - [dorjeduck/llm.mojo](https://github.com/dorjeduck/llm.mojo) \u003cimg src=\"https://img.shields.io/github/stars/dorjeduck/llm.mojo?style=social\"/\u003e : port of Andrjey Karpathy's llm.c to Mojo.\r\n\r\n\r\n        - ##### Rust Implementation\r\n\r\n            - [Candle](https://github.com/huggingface/candle) \u003cimg src=\"https://img.shields.io/github/stars/huggingface/candle?style=social\"/\u003e : Minimalist ML framework for Rust.\r\n\r\n            - [Safetensors](https://github.com/huggingface/safetensors) \u003cimg src=\"https://img.shields.io/github/stars/huggingface/safetensors?style=social\"/\u003e : Simple, safe way to store and distribute tensors. [huggingface.co/docs/safetensors](https://huggingface.co/docs/safetensors/index)\r\n\r\n            - [Tokenizers](https://github.com/huggingface/tokenizers) \u003cimg src=\"https://img.shields.io/github/stars/huggingface/tokenizers?style=social\"/\u003e : 💥 Fast State-of-the-Art Tokenizers optimized for Research and Production. [huggingface.co/docs/tokenizers](https://huggingface.co/docs/tokenizers/index)\r\n\r\n            - [Burn](https://github.com/burn-rs/burn) \u003cimg src=\"https://img.shields.io/github/stars/burn-rs/burn?style=social\"/\u003e : Burn - A Flexible and Comprehensive Deep Learning Framework in Rust. [burn-rs.github.io/](https://burn-rs.github.io/)\r\n\r\n            - [dfdx](https://github.com/coreylowman/dfdx) \u003cimg src=\"https://img.shields.io/github/stars/coreylowman/dfdx?style=social\"/\u003e : Deep learning in Rust, with shape checked tensors and neural networks.\r\n\r\n            - [luminal](https://github.com/jafioti/luminal) \u003cimg src=\"https://img.shields.io/github/stars/jafioti/luminal?style=social\"/\u003e : Deep learning at the speed of light. [www.luminalai.com/](https://www.luminalai.com/)\r\n\r\n            - [crabml](https://github.com/crabml/crabml) \u003cimg src=\"https://img.shields.io/github/stars/crabml/crabml?style=social\"/\u003e : crabml is focusing on the reimplementation of GGML using the Rust programming language.\r\n\r\n            - [TensorFlow Rust](https://github.com/tensorflow/rust) \u003cimg src=\"https://img.shields.io/github/stars/tensorflow/rust?style=social\"/\u003e : Rust language bindings for TensorFlow.\r\n\r\n            - [tch-rs](https://github.com/LaurentMazare/tch-rs) \u003cimg src=\"https://img.shields.io/github/stars/LaurentMazare/tch-rs?style=social\"/\u003e : Rust bindings for the C++ api of PyTorch.\r\n\r\n            - [rustai-solutions/candle_demo_openchat_35](https://github.com/rustai-solutions/candle_demo_openchat_35) \u003cimg src=\"https://img.shields.io/github/stars/rustai-solutions/candle_demo_openchat_35?style=social\"/\u003e : candle_demo_openchat_35.\r\n\r\n            - [llama2.rs](https://github.com/srush/llama2.rs) \u003cimg src=\"https://img.shields.io/github/stars/srush/llama2.rs?style=social\"/\u003e : A fast llama2 decoder in pure Rust.\r\n\r\n            - [Llama2-burn](https://github.com/Gadersd/llama2-burn) \u003cimg src=\"https://img.shields.io/github/stars/Gadersd/llama2-burn?style=social\"/\u003e : Llama2 LLM ported to Rust burn.\r\n\r\n            - [gaxler/llama2.rs](https://github.com/gaxler/llama2.rs) \u003cimg src=\"https://img.shields.io/github/stars/gaxler/llama2.rs?style=social\"/\u003e : Inference Llama 2 in one file of pure Rust 🦀\r\n\r\n            - [whisper-burn](https://github.com/Gadersd/whisper-burn) \u003cimg src=\"https://img.shields.io/github/stars/Gadersd/whisper-burn?style=social\"/\u003e : A Rust implementation of OpenAI's Whisper model using the burn framework.\r\n\r\n            - [stable-diffusion-burn](https://github.com/Gadersd/stable-diffusion-burn) \u003cimg src=\"https://img.shields.io/github/stars/Gadersd/stable-diffusion-burn?style=social\"/\u003e : Stable Diffusion v1.4 ported to Rust's burn framework.\r\n\r\n            - [coreylowman/llama-dfdx](https://github.com/coreylowman/llama-dfdx) \u003cimg src=\"https://img.shields.io/github/stars/coreylowman/llama-dfdx?style=social\"/\u003e : [LLaMa 7b](https://ai.facebook.com/blog/large-language-model-llama-meta-ai/) with CUDA acceleration implemented in rust. Minimal GPU memory needed!\r\n\r\n            - [tazz4843/whisper-rs](https://github.com/tazz4843/whisper-rs) \u003cimg src=\"https://img.shields.io/github/stars/tazz4843/whisper-rs?style=social\"/\u003e : Rust bindings to [whisper.cpp](https://github.com/ggerganov/whisper.cpp).\r\n\r\n            - [rustformers/llm](https://github.com/rustformers/llm) \u003cimg src=\"https://img.shields.io/github/stars/rustformers/llm?style=social\"/\u003e : Run inference for Large Language Models on CPU, with Rust 🦀🚀🦙.\r\n\r\n            - [Chidori](https://github.com/ThousandBirdsInc/chidori) \u003cimg src=\"https://img.shields.io/github/stars/ThousandBirdsInc/chidori?style=social\"/\u003e : A reactive runtime for building durable AI agents. [docs.thousandbirds.ai](https://docs.thousandbirds.ai/).\r\n\r\n            - [llm-chain](https://github.com/sobelio/llm-chain) \u003cimg src=\"https://img.shields.io/github/stars/sobelio/llm-chain?style=social\"/\u003e : llm-chain is a collection of Rust crates designed to help you work with Large Language Models (LLMs) more effectively. [llm-chain.xyz](https://llm-chain.xyz/)\r\n\r\n            - [Abraxas-365/langchain-rust](https://github.com/Abraxas-365/langchain-rust) \u003cimg src=\"https://img.shields.io/github/stars/Abraxas-365/langchain-rust?style=social\"/\u003e : 🦜️🔗LangChain for Rust, the easiest way to write LLM-based programs in Rust.\r\n\r\n            - [Atome-FE/llama-node](https://github.com/Atome-FE/llama-node) \u003cimg src=\"https://img.shields.io/github/stars/Atome-FE/llama-node?style=social\"/\u003e : Believe in AI democratization. llama for nodejs backed by llama-rs and llama.cpp, work locally on your laptop CPU. support llama/alpaca/gpt4all/vicuna model. [www.npmjs.com/package/llama-node](https://www.npmjs.com/package/llama-node)\r\n\r\n            - [Noeda/rllama](https://github.com/Noeda/rllama) \u003cimg src=\"https://img.shields.io/github/stars/Noeda/rllama?style=social\"/\u003e : Rust+OpenCL+AVX2 implementation of LLaMA inference code.\r\n\r\n            - [lencx/ChatGPT](https://github.com/lencx/ChatGPT) \u003cimg src=\"https://img.shields.io/github/stars/lencx/ChatGPT?style=social\"/\u003e : 🔮 ChatGPT Desktop Application (Mac, Windows and Linux). [NoFWL](https://app.nofwl.com/).\r\n\r\n            - [Synaptrix/ChatGPT-Desktop](https://github.com/Synaptrix/ChatGPT-Desktop) \u003cimg src=\"https://img.shields.io/github/stars/Synaptrix/ChatGPT-Desktop?style=social\"/\u003e : Fuel your productivity with ChatGPT-Desktop - Blazingly fast and supercharged!\r\n\r\n            - [Poordeveloper/chatgpt-app](https://github.com/Poordeveloper/chatgpt-app) \u003cimg src=\"https://img.shields.io/github/stars/Poordeveloper/chatgpt-app?style=social\"/\u003e : A ChatGPT App for all platforms. Built with Rust + Tauri + Vue + Axum.\r\n\r\n            - [mxismean/chatgpt-app](https://github.com/mxismean/chatgpt-app) \u003cimg src=\"https://img.shields.io/github/stars/mxismean/chatgpt-app?style=social\"/\u003e : Tauri 项目：ChatGPT App.\r\n\r\n            - [sonnylazuardi/chat-ai-desktop](https://github.com/sonnylazuardi/chat-ai-desktop) \u003cimg src=\"https://img.shields.io/github/stars/sonnylazuardi/chat-ai-desktop?style=social\"/\u003e : Chat AI Desktop App. Unofficial ChatGPT desktop app for Mac \u0026 Windows menubar using Tauri \u0026 Rust.\r\n\r\n            - [yetone/openai-translator](https://github.com/yetone/openai-translator) \u003cimg src=\"https://img.shields.io/github/stars/yetone/openai-translator?style=social\"/\u003e : The translator that does more than just translation - powered by OpenAI.\r\n\r\n            - [m1guelpf/browser-agent](https://github.com/m1guelpf/browser-agent) \u003cimg src=\"https://img.shields.io/github/stars/m1guelpf/browser-agent?style=social\"/\u003e : A browser AI agent, using GPT-4. [docs.rs/browser-agent](https://docs.rs/browser-agent/latest/browser_agent/)\r\n\r\n            - [sigoden/aichat](https://github.com/sigoden/aichat) \u003cimg src=\"https://img.shields.io/github/stars/sigoden/aichat?style=social\"/\u003e : Using ChatGPT/GPT-3.5/GPT-4 in the terminal.\r\n\r\n            - [uiuifree/rust-openai-chatgpt-api](https://github.com/uiuifree/rust-openai-chatgpt-api) \u003cimg src=\"https://img.shields.io/github/stars/uiuifree/rust-openai-chatgpt-api?style=social\"/\u003e : \"rust-openai-chatgpt-api\" is a Rust library for accessing the ChatGPT API, a powerful NLP platform by OpenAI. The library provides a simple and efficient interface for sending requests and receiving responses, including chat. It uses reqwest and serde for HTTP requests and JSON serialization.\r\n\r\n            - [1595901624/gpt-aggregated-edition](https://github.com/1595901624/gpt-aggregated-edition) \u003cimg src=\"https://img.shields.io/github/stars/1595901624/gpt-aggregated-edition?style=social\"/\u003e : 聚合ChatGPT官方版、ChatGPT免费版、文心一言、Poe、chatchat等多平台，支持自定义导入平台。\r\n\r\n            - [Cormanz/smartgpt](https://github.com/Cormanz/smartgpt) \u003cimg src=\"https://img.shields.io/github/stars/Cormanz/smartgpt?style=social\"/\u003e : A program that provides LLMs with the ability to complete complex tasks using plugins.\r\n\r\n            - [femtoGPT](https://github.com/keyvank/femtoGPT) \u003cimg src=\"https://img.shields.io/github/stars/keyvank/femtoGPT?style=social\"/\u003e : femtoGPT is a pure Rust implementation of a minimal Generative Pretrained Transformer. [discord.gg/wTJFaDVn45](https://github.com/keyvank/femtoGPT)\r\n\r\n            - [shafishlabs/llmchain-rs](https://github.com/shafishlabs/llmchain-rs) \u003cimg src=\"https://img.shields.io/github/stars/shafishlabs/llmchain-rs?style=social\"/\u003e : 🦀Rust + Large Language Models - Make AI Services Freely and Easily. Inspired by LangChain.\r\n\r\n            - [flaneur2020/llama2.rs](https://github.com/flaneur2020/llama2.rs) \u003cimg src=\"https://img.shields.io/github/stars/flaneur2020/llama2.rs?style=social\"/\u003e : An rust reimplementatin of [https://github.com/karpathy/llama2.c](https://github.com/karpathy/llama2.c).\r\n\r\n            - [Heng30/chatbox](https://github.com/Heng30/chatbox) \u003cimg src=\"https://img.shields.io/github/stars/Heng30/chatbox?style=social\"/\u003e : A Chatbot for OpenAI ChatGPT. Based on Slint-ui and Rust.\r\n\r\n            - [fairjm/dioxus-openai-qa-gui](https://github.com/fairjm/dioxus-openai-qa-gui) \u003cimg src=\"https://img.shields.io/github/stars/fairjm/dioxus-openai-qa-gui?style=social\"/\u003e : a simple openai qa desktop app built with dioxus.\r\n\r\n            - [purton-tech/bionicgpt](https://github.com/purton-tech/bionicgpt) \u003cimg src=\"https://img.shields.io/github/stars/purton-tech/bionicgpt?style=social\"/\u003e : Accelerate LLM adoption in your organisation. Chat with your confidential data safely and securely. [bionic-gpt.com](https://bionic-gpt.com/)\r\n\r\n            - [InfiniTensor/transformer-rs](https://github.com/InfiniTensor/transformer-rs) \u003cimg src=\"https://img.shields.io/github/stars/InfiniTensor/transformer-rs?style=social\"/\u003e : 从 [YdrMaster/llama2.rs](https://github.com/YdrMaster/llama2.rs) 发展来的手写 transformer 模型项目。\r\n\r\n\r\n        - #### Zig Implementation\r\n\r\n            - [llama2.zig](https://github.com/cgbur/llama2.zig) \u003cimg src=\"https://img.shields.io/github/stars/cgbur/llama2.zig?style=social\"/\u003e : Inference Llama 2 in one file of pure Zig.\r\n\r\n            - [renerocksai/gpt4all.zig](https://github.com/renerocksai/gpt4all.zig) \u003cimg src=\"https://img.shields.io/github/stars/renerocksai/gpt4all.zig?style=social\"/\u003e : ZIG build for a terminal-based chat client for an assistant-style large language model with ~800k GPT-3.5-Turbo Generations based on LLaMa.\r\n\r\n            - [EugenHotaj/zig_inference](https://github.com/EugenHotaj/zig_inference) \u003cimg src=\"https://img.shields.io/github/stars/EugenHotaj/zig_inference?style=social\"/\u003e : Neural Network Inference Engine in Zig.\r\n\r\n\r\n        - ##### Go Implementation\r\n\r\n            - [Ollama](https://github.com/ollama/ollama/) \u003cimg src=\"https://img.shields.io/github/stars/ollama/ollama?style=social\"/\u003e : Get up and running with Llama 2, Mistral, Gemma, and other large language models. [ollama.com](https://ollama.com/)\r\n\r\n\r\n            - [go-skynet/LocalAI](https://github.com/go-skynet/LocalAI) \u003cimg src=\"https://img.shields.io/github/stars/go-skynet/LocalAI?style=social\"/\u003e : 🤖 Self-hosted, community-driven, local OpenAI-compatible API. Drop-in replacement for OpenAI running LLMs on consumer-grade hardware. Free Open Source OpenAI alternative. No GPU required. LocalAI is an API to run ggml compatible models: llama, gpt4all, rwkv, whisper, vicuna, koala, gpt4all-j, cerebras, falcon, dolly, starcoder, and many other. [localai.io](https://localai.io/)\r\n\r\n\r\n\r\n\r\n    - #### LLM Quantization Framework\r\n      ##### LLM量化框架\r\n\r\n        - [GPTQ](https://github.com/IST-DASLab/gptq) \u003cimg src=\"https://img.shields.io/github/stars/IST-DASLab/gptq?style=social\"/\u003e :  \"GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers\". (**[ICLR 2023](https://arxiv.org/abs/2210.17323)**).\r\n\r\n        - [SmoothQuant](https://github.com/mit-han-lab/smoothquant) \u003cimg src=\"https://img.shields.io/github/stars/mit-han-lab/smoothquant?style=social\"/\u003e :  \"SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models\". (**[ICML 2023](https://arxiv.org/abs/2211.10438)**).\r\n\r\n        - [AWQ](https://github.com/mit-han-lab/llm-awq) \u003cimg src=\"https://img.shields.io/github/stars/mit-han-lab/llm-awq?style=social\"/\u003e :  \"AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration\". (**[MLSys 2024](https://arxiv.org/abs/2306.00978)**).\r\n\r\n\r\n\r\n\r\n    - #### Application Development Platform\r\n      ##### 应用程序开发平台\r\n\r\n        - [LangChain](https://github.com/langchain-ai/langchain) \u003cimg src=\"https://img.shields.io/github/stars/hwchase17/langchain?style=social\"/\u003e :  🦜️🔗 LangChain. ⚡ Building applications with LLMs through composability ⚡ [python.langchain.com](https://python.langchain.com/docs/get_started/introduction.html)\r\n\r\n        - [Dify](https://github.com/langgenius/dify) \u003cimg src=\"https://img.shields.io/github/stars/langgenius/dify?style=social\"/\u003e : Dify is an open-source LLM app development platform. Dify's intuitive interface combines AI workflow, RAG pipeline, agent capabilities, model management, observability features and more, letting you quickly go from prototype to production. [dify.ai](https://dify.ai/)\r\n\r\n        - [Lobe Chat](https://github.com/lobehub/lobe-chat) \u003cimg src=\"https://img.shields.io/github/stars/lobehub/lobe-chat?style=social\"/\u003e : 🤯 Lobe Chat - an open-source, modern-design AI chat framework. Supports Multi AI Providers( OpenAI / Claude 3 / Gemini / Ollama / Qwen / DeepSeek), Knowledge Base (file upload / knowledge management / RAG ), Multi-Modals (Vision/TTS/Plugins/Artifacts). One-click FREE deployment of your private ChatGPT/ Claude application. [chat-preview.lobehub.com](https://chat-preview.lobehub.com/)\r\n\r\n        - [AutoChain](https://github.com/Forethought-Technologies/AutoChain) \u003cimg src=\"https://img.shields.io/github/stars/Forethought-Technologies/AutoChain?style=social\"/\u003e :  AutoChain: Build lightweight, extensible, and testable LLM Agents. [autochain.forethought.ai](https://autochain.forethought.ai/)\r\n\r\n        - [Auto-GPT](https://github.com/Significant-Gravitas/Auto-GPT) \u003cimg src=\"https://img.shields.io/github/stars/Significant-Gravitas/Auto-GPT?style=social\"/\u003e : Auto-GPT: An Autonomous GPT-4 Experiment. Auto-GPT is an experimental open-source application showcasing the capabilities of the GPT-4 language model. This program, driven by GPT-4, chains together LLM \"thoughts\", to autonomously achieve whatever goal you set. As one of the first examples of GPT-4 running fully autonomously, Auto-GPT pushes the boundaries of what is possible with AI. [agpt.co](https://news.agpt.co/)\r\n\r\n        - [LiteChain](https://github.com/rogeriochaves/litechain) \u003cimg src=\"https://img.shields.io/github/stars/rogeriochaves/litechain?style=social\"/\u003e : Build robust LLM applications with true composability 🔗. [rogeriochaves.github.io/litechain/](https://rogeriochaves.github.io/litechain/)\r\n\r\n        - [Open-Assistant](https://github.com/LAION-AI/Open-Assistant) \u003cimg src=\"https://img.shields.io/github/stars/LAION-AI/Open-Assistant?style=social\"/\u003e : OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so. [open-assistant.io](https://open-assistant.io/)\r\n\r\n        - [om-ai-lab/OmAgent](https://github.com/om-ai-lab/OmAgent) \u003cimg src=\"https://img.shields.io/github/stars/om-ai-lab/OmAgent?style=social\"/\u003e : Build multimodal language agents for fast prototype and production. [om-agent.com](https://om-agent.com/)\r\n\r\n\r\n\r\n\r\n    - #### RAG Framework\r\n      ##### 检索增强生成框架\r\n\r\n        - [LlamaIndex](https://github.com/run-llama/llama_index) \u003cimg src=\"https://img.shields.io/github/stars/run-llama/llama_index?style=social\"/\u003e : LlamaIndex is a data framework for your LLM applications. [docs.llamaindex.ai](https://docs.llamaindex.ai/)\r\n\r\n        - [Embedchain](https://github.com/embedchain/embedchain) \u003cimg src=\"https://img.shields.io/github/stars/embedchain/embedchain?style=social\"/\u003e : The Open Source RAG framework. [docs.embedchain.ai](https://docs.embedchain.ai/)\r\n\r\n        - [QAnything](https://github.com/netease-youdao/QAnything) \u003cimg src=\"https://img.shields.io/github/stars/netease-youdao/QAnything?style=social\"/\u003e : Question and Answer based on Anything. [qanything.ai](https://qanything.ai/)\r\n\r\n        - [R2R](https://github.com/SciPhi-AI/R2R) \u003cimg src=\"https://img.shields.io/github/stars/SciPhi-AI/R2R?style=social\"/\u003e : A framework for rapid development and deployment of production-ready RAG systems. [docs.sciphi.ai](https://docs.sciphi.ai/)\r\n\r\n        - [langchain-ai/rag-from-scratch](https://github.com/langchain-ai/rag-from-scratch) \u003cimg src=\"https://img.shields.io/github/stars/langchain-ai/rag-from-scratch?style=social\"/\u003e : Retrieval augmented generation (RAG) comes is a general methodology for connecting LLMs with external data sources. These notebooks accompany a video series will build up an understanding of RAG from scratch, starting with the basics of indexing, retrieval, and generation.\r\n\r\n    - #### Vector Database\r\n      ##### 向量数据库\r\n\r\n        - [Qdrant](https://github.com/milvus-io/milvus) \u003cimg src=\"https://img.shields.io/github/stars/milvus-io/milvus?style=social\"/\u003e : Milvus is an open-source vector database built to power embedding similarity search and AI applications. Milvus makes unstructured data search more accessible, and provides a consistent user experience regardless of the deployment environment. [milvus.io](https://milvus.io/)\r\n\r\n        - [Qdrant](https://github.com/qdrant/qdrant) \u003cimg src=\"https://img.shields.io/github/stars/qdrant/qdrant?style=social\"/\u003e : Qdrant - Vector Database for the next generation of AI applications. Also available in the cloud [https://cloud.qdrant.io/](https://cloud.qdrant.io/). [qdrant.tech](https://qdrant.tech/)\r\n\r\n\r\n\r\n\r\n    - #### Memory Management\r\n      ##### 内存管理\r\n\r\n        - [microsoft/vattention](https://github.com/microsoft/vattention) \u003cimg src=\"https://img.shields.io/github/stars/microsoft/vattention?style=social\"/\u003e : Dynamic Memory Management for Serving LLMs without PagedAttention.\r\n\r\n\r\n\r\n\r\n  - ### Awesome List\r\n\r\n    - [deepseek-ai/awesome-deepseek-integration](https://github.com/deepseek-ai/awesome-deepseek-integration) \u003cimg src=\"https://img.shields.io/github/stars/deepseek-ai/awesome-deepseek-integration?style=social\"/\u003e : Integrate the DeepSeek API into popular softwares. Access [DeepSeek Open Platform](https://platform.deepseek.com/) to get an API key.\r\n\r\n    - [Hannibal046/Awesome-LLM](https://github.com/Hannibal046/Awesome-LLM) \u003cimg src=\"https://img.shields.io/github/stars/Hannibal046/Awesome-LLM?style=social\"/\u003e :","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/coderonion%2Fawesome-llm-and-aigc/projects"}