An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with llm-training

A curated list of projects in awesome lists tagged with llm-training .

https://github.com/liguodongiot/llm-action

本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

llm llm-inference llm-serving llm-training llmops

Last synced: 15 May 2025

https://github.com/skypilot-org/skypilot

SkyPilot: Run AI and batch jobs on any infra (Kubernetes or 16+ clouds). Get unified execution, cost savings, and high GPU availability via a simple interface.

cloud-computing cloud-management cost-management cost-optimization data-science deep-learning distributed-training finops gpu hyperparameter-tuning job-queue job-scheduler llm-serving llm-training machine-learning ml-infrastructure ml-platform multicloud spot-instances tpu

Last synced: 02 Apr 2026

https://github.com/linkedin/liger-kernel

Efficient Triton Kernels for LLM Training

finetuning gemma2 llama llama3 llm-training llms mistral phi3 triton triton-kernels

Last synced: 13 May 2025

https://github.com/h2oai/h2o-llmstudio

H2O LLM Studio - a framework and no-code GUI for fine-tuning LLMs. Documentation: https://docs.h2o.ai/h2o-llmstudio/

ai chatbot chatgpt fedramp fine-tuning finetuning generative generative-ai gpt llama llama2 llm llm-training

Last synced: 07 Apr 2026

https://github.com/internlm/xtuner

An efficient, flexible and full-featured toolkit for fine-tuning LLM (InternLM2, Llama3, Phi3, Qwen, Mistral, ...)

agent baichuan chatbot chatglm2 chatglm3 conversational-ai internlm large-language-models llama2 llama3 llava llm llm-training mixtral msagent peft phi3 qwen supervised-finetuning

Last synced: 30 Oct 2025

https://github.com/InternLM/xtuner

An efficient, flexible and full-featured toolkit for fine-tuning LLM (InternLM2, Llama3, Phi3, Qwen, Mistral, ...)

agent baichuan chatbot chatglm2 chatglm3 conversational-ai internlm large-language-models llama2 llama3 llava llm llm-training mixtral msagent peft phi3 qwen supervised-finetuning

Last synced: 20 Mar 2025

https://github.com/linkedin/Liger-Kernel

Efficient Triton Kernels for LLM Training

finetuning gemma2 llama llama3 llm-training llms mistral phi3 triton triton-kernels

Last synced: 21 Aug 2025

https://github.com/databricks/dbrx

Code examples and resources for DBRX, a large language model developed by Databricks

databricks gen-ai generative-ai llm llm-inference llm-training mosaic-ai

Last synced: 25 Oct 2025

https://github.com/moonshotai/moba

MoBA: Mixture of Block Attention for Long-Context LLMs

flash-attention llm llm-serving llm-training moe pytorch transformer

Last synced: 14 May 2025

https://github.com/MoonshotAI/MoBA

MoBA: Mixture of Block Attention for Long-Context LLMs

flash-attention llm llm-serving llm-training moe pytorch transformer

Last synced: 31 Mar 2025

https://github.com/intelligent-machine-learning/dlrover

DLRover: An Automatic Distributed Deep Learning System

distributed-training hacktoberfest k8s llm-training

Last synced: 29 Dec 2025

https://github.com/openlake-project/openlake

High performance storage system for LLM Inference and GPU Training. Feed your GPUs at blazing fast speeds

blackwell gpt gpu high-performance llm llm-training model-serving rdma rust storage throughput

Last synced: 14 Jun 2026

https://github.com/tingaicompass/AI-Compass

“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。

agent ai llm llm-inference llm-training nlp rl rlhf

Last synced: 02 Sep 2026

https://github.com/arahim3/mlx-tune

Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, and Vision fine-tuning — natively on MLX. Unsloth-compatible API.

apple-silicon deep-learning huggingface large-language-models llm llm-finetuning llm-training local-llm lora machine-learning macos metal mlx mlx-framework on-device-ai peft transformers unsloth vision-language-model

Last synced: 01 Apr 2026

https://github.com/volcengine/vescale

A PyTorch Native LLM Training Framework

llm-training pytorch

Last synced: 03 Jul 2025

https://github.com/volcengine/veScale

A PyTorch Native LLM Training Framework

llm-training pytorch

Last synced: 30 Jul 2025

https://github.com/open-sciencelab/GraphGen

GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation

ai4science data-generation data-synthesis graphgen knowledge-graph llama-factory llm llm-training pretrain pretraining qa question-answering qwen sft sft-data xtuner

Last synced: 29 Nov 2025

https://github.com/feifeibear/long-context-attention

USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference

attention-is-all-you-need deepspeed-ulysses llm-inference llm-training pytorch ring-attention

Last synced: 14 May 2025

https://github.com/flagai-open/aquila2

The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.

llm llm-inference llm-training

Last synced: 15 May 2025

https://github.com/FlagAI-Open/Aquila2

The official repo of Aquila2 series proposed by BAAI, including pretrained & chat large language models.

llm llm-inference llm-training

Last synced: 08 Apr 2025

https://github.com/internlm/internevo

InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.

910b deepspeed-ulysses flash-attention gemma internlm internlm2 llama3 llava llm-framework llm-training multi-modal pipeline-parallelism pytorch ring-attention sequence-parallelism tensor-parallelism transformers-models zero3

Last synced: 07 Oct 2025

https://github.com/InternLM/InternEvo

InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.

910b deepspeed-ulysses flash-attention gemma internlm internlm2 llama3 llava llm-framework llm-training multi-modal pipeline-parallelism pytorch ring-attention sequence-parallelism tensor-parallelism transformers-models zero3

Last synced: 27 Mar 2025

https://github.com/armbues/SiLLM

SiLLM simplifies the process of training and running Large Language Models (LLMs) on Apple Silicon by leveraging the MLX framework.

apple-silicon dpo large-language-models llm llm-inference llm-training lora mlx

Last synced: 18 Jul 2025

https://github.com/shivendrra/smalllanguagemodel

a LLM cookbook, for building your own from scratch, all the way from gathering data to training a model

bert-model decoder-model gpt llm-cookbook llm-training llms machine-learning neural-networks transformer

Last synced: 12 Apr 2025

https://github.com/shivendrra/SmallLanguageModel

a LLM cookbook, for building your own from scratch, all the way from gathering data to training a model

bert-model decoder-model gpt llm-cookbook llm-training llms machine-learning neural-networks transformer

Last synced: 14 Mar 2025

https://github.com/itachi-uchiha581/auto-data

Auto Data is a library designed for quick and effortless creation of datasets tailored for fine-tuning Large Language Models (LLMs).

ai data finetuning-large-language-models finetuning-llms generative-ai llm llm-training python python3

Last synced: 20 Sep 2025

https://github.com/hkust-nlp/dart-math

[NeurIPS'24] Official code for *🎯DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving*

deep-learning llm llm-evaluation llm-inference llm-training mathematics nlp

Last synced: 10 Jul 2025

https://github.com/simplifine-llm/Simplifine

🚀 Easy, open-source LLM finetuning with one-line commands, seamless cloud integration, and popular optimization frameworks. ✨

ai cloud fine-tuning fine-tuning-llm finetuning-llms gpt instruction-tuning large-language-models llama llama3 llm llm-training lora mistral moe open-source peft phi qwen

Last synced: 29 Jul 2025

https://github.com/dsdanielpark/open-llm-datasets

Repository for organizing datasets and papers used in Open LLM.

datasets large-language-models llm llm-datasets llm-training natural-language-processing

Last synced: 05 Mar 2026

https://github.com/Simplifine-gamedev/Simplifine

🚀 Easy, open-source LLM finetuning with one-line commands, seamless cloud integration, and popular optimization frameworks. ✨

ai cloud fine-tuning fine-tuning-llm finetuning-llms gpt instruction-tuning large-language-models llama llama3 llm llm-training lora mistral moe open-source peft phi qwen

Last synced: 31 Oct 2025

https://github.com/tatevkaren/babygpt-build_gpt_from_scratch

BabyGPT: Build Your Own GPT Large Language Model from Scratch Pre-Training Generative Transformer Models: Building GPT from Scratch with a Step-by-Step Guide to Generative AI in PyTorch and Python

attention-is-all-you-need dropout-layers gpt gpt-3 large-language-models layer-normalization llm-training llms multi-head-self-attention neural-networks python pytorch residual-connections transformers

Last synced: 10 Apr 2025

https://github.com/amazon-science/Cyber-Zero

Cyber-Zero: Training Cybersecurity Agents Without Runtime

agent ctf cybersecurity large-language-models llm llm-training offensive-security

Last synced: 22 Jun 2026

https://github.com/microsoft/llf-bench

A benchmark for evaluating learning agents based on just language feedback

large-language-models llm llm-training llms machine-learning natural-language-processing reinforcement-learning

Last synced: 30 Oct 2025

https://github.com/lum1104/mer-factory

🚀 Pre-process, annotate, evaluate, and train your Affect Computing (e.g., Multimodal Emotion Recognition, Sentiment Analysis) datasets ALL within MER-Factory! (LangGraph Based Agent Workflow)

affective-computing collaborate llm-training multimodal-emotion-recognition sentiment-analysis

Last synced: 12 Mar 2026

https://github.com/smerkyg/gptcore

Fast modular code to create and train cutting edge LLMs

gpt llm-training llms machine-learning

Last synced: 12 May 2025

https://github.com/erogol/blagpt

Experimental playground for benchmarking language model (LM) architectures, layers, and tricks on smaller datasets. Designed for flexible experimentation and exploration.

attention-mechanisms deep-learning diffusion-llm dllm gpt gpt-2 hymba large-language-models llm llm-training machine-learning position-embedding pytorch transformers

Last synced: 07 Jul 2025

https://github.com/lenguajenatural-ai/autotransformers

A Python package for automatically training and comparing language models.

language-model llm-training llms nlp

Last synced: 14 Jan 2026

https://github.com/microsoft/LLF-Bench

A benchmark for evaluating learning agents based on just language feedback

large-language-models llm llm-training llms machine-learning natural-language-processing reinforcement-learning

Last synced: 18 Apr 2025

https://github.com/microsoft/mathoctopus

This repository contains resources for accessing the official benchmarks, codes, and checkpoints of the paper: "[**Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations**]".

llm-training

Last synced: 20 Oct 2025

https://github.com/microsoft/MathOctopus

This repository contains resources for accessing the official benchmarks, codes, and checkpoints of the paper: "[**Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations**]".

llm-training

Last synced: 16 Oct 2025

https://github.com/google/litmus

Litmus is a comprehensive LLM testing and evaluation tool designed for GenAI Application Development. It provides a robust platform with a user-friendly UI for streamlining the process of building and assessing the performance of your LLM-powered applications.

api apitesting cicd devops llm llm-evaluation llm-security llm-training llmops testing testing-tools

Last synced: 14 Jan 2026

https://github.com/amazon-science/llm-code-preference

Training and Benchmarking LLMs for Code Preference.

code-generation llm-evaluation llm-training llms-benchmarking

Last synced: 07 Oct 2025

https://github.com/wassemgtk/llm.scala

Extensible implementation of a Language Model (LLM) training framework in Scala.

gpt llm llm-training transformer transformers-library

Last synced: 19 Apr 2025

https://github.com/mofheka/llama-megatron

A LLaMA1/LLaMA12 Megatron implement.

llama llama2 llm llm-training megatron megatron-lm pytorch

Last synced: 16 Jun 2025

https://github.com/arcee-ai/arcee-python

The Arcee client for executing domain-adpated language model routines https://pypi.org/project/arcee-py/

ai llm llm-inference llm-training llmops

Last synced: 17 Mar 2026

https://github.com/ai4sd/number-token-loss

PyPI package for number token loss.

language-models llm llm-training reasoning

Last synced: 08 Apr 2026

https://github.com/armbues/SiLLM-examples

Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon

apple-silicon dpo large-language-models llm llm-inference llm-training lora mlx

Last synced: 10 Apr 2025

https://github.com/ubos-tech/ai-chatbot-starter-kit

AI Chatbot Starter Kit: An open-source, extensible framework for rapidly developing custom AI chatbots with integrations for popular data sources, messaging platforms, LLM models, and CRM systems. Ideal for developers looking for a minimal boilerplate solution.

chatgpt chatgpt-api chatgpt-bot facebook-bot instagram-chatbot llm llm-agent llm-training low-code lowcode-editor node-red openai openai-api pinecone support-bot telegram-chat-bot ubos-tech whatsapp-bot

Last synced: 15 May 2025

https://github.com/aklinker1/vitepress-knowledge

Free, self-hosted LLM chatbot trained on your VitePress website.

ai llm-training vitepress vitepress-plugin

Last synced: 30 Dec 2025

https://github.com/levitation-opensource/manipulative-expression-recognition

MER is a software that identifies and highlights manipulative communication in text from human conversations and AI-generated responses. MER benchmarks language models for manipulative expressions, fostering development of transparency and safety in AI. It also supports manipulation victims by detecting manipulative patterns in human communication.

benchmarking conversation-analysis conversation-analytics expression-recognition fraud-detection fraud-prevention human-computer-interaction human-robot-interaction llm llm-security llm-test llm-training manipulation misinformation prompt-engineering prompt-injection psychometrics sentiment-analysis sentiment-classification transparency

Last synced: 28 Jan 2026

https://github.com/moinulmoin/free-llmstxt-generator

converts webpage content into Markdown format, optimized for LLM training and context

crawling llm-context llm-training llmstxt markdown

Last synced: 23 Mar 2025

https://github.com/epistates/pmetal

Powdered Metal — High-performance LLM fine-tuning framework for Apple Silicon

ai llm-inference llm-training metal mlx

Last synced: 08 May 2026

https://github.com/kvignesh1420/cot-icl-lab

[ACL 2025] CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations

chain-of-thought graphs in-context-learning llm-inference llm-training transformers

Last synced: 01 Jul 2025

https://github.com/rs-py/howtofinetunellama3.1

Quick tutorial showing how to fine-tune Llama3.1 with nothing but free tools and text data. All code included in ipynb. For a step by step walkthrough take a look at the tutorial below on medium.

fine-tuning finetuning huggingface llama3 llm llm-training

Last synced: 24 Apr 2025

https://github.com/angrysky56/llada_gui_new

GUI for LLaDA Diffusion LLM with Quantization for low end GPU and CPU options. Now with Prototype Training and Vector DB.

gui inference llada llm-training

Last synced: 24 Apr 2026

https://github.com/endevsols/long-trainer

Introducing LongTrainer, a sophisticated extension of the LangChain framework designed specifically for managing multiple bots and providing isolated, context-aware chat sessions. Ideal for developers and businesses looking to integrate complex conversational AI into their systems, LongTrainer simplifies the deployment and customization of LLMs.

gpt langchain langchain-python llm-training longtrainer openai rag

Last synced: 26 Feb 2026

https://github.com/mewmix/gh_llm_loader

clone GitHub repositories and prepare their data for ingestion for LLMs.

context data data-structures github llm llm-training python

Last synced: 19 Sep 2025

https://github.com/airmomo/tpo-llm-webui

TPO 是一个优化 LLM 输出文本的框架,通过迭代反馈和优化提示的方式来“微调模型”,而非直接调整模型的参数,使模型在推理过程中与人类偏好对齐以生成更好的结果。本项目提供了一个友好的 WebUI 来加载模型,实时优化基础模型并展示最佳结果。

llm llm-training tpo web-ui

Last synced: 24 Dec 2025

https://github.com/altunenes/rustysozluk

Efficiently fetch and perform sentiment analysis (Turkish Only) on eksisozluk.com entries using Rust

duyguanalizi eksi-sozluk eksisozluk llm-datasets llm-training reqwest rust rust-lang rust-scraping scraper sentiment-analysis turkish webscraping

Last synced: 26 Sep 2025

https://github.com/prismadic/tractor-beam

high-efficiency text & file scraper with smart tracking, client/server networking for building language model datasets fast

botnet cluster data file-downloader llm llm-finetuning llm-training mass-downloader scraping

Last synced: 07 May 2025

https://github.com/puneetkakkar/bitnet-1.58b

Bitnet 1.58b: This project implements the innovative 1-bit LLM architecture described in recent whitepapers, focusing on efficient training, inference, and open-source collaboration.

1-bit-quantization deep-learning large-language-models llm-training llms machine-learning nlp pytorch research-and-development

Last synced: 01 Sep 2025

https://github.com/stoyan-stoyanov/transformers-calculator

Transformer Calculator - Estimate training time for transformer models.

llm llm-training transformers

Last synced: 18 Jun 2025

https://github.com/holasoymalva/llm-glossary

Glossary of key concepts in Large Language Models (LLMs), Artificial Intelligence (AI), Natural Language Processing (NLP), and Machine Learning (ML). Clear definitions and resources for developers, researchers, and AI enthusiasts exploring generative language technologies.

ai artificial-intelligence artificial-intelligence-algorithms artificial-intelligence-framework artificial-intelligence-models generative-ai generative-ai-model glossary hacktoberfest large-language-models llm llm-agents llm-evaluation llm-framework llm-inference llm-tools llm-training llms machine-learning-models

Last synced: 24 Jun 2026

https://github.com/partiql/partiql-beamline

Beamline is a tool for fast data generation for your AI/LLM/ML model training, simulation, and testing use-cases. It generates reproducible pseudo-random data using a stochastic approach and probability distributions, meaning you can create realistic datasets that follow specific mathematical patterns.

fuzz-testing llm-training ml-training model-training-and-tuning query-generator simulation stochastic-processes testing

Last synced: 14 Aug 2026

https://github.com/Rs-py/HowToFineTuneLlama3.1

Quick tutorial showing how to fine-tune Llama3.1 with nothing but free tools and text data. All code included in ipynb. For a step by step walkthrough take a look at the tutorial below on medium.

fine-tuning finetuning huggingface llama3 llm llm-training

Last synced: 11 Sep 2025

https://github.com/maris205/llama-gene

A General-purpose Gene Task Large Language Model Based on Instruction Fine-tuning

genetic-algorithm llama llm-training

Last synced: 10 Apr 2025

https://github.com/firojalam/llamalens

This repository contains the resources, code, and documentation for LlamaLens, a specialized multilingual large language model (LLM) designed to analyze news and social media content effectively. LlamaLens supports multiple languages, including Arabic, English, and Hindi, and is tailored for diverse tasks such as sentiment analysis, misinformation.

arabic downstream-tasks emotion-detection english hindi llm llm-inference llm-training newsmedia sentiment-classification social-media

Last synced: 04 Aug 2025

https://github.com/ernie-research/ma-rlhf

[ICLR'25] MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

llm-training ma-rlhf ppo rlhf

Last synced: 19 Sep 2025

https://github.com/auxten/llm-cluster-viz

A 3D visualizaion of LLM training and inferencing on modern cluster

gb200 gpu inference llm llm-training nvidia nvl72 react three-js visualization webgl

Last synced: 13 Aug 2026

https://github.com/thetwopct/folder2txt

Convert local folder contents into a single text file with ease - perfect for analysis, documentation, or AI/LLM training purposes.

ai-tools ai-training cli llm-training text

Last synced: 29 May 2026

https://github.com/es7/introduction-to-llms

In this repository I have explained the application of Large Language Models (LLMs). Starting from how to use LLMs in our own application till how to build a LLM.

computer-vision deep-learning huggingface llm llm-framework llm-inference llm-training machine-learning natural-language-processing prompt-engineering prompt-learning

Last synced: 18 Jul 2026