An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with trl

A curated list of projects in awesome lists tagged with trl .

https://github.com/argilla-io/notus

Notus is a collection of fine-tuned LLMs using SFT, DPO, SFT+DPO, and/or any other RLHF techniques, while always keeping a data-first approach

alignment-handbook dpo fine-tuning lm-alignment preference-data trl zephyr

Last synced: 17 Jun 2025

https://github.com/ssbuild/llm_rlhf

realize the reinforcement learning training for gpt2 llama bloom and so on llm model

llm llm-rlhf lora reward rlhf trl trlx

Last synced: 24 Apr 2025

https://github.com/akshint0407/nano-r1

This project demonstrates the process of fine-tuning the Qwen2.5-3B-Instruct model using GRPO (Generalized Reward Policy Optimization) on the GSM8K dataset.

adapters grpo huggingface python qwen2-5 safetensors text-generation-inference transformer trl unsloth

Last synced: 28 Apr 2026

https://github.com/yancotta/post_training_llms

Different post-training techniques for LLMs, including: SFT, DPO and Online RL

alignment dpo fine-tuning huggingface huggingface-transformers llm pytorch reinforcement-learning sft trl

Last synced: 05 Oct 2025

https://github.com/rasyosef/phi-2-sft-and-dpo

Notebooks to create an instruction following version of Microsoft's Phi 2 LLM with Supervised Fine Tuning and Direct Preference Optimization (DPO)

direct-preference-optimization huggingface llm pytorch supervised-finetuning transformers trl

Last synced: 14 May 2026

https://github.com/mikesterner87/nano-r1

This project demonstrates the process of fine-tuning the Qwen2.5-3B-Instruct model using GRPO (Generalized Reward Policy Optimization) on the GSM8K dataset.

adapters build grpo huggingface nanopi nanopi-r1 nanopi-r1s openwrt python safetensors text-generation-inference transformer trl unsloth

Last synced: 03 Sep 2025

https://github.com/rasyosef/phi-1_5-instruct

Notebooks to create an instruction following version of Microsoft's Phi 1.5 LLM with Supervised Fine Tuning and Direct Preference Optimization (DPO)

direct-preference-optimization llm pytorch supervised-finetuning transformers trl

Last synced: 05 Oct 2025

https://github.com/longern/pd-trainer

Fine-tune your LLM using minimal data and computing power with the Prompt Distillation trainer.

fine-tuning knowledge-distillation llm prompt-distillation trl

Last synced: 08 Sep 2025

https://github.com/jorahn/jorahn

Profile for Jonathan Rahn — AI Lab Lead at Drees & Sommer. Chess-reasoning LMs (policy + world model), RL with verifiable rewards.

chess-ai cot policy rlvr self-play transformers trl world-models

Last synced: 19 Sep 2025

https://github.com/sofiakhutsieva/llm_experiments

Эксперименты с LLM (инференс, rag, дообучение)

langchain llamacpp llm mistral peft rag trl

Last synced: 21 Jan 2026

https://github.com/shekswess/tiny-reasoning-language-model

Code repository dedicated to experimenting and research with tiny reasoning language model

llm post-training reasoning research sft slm transformers trl

Last synced: 11 Oct 2025

https://github.com/sivasaiyadav8143/llm-finetuning-playbook

A hands-on playbook for fine-tuning LLMs (TinyLlama, Gemma-2B) with LoRA/QLoRA, SFT, and DPO on domain-specific datasets (pharma, legal). Demonstrates end-to-end pipelines for classification, reasoning, and preference alignment—all on free Colab GPUs. Perfect for learning parameter-efficient fine-tuning and showcasing applied LLM skills.

dpo fine-tuning huggingface llm lora nlp peft qlora tinyllama transformers trl unsloth

Last synced: 06 Jul 2026