Projects in Awesome Lists tagged with trl
A curated list of projects in awesome lists tagged with trl .
https://github.com/argilla-io/notus
Notus is a collection of fine-tuned LLMs using SFT, DPO, SFT+DPO, and/or any other RLHF techniques, while always keeping a data-first approach
alignment-handbook dpo fine-tuning lm-alignment preference-data trl zephyr
Last synced: 17 Jun 2025
https://github.com/akshint0407/nano-r1
This project demonstrates the process of fine-tuning the Qwen2.5-3B-Instruct model using GRPO (Generalized Reward Policy Optimization) on the GSM8K dataset.
adapters grpo huggingface python qwen2-5 safetensors text-generation-inference transformer trl unsloth
Last synced: 28 Apr 2026
https://github.com/yancotta/post_training_llms
Different post-training techniques for LLMs, including: SFT, DPO and Online RL
alignment dpo fine-tuning huggingface huggingface-transformers llm pytorch reinforcement-learning sft trl
Last synced: 05 Oct 2025
https://github.com/rasyosef/phi-2-sft-and-dpo
Notebooks to create an instruction following version of Microsoft's Phi 2 LLM with Supervised Fine Tuning and Direct Preference Optimization (DPO)
direct-preference-optimization huggingface llm pytorch supervised-finetuning transformers trl
Last synced: 14 May 2026
https://github.com/mikesterner87/nano-r1
This project demonstrates the process of fine-tuning the Qwen2.5-3B-Instruct model using GRPO (Generalized Reward Policy Optimization) on the GSM8K dataset.
adapters build grpo huggingface nanopi nanopi-r1 nanopi-r1s openwrt python safetensors text-generation-inference transformer trl unsloth
Last synced: 03 Sep 2025
https://github.com/rasyosef/phi-1_5-instruct
Notebooks to create an instruction following version of Microsoft's Phi 1.5 LLM with Supervised Fine Tuning and Direct Preference Optimization (DPO)
direct-preference-optimization llm pytorch supervised-finetuning transformers trl
Last synced: 05 Oct 2025
https://github.com/longern/pd-trainer
Fine-tune your LLM using minimal data and computing power with the Prompt Distillation trainer.
fine-tuning knowledge-distillation llm prompt-distillation trl
Last synced: 08 Sep 2025
https://github.com/jorahn/jorahn
Profile for Jonathan Rahn — AI Lab Lead at Drees & Sommer. Chess-reasoning LMs (policy + world model), RL with verifiable rewards.
chess-ai cot policy rlvr self-play transformers trl world-models
Last synced: 19 Sep 2025
https://github.com/shekswess/tiny-reasoning-language-model
Code repository dedicated to experimenting and research with tiny reasoning language model
llm post-training reasoning research sft slm transformers trl
Last synced: 11 Oct 2025
https://github.com/sivasaiyadav8143/llm-finetuning-playbook
A hands-on playbook for fine-tuning LLMs (TinyLlama, Gemma-2B) with LoRA/QLoRA, SFT, and DPO on domain-specific datasets (pharma, legal). Demonstrates end-to-end pipelines for classification, reasoning, and preference alignment—all on free Colab GPUs. Perfect for learning parameter-efficient fine-tuning and showcasing applied LLM skills.
dpo fine-tuning huggingface llm lora nlp peft qlora tinyllama transformers trl unsloth
Last synced: 06 Jul 2026