Projects in Awesome Lists tagged with rlhf
A curated list of projects in awesome lists tagged with rlhf .
https://github.com/hiyouga/LLaMAFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
agent ai deepseek fine-tuning gemma gpt instruction-tuning large-language-models llama llama3 llm lora moe nlp peft qlora quantization qwen rlhf transformers
Last synced: 18 Aug 2026
https://github.com/hiyouga/llama-factory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
agent ai chatglm fine-tuning gpt instruction-tuning language-model large-language-models llama llama3 llm lora mistral moe peft qlora quantization qwen rlhf transformers
Last synced: 02 Jan 2026
https://github.com/hiyouga/LLaMA-Factory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
agent ai chatglm fine-tuning gpt instruction-tuning language-model large-language-models llama llama3 llm lora mistral moe peft qlora quantization qwen rlhf transformers
Last synced: 14 Mar 2025
https://github.com/laion-ai/open-assistant
OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
ai assistant chatgpt discord-bot language-model machine-learning nextjs python rlhf
Last synced: 14 May 2025
https://github.com/LAION-AI/Open-Assistant
OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
ai assistant chatgpt discord-bot language-model machine-learning nextjs python rlhf
Last synced: 15 Mar 2025
https://github.com/rucaibox/llmsurvey
The official GitHub page for the survey paper "A Survey of Large Language Models".
chain-of-thought chatgpt in-context-learning instruction-tuning large-language-models llm llms natural-language-processing pre-trained-language-models pre-training rlhf
Last synced: 11 May 2025
https://github.com/RUCAIBox/LLMSurvey
The official GitHub page for the survey paper "A Survey of Large Language Models".
chain-of-thought chatgpt in-context-learning instruction-tuning large-language-models llm llms natural-language-processing pre-trained-language-models pre-training rlhf
Last synced: 14 Mar 2025
https://github.com/ymcui/Chinese-LLaMA-Alpaca-2
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
64k alpaca alpaca-2 alpaca2 flash-attention large-language-models llama llama-2 llama2 llm nlp rlhf yarn
Last synced: 24 Mar 2025
https://github.com/ymcui/chinese-llama-alpaca-2
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
64k alpaca alpaca-2 alpaca2 flash-attention large-language-models llama llama-2 llama2 llm nlp rlhf yarn
Last synced: 14 May 2025
https://github.com/internlm/internlm
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
chatbot chinese fine-tuning-llm flash-attention gpt large-language-model llm long-context pretrained-models rlhf
Last synced: 14 May 2025
https://github.com/InternLM/InternLM
Official release of InternLM2 7B and 20B base and chat models. 200K context support
chatbot chinese fine-tuning-llm flash-attention gpt large-language-model llm long-context pretrained-models rlhf
Last synced: 16 Mar 2025
https://github.com/huggingface/alignment-handbook
Robust recipes to align language models with human and AI preferences
Last synced: 12 May 2025
https://github.com/Kiln-AI/kiln
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
ai chain-of-thought collaboration dataset-generation evals evaluation evaluation-framework fine-tuning machine-learning macos mcp ml ollama openai prompt prompt-engineering python rlhf synthetic-data windows
Last synced: 16 Aug 2026
https://github.com/Gen-Verse/OpenClaw-RL
OpenClaw-RL: Train any agent simply by talking
async coding grpo gui-application memory-systems on-policy-distillation open-claw openclaw-skills rlhf sglang skill-learning slime tinker
Last synced: 01 May 2026
https://github.com/argilla-io/argilla
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
active-learning ai annotation-tool developer-tools gpt-4 human-in-the-loop langchain llm machine-learning mlops natural-language-processing nlp rlhf text-annotation text-labeling weak-supervision weakly-supervised-learning
Last synced: 13 May 2025
https://github.com/hiyouga/chatglm-efficient-tuning
Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调
alpaca chatglm chatglm2 chatgpt fine-tuning huggingface language-model lora peft pytorch qlora rlhf transformers
Last synced: 29 Sep 2025
https://github.com/hiyouga/ChatGLM-Efficient-Tuning
Fine-tuning ChatGLM-6B with PEFT | 基于 PEFT 的高效 ChatGLM 微调
alpaca chatglm chatglm2 chatgpt fine-tuning huggingface language-model lora peft pytorch qlora rlhf transformers
Last synced: 29 Mar 2025
https://github.com/pku-alignment/align-anything
Align Anything: Training All-modality Model with Feedback
chameleon dpo large-language-models multimodal rlhf vision-language-model
Last synced: 14 May 2025
https://github.com/kiln-ai/kiln
The easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets.
ai chain-of-thought collaboration dataset-generation evals evaluation fine-tuning machine-learning macos ml ollama openai prompt prompt-engineering python rlhf synthetic-data windows
Last synced: 23 Apr 2025
https://github.com/transformerlab/transformerlab-app
Open Source Application for Advanced LLM Engineering: interact, train, fine-tune, and evaluate large language models on your own computer.
electron llama llms lora mlx rlhf transformers
Last synced: 12 Feb 2026
https://github.com/docta-ai/docta
A Doctor for your data
data data-centric-ai data-centric-machine-learning data-curation data-diagnosis language-model rlhf
Last synced: 13 May 2025
https://github.com/PKU-Alignment/align-anything
Align Anything: Training All-modality Model with Feedback
chameleon dpo large-language-models multimodal rlhf vision-language-model
Last synced: 01 Apr 2025
https://github.com/Docta-ai/docta
A Doctor for your data
data data-centric-ai data-centric-machine-learning data-curation data-diagnosis language-model rlhf
Last synced: 26 Mar 2025
https://github.com/argilla-io/distilabel
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
ai huggingface llms openai python rlaif rlhf synthetic-data synthetic-dataset-generation
Last synced: 11 Apr 2025
https://github.com/tatsu-lab/alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
deep-learning evaluation foundation-models instruction-following large-language-models leaderboard nlp rlhf
Last synced: 13 May 2025
https://tatsu-lab.github.io/alpaca_eval/
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
deep-learning evaluation foundation-models instruction-following large-language-models leaderboard nlp rlhf
Last synced: 23 Mar 2025
https://github.com/thudm/webglm
WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)
Last synced: 14 May 2025
https://github.com/THUDM/WebGLM
WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)
Last synced: 24 Mar 2025
https://github.com/natolambert/rlhf-book
Textbook on reinforcement learning from human feedback
Last synced: 08 Feb 2026
https://github.com/pku-alignment/safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
ai-safety alpaca beaver datasets deepspeed gpt large-language-models llama llm llms reinforcement-learning reinforcement-learning-from-human-feedback rlhf safe-reinforcement-learning safe-reinforcement-learning-from-human-feedback safe-rlhf safety transformer transformers vicuna
Last synced: 16 May 2025
https://github.com/PKU-Alignment/safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
ai-safety alpaca beaver datasets deepspeed gpt large-language-models llama llm llms reinforcement-learning reinforcement-learning-from-human-feedback rlhf safe-reinforcement-learning safe-reinforcement-learning-from-human-feedback safe-rlhf safety transformer transformers vicuna
Last synced: 09 May 2025
https://github.com/zai-org/ImageReward
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
diffusion-models generative-model human-preferences rlhf
Last synced: 29 Dec 2025
https://github.com/THUDM/ImageReward
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
diffusion-models generative-model human-preferences rlhf
Last synced: 28 Mar 2025
https://github.com/openlmlab/moss-rlhf
Secrets of RLHF in Large Language Models Part I: PPO
Last synced: 16 May 2025
https://github.com/OpenLMLab/MOSS-RLHF
Secrets of RLHF in Large Language Models Part I: PPO
Last synced: 29 Mar 2025
https://openlmlab.github.io/MOSS-RLHF/
Secrets of RLHF in Large Language Models Part I: PPO
Last synced: 25 Mar 2025
https://github.com/RLHFlow/RLHF-Reward-Modeling
Recipes to train reward model for RLHF.
Last synced: 07 May 2025
https://github.com/tingaicompass/AI-Compass
“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。
agent ai llm llm-inference llm-training nlp rl rlhf
Last synced: 02 Sep 2026
https://github.com/alibaba/ROLL
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
Last synced: 15 Jun 2025
https://github.com/Kiln-AI/Kiln
The easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets.
ai chain-of-thought collaboration dataset-generation fine-tuning machine-learning macos ml ollama openai prompt prompt-engineering python rlhf synthetic-data windows
Last synced: 06 Oct 2025
https://github.com/xtreme1-io/xtreme1
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
3d-annotation annotation annotation-tool computer-vision image-annotation image-classification image-labelling-tool labeling-tool multimodal point-cloud rlhf
Last synced: 20 Mar 2025
https://github.com/GaryYufei/AlignLLMHumanSurvey
Aligning Large Language Models with Human: A Survey
awesome chatgpt chinese-llama gpt-4 large-language-models llama llama2 llms rlhf supervised-finetuning survey
Last synced: 11 May 2025
https://github.com/garyyufei/alignllmhumansurvey
Aligning Large Language Models with Human: A Survey
awesome chatgpt chinese-llama gpt-4 large-language-models llama llama2 llms rlhf supervised-finetuning survey
Last synced: 01 Oct 2025
https://github.com/verl-project/verl-omni
Multimodal RL training framework for diffusion & omni models
diffusion-models flow-matching grpo multimodal qwen reinforcement-learning rlhf vllm
Last synced: 02 Aug 2026
https://github.com/allenai/reward-bench
RewardBench: the first evaluation tool for reward models.
Last synced: 11 Sep 2025
https://github.com/jerry1993-tech/Cornucopia-LLaMA-Fin-Chinese
聚宝盆(Cornucopia): 中文金融系列开源可商用大模型,并提供一套高效轻量化的垂直领域LLM训练框架(Pretraining、SFT、RLHF、Quantize等)
chinese finance large-language-models llama nlp qa rlhf sft text-generation transformers
Last synced: 01 Apr 2025
https://github.com/voidful/textrl
Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
chatgpt controlled-nlg gpt-2 gpt-3 language-model nlg nlp pytorch reinforcement-learning rlhf
Last synced: 16 May 2025
https://github.com/voidful/TextRL
Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
chatgpt controlled-nlg gpt-2 gpt-3 language-model nlg nlp pytorch reinforcement-learning rlhf
Last synced: 29 Mar 2025
https://github.com/RLHFlow/Online-RLHF
A recipe for online RLHF and online iterative DPO.
Last synced: 13 Jul 2026
https://github.com/mindspore-courses/step_into_llm
MindSpore online courses: Step into LLM
bert chatglm chatglm2 chatgpt codegeex gpt gpt2 instruction-tuning large-language-models llama llama2 llm mindspore moe natural-language-processing nlp parallel-computing peft prompt-tuning rlhf
Last synced: 15 May 2025
https://github.com/cambioml/pykoi-rlhf-finetuned-transformers
pykoi: Active learning in one unified interface
ai chatbot feedback language-model llm machine-learning rlhf
Last synced: 11 Sep 2025
https://github.com/CambioML/pykoi-rlhf-finetuned-transformers
pykoi: Active learning in one unified interface
ai chatbot feedback language-model llm machine-learning rlhf
Last synced: 04 Apr 2025
https://github.com/rlhflow/online-rlhf
A recipe for online RLHF and online iterative DPO.
Last synced: 08 Apr 2025
https://github.com/sail-sg/oat
🌾 OAT: A research-friendly framework for LLM online alignment, including preference learning, reinforcement learning, etc.
alignment distributed-rl distributed-training dpo dueling-bandits grpo llm llm-aligment llm-exploration online-alignment online-rl ppo r1-zero reasoning rlhf thompson-sampling
Last synced: 08 May 2025
https://github.com/WangRongsheng/MedQA-ChatGLM
🛰️ 基于真实医疗对话数据在ChatGLM上进行LoRA、P-Tuning V2、Freeze、RLHF等微调,我们的眼光不止于医疗问答
chatglm-6b chatgpt dataset fine-tuning freeze huggingface large-language-models llms lora medical rlhf transformer
Last synced: 20 Apr 2025
https://github.com/goekdeniz-guelmez/mlx-lm-lora
Train Large Language Models on MLX.
apple deep-learning dpo fine finetuning-llms ml rlhf supervised-machine-learning training
Last synced: 09 Mar 2026
https://github.com/haoliuhl/chain-of-hindsight
Simple next-token-prediction for RLHF
large-language-models learning-from-human-feedback rlhf
Last synced: 03 Apr 2025
https://github.com/mihirp1998/VADER
Video Diffusion Alignment via Reward Gradients. We improve a variety of video diffusion models such as VideoCrafter, OpenSora, ModelScope and StableVideoDiffusion by finetuning them using various reward models such as HPS, PickScore, VideoMAE, VJEPA, YOLO, Aesthetics etc.
alignment diffusion reinforcement-learning reinforcement-learning-human-feedback rl rlhf vader video-diffusion video-diffusion-alignment
Last synced: 28 Mar 2025
https://github.com/jackaduma/vicuna-lora-rlhf-pytorch
A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Vicuna architecture. Basically ChatGPT but with Vicuna
chatgpt finetune gpt llama llm lora peft ppo pytorch reward-models rlhf vicuna vicuna-7b
Last synced: 13 Apr 2025
https://github.com/jianzhnie/open-r1
The open source implementation of DeepSeek-R1. 开源复现 DeepSeek-R1
deepseek-r1 deepseek-v3 grpo llm rlhf
Last synced: 04 Apr 2025
https://github.com/thudm/visionreward
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
Last synced: 21 Jun 2025
https://github.com/harderthenharder/rlloggingboard
A visuailzation tool to make deep understaning and easier debugging for RLHF training.
Last synced: 07 May 2025
https://github.com/tomekkorbak/pretraining-with-human-feedback
Code accompanying the paper Pretraining Language Models with Human Preferences
ai-alignment ai-safety decision-transformers gpt language-models pretraining reinforcement-learning rlhf
Last synced: 07 May 2025
https://github.com/jianzhnie/Open-R1
The open source implementation of ChatGPT, Alpaca, Vicuna and RLHF Pipeline. 从0开始实现一个ChatGPT.
chatgpt gpt llama llm lora peft ppo rlhf stanford-alpaca
Last synced: 05 Oct 2025
https://github.com/pku-alignment/aligner
[NeurIPS 2024 Oral] Aligner: Efficient Alignment by Learning to Correct
aisafety aligner alignment interpretability llm mecinterp rlhf weak-to-strong
Last synced: 07 May 2025
https://github.com/xrsrke/instructgoose
Implementation of Reinforcement Learning from Human Feedback (RLHF)
chatgpt human-feedback instructgpt reinforcement-learning rlhf
Last synced: 09 Apr 2025
https://github.com/xrsrke/instructGOOSE
Implementation of Reinforcement Learning from Human Feedback (RLHF)
chatgpt human-feedback instructgpt reinforcement-learning rlhf
Last synced: 29 Mar 2025
https://github.com/rlhflow/rlhf-reward-modeling
A recipe to train reward models for RLHF.
Last synced: 21 Aug 2025
https://github.com/pku-alignment/beavertails
BeaverTails is a collection of datasets designed to facilitate research on safety alignment in large language models (LLMs).
ai-safety beaver datasets gpt human-feedback human-feedback-data language-model large-language-model llama llm llms rlhf safe-rlhf safety
Last synced: 09 Aug 2025
https://github.com/liziniu/ReMax
Code for Paper (ReMax: A Simple, Efficient and Effective Reinforcement Learning Method for Aligning Large Language Models)
large-language-models policy-gradient reinforcement-learning rlhf
Last synced: 09 May 2025
https://github.com/jackaduma/chatglm-lora-rlhf-pytorch
A full pipeline to finetune ChatGLM LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the ChatGLM architecture. Basically ChatGPT but with ChatGLM
chatglm chatglm-6b chatgpt deepspeed finetune gpt llama llm lora peft ppo pytorch reward-models rlhf
Last synced: 27 Apr 2025
https://github.com/modelscope/trinity-rft
Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).
Last synced: 14 Jun 2025
https://github.com/l294265421/alpaca-rlhf
Finetuning LLaMA with RLHF (Reinforcement Learning with Human Feedback) based on DeepSpeed Chat
alpaca chatgpt language-model large-language-models llama llm reinforcement-learning rlhf
Last synced: 29 Mar 2025
https://github.com/niutrans/vision-llm-alignment
This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.
alignment dpo llama3-vision llava llm mllm multi-model ppo reward rlhf sft vision
Last synced: 06 Apr 2025
https://github.com/NiuTrans/Vision-LLM-Alignment
This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.
alignment dpo llama3-vision llava llm mllm multi-model ppo reward rlhf sft vision
Last synced: 07 May 2025
https://github.com/log10-io/log10
Python client library for improving your LLM app accuracy
agents ai anthropic artificial-intelligence autonomous-agents debugging evaluations feedback fine-tuning llmops llms logging monitoring openai python rlhf
Last synced: 11 Apr 2025
https://github.com/nlp-uoregon/Okapi
Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback
bloom chatbot dataset instruction-tuning language-model large-language-models llama multilingual natural-language-processing nlp question-answering reinforcement-learning reinforcement-learning-from-human-feedback rlhf
Last synced: 16 Oct 2025
https://github.com/nlp-uoregon/okapi
Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback
bloom chatbot dataset instruction-tuning language-model large-language-models llama multilingual natural-language-processing nlp question-answering reinforcement-learning reinforcement-learning-from-human-feedback rlhf
Last synced: 28 Apr 2025
https://github.com/rkinas/rlhf_thinking_model
This repository serves as a collection of research notes and resources on training large language models (LLMs) and Reinforcement Learning from Human Feedback (RLHF). It focuses on the latest research, methodologies, and techniques for fine-tuning language models.
Last synced: 09 Apr 2025
https://github.com/opening-up-chatgpt/opening-up-chatgpt.github.io
Tracking instruction-tuned LLM openness. Paper: Liesenfeld, Andreas, Alianda Lopez, and Mark Dingemanse. 2023. “Opening up ChatGPT: Tracking Openness, Transparency, and Accountability in Instruction-Tuned Text Generators.” In Proceedings of the 5th International Conference on Conversational User Interfaces. doi:10.1145/3571884.3604316.
chatgpt chatgpt-free llm open-source rlhf transparency
Last synced: 15 Apr 2025
https://github.com/cogment/cogment-verse
Research platform for Human-in-the-loop learning (HILL) & Multi-Agent Reinforcement Learning (MARL)
cogment human-in-the-loop-learning reinforcement-learning rlhf
Last synced: 13 Oct 2025
https://github.com/llmsresearch/llm-flashcards
Visual knowledge bank for understanding large language models, with 180 concept cards from tokenization to deployment.
agents ai anki attention deep-learning fine-tuning flashcards gpt interview-preparation large-language-models llm llm-resources machine-learning nlp prompt-engineering rag rlhf study-notes transformers
Last synced: 13 Jul 2026
https://github.com/rlhflow/directional-preference-alignment
Directional Preference Alignment
ai-alignment large-language-models rlhf
Last synced: 09 Mar 2026
https://github.com/jackaduma/alpaca-lora-rlhf-pytorch
A full pipeline to finetune Alpaca LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Alpaca architecture. Basically ChatGPT but with Alpaca
alpaca chatgpt deepspeed finetune gpt llama llm lora peft ppo pytorch reward-models rlhf
Last synced: 16 Jun 2025
https://github.com/RLHFlow/Directional-Preference-Alignment
Directional Preference Alignment
ai-alignment large-language-models rlhf
Last synced: 13 Jul 2026
https://github.com/sail-sg/dice
Official implementation of Bootstrapping Language Models via DPO Implicit Rewards
alignment large-language-models preference-learning rlhf
Last synced: 16 Jul 2025
https://github.com/holarissun/rewardmodelingbeyondbradleyterry
official implementation of ICLR'2025 paper: Rethinking Bradley-Terry Models in Preference-based Reward Modeling: Foundations, Theory, and Alternatives
inverse-reinforcement-learning large-language-models largelanguagemodels llm-aligment llmalignment reward reward-modeling reward-models rlhf
Last synced: 19 Sep 2025
https://github.com/quentinwach/image-ranker
Rank images using TrueSkill by comparing them against each other in the browser. 🖼📊
data-analysis deep-learning fine-tuning finetune flux generative-ai image-analysis image-annotation image-annotation-tool image-classification image-rank rank ranking ranking-algorithm reinforcement-learning rlhf stable-diffusion trueskill trueskill-algorithm web-ui
Last synced: 20 Sep 2025
https://github.com/jackfsuia/nanorlhf
RLHF experiments on a single A100 40G GPU. Support PPO, GRPO, REINFORCE, RAFT, RLOO, ReMax, DeepSeek R1-Zero reproducing.
Last synced: 23 Mar 2025
https://github.com/astorfi/llm-alignment-project
A comprehensive template for aligning large language models (LLMs) using Reinforcement Learning from Human Feedback (RLHF), transfer learning, and more. Build your own customizable LLM alignment solution with ease.
ai alignment deep-learning generative-ai large-language-models llms machine-learning rlhf template
Last synced: 22 Jul 2025
https://github.com/vicgalle/zero-shot-reward-models
ZYN: Zero-Shot Reward Models with Yes-No Questions
llm reinforcement-learning reward-models rlaif rlhf trlx zero-shot
Last synced: 05 Mar 2025
https://github.com/wschella/llm-reliability
Code for the paper "Larger and more instructable language models become less reliable"
bloom evaluation gpt llama llm reliability rlhf scaling supervision
Last synced: 12 Apr 2025
https://github.com/ssbuild/chatglm_rlhf
chatglm_rlhf_finetuning
chat chatglm finetuning lora qlora reward rlhf
Last synced: 17 Oct 2025
https://github.com/dannylee1020/openpo
Building synthetic data for preference tuning
ai ai-feedback dpo evaluation finetuning huggingface llm llm-evaluation python rlaif rlhf synthetic-data synthetic-data-generation
Last synced: 11 Sep 2025
https://github.com/general-preference/general-preference-model
Official implementation of ICML 2025 paper "Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment" (https://arxiv.org/abs/2410.02197)
alignment large-language-models preference-modeling preference-optimization rlhf
Last synced: 19 Sep 2025