{"id":131535,"url":"https://github.com/thinkwee/awesomeopd","name":"awesomeopd","description":"Awesome List for On-Policy Distillation","projects_count":361,"last_synced_at":"2026-07-26T06:00:26.681Z","repository":{"id":354455499,"uuid":"1222978157","full_name":"thinkwee/AwesomeOPD","owner":"thinkwee","description":"Awesome List for On-Policy Distillation","archived":false,"fork":false,"pushed_at":"2026-06-23T19:48:24.000Z","size":4160,"stargazers_count":724,"open_issues_count":0,"forks_count":16,"subscribers_count":10,"default_branch":"main","last_synced_at":"2026-07-07T11:04:10.565Z","etag":null,"topics":["awesome-list","distillation","large-language-model","on-policy-distillation"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/thinkwee.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-04-27T22:30:47.000Z","updated_at":"2026-07-07T08:43:03.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/thinkwee/AwesomeOPD","commit_stats":null,"previous_names":["thinkwee/awesomeopd"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/thinkwee/AwesomeOPD","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thinkwee%2FAwesomeOPD","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thinkwee%2FAwesomeOPD/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thinkwee%2FAwesomeOPD/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thinkwee%2FAwesomeOPD/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/thinkwee","download_url":"https://codeload.github.com/thinkwee/AwesomeOPD/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thinkwee%2FAwesomeOPD/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35902828,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-26T02:00:06.503Z","response_time":89,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2026-05-17T07:50:13.254Z","updated_at":"2026-07-26T06:00:26.682Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["📚 Surveys, Foundations \u0026 Position Papers","⚡ Speculative-Decoding Distillation","🎭 OPD with Black-Box / Outcome-Based Teachers","🛠️ Frameworks \u0026 Toolkits","🖼️ Multimodal OPD (VLM, Video, Audio, Image)","🏭 Industrial / Production Model Reports","🤝 OPD-RL Hybrids — Inside-RL OPD","♻️ Self-Distillation with Privileged Context — OPSD","🔬 OPD with Larger External Teachers — White-Box","🌟 Curator's Picks — where to start","🤖 Agent \u0026 Embodied OPD (by application)","🧠 Reasoning OPD (by application)","Star History","Updates"],"sub_categories":["🔁 Iterative Self-Bootstrapping"],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"banner.png\" alt=\"banner Logo\" width=\"800\"\u003e\n\u003c/div\u003e\n\n\n\u003cdiv align=\"center\"\u003e\n\n![Surveys](https://img.shields.io/badge/Surveys_\u0026_Position-22-4E6813?style=for-the-badge)\n![White-Box](https://img.shields.io/badge/White--Box_OPD-51-BFA2DB?style=for-the-badge)\n![Black-Box](https://img.shields.io/badge/Black--Box_OPD-9-845C40?style=for-the-badge)\n\u003cbr\u003e\n![OPSD](https://img.shields.io/badge/OPSD-72-A259FF?style=for-the-badge)\n![Iterative](https://img.shields.io/badge/Iterative_Self--Bootstrapping-3-50C878?style=for-the-badge)\n![OPD-RL](https://img.shields.io/badge/OPD--RL_Hybrids-50-9B59B6?style=for-the-badge)\n\u003cbr\u003e\n![Reasoning](https://img.shields.io/badge/Reasoning_OPD-3-FF69B4?style=for-the-badge)\n![Multimodal](https://img.shields.io/badge/Multimodal_OPD-34-2ECC71?style=for-the-badge)\n![Agent](https://img.shields.io/badge/Agent_\u0026_Embodied-32-1F4CAD?style=for-the-badge)\n\u003cbr\u003e\n![SpecDec](https://img.shields.io/badge/Speculative_Decoding-15-D89F7B?style=for-the-badge)\n![Frameworks](https://img.shields.io/badge/Frameworks-16-FA5A4C?style=for-the-badge)\n![Industrial](https://img.shields.io/badge/Production_Reports-23-ffc884?style=for-the-badge)\n\n\u003c/div\u003e\n\n# When LLMs Distill On-Policy\n\n**AwesomeOPD** is an awesome list summarising **open-source repositories and papers** for training LLMs (and VLMs / agents / draft models) with **On-Policy Distillation (OPD)** and **On-Policy Self-Distillation (OPSD)**:\n - 🎯 **OPD = C1 + C2.** `C1`: student samples its own trajectories `y ~ π_student(·|x)` during training. `C2`: teacher provides per-token / sequence supervision on those student samples. Methods that only partially satisfy are flagged in **📝 Strictness notes** per section.\n - 🪞 **OPSD** = special case where teacher *is the same model*, conditioned on privileged context (verified trace / answer / \"be concise\" prefix / longer context) or an earlier checkpoint.\n - 🚀 Each entry is annotated along four design axes — **teacher source** (external · same model with privileged context · earlier checkpoint · multi-teacher · discriminator), **supervision signal** (logits / top-k / sequence reward / verbal score / discriminator / verifier / feature), **rollout consumption** (all / selected / truncated / replaced / as PG samples), and **pipeline slot** (cold-start / mid / RL-replacement / inside-RL / inter-stage / compression / continual-anchor).\n - ⚠️ Built by reading paper PDFs, project pages, and source code with LLM coding agents; manually reviewed but errors possible. PRs welcome.\n - 📌 If you find this repository helpful for your research, please cite it via the **\"Cite this repository\"** button in the right sidebar of the GitHub page.\n - 📅 Last updated: 2026-07-23\n\nTaxonomy:\n - **📚 Surveys, Foundations \u0026 Position Papers** — meta-references and seed papers (GKD, MiniLLM, Thinking Machines blog, Tencent / THUNLP surveys)\n - **🔬 White-Box** — logit-based OPD on student rollouts with an external teacher\n - **🎭 Black-Box** — discriminator / verbal / preference, no teacher logits\n - **♻️ OPSD** — privileged-context self-distillation (same model, different conditioning)\n - **🔁 Iterative Self-Bootstrapping** — same model as previous-checkpoint teacher\n - **🤝 OPD-RL Hybrids** — inside-RL OPD: KL-as-reward, RL+OPD fusion\n - **🧠 Reasoning / 🖼️ Multimodal / 🤖 Agent \u0026 Embodied** — by application; cuts across all teacher-source categories\n - **⚡ Speculative-Decoding Distillation** — drafter distillation; \"student\" is a draft model\n - **🛠️ Frameworks \u0026 Toolkits** — what to actually run\n - **🏭 Industrial / Production Reports** — what the labs ship\n\nShorthand: **FKL** = forward KL · **RKL** = reverse KL · **JSD** = Jensen–Shannon · **Skew-KL** / **AKL** = skewed / adaptive KL · `📄 paper-only` = no public code yet.\n\n## Updates\n\n\u003cdetails\u003e\n\u003csummary\u003e📢 click to expand\u003c/summary\u003e\n\n- **2026-07-23 (c)** — **independent re-audit of all 191 entries added in (a) and (b).** Every entry was re-read against its PDF by a second reviewer with no access to the first pass's reasoning; **189/191 verdicts were reproduced (98.9%)**.\n  - *Removed as out of scope* (2): **CollectionLoRA** ([2605.25378](https://arxiv.org/abs/2605.25378)) — single-step LoRA image editing, no sequence model; **FA-OPD** ([2605.27095](https://arxiv.org/abs/2605.27095)) — MLP policy on Gym/D4RL continuous control, no language or generative-sequence component. The flow-matching entries that *are* retained (DiffusionOPD, D-OPSD, OPSD-V, dOPSD, AnyFlow, WMSD, π-Flow, Qwen-Image-2.0-RL) generate a supervised *trajectory*; these two do not.\n  - *Corrected*: a transcription error in **Multi-Rollout OPD** (AIME25 mean@8 41.21 → **25.41**); a misattributed ablation in **OPSD Compresses RLVR** (−26.8% → −29.22% at −1.03 pp); conflated settings in **ADWIN**; a worst-case-only figure in **GeoSD**; teacher count in **MAD-OPD**; a loose \"doubles\" in **Tool-Call Boundary Drift**.\n  - *Metadata*: an unverifiable venue tag on **Revisiting OPD**; four wrong or over-specific Org cells (**HPD** listed a GitHub username; **Draft-OPD** listed an affiliation absent from the paper; **CoPD** \"JD Explore\" → JD.COM; **EffOPD** \"Tencent Hunyuan\" → Tencent); ten Date cells normalised to the arXiv-ID month; **ShortOPD** refiled from White-Box to OPSD (its teacher is its own pre-compression checkpoint, not an external model); two new acronym collisions documented (**PBSD** ×2, **COPD**/**CoPD**); all badge counts recomputed.\n  - *Checked and confirmed correct*: the eight cross-listed arXiv IDs are the intended cross-section duplicates; `HJSang/OPSD_OnPolicyDistillation` genuinely hosts three separate papers (PACED, TIP, Sparse-to-Dense); no dead repo links.\n- **2026-07-23 (b)** — **backfill sweep of 2026-04-25 → 2026-06-20: 152 papers read in full, 120 added, 32 rejected.** This period predates the previous update and had never been covered, so the list was missing roughly half the field's output during its busiest months. Highlights: *Decoupling KL and Trajectories* (the cleanest formal taxonomy of SFT/DAgger/offline-RL/OPD), Apple's *Unmasking OPD*, three independent parameter-geometry studies, the prefix-truncation family (Prune-OPD, ESR, ADWIN, Truncated OPD), cross-tokenizer OPD (SimCT, Tokenizer Barrier), Meta's logit-free *OmniOPD*, and two production reports with genuine MOPD (Kwai Keye-VL-2.0, Nemotron 3 Ultra). New acronym-collision and contested-finding notes added to the relevant strictness sections. Backfill entries carry their mechanism summary inline in the table rather than a separate technical-details row.\n- **2026-07-23 (a)** — **large sweep (~60 entries), every paper PDF read and every repo link checked.**\n  - *Surveys*: Formula-Driven Survey, NAIL, When Does Online IL Help, Demystifying OPD\n  - *White-Box*: IW-OPD, DEAR, ReNIO, PG-OPD, SEAD, DOPD, MOPD (Xiaomi), Blockwise Drift Gating, Direct-OPD, OPD², TOP-D, COPD, ShortOPD, TOPD\n  - *OPSD*: CaOPD, Vision-OPD, D-OPSD, PW-OPSD, OPSD Predictive Law, SDSD Diversity, PHF, Visual-OPSD, Denser ≠ Better, Purified OPSD, DemoPSD, Rethinking OPSD, GeoSD, AD-OPSD, dOPSD, PromptSD; BIRD → Iterative Self-Bootstrapping\n  - *OPD-RL*: OPID, ATOD, PASS, CRAFT, UCOB, DRIFT, RG-OPD, CPO, Distilled RL, CriPO, H²SD, CADENCE\n  - *Multimodal*: DiffusionOPD, V-Zero, H-OPD, OPSD-V, Med-OPD, OPD-IAD\n  - *Agent*: KbSD, Two-Phase Distillation, SEED, ReOPD, UI-MOPD, TurnOPD, FA-SD Failure, Tool-Call Boundary Drift, ROAD-VLA\n  - *SpecDec*: TIGER, AdaFlash · *Frameworks*: Spider, AsyncOPD, EasyOPD\n  - *Industrial*: Solar Open 2, KAT-Coder-V2.5, OvisOCR2, Mach-Mind-4-Flash, Agents-A1, Qwen-Image-2.0-RL, Audex\n  - *Maintenance*: TCOD upgraded to its released repo (COLM 2026); Draft-OPD relinked to the canonical `Simplified-Reasoning` org; venue tags added to HPD (ICML 2026) and Revisiting OPD (COLM 2026); all badge counts recomputed (several were stale)\n  - *Verified and excluded*: TREK, REGEN, A Few Teacher Steps, Prefix-GRPO, HyperDFlash, SOPD-SocialNav, DataFlex-RL — each fails C1 or C2 on a full read; rationale in the relevant section's strictness notes\n- **2026-06-23** — add FiRe-OPD (White-Box; hard trajectory filtering + soft token reweighting)\n- **2026-06-19** — add d-OPSD, RLCSD, SSOPD (OPSD); TGPO, OPD+ (OPD-RL); Flow-OPD, Decomposed-OPD/VGS (Multimodal); ROPD (Black-Box); BRTS, OPRD (White-Box); Draft-OPD (SpecDec); SDAR (Agent); *The Many Faces of OPD* (Surveys)\n- **2026-06-13** — add TRD (White-Box; trajectory-level refinement) + Li Jiang's OPD reflection blog\n- **2026-06-08** — add EMPO² (OPSD; cross-listed into Agent)\n- **2026-06-07** — add SPOT\n- **2026-05-18** — add COPSD, MSD\n- **2026-05-15** — add TCOD, Healthcare AI GYM, HyperEyes (and cross-list Skill-SD into Agent)\n- **2026-05-14** — add CORD\n- **2026-05-13** — add Uni-OPD; add HY-MT, Baichuan-M3, KAT-Coder-V2, HY-Embodied, Qwen3.5-Omni\n- **2026-05-11** — add Skill-SD\n- **2026-04-30** — add π-Play\n- **2026-04-29** — add SD-Zero, *Why Does Self-Distillation (Sometimes) Degrade Reasoning?*\n- **2026-04-28** — initial release; add NPO, VLA-OPD, KDFlow, HPD, DeepSeek-V4\n\n\u003c/details\u003e\n\n---\n\n## 📚 Surveys, Foundations \u0026 Position Papers\n\n| Resource | 🌟 Stars | Date | Org | Paper / Link | Title / Notes |\n| :----: | :----: | :----: |  :----: | :----: | :---- |\n| [GKD](https://arxiv.org/abs/2306.13649) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2306.13649) | 2023.06 | Google DeepMind (Agarwal et al.) | [arXiv 2306.13649](https://arxiv.org/abs/2306.13649) — implemented in [TRL `GKDTrainer`](https://github.com/huggingface/trl/blob/main/trl/experimental/gkd/gkd_trainer.py) | **GKD: On-Policy Distillation of Language Models — Learning from Self-Generated Mistakes** (Seminal · ICLR 2024) |\n| [Blog](https://thinkingmachines.ai/blog/on-policy-distillation/) | [![Blog](https://img.shields.io/badge/blog_post-3.2k_cookbook-blue?style=for-the-badge)](https://thinkingmachines.ai/blog/on-policy-distillation/) | 2025.10 | Thinking Machines Lab (Kevin Lu et al.) | [Blog](https://thinkingmachines.ai/blog/on-policy-distillation/) · [tinker-cookbook](https://github.com/thinking-machines-lab/tinker-cookbook) | **Thinking Machines Lab — On-Policy Distillation (blog)** |\n| [tinker-cookbook](https://github.com/thinking-machines-lab/tinker-cookbook) | \u003cimg src=\"https://img.shields.io/github/stars/thinking-machines-lab/tinker-cookbook?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2025.10 | Thinking Machines Lab | — | Reference impl. of the OPD recipe on the Tinker SDK |\n| [revisiting_opd](https://github.com/hhh675597/revisiting_opd) | \u003cimg src=\"https://img.shields.io/github/stars/hhh675597/revisiting_opd?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | CASIA (Fu et al.) | [arXiv 2603.25562](https://arxiv.org/abs/2603.25562) | Revisiting OPD: Failure Modes \u0026 Simple Fixes |\n| [Tencent OPD Survey](https://arxiv.org/abs/2604.00626) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.00626) | 2026.04 | Tencent (Mingyang Song \u0026 Mao Zheng) | [arXiv 2604.00626](https://arxiv.org/abs/2604.00626) | **A Survey of On-Policy Distillation for LLMs** |\n| [OPD](https://github.com/thunlp/OPD) | \u003cimg src=\"https://img.shields.io/github/stars/thunlp/OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | Tsinghua THUNLP | [arXiv 2604.13016](https://arxiv.org/abs/2604.13016) | **Rethinking On-Policy Distillation: Phenomenology, Mechanism \u0026 Recipe** |\n| [Lightning OPD](https://arxiv.org/abs/2604.13010) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.13010) | 2026.04 | Wu, Han, Cai | [arXiv 2604.13010](https://arxiv.org/abs/2604.13010) | **Lightning OPD: Efficient Post-Training with Offline OPD** |\n| [OPSD Survey](https://arxiv.org/abs/2605.18141) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.18141) | 2026.05 | Academic | [arXiv 2605.18141](https://arxiv.org/abs/2605.18141) | **A Brief Overview: On-Policy Self-Distillation in LLMs** |\n| [Many Faces of OPD](https://arxiv.org/abs/2605.11182) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.11182) | 2026.05 | UIUC (Ge Liu's ULab) | [arXiv 2605.11182](https://arxiv.org/abs/2605.11182) | **The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixes** — diagnoses when OPD/OPSD succeed or fail (distribution mismatch, optimization instability, PI-free policy limits) \u0026 proposes fixes; companion to Revisiting OPD \u0026 THUNLP Rethinking |\n| [Blog](https://louieworth.github.io/blog/opd_reflection/) | [![Blog](https://img.shields.io/badge/blog_post-reflection-blue?style=for-the-badge)](https://louieworth.github.io/blog/opd_reflection/) | 2026.06 | Li Jiang | [Blog](https://louieworth.github.io/blog/opd_reflection/) · [arXiv 2606.08432](https://arxiv.org/abs/2606.08432) | **On-Policy Distillation: Promise, Pitfalls, and Prospects** — reflection on OPD's promise, three failure mechanisms (local teacher noise, coverage decay, myopic gradients) \u0026 prospects; companion to TRD |\n| [Formula-Driven Survey](https://arxiv.org/abs/2606.22793) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.22793) | 2026.06 | Tsinghua (Bowen Zhang) | [arXiv 2606.22793](https://arxiv.org/abs/2606.22793) | **A Formula-Driven Survey and Research Agenda for OPD** — organises OPD as a *feedback-to-update path* rather than by KL direction; seven formula-derived variables + an explicit evidence-tier table (E0 papers → E3 own hypotheses) |\n| [NAIL](https://github.com/plau666/NAIL) | \u003cimg src=\"https://img.shields.io/github/stars/plau666/NAIL?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | Columbia (Sriraman, Liu, Hsu, Block) | [arXiv 2606.30923](https://arxiv.org/abs/2606.30923) | **Behavior Cloning is Not All You Need: The Optimality of OPD for Noisy Expert Feedback** — proves offline BC needs samples *exponential* in horizon under a noisy expert (and that this is necessary for **any** offline IL), while OPD is polynomial; proposes **NAIL** |\n| [Online IL Realizability](https://arxiv.org/abs/2606.30445) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.30445) | 2026.06 | Tsinghua / CMU / Berkeley / Harvard | [arXiv 2606.30445](https://arxiv.org/abs/2606.30445) | **When Does Online Imitation Learning Help in LLM Post-Training?** — challenges the error-accumulation story: under **realizability** OPD buys nothing over SFT; **non-realizability** is the real source of online gains |\n| [Demystifying OPD](https://arxiv.org/abs/2607.13399) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.13399) | 2026.07 | CUHK / Tencent AI Lab | [arXiv 2607.13399](https://arxiv.org/abs/2607.13399) | **Demystifying OPD: Roles, Pathologies, and Regulations** — OPD as *exploration catalyst*, not ceiling-raiser; diagnoses Student–Teacher Mismatch (the strongest teacher can be the worst) + Length Exploitation, and fixes both with clipping / log-compression |\n| [Unmasking OPD](https://arxiv.org/abs/2605.10889) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.10889) | 2026.05 | Apple | [arXiv 2605.10889](https://arxiv.org/abs/2605.10889) | **Unmasking OPD: Where It Helps, Where It Hurts, and Why** — training-free diagnostic that derives an *ideal* per-token gradient from empirical success probabilities and scores GKD / MiniLLM / Dr.GRPO objectives by cosine alignment to it. Finds guidance is far better aligned on **incorrect** rollouts than correct ones |\n| [Decoupling KL \u0026 Trajectories](https://arxiv.org/abs/2605.16826) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.16826) | 2026.05 | EIT Ningbo / HK PolyU / HKUST / SJTU | [arXiv 2605.16826](https://arxiv.org/abs/2605.16826) | **A Unified Perspective for SFT, DAgger, Offline RL, and OPD** — decomposes distillation along *prefix source* × *KL direction*, showing the four cells are exactly SFT / on-policy SFT / offline-RL distillation / OPD, and proves student-prefix + reverse KL equals a dense-reward REINFORCE objective |\n| [States, Not Tokens](https://arxiv.org/abs/2605.22731) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.22731) | 2026.05 | Independent (Dong Nie) | [arXiv 2605.22731](https://arxiv.org/abs/2605.22731) | **Post-Training is About States, Not Tokens** — recasts SFT/RL/OPD along *state source* × *signal source*; shows continuation-OPD from a deliberately **degraded** teacher still lifts the student past that teacher on all three axes |\n| [Geometry of OPD](https://arxiv.org/abs/2606.07082) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.07082) | 2026.06 | HKUST / Brown / ZJU / HK PolyU / USTC / BUPT | [arXiv 2606.07082](https://arxiv.org/abs/2606.07082) | **On the Geometry of On-Policy Distillation** — locates OPD between SFT and RLVR in parameter space (update sparsity 51.6% vs 8.1% / 77.2%) and identifies early **subspace locking** into a persistent low-rank update channel |\n| [Dense Supervision, Sparse Updates](https://github.com/SydCS/OPD-Param-Analysis) | \u003cimg src=\"https://img.shields.io/github/stars/SydCS/OPD-Param-Analysis?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | Nanjing Univ. / Amap Alibaba | [arXiv 2606.13657](https://arxiv.org/abs/2606.13657) | **On the Sparsity and Geometry of OPD** — OPD updates are tiny (0.036–0.142% relative Frobenius norm) and 67–90% coordinate-sparse, concentrated in FFN; **training only the discovered sparse mask nearly recovers full OPD** |\n| [Rock Tokens](https://github.com/YuxuanJiang1/Rock-Token) | \u003cimg src=\"https://img.shields.io/github/stars/YuxuanJiang1/Rock-Token?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | UMBC / Case Western / ASU / VU Amsterdam | [arXiv 2605.09253](https://arxiv.org/abs/2605.09253) | **Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in OPD** — tokens keeping high KL after saturation are mostly structural scaffolding whose gradients Adam neutralizes, and are causally near-irrelevant to accuracy |\n| [OPSD Compresses RLVR](https://arxiv.org/abs/2605.06188) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.06188) | 2026.05 | Yonsei | [arXiv 2605.06188](https://arxiv.org/abs/2605.06188) | **OPSD Compresses What RLVR Teaches** — separating correct-only from incorrect-only rollouts shows OPSD mainly *compresses* already-correct traces (−29.22% length at −1.03 pp accuracy, Qwen3-8B) rather than repairing failures; proposes it as a post-RL compaction stage |\n| [Extrapolation Cliff](https://arxiv.org/abs/2605.08737) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.08737) | 2026.05 | NTU Singapore | [arXiv 2605.08737](https://arxiv.org/abs/2605.08737) | **The Extrapolation Cliff in OPD of Near-Deterministic Structured Outputs** — derives a closed-form clip-safety threshold λ\\* beyond which reward-extrapolated OPD collapses output validity; the quantitative bound on G-OPD/ExOPD-style λ\u003e1 |\n\n\u003cdetails\u003e\n\u003csummary\u003e📋 Click to view technical details\u003c/summary\u003e\n\n| Resource | Loss / Divergence | Data | Teacher Access | Granularity | Notes |\n| :----: | :----: | :----: | :----: | :----: | :---- |\n| GKD (Agarwal) | Generalised JSD (FKL/RKL configurable) | Mixed (`λ` interpolates teacher↔student) | White-box | Token | The seminal paper that named OPD; introduced student-self-rollout supervision. |\n| Thinking Machines blog | Reverse KL (student‖teacher) | Student rollouts | White-box | Token | \"Swap KL ref model for stronger teacher\" recipe; one-line addition to RL trainer. Replicates Qwen3 result at ~1/10 RL cost. |\n| Revisiting OPD | Truncated reverse KL + top-p sampling + special-token masking | Student | White-box | Token (filtered) | Diagnoses 3 failure modes: imbalanced one-token signal, unreliable prefix guidance, tokenizer mismatch. |\n| Tencent OPD Survey | (survey) | (survey) | (survey) | (survey) | Catalogues 50+ methods; useful as a reference index. |\n| THUNLP Rethinking OPD | Reverse KL with progressive top-K alignment | Student | White-box | Token | Identifies two success conditions: compatible thinking patterns + genuinely new teacher capability. Recipe = **off-policy cold-start + teacher-aligned prompt selection**. |\n| Lightning OPD | Cached teacher log-probs over SFT rollouts (offline OPD) | Student (cached) | White-box | Token | Introduces \"teacher consistency\" — same teacher must be used for SFT and OPD or else gradient bias. Eliminates the live teacher server. |\n| OPSD Survey|(survey) | (survey) | (survey) | (survey) | Categorize eight designs; useful as a reference index. |\n| Many Faces of OPD | Reverse KL (FKL/RKL comparison) | Student (self-sampled trajectories) | White-box / self | Token | **Diagnostic study** (not a training method). Identifies three failure mechanisms — distribution mismatch, optimization instability, and a PI-free policy limit specific to OPSD — and proposes stop-gradient objectives + stabilized training as fixes. Covers both OPD and OPSD. |\n| Formula-Driven Survey | (survey) | (survey) | (survey) | (survey) | Refuses the \"KL direction × teacher access\" taxonomy; models OPD as feedback→update with seven variables: state distribution, feedback source, comparison support Ω_t (sampled-token / top-k / full-vocab), temporal credit A_t, gate-or-weight w_t, vocabulary routing, update route (direct-loss vs policy-gradient score-function). Carves out OPD-hybrids and OPD-adjacent (RLVR, teacher-forced KD, SFT) as out of scope. Proposes two unimplemented designs (GAE-OPD, CR-OPD), self-labelled as E3 hypotheses. |\n| NAIL (Behavior Cloning Is Not All You Need) | Forward-KL and reverse-KL variants, on a rollout distribution *distinct* from the scored policy | Student | White-box | Token | **Theory + method.** Noisy-expert model π*_η = (1−η)π* + η·ν. Offline BC needs n ≳ (1−η)^{−(H+2)}log\\|Π\\|/ε — exponential in horizon and *necessary for any offline IL algorithm*, unlike the horizon-free clean-expert result; the OPD variant is polynomial in H. Key design claim: the loss must distinguish the **rollout** distribution (greedy) from the **scored** policy (temp 1) — which standard OPD does not. NAIL ≈ BC/OPD at low noise, far better at high noise (GSM8K + modular addition). |\n| When Does Online IL Help | Reverse KL + entropy under the student (analysis only) | Student | White-box / self | Sequence (contextual bandit, H=1) | **Position/theory paper.** Under *realizability* (student class can represent the expert) SFT on expert samples fully matches expert performance and OPD adds neither accuracy nor speed — replicated on Countdown, GSM8K (Llama-3.2-3B) and DeepScaleR (R1-Distill-Qwen-1.5B). Under *misspecification*, discrepancy-based (TV/Hellinger) bounds go vacuous, and an information-theoretic lower bound with coverage coefficient C^e_∞ applies even at H=1. Notes strong-to-weak OPD lives in the non-realizable regime — which is why it works. |\n| Demystifying OPD | Reverse KL as policy gradient, per-token Δℓ_t = log π_T − log π_θ, PPO-clipped | Student | White-box | Token | **Analysis + method.** (1) OPD steers the student toward correct paths without expanding the capability ceiling (pass@1024 converges to the base model's); prompt diversity beats per-prompt rollout depth. (2) *Student–Teacher Mismatch*: Qwen3-4B-GRPO, the strongest teacher, is the worst teacher (student stuck ~2% AIME25); an Informativeness metric I = E[Δℓ̄\\|r=1] − E[Δℓ̄\\|r=0] predicts this in advance. (3) *Length Exploitation*: token-mean advantage lets the student wash out negative advantage by padding or truncate early on a favourable prefix. Fixes = Hard Clipping + Soft Log-Scale Compression, zero extra compute. Regulated OPD with a 1.7B/4B teacher beats prior OPD from a 30B teacher (AIME24 45.2 vs Uni-OPD 35.2 / G-OPD 37.3). |\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e📝 \u003cb\u003eStrictness notes\u003c/b\u003e (against the strict OPD definition \u003ccode\u003eC1: student samples its own trajectories during training\u003c/code\u003e + \u003ccode\u003eC2: teacher provides supervision on those samples\u003c/code\u003e)\u003c/summary\u003e\n\n- **Lightning OPD** — ⚠️ partially satisfies C1: teacher log-probs are pre-computed *once* over SFT rollouts and reused during training; student doesn't actively sample during the OPD step. Authors call this \"offline OPD\" explicitly. Listed in OPD because the data is past-student-generated rollouts, not teacher-generated.\n\n\u003c/details\u003e\n\n---\n\n## 🔬 OPD with Larger External Teachers — White-Box\n\nWhite-box methods use **teacher logits / log-probabilities** to supervise the student on **student-generated rollouts**. Each entry below has been verified to (a) train on student rollouts and (b) operate at the token level.\n\nMethods that turned out to be RL-style on verification have been moved to [OPD-RL Hybrids](#-opd-rl-hybrids); off-policy / pure-loss-function / pretraining-side methods are excluded from this list.\n\n| Resource | 🌟 Stars | Date | Org | Paper Link | Title / Notes |\n| :----: | :----: | :----: |  :----: | :----: | :---- |\n| [LMOps `/minillm`](https://github.com/microsoft/LMOps/tree/main/minillm) | \u003cimg src=\"https://img.shields.io/github/stars/microsoft/LMOps?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2023.06 | Microsoft / Tsinghua | [arXiv 2306.08543](https://arxiv.org/abs/2306.08543) | MiniLLM (ICLR 2024) |\n| [distillm](https://github.com/jongwooko/distillm) | \u003cimg src=\"https://img.shields.io/github/stars/jongwooko/distillm?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2024.02 | KAIST / Microsoft | [arXiv 2402.03898](https://arxiv.org/abs/2402.03898) | DistiLLM (ICML 2024) |\n| [google-research `/speculative_kd`](https://github.com/google-research/google-research/tree/master/speculative_kd) | \u003cimg src=\"https://img.shields.io/github/stars/google-research/google-research?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2024.10 | UCSB / Google | [arXiv 2410.11325](https://arxiv.org/abs/2410.11325) | Speculative KD (ICLR 2025) |\n| [distillm-2](https://github.com/jongwooko/distillm-2) | \u003cimg src=\"https://img.shields.io/github/stars/jongwooko/distillm-2?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2025.03 | KAIST / Microsoft | [arXiv 2503.07067](https://arxiv.org/abs/2503.07067) | DistiLLM-2 (ICML 2025 Oral) |\n| [DSKDv2](https://github.com/songmzhang/DSKDv2) | \u003cimg src=\"https://img.shields.io/github/stars/songmzhang/DSKDv2?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2025.04 | BJTU | [arXiv 2504.11426](https://arxiv.org/abs/2504.11426) | DSKDv2 — cross-tokenizer; supports on-policy mode |\n| [Constrained OPD](https://arxiv.org/abs/2509.22921) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2509.22921) | 2025.09 | Huawei Noah's Ark | [arXiv 2509.22921](https://arxiv.org/abs/2509.22921) | Constrained OPD (CMDP) |\n| [AdaSwitch](https://arxiv.org/abs/2510.07842) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2510.07842) | 2025.10 | RUC / Baidu | [arXiv 2510.07842](https://arxiv.org/abs/2510.07842) | AdaSwitch (on-/off-policy switching) |\n| [Veto](https://arxiv.org/abs/2601.07155) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2601.07155) | 2026.01 | SNU | [arXiv 2601.07155](https://arxiv.org/abs/2601.07155) | Veto (Stable OPD) — ACL 2026 Findings |\n| [G-OPD](https://github.com/RUCBM/G-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/RUCBM/G-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.02 | RUC / Tencent | [arXiv 2602.12125](https://arxiv.org/abs/2602.12125) | G-OPD |\n| [Fast OPD](https://arxiv.org/abs/2602.15260) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2602.15260) | 2026.02 | Industrial | [arXiv 2602.15260](https://arxiv.org/abs/2602.15260) | Fast OPD (prefix-truncated) |\n| [Entropy-Aware OPD](https://arxiv.org/abs/2603.07079) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2603.07079) | 2026.03 | KAIST / IBM | [arXiv 2603.07079](https://arxiv.org/abs/2603.07079) | Entropy-Aware OPD |\n| [REOPOLD](https://arxiv.org/abs/2603.11137) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2603.11137) | 2026.03 | KAIST / Microsoft | [arXiv 2603.11137](https://arxiv.org/abs/2603.11137) | REOPOLD (Relaxed OPD) — code soon |\n| [OPSD_OnPolicyDistillation](https://github.com/HJSang/OPSD_OnPolicyDistillation) | \u003cimg src=\"https://img.shields.io/github/stars/HJSang/OPSD_OnPolicyDistillation?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | LinkedIn | [arXiv 2603.11178](https://arxiv.org/abs/2603.11178) | PACED — frontier curriculum self-distill |\n| [TSD-KD](https://github.com/kmswin1/TSD-KD) | \u003cimg src=\"https://img.shields.io/github/stars/kmswin1/TSD-KD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | Korea Univ. | [arXiv 2603.13260](https://arxiv.org/abs/2603.13260) | TSD-KD — token-selective dual KD (ICLR 2026) |\n| [SCOPE](https://github.com/machine981/SCOPE) | \u003cimg src=\"https://img.shields.io/github/stars/machine981/SCOPE?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | USTC / Meituan / Fudan | [arXiv 2604.10688](https://arxiv.org/abs/2604.10688) | SCOPE — signal-calibrated dual-path |\n| [OPSD_OnPolicyDistillation](https://github.com/HJSang/OPSD_OnPolicyDistillation) | \u003cimg src=\"https://img.shields.io/github/stars/HJSang/OPSD_OnPolicyDistillation?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | Meta / LinkedIn | [arXiv 2604.14084](https://arxiv.org/abs/2604.14084) | TIP — Token Importance, shares LinkedIn OPSD repo with PACED |\n| [Hybrid-Policy-Distillation](https://github.com/zwhong714/Hybrid-Policy-Distillation) | \u003cimg src=\"https://img.shields.io/github/stars/zwhong714/Hybrid-Policy-Distillation?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | SJTU / Shanghai Innovation Institute / Tencent | [arXiv 2604.20244](https://arxiv.org/abs/2604.20244) | HPD — Hybrid Policy Distillation; LlamaFactory + verl backends (ICML 2026). ⚠️ headline Tables 2–4 use a *lightweight approximation* of on-policy sampling that avoids full-sequence rollouts; true on-policy results appear only in §5.4/Table 5 |\n| [BRTS](https://github.com/BWGZK-keke/BRTS) | \u003cimg src=\"https://img.shields.io/github/stars/BWGZK-keke/BRTS?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | JHU (Patel group) | [arXiv 2605.09725](https://arxiv.org/abs/2605.09725) | **BRTS — Best-of-N Teacher Rollout Selection**; augments student-context OPD with a curated teacher-context branch (correctness-first, then student-alignment) to cut single-rollout teacher variance |\n| [FiRe-OPD](https://github.com/YuYingLi0/FiRe-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/YuYingLi0/FiRe-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | THU / HKUST / Meituan (Li et al.) | [arXiv 2606.02684](https://arxiv.org/abs/2606.02684) | **FiRe-OPD — Filter, then Reweight**; decouples *optimization granularity* — **hard** trajectory-level filtering (drop bottom-p% rollouts by teacher log-prob) + **soft** token-level reweighting (teacher-confidence × student-confusion), arguing soft weighting beats hard token selection (cf. TIP); verl-based, with a multi-teacher math+code variant |\n| [trd](https://github.com/louieworth/trd) | \u003cimg src=\"https://img.shields.io/github/stars/louieworth/trd?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | McGill / Mila / UT Austin (Jiang et al.) | [arXiv 2606.08432](https://arxiv.org/abs/2606.08432) | **TRD — Trajectory-Refined Distillation**; diagnoses *prefix failure* of dense per-token OPD, refines student rollouts at trajectory level before distilling; verl-based, also applies to OPSD |\n| [OPRD](https://github.com/ShenzhiYang2000/OPRD) | \u003cimg src=\"https://img.shields.io/github/stars/ShenzhiYang2000/OPRD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | ZJU / Ant Group | [arXiv 2606.06021](https://arxiv.org/abs/2606.06021) | **OPRD — On-Policy Representation Distillation**; first OPD to supervise in *hidden-state space* (aligns teacher/student representations across layers on student rollouts, bypassing the LM head) rather than logits; built on the THUNLP OPD stack |\n| [IW-OPD](https://github.com/YannX1e/Importance-Weighted-On-Policy-Distillation) | \u003cimg src=\"https://img.shields.io/github/stars/YannX1e/Importance-Weighted-On-Policy-Distillation?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | Xidian / Georgia Tech / Amazon AGI SF Lab | [arXiv 2606.22600](https://arxiv.org/abs/2606.22600) · [project](https://yannx1e.github.io/IW-OPD/) | **IW-OPD — On the Position Bias of OPD**; supervising only the *prefix* 30% of tokens matches full-token OPD while suffix-30% barely learns; reweights the OPD advantage by a prefix-importance term derived from a trust-region argument |\n| [DEAR](https://arxiv.org/abs/2606.22830) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.22830) | 2026.06 | Meituan LongCat / Nanjing Univ. / TJUNLP | [arXiv 2606.22830](https://arxiv.org/abs/2606.22830) | **DEAR — Finding the Evidence**; argues entropy-selective OPD (cf. TIP) captures only *decision* tokens and structurally misses low-entropy, high-divergence **evidence** tokens where the student is confident yet wrong |\n| [ReNIO](https://github.com/BDML-lab/ReNIO) | \u003cimg src=\"https://img.shields.io/github/stars/BDML-lab/ReNIO?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | ECNU / Shanghai Innovation Institute | [arXiv 2606.23104](https://arxiv.org/abs/2606.23104) | **ReNIO — Reweighting Negative Trajectory Importance**; finds training on *incorrect* student outputs beats correct-only under both OPD and OPSD, then approximates correctness with a prefix-computable pivotal-token proxy |\n| [PG-OPD](https://arxiv.org/abs/2606.21994) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.21994) | 2026.06 | TeleAI / SJTU | [arXiv 2606.21994](https://arxiv.org/abs/2606.21994) | **Prefix-Guided OPD — Mining Golden Trajectories from Rollouts**; probes candidates at a short prefix by teacher–student top-k overlap and only continues the promising ones to full length (up to +4.80 avg, 2.46× wall-clock) |\n| [SEAD](https://arxiv.org/abs/2606.28562) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.28562) | 2026.06 | Capital One | [arXiv 2606.28562](https://arxiv.org/abs/2606.28562) | **SEAD — Competence-Aware OPD**; entropy zones the tokens (skip / RKL / FKL), cosine-anneals FKL→RKL, and adds the **first prompt-level curriculum for OPD** |\n| [DOPD](https://arxiv.org/abs/2606.30626) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.30626) | 2026.06 | JD Explore / NUS / MMLab CUHK / PKU | [arXiv 2606.30626](https://arxiv.org/abs/2606.30626) | **DOPD — Dual On-policy Distillation**; names **\"privilege illusion\"** (part of a privileged teacher's gap is information asymmetry, not transferable capability) and routes each token among four regimes by a privilege-advantage gap |\n| [MOPD](https://arxiv.org/abs/2606.30406) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.30406) | 2026.06 | Xiaomi LLM-Core / PKU / HKU / RUC | [arXiv 2606.30406](https://arxiv.org/abs/2606.30406) | **MOPD — Multi-Teacher OPD for Capability Integration**; the method paper behind [MiMo-V2-Flash](#-industrial--production-model-reports)'s MOPD stage. Shows **same-origin teachers are essential** — a stronger but distributionally distant teacher diverges outright |\n| [Blockwise Drift Gating](https://arxiv.org/abs/2606.24084) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.24084) | 2026.06 | Independent (Zheng \u0026 Jiang) | [arXiv 2606.24084](https://arxiv.org/abs/2606.24084) | Blockwise Policy-Drift Gating — student-only old↔current drift gate for OPD under rollout reuse (author-labelled *preliminary study*) |\n| [Direct-OPD](https://github.com/BytedTsinghua-SIA/Direct-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/BytedTsinghua-SIA/Direct-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.07 | Tsinghua AIR SIA-Lab / ByteDance Seed / PKU | [arXiv 2607.05394](https://arxiv.org/abs/2607.05394) · [project](https://bytedtsinghua-sia.github.io/Direct-OPD/) | **Direct-OPD — Weak-to-Strong Generalization**; distills the teacher's *implicit reward* (post-RL minus pre-RL log-ratio) instead of its distribution, so a **weaker** post-RL teacher can improve a stronger student. Qwen3-1.7B AIME24 48.3→58.3 in ~4 h on 8×A100 |\n| [OPD²](https://github.com/naver-ai/opd2) | \u003cimg src=\"https://img.shields.io/github/stars/naver-ai/opd2?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.07 | NAVER AI Lab | [arXiv 2607.15161](https://arxiv.org/abs/2607.15161) | **On-Policy Delta Distillation**; replaces the teacher–student log-ratio reward with a **delta signal** (teacher minus *its own base model*), plus a sign-agreement gate. Qwen3-1.7B math avg 34.8→54.6 |\n| [TOP-D](https://arxiv.org/abs/2607.04751) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.04751) | 2026.07 | HKUST(GZ) / Microsoft | [arXiv 2607.04751](https://arxiv.org/abs/2607.04751) | **Trust Region Policy Distillation**; interpolates teacher and student *in probability space* so the token reward is lower-bounded and gradient variance is **provably bounded**. AIME24 avg@32 50.42 vs 24.58 for standard OPD |\n| [COPD](https://arxiv.org/abs/2607.19046) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.19046) | 2026.07 | SJTU / Alibaba Qwen | [arXiv 2607.19046](https://arxiv.org/abs/2607.19046) | **Contrastive OPD**; the teacher re-scores the student's tokens twice — under a \"light-thinking\" and a \"heavy-thinking\" prefix — and the *contrast* becomes the advantage. +2.1 pp with 57% shorter responses; ships a **COPSD** self-distillation variant |\n| [TOPD](https://arxiv.org/abs/2607.16872) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.16872) | 2026.07 | CASIA / UCAS | [arXiv 2607.16872](https://arxiv.org/abs/2607.16872) | **Trace-Based OPD for masked diffusion LMs**; supervises only *trace-aligned* denoising decisions (commitments surviving into the final answer), avoiding the backward-reconstruction states random-mask supervision creates. MATH500 +5.7 with 4× fewer rollout rounds |\n| [AOPD](https://arxiv.org/abs/2605.06387) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.06387) | 2026.05 | HUST / PKU / Meituan | [arXiv 2605.06387](https://arxiv.org/abs/2605.06387) | **Asymmetric OPD**; splits the OPD gradient into positive (exploit) and non-positive (imitate) token regions, replacing the noisy policy gradient with truncated forward KL only on the latter. Preserves math accuracy after continual tool-use training (+0.61 vs −7.49 for OPD) |\n| [SimCT](https://github.com/sunjie279/SimCT-) | \u003cimg src=\"https://img.shields.io/github/stars/sunjie279/SimCT-?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | Tencent Hunyuan / USTC / Shanghai Innovation Institute | [arXiv 2605.07711](https://arxiv.org/abs/2605.07711) | **SimCT — Recovering Lost Supervision for Cross-Tokenizer OPD**; builds a common supervision space of shared-vocab tokens ∪ *minimal aligned units* (finest jointly-tokenizable spans) so the unchanged OPD loss applies across tokenizers |\n| [Prune-OPD](https://github.com/yangzhch6/Prune-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/yangzhch6/Prune-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | HKUST(GZ) / MBZUAI / UC Merced / SYSU | [arXiv 2605.07804](https://arxiv.org/abs/2605.07804) | **Prune-OPD** — monitors per-position teacher/student top-k overlap, attenuates rewards on prefix drift and truncates once supervision is locally unexploitable; **−37.6–68.0% training time**. The mechanism [KAT-Coder-V2.5](#-industrial--production-model-reports) credits for its drift-aware truncation |\n| [ESR](https://arxiv.org/abs/2605.27028) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.27028) | 2026.05 | UCLA / BIGAI | [arXiv 2605.27028](https://arxiv.org/abs/2605.27028) | **Less is More: Early Stopping Rollout for OPD**; names **Off-Policy Teacher Decay** — conditioning the teacher on the student’s drifting prefix collapses its late-token distribution toward student-level autocomplete. Beats full-rollout OPD in every cell at up to 24× lower cost |\n| [ADWIN](https://arxiv.org/abs/2605.28396) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.28396) | 2026.05 | PKU / Tencent | [arXiv 2605.28396](https://arxiv.org/abs/2605.28396) | **ADWIN — Adaptive Windows for Horizon-Aware OPD**; truncates by a prefix-admissibility criterion based on **gradient-cosine alignment** between short-prefix and full-rollout updates, audited by delayed full-rollout probes. 59.3→60.9 avg at ~3.4× fewer FLOPs single-task (the 4.1× figure is from the separate strong-to-weak setting) |\n| [Truncated OPD](https://arxiv.org/abs/2605.31490) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.31490) | 2026.05 | CASIA / Meituan | [arXiv 2605.31490](https://arxiv.org/abs/2605.31490) | **Are Full Rollouts Necessary for OPD?** — proves sequence-level OPD accumulates noisy future signal at **O(T³) MSE** while token-level does not; truncating to 10% of the horizon matches full-rollout OPD while cutting training time **82%** |\n| [Prefix Teach, Suffix Fade](https://arxiv.org/abs/2605.13643) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.13643) | 2026.05 | ZJU / Meituan LongCat / Jilin | [arXiv 2605.13643](https://arxiv.org/abs/2605.13643) | **Local Teachability Collapse in Strong-to-Weak OPD**; cuts dense supervision at a **BIC-detected change-point** in the teacher’s top-1/top-2 margin over the student’s reachable candidates. Avg 36.8→40.1 |\n| [KAT](https://arxiv.org/abs/2606.09471) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.09471) | 2026.06 | HKUST(GZ) / HKUST / HK PolyU / EIT | [arXiv 2606.09471](https://arxiv.org/abs/2606.09471) | **Escaping the KL Agreement Trap in OPD** — low KL is not evidence of health: the student drifts into a corrupted prefix and the teacher *locally agrees*, silencing the corrective signal. Sliding-window detection cuts rollout length 59.7% at 2.4× speedup |\n| [TA-OPD](https://github.com/wyy-code/TA-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/wyy-code/TA-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | HK PolyU / InfiX.ai | [arXiv 2605.26844](https://arxiv.org/abs/2605.26844) | **Not All Disagreement Is Learnable — Token Teachability in OPD**; supervises only a support-aligned \"teachable\" subset, matching or beating full-token OPD with **5–10% of tokens** |\n| [TRB](https://arxiv.org/abs/2605.31159) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.31159) | 2026.05 | T-Tech | [arXiv 2605.31159](https://arxiv.org/abs/2605.31159) | **Trust-Region Behavior Blending**; leaves the OPD loss untouched and instead controls the *rollout behaviour policy* μ ∝ π_S^(1−β)·π_T^β under a trust-region constraint, annealing back to pure on-policy sampling |\n| [TrOPD](https://github.com/Xingrun-Xing2/TrOPD) | \u003cimg src=\"https://img.shields.io/github/stars/Xingrun-Xing2/TrOPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | Samsung Research Beijing / Oxford / PKU | [arXiv 2606.01249](https://arxiv.org/abs/2606.01249) | **Trust Region On-Policy Distillation**; partitions student tokens into a teacher-verifiable trust region (reverse KL) vs outliers (top-k forward KL), plus annealed off-policy guidance. AIME25 44.06 vs OPD 40.72 |\n| [Near-Future Guidance](https://arxiv.org/abs/2606.00305) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.00305) | 2026.06 | UMBC | [arXiv 2606.00305](https://arxiv.org/abs/2606.00305) | **Bridging Reasoning Trajectories in OPD via Near-Future Guidance**; shows per-token KL is a weak proxy for real trajectory divergence (Pearson **r=0.126**) and adds an optimal-transport-aligned target spanning several future positions |\n| [PowerOPD](https://arxiv.org/abs/2606.17199) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.17199) | 2026.06 | EIT Ningbo / HK PolyU / SJTU / Waterloo | [arXiv 2606.17199](https://arxiv.org/abs/2606.17199) | **Stabilizing OPD with Bounded Power Transformation**; the vanilla reward log(π_T/π_θ) is *unbounded* and drives instability, so it is replaced by a sign-consistent Box-Cox transform bounded in [−1,1]. +6.37 Avg@8 at 59% less wall-clock |\n| [NPD](https://arxiv.org/abs/2605.05940) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.05940) | 2026.05 | Huawei / Tianjin Univ. | [arXiv 2605.05940](https://arxiv.org/abs/2605.05940) | **Near-Policy Distillation**; decouples generation from updates so teacher top-k logits can be computed by packed parallel prefill — **8.1× speedup**, at the cost of deliberate policy lag |\n| [RWOPD](https://arxiv.org/abs/2605.13501) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.13501) | 2026.05 | NUS | [arXiv 2605.13501](https://arxiv.org/abs/2605.13501) | **Reward-Weighted OPD with an Open Property-Equivalence Verifier**; forward KL on rollouts passing a SymbiYosys+Z3 equivalence check, for NL→SystemVerilog Assertions. **Beats its own 14B teacher and 671B general baselines** |\n| [ProteinOPD](https://github.com/THU-AI4S/ProteinOPD) | \u003cimg src=\"https://img.shields.io/github/stars/THU-AI4S/ProteinOPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | Tsinghua / IDEA / HKUST(GZ) / NTU | [arXiv 2605.10189](https://arxiv.org/abs/2605.10189) | **ProteinOPD** — multi-teacher OPD for protein language models; distills a normalized **product-of-experts** consensus over preference-specific teachers via token-level generalized JSD. PPL −83.7%, thermostability +54.2% |\n| [MAD-OPD](https://github.com/chiefovoavicii/MAD-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/chiefovoavicii/MAD-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | HUST / Alibaba | [arXiv 2605.01347](https://arxiv.org/abs/2605.01347) | **Breaking the Ceiling in OPD via Multi-Agent Debate**; two larger teachers (K=2) *debate* across rounds to form the token-level target, with task-adaptive divergence. A 4B student **beats its own 14B teacher** on LiveCodeBench-v6 |\n\n\u003cdetails\u003e\n\u003csummary\u003e📋 Click to view technical details\u003c/summary\u003e\n\n| Method | Loss / Divergence | Data | Granularity | Domain | Notes |\n| :----: | :----: | :----: | :----: | :----: | :---- |\n| MiniLLM | Reverse KL via policy gradient | Student | Sequence (PG) | General | The seminal \"OPD\" recipe by Yuxian Gu et al.; predates GKD by days. Mode-seeking. |\n| DistiLLM | Skewed-KL (mix of FKL/RKL) | Mixed (adaptive off→on, with student samples) | Token | General | Skew parameter `α` interpolates between FKL and RKL; importance-reweighted student samples. |\n| Speculative KD (Xu) | Interleaved propose-and-correct (gated KL) | Student-proposed, teacher-corrected | Token | General | Bridges teacher-student gap via interleaved sampling. |\n| DistiLLM-2 | Contrastive: Skew-FKL on teacher data + Skew-RKL on student data | Mixed | Token | General | Asymmetric losses on each data source; ICML 2025 oral. |\n| DSKDv2 | KL in dual aligned space; explicit on-policy mode | Student | Token | Cross-tokenizer | Cross-vocabulary distillation; supports both on/off-policy. |\n| Constrained OPD | KL-constrained CMDP | Student | Token | General | Hard KL constraint instead of soft penalty. Borderline OPD-RL. |\n| AdaSwitch | Adaptive on/off-policy switching | Mixed | Token | General | Switches between teacher-data and student-rollout based on divergence threshold. |\n| Veto | Logit-space geometric bridge with adaptive gradient veto | Student | Token | General | Adaptive Target Reformulation. |\n| G-OPD / ExOPD | Reverse KL + scaled reward extrapolation | Student | Token | General | Generalises OPD as KL-constrained RL; allows reward scale \u003e 1 to \"exceed\" the teacher. |\n| Fast OPD | Prefix-truncated distillation reducing FLOPs | Student | Token (truncated) | Reasoning | 2× to 47× speedup via reasoning-prefix truncation. |\n| Entropy-Aware OPD | Switch between FKL and RKL based on teacher entropy | Student | Token | Reasoning | When teacher entropy high → FKL; low → RKL. |\n| REOPOLD | Mixture-based reward clipping + entropy-based dynamic sampling | Student | Token | Reasoning | \"Relaxed OPD\"; views OPD as policy optimisation with teacher-student log-ratio reward. |\n| PACED | Frontier curriculum at student competence boundary | Student | Token | General | Self-distill style (privileged-context / earlier-checkpoint); difficulty weighting `w(p)=p(1−p)`. |\n| TSD-KD | Indirect (student-propose / teacher re-rank) + direct selective logit KD | Mixed | Token (selected) | General | Hybrid; partial OPD + partial preference. |\n| SCOPE | Teacher-PPL-weighted KL on incorrect rollouts; student-PPL-weighted MLE on correct | Student | Token | Reasoning | Signal-Calibrated OPD with Dual-Path Adaptive Weighting; verifier-routing. |\n| TIP | Top-50% high-entropy student tokens carry the OPD signal | Student (selected) | Token (filtered) | Reasoning | ~47% memory savings; only entropy-high student tokens trained. |\n| HPD | Reweighted log-likelihood unifying FKL + RKL | Mixed (off-policy + lightweight approximate on-policy sampling) | Token | General | Unifies KD as token-level reweighted likelihood; lightweight on-policy sampling preserves training efficiency. |\n| TRD | Trajectory-level refinement of student rollouts, then distillation | Student (refined) | Trajectory → Token | Reasoning | Argues dense per-token supervision causes *prefix failure*; revises problematic student predictions at the trajectory level before distilling. Generalises to on-policy self-distillation. |\n| BRTS | Token-level KL on student rollouts + auxiliary teacher-context loss | Student + Best-of-N-selected teacher rollouts | Token | Reasoning (AIME/AMC) | Selection waterfall = correctness → student-alignment → ground-truth-guided recovery; the curated teacher trajectory replaces a single high-variance teacher rollout. |\n| OPRD | Layer-wise representation alignment (deterministic; avoids Monte-Carlo KL variance over large vocab) | Student | Hidden-state (per-layer) | Reasoning (competition math) | Lifts distillation from output space into hidden-state space — \"bypasses the LM head entirely\". Teacher access is white-box (representations rather than logits): the first **feature-based** OPD. |\n| FiRe-OPD | RKL with hard trajectory filtering + soft token reweighting | Student (filtered) | Trajectory (hard) → Token (soft) | Reasoning + Code | Decouples granularity: hard filtering wins at the trajectory level, soft weighting beats hard selection at the token level. Weight `w_t=(1+α·c^T_t)(1+β·c^S_t)` from teacher confidence × student confusion, per-trajectory normalised. +6.25 AIME24 (Qwen3-4B, 30B teacher); multi-teacher math+code variant. |\n| IW-OPD | Per-token OPD advantage × detached, normalised **prefix-importance** weight | Student | Token (position-reweighted) | Reasoning + code | Diagnoses a *position bias*: student rollouts drift out of the teacher's distribution as they lengthen, so late tokens carry degraded supervision. Derived as the closed-form optimum of `min_q D_KL(q‖π_T) s.t. D_KL(q‖π_θ) ≤ ρ`, changed of measure back to π_θ. Stabilised with log-space scaling, *unsigned* accumulation, within-sample normalisation. No extra teacher forward passes. **+6.9 AIME-2025 at step 10**; gains grow as the student shrinks (+1.0/+1.2/+1.9 for 4B/1.7B/0.6B) and as the compression ratio rises (+4.0% at 1.0× → +14.9% at 6.7×). Composable with ExOPD. |\n| DEAR | Teacher–student log-prob gap advantage on a selected token subset (D ∪ E, ~36% of tokens) | Student (selected) | Token (filtered) | Reasoning + code | **Direct methodological rebuttal to [TIP](#-opd-with-larger-external-teachers--white-box).** Entropy selectors find *decisions* (where to branch) but the substantive knowledge sits in *evidence* tokens — low-entropy, high-divergence positions where the student is confident yet wrong, structurally unreachable by any entropy criterion. Stage 2 scores non-decision tokens by hidden-state cosine similarity to decision anchors × normalised divergence. Gradient-mass coverage: 39.1% (decision-only) / 35.9% (random) / **75.8% (DEAR)**. Up to +2.5 pp competition math, +5.7 pp code. |\n| ReNIO | Standard token-level OPD divergence × a per-sample weight from clipped \"pivotal tokens\" | Student (reweighted) | Sample-level weight on token loss | Math + code | Controlled filtering shows incorrect-only training beats correct-only under both OPD (+2.59) and OPSD (+2.50), and yields longer responses with more reflection markers. Because correctness labels need full answer-bearing rollouts (defeating OPD's short-prefix cost advantage), the weight is built from prefix-computable log-ratios instead. Evaluated in both external-teacher OPD and teacher-free OPSD (teacher = the model's own initial parameters) modes; the OPSD tables are the stronger half. |\n| PG-OPD | Unchanged reverse KL; the contribution is rollout budgeting | Student (prefix-screened) | Token | Math reasoning | All K candidates decode to a fixed prefix; teacher–student top-k overlap over the first R probe tokens scores each; only high scorers (plus a guaranteed per-prompt best) continue to full length, while prefix tokens of *all* candidates still supervise. Up to +4.80 avg over OPD (59.40→64.20) at 2.46× wall-clock; beats PRUNE-OPD. |\n| SEAD | Zone-dependent: skip (both low-entropy) / RKL (teacher confident, student uncertain) / FKL (teacher uncertain) | Student | Token (zoned) + prompt (curriculum) | Math reasoning | Three scales of competence-adaptivity. Token: joint top-k entropy assigns ~50% of positions zero gradient, ~40% RKL to sharpen, ~10% FKL to preserve multi-path diversity. Temporal: α cosine-anneals 0.8→0.0, a continuous exploration→refinement transition. Prompt: a competence-gated curriculum admits prompts as the student's measured pass rate allows. 64.0 avg vs 59.2 vanilla OPD; a 2³ factorial shows curriculum alone is the single strongest factor (+4.20) while zones + annealing are super-additive. |\n| DOPD | Four-way routed: Top-K RKL→teacher / stop-grad self-anchor / full-vocab JS→teacher / RKL→privileged student | Student (non-privileged rollout) | Token (routed) | Reasoning + VLM | The privileged-teacher trick raises the ceiling but conflates two gaps: the *transferable capability gap* the student is meant to close, and the *information-asymmetry gap* it can only mimic. Uniform distillation therefore teaches privileged shortcuts and collapses entropy. A per-token gap A = \\|log π_T − log π_S\\| measured under identical privileged conditions selects the regime. Qwen3-8B→1.7B avg 51.4 vs 43.9 vanilla OPD, recovering 89.8% of the teacher–student gap. |\n| MOPD (Xiaomi) | Per-token reverse KL; PG form with clipped Â = sg[log π_teacher − log π_student], or top-k (k=64) with a bias-correction term | Student | Token | Multi-domain (Math, IF, SWE, Code, Tool Use) | The canonical **inter-stage consolidation** recipe: one shared SFT checkpoint → fully parallel per-domain RL experts → student re-initialised from that same SFT checkpoint and distilled from domain-routed teachers on its own rollouts. Teachers run as standalone prefill services so teacher cost hides behind student sampling. Normalised score 0.9373 vs 0.8818 (Mix-RL) / 0.8241 (off-policy finetune) / 0.8574 (param-merge). **Same-origin teachers are essential** — substituting the stronger but distant Qwen3-235B-A22B collapses it to 0.60 (PG) / −1.19 (top-k) with divergence at step 18. Iter-2 reaches 0.986. |\n| Blockwise Drift Gating | Existing OPD loss × detached gate g = exp(−τ\\|s\\|) from block-aggregated old↔current drift | Student (rollout-reuse) | Block (64-token / newline span) | Math reasoning | Targets the PPO-style setting where one rollout is reused across epochs, so π_θ drifts from the π_old that generated it. Teacher targets, support and rollout policy are all untouched. LSM+Block64 is the best trained student at 53.3 avg (vs LSM 49.8, base 40.1, teacher 65.2). |\n| Direct-OPD | Teacher **log-ratio** r_t = log π_T − log π_T_ref (the teacher's implicit reward), Rao–Blackwellised over top-k, with adaptive α anchoring to the student's init | Student | Token | Reasoning | Weak-to-strong: run cheap RL on a *small* model, then transfer what that run learned to a bigger target. Vanilla OPD toward a weak post-RL teacher actively degrades a stronger student (R1-Distill-7B 56.7→~50); reading the pre-RL→post-RL *shift* instead improves it. The objective is itself KL-regularised RL anchored at the student's own initialisation. Qwen3-1.7B AIME24 48.3→58.3 (~4 h, 8×A100, vs Polaris direct RL at 32×A100 for a week); shifts compose (48.3→58.3→63.8). Transfer works **without** rising teacher–student top-k overlap, so it is not progressive imitation. |\n| OPD² (Delta Distillation) | Delta reward R^Δ = log π*(y_t) − log π*_base(y_t), both advantages top-k-centred, with a sign-agreement gate | Student | Token | Math + science + code | Isolates *what post-training added* to the teacher rather than the teacher's absolute distribution, so the student is not pulled toward capabilities it already shares with the base model. The sign-agreement gate zeroes updates where the delta and ordinary OPD advantages disagree, fixing the convergence point. Implemented on TRL GRPOTrainer, 100 steps. Qwen3-1.7B math avg 34.8→54.6 (vs 51.0 OPD, 51.4 ExOPD); Qwen3-8B AIME24 76.2. |\n| TOP-D | Probability-space teacher–student interpolation collapsing to r̃ = log(αρ + 1 − α); GRPO-style clipped surrogate with token-level normalisation | Student (rollouts reused across mini-batch epochs) | Token | Math reasoning | The smoothed reward is strictly lower-bounded, so gradient variance is bounded (Thm 4.2) instead of exploding as π*→0 — at zero extra compute over standard OPD. Comes with a global convergence bound and a monotonic-improvement bound. Qwen3-8B-Base ← Qwen3-30B-A3B: AIME24 avg@32 **50.42 vs 24.58** standard OPD, AIME25 +10.73, AIME26 +18.64; vs GRPO 30.10 / DAPO 32.92. |\n| COPD | Log-likelihood **contrast** A_t = ℓ^LT − ℓ^HT between two teacher scorings of the same student tokens, clipped and used as a PPO-style advantage | Student | Token | Multimodal reasoning | Instead of matching a distribution, it asks the teacher which of two *reasoning modes* the token belongs to. No accuracy reward, no verifier, no length penalty. vs ExOPD on Qwen3-VL-8B→2B: +2.1 pp with **57.0% shorter** responses (Acc@1K 26.7→64.4) at 60 GPU-h vs 132 (OPD) / 150 (ExOPD). The **COPSD** variant swaps in a frozen snapshot of the student itself: +3.5 pp, −63.8% length. |\n| ShortOPD | Generalized JSD (α=0.5) over top-100 logits + aggregated tail mass | Student (adaptive horizon) | Token | Compression (post-pruning recovery) | Pruning Qwen3-4B-Instruct by 4/36 layers collapses greedy GSM8K 88.1→49.0 but pass@64 recovers to 91.2 — correct trajectories are demoted, not erased, so recovery needs on-policy states and dense targets with the frozen pre-compression parent as teacher (no labels, verifier, or external teacher). At 25% pruning 55–75% of early rollouts end in repetitive suffixes carrying ~35× less signal, so EMAs of repetition / truncation / effective length adapt the per-step horizon. Avg over 8 tasks 5.71→48.46 = **64.5% of the dense teacher**, +17.9 over the best off-policy baseline, 8.5 h vs 35.9 h. |\n| TOPD (masked dLLM) | Token-level reverse KL via a sampled-token score-function estimator, on trace-aligned decisions only | Student (own denoising trajectory) | Token (trace-aligned) | Reasoning (diffusion LMs) | Random-mask supervision — the default in dLLM RL/SFT pipelines — can reveal later answer tokens while hiding earlier reasoning ones, creating backward-reconstruction states the student never visits at inference. TOPD instead rolls out the real low-confidence-remasking decoder and keeps only commitments that survive into the final answer. SDAR-4B-Chat ← TraDo-8B-Instruct: MATH500 70.2→75.9, matching TraceRL-trained TraDo-4B with **4× fewer rollout rounds**. Ablations: on-policy 75.2 \u003e off-policy 74.5 \u003e semi-AR SFT 73.5; trace-aligned 75.2 \u003e random-mask 74.3; RKL \u003e JSD \u003e FKL. |\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e📝 \u003cb\u003eStrictness notes\u003c/b\u003e\u003c/summary\u003e\n\n- **BRTS** — ⚠️ Partially dilutes C1: the primary student-context leg is strict OPD (student trains on its own rollouts), but the auxiliary teacher-context branch supervises on *teacher*-generated (off-policy) trajectories. Listed because the student-context leg is the core objective and the teacher branch only stabilises it.\n- **OPRD** — ⚠️ Not logit-based: C1 ✓ / C2 ✓ on student rollouts, but supervision is **feature/representation-space** (hidden states across layers), not next-token logits. Listed in White-Box because teacher access is white-box; flagged here because the \"feature\" supervision signal sits outside the section's default logit-matching form.\n- **Direct-OPD** — ⚠️ Two departures from the section default. The supervision is a teacher **log-ratio** (implicit reward), not the teacher distribution, so it is arguably an OPD/RL hybrid; and the teacher is *smaller* than the student (weak-to-strong), which inverts the usual strong-to-weak setting. C1 ✓ / C2 ✓ — the teacher is still queried on the student's own visited prefixes.\n- **OPD²** — ⚠️ Stronger access assumption than ordinary white-box OPD: three models must be loaded (student, teacher, **and the teacher's pre-post-training base checkpoint**), which is not available for most released teachers.\n- **COPD** — ⚠️ C2 is satisfied at token granularity (teacher log-probs on student tokens) but the loss is **not** a KL/distribution-matching objective — the teacher signal is converted into an RL advantage. Listed in White-Box rather than OPD-RL Hybrids because there is no reward model or verifier anywhere in the objective. Its COPSD variant is squarely OPSD.\n- **TOP-D** — ⚠️ C1 slightly relaxed: the internal trust-region iterations deliberately reuse rollouts across mini-batch epochs, so updates are *near*-on-policy. The paper's own \"w/o off-policy\" ablation is the strict variant.\n- **Blockwise Drift Gating** — ⚠️ Authors describe it as \"a preliminary empirical study\": one student, one teacher, one dataset, **no repeated seeds**, and AIME sets with tiny sample counts — a 1.7-point pass@8 delta on 4 benchmarks is plausibly within noise. Also a pure loss-weighting heuristic on an existing OPD loss, not a new supervision mechanism.\n- **ShortOPD** — teacher is the student's own *uncompressed parent*, so it sits between the compression slot and OPSD. C1 ✓ / C2 ✓ otherwise textbook.\n- **⚠️ Acronym collisions.** This field has reused several abbreviations for unrelated work; always resolve by arXiv ID:\n  - **TOPD** = Trace-Based OPD for masked dLLMs ([2607.16872](https://arxiv.org/abs/2607.16872)) · Near-Future-Guidance trajectory OPD ([2606.00305](https://arxiv.org/abs/2606.00305)) · Truncated OPD ([2605.31490](https://arxiv.org/abs/2605.31490)).\n  - **MOPD** = Multi-*Teacher* OPD, Xiaomi ([2606.30406](https://arxiv.org/abs/2606.30406)) · Multi-*Rollout* OPD, Microsoft/CMU/Purdue ([2605.12652](https://arxiv.org/abs/2605.12652)) · and generically for multi-teacher consolidation in most 2026 production reports.\n  - **COPSD** = *Crosslingual* OPSD ([2605.09548](https://arxiv.org/abs/2605.09548)) · *Constitutional* On-Policy Safe Distillation ([2606.03089](https://arxiv.org/abs/2606.03089)).\n  - **D-OPSD / d-OPSD / dOPSD** = step-distilled image diffusion ([2605.05204](https://arxiv.org/abs/2605.05204)) · dLLM self-future ([2606.18195](https://arxiv.org/abs/2606.18195)) · dLLM peek-ahead ([2607.04428](https://arxiv.org/abs/2607.04428)).\n  - **PBSD** = *Preference-Based* Self-Distillation ([2605.05040](https://arxiv.org/abs/2605.05040), OPSD) · *Posterior-Bayesian* Self-Distillation ([2606.09348](https://arxiv.org/abs/2606.09348), Agent) — unrelated mechanisms, same acronym.\n  - **COPD** ([2607.19046](https://arxiv.org/abs/2607.19046), SJTU/Qwen — *contrastive* light-vs-heavy-thinking prefixes) and **CoPD** ([2604.27083](https://arxiv.org/abs/2604.27083), JD.COM — *co-evolving* sibling branches) differ only in capitalisation.\n  - **TrOPD** ([2606.01249](https://arxiv.org/abs/2606.01249), Samsung/Oxford/PKU) and **TOP-D** ([2607.04751](https://arxiv.org/abs/2607.04751), Microsoft/HKUST-GZ) are different papers with different mechanisms — verified separately.\n- **Teacher decay is a recurring, independently-rediscovered phenomenon.** Supervision quality degrades as the student's prefix lengthens, named separately as *Off-Policy Teacher Decay* ([ESR](https://arxiv.org/abs/2605.27028)), *Supervision Fidelity Decay* ([LGR](https://arxiv.org/abs/2605.30833)), *local teachability collapse* ([2605.13643](https://arxiv.org/abs/2605.13643)) and depth-inverted discriminability ([TurnOPD](https://arxiv.org/abs/2607.05804)). [KAT](https://arxiv.org/abs/2606.09471) adds the sharper warning that **low KL is not evidence of health** — the teacher may simply be agreeing with a corrupted prefix. The prefix-truncation family (Fast OPD, Prune-OPD, PG-OPD, ESR, ADWIN, Truncated OPD) is the practical response.\n- **Whether incorrect or correct rollouts carry the signal is genuinely contested.** [ReNIO](https://arxiv.org/abs/2606.23104) and [Apple's diagnostic](https://arxiv.org/abs/2605.10889) find supervision is better aligned on *incorrect* rollouts; [Yonsei's compaction study](https://arxiv.org/abs/2605.06188) finds OPSD mainly compresses *already-correct* traces and barely repairs failures. Both were verified by full reads; the disagreement is real, not a listing error.\n- **TOPD / diffusion-LM entries** — C1 ✓ (the student rolls out its own denoising trajectory with the real inference-time decoder) and C2 ✓ at token level, but \"trajectory\" means a denoising path over masked positions rather than a left-to-right generation, so per-step semantics differ from the autoregressive entries.\n\n\u003c/details\u003e\n\n---\n\n## 🎭 OPD with Black-Box / Outcome-Based Teachers\n\nWhen the teacher is **API-only** (no logits), OPD uses scalar rewards, verbal scores, preferences, or adversarial discriminators — all evaluated on **student rollouts**. Entries that turned out to use static teacher data only (Lion, SuperCorrect, DAIL, SODA) are excluded from this list.\n\n| Resource | 🌟 Stars | Date | Org | Paper Link | Title / Notes |\n| :----: | :----: | :----: |  :----: | :----: | :---- |\n| [ORPO-Distill](https://arxiv.org/abs/2509.25100) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2509.25100) | 2025.09 | Industrial | [arXiv 2509.25100](https://arxiv.org/abs/2509.25100) | ORPO-Distill |\n| [LMOps `/gad`](https://github.com/microsoft/LMOps) | \u003cimg src=\"https://img.shields.io/github/stars/microsoft/LMOps?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2025.11 | Microsoft Research | [arXiv 2511.10643](https://arxiv.org/abs/2511.10643) · [project](https://ytianzhu.github.io/Generative-Adversarial-Distillation/) | GAD — Black-Box OPD |\n| [OVD](https://arxiv.org/abs/2601.21968) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2601.21968) | 2026.01 | HKU / Huawei | [arXiv 2601.21968](https://arxiv.org/abs/2601.21968) | OVD (On-policy Verbal Distillation) — project page `OVD.github.io` 404s |\n| [SPoT](https://github.com/Visual-AI/SPoT) | \u003cimg src=\"https://img.shields.io/github/stars/Visual-AI/SPoT?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | Visual-AI | [arXiv 2603.01683](https://arxiv.org/abs/2603.01683) | **SPOT: Surgical Post-Training** — black-box oracle edits student failures into proximal rollouts |\n| [SODA](https://arxiv.org/pdf/2604.03873) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/pdf/2604.03873) | 2026.04 | Academic | [arXiv 2604.03873](https://arxiv.org/pdf/2604.03873) | SODA — Semi On-Policy Black-Box Distillation |\n| [ROPD](https://github.com/Peregrine123/ROPD_official) | \u003cimg src=\"https://img.shields.io/github/stars/Peregrine123/ROPD_official?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | NUS / USTC / Tencent | [arXiv 2605.07396](https://arxiv.org/abs/2605.07396) | **ROPD — Rubric-based On-Policy Distillation**; induces prompt-specific rubrics from teacher–student contrasts, then scores student rollouts by those rubrics (logit-free / black-box); up to 10× sample efficiency |\n| [PRISM](https://github.com/XIAO4579/PRISM) | \u003cimg src=\"https://img.shields.io/github/stars/XIAO4579/PRISM?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | HKUST(GZ) / Tsinghua / NTU / RUC / USTC / UCAS | [arXiv 2604.28123](https://arxiv.org/abs/2604.28123) | **PRISM — Pre-alignment via Black-Box OPD for Multimodal RL**; an adversarial OPD stage inserted *between SFT and RLVR*, scoring student rollouts with a **Mixture-of-Experts discriminator** (separate perception and reasoning experts, Bradley–Terry loss). +4.4 / +6.0 avg over SFT→RLVR at 4B / 8B |\n| [OmniOPD](https://arxiv.org/abs/2606.01476) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.01476) | 2026.06 | Meta AI | [arXiv 2606.01476](https://arxiv.org/abs/2606.01476) | **OmniOPD — Logit-Free OPD via Speculative Verification**; an entropy-driven scheduler picks uncertain chunks of the student’s rollout, a black-box teacher generates Monte-Carlo continuations, and they are scored by **semantic similarity** rather than logits. **Beats white-box OPD with the same teacher family (69.08 vs 64.16)** |\n| [ExpRL](https://github.com/violetxi/ExpRL) | \u003cimg src=\"https://img.shields.io/github/stars/violetxi/ExpRL?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | Stanford / CMU | [arXiv 2606.17024](https://arxiv.org/abs/2606.17024) | **ExpRL — Exploratory RL for LLM Mid-Training**; an LLM judge scores the student’s own rollouts and prefixes against a *hidden* reference solution under a fixed rubric, giving dense process-level reward before sparse RL ⚠️ see strictness note |\n\n\n\n\u003cdetails\u003e\n\u003csummary\u003e📋 Click to view technical details\u003c/summary\u003e\n\n| Method | Feedback Signal | Data | Granularity | Domain | Notes |\n| :----: | :----: | :----: | :----: | :----: | :---- |\n| ORPO-Distill | Student-Generated Outputs (SGO) + ORPO contrastive | Mixed (student-generated negatives, teacher positives) | Sequence | Cross-architecture | \"Mixed-policy strategy utilizing student-generated outputs\"; NeurIPS 2025 WS. |\n| GAD (Generative Adversarial Distillation) | Discriminator (on-policy reward model) | Student | Sequence | General | A trained discriminator distinguishes student outputs from teacher (e.g. GPT-5) responses; minimax game makes the discriminator co-evolve into an on-policy reward model. Qwen2.5-14B student becomes comparable to GPT-5-Chat on LMSYS. |\n| OVD | Verbal scores (0–9) on student trajectories | Student | Sequence | General | Replaces token-level logit matching with verbal scoring; +25.7% over baselines. |\n| SPOT | Black-box Oracle step edits + BCE reward objective | Student rollouts, Oracle-rectified | Step / sequence | Math reasoning | Minimal edits keep samples proximal to the student distribution, targeting reasoning gains with knowledge retention. |\n| SODA | DPO: teacher responses as preferred vs. base student (q₀) zero-shot responses as rejected | Mixed | Sequence | Cross-architecture | \"Semi on-policy\" paradigm: captures student-specific inferior behaviors from a one-time static snapshot of q₀, eliminating the need for dynamic rollouts or adversarial training. 10× faster and 27% less peak GPU memory than GAD. Outperforms GAD on 15/16 benchmarks |\n| ROPD | Prompt-specific rubric scores (Rubricator + Verifier) | Student | Sequence (rubric-weighted reward) | General | Black-box-compatible alternative to logit OPD: a Rubricator contrasts teacher vs. student responses to induce prompt-specific rubrics, a Verifier scores each student rollout into a weighted pass-rate reward. Needs only teacher-generated responses, no logits. |\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e📝 \u003cb\u003eStrictness notes\u003c/b\u003e\u003c/summary\u003e\n\n- **ExpRL** — ⚠️ The paper presents itself as RL *mid-training*, not distillation, and explicitly benchmarks against (and beats) a true OPSD baseline. It qualifies here only under the black-box reading: an LLM judge scores the student's own rollouts against a hidden reference under a fixed rubric. No teacher logits and no KL term anywhere.\n- **OmniOPD** — the teacher may be fully API-only; supervision is a continuous semantic-similarity score over Monte-Carlo teacher continuations at student-chosen chunks, not a distribution. Notable for **outperforming white-box OPD with the same teacher family**.\n- **Excluded from this section after a full read:** [ZPPO](https://arxiv.org/abs/2606.18216) (NVIDIA) — its own subtitle, *\"Teacher in Prompts, Not Gradients\"*, states the disqualifier exactly: the teacher rewrites the prompt, the student resamples, and training is plain GRPO on binary reward. The teacher never scores a student token. It is the cleanest illustration of where C2 draws the line.\n\n\u003c/details\u003e\n\n---\n\n## ♻️ Self-Distillation with Privileged Context — OPSD\n\n**Same model = teacher = student**, but the teacher is conditioned on something the student doesn't see (verified trace, ground-truth answer, \"be concise\" prefix, longer context, document, …). The gap exists *because of the conditioning*, not weights.\n\nSeveral entries previously listed here turned out on verification to use static teacher data or a fixed self-rewritten dataset rather than student rollouts; those have been excluded. SPIN was reclassified to [Iterative Self-Bootstrapping](#-iterative-self-bootstrapping).\n\n| Resource | 🌟 Stars | Date | Org | Paper Link | Title / Notes |\n| :----: | :----: | :----: |  :----: | :----: | :---- |\n| [OPSD](https://github.com/siyan-zhao/OPSD) | \u003cimg src=\"https://img.shields.io/github/stars/siyan-zhao/OPSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.01 | UCLA / Meta FAIR | [arXiv 2601.18734](https://arxiv.org/abs/2601.18734) · [blog](https://siyan-zhao.github.io/blog/2026/opsd/) | OPSD — Self-Distilled Reasoner |\n| [Self-Distillation](https://github.com/idanshen/Self-Distillation) | \u003cimg src=\"https://img.shields.io/github/stars/idanshen/Self-Distillation?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.01 | MIT / ETH | [arXiv 2601.19897](https://arxiv.org/abs/2601.19897) | SDFT-Continual |\n| [mtp-lm](https://github.com/jwkirchenbauer/mtp-lm) | \u003cimg src=\"https://img.shields.io/github/stars/jwkirchenbauer/mtp-lm?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.02 | UMD / LLNL | [arXiv 2602.06019](https://arxiv.org/abs/2602.06019) | MTP Self-Distill |\n| [LMOps `/opcd`](https://github.com/microsoft/LMOps) | \u003cimg src=\"https://img.shields.io/github/stars/microsoft/LMOps?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.02 | Microsoft Research | [arXiv 2602.12275](https://arxiv.org/abs/2602.12275) | OPCD — On-Policy Context Distillation |\n| [GATES](https://arxiv.org/abs/2602.20574) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2602.20574) | 2026.02 | UMD | [arXiv 2602.20574](https://arxiv.org/abs/2602.20574) | GATES (Self-Distillation under Privileged Context) |\n| [EMPO²](https://agent-lightning.github.io/posts/empo2/) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2602.23008) | 2026.02 | Microsoft Research | [arXiv 2602.23008](https://arxiv.org/abs/2602.23008) · [code](https://github.com/microsoft/agent-lightning/tree/main/contrib/recipes/envs) · [blog](https://agent-lightning.github.io/posts/empo2/) | **EMPO²** — memory-tip-conditioned online self-distillation for exploratory LLM agents (ICLR 2026; cross-listed into Agent) |\n| [CRISP_Reasoning_Compression](https://github.com/HJSang/CRISP_Reasoning_Compression) | \u003cimg src=\"https://img.shields.io/github/stars/HJSang/CRISP_Reasoning_Compression?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | LinkedIn | [arXiv 2603.05433](https://arxiv.org/abs/2603.05433) | OPSDC / CRISP |\n| [LMOps `/oel`](https://github.com/microsoft/LMOps) | \u003cimg src=\"https://img.shields.io/github/stars/microsoft/LMOps?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | Microsoft Research | [arXiv 2603.16856](https://arxiv.org/abs/2603.16856) | OEL — Online Experiential Learning |\n| [self-distillation-analysis](https://github.com/beanie00/self-distillation-analysis) | \u003cimg src=\"https://img.shields.io/github/stars/beanie00/self-distillation-analysis?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.03 | MSR / KAIST / SNU | [arXiv 2603.24472](https://arxiv.org/abs/2603.24472) | **Why Does Self-Distillation (Sometimes) Degrade Reasoning?** — diagnostic study of OPSD failure modes |\n| [ml-ssd](https://github.com/apple/ml-ssd) | \u003cimg src=\"https://img.shields.io/github/stars/apple/ml-ssd?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | Apple MLR | [arXiv 2604.01193](https://arxiv.org/abs/2604.01193) | Apple — Embarrassingly Simple Self-Distillation |\n| [Skill-SD](https://skill-sd.github.io/) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.10674) | 2026.04 | UCAS / CUHK / USTC / vivo AI Lab | [arXiv 2604.10674](https://arxiv.org/abs/2604.10674) | **Skill-SD** — skill-conditioned OPSD for multi-turn LLM agents |\n| [SD-Zero](https://arxiv.org/abs/2604.12002) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.12002) | 2026.04 | Princeton / Toronto / CMU | [arXiv 2604.12002](https://arxiv.org/abs/2604.12002) | **SD-Zero** — Self-Revision turns binary rewards into dense supervision |\n| [π-Play](https://arxiv.org/abs/2604.14054) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.14054) | 2026.04 | CASIA / UCAS / Meituan | [arXiv 2604.14054](https://arxiv.org/abs/2604.14054) | **π-Play** — multi-agent self-play turns the question-construction path into privileged context for OPSD on search agents |\n| [OPSDL](https://arxiv.org/abs/2604.17535) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.17535) | 2026.04 | Baidu | [arXiv 2604.17535](https://arxiv.org/abs/2604.17535) | OPSDL (Long-Context Self-Distillation) |\n| [MSD](https://arxiv.org/abs/2605.02971) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.02971) | 2026.05 | Tongji / Shanghai AI Lab | [arXiv 2605.02971](https://arxiv.org/abs/2605.02971) | **MSD** — multilingual safety OPSD; teacher conditioned on English query translation + CoT instruction; DPSW weights safety-critical tokens |\n| [COPSD](https://github.com/cisnlp/COPSD) | \u003cimg src=\"https://img.shields.io/github/stars/cisnlp/COPSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | LMU Munich / MCML | [arXiv 2605.09548](https://arxiv.org/abs/2605.09548) | **COPSD** — crosslingual OPSD; teacher sees English problem translation + reference solution, student rolls out in low-resource language (17 African languages) |\n| [SGSD](https://github.com/walawalagoose/SGSD) | \u003cimg src=\"https://img.shields.io/github/stars/walawalagoose/SGSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | THU | [arXiv 2605.28791](https://arxiv.org/pdf/2605.28791) | **SGSD** — Skill-Conditional Gated SD |\n| [CODE](https://github.com/CrashBugger/CODE) | \u003cimg src=\"https://img.shields.io/github/stars/CrashBugger/CODE?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | USTC | [arXiv 2605.28303](https://arxiv.org/pdf/2605.28303v1) | **CODE** — OPSD on Knowledge Editing + Casual Editing |\n| [SSOPD](https://arxiv.org/abs/2605.17497) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.17497) | 2026.05 | THU / Beihang | [arXiv 2605.17497](https://arxiv.org/abs/2605.17497) | **SSOPD — Self-Supervised OPSD**; privileged context is the model's *own shortest correct completion* within a GRPO group (no external traces), distilled into prefixes of the longest wrong completion |\n| [RLCSD](https://github.com/THU-BPM/RLCSD) | \u003cimg src=\"https://img.shields.io/github/stars/THU-BPM/RLCSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | THU (BPM) / Alibaba Tongyi | [arXiv 2606.11709](https://arxiv.org/abs/2606.11709) | **RLCSD — Contrastive OPSD**; cancels *privilege-induced style drift* by contrasting the teacher–student gap under a correct hint vs. a wrong hint; verl-based |\n| [d-OPSD](https://github.com/xingzhejun/d-opsd-code) | \u003cimg src=\"https://img.shields.io/github/stars/xingzhejun/d-opsd-code?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | THU / TUM / NTU / UT Austin | [arXiv 2606.18195](https://arxiv.org/abs/2606.18195) | **d-OPSD — first OPSD for diffusion LLMs**; self-generated answers as *suffix* conditioning (\"self future-experience\"); step-level (not token-level) divergence aligned to the denoising process |\n| [CaOPD](https://github.com/SalesforceAIResearch/CaOPD) | \u003cimg src=\"https://img.shields.io/github/stars/SalesforceAIResearch/CaOPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.04 | Salesforce AI Research | [arXiv 2604.16830](https://arxiv.org/abs/2604.16830) | **CaOPD — The Illusion of Certainty**; proves privileged conditioning makes the teacher's confidence a *non-identifiable* and upward-biased target, so OPD reliably buys accuracy at the cost of severe overconfidence; fixes it by rewriting confidence targets to the free empirical rollout success rate |\n| [Vision-OPD](https://github.com/VisionOPD/Vision-OPD) | \u003cimg src=\"https://img.shields.io/github/stars/VisionOPD/Vision-OPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | ISCAS / UCAS / Xiaohongshu | [arXiv 2605.18740](https://arxiv.org/abs/2605.18740) | **Vision-OPD** — regional-to-global OPSD for MLLM fine detail; teacher sees a 2×-upscaled evidence crop, student sees the full image. 9B beats Gemini-3.1-Pro on the fine-grained suite using 6.2K *fully synthetic* triplets, no GT labels or verifier |\n| [D-OPSD](https://github.com/vvvvvjdy/D-OPSD) | \u003cimg src=\"https://img.shields.io/github/stars/vvvvvjdy/D-OPSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | HKUST / Z-Image Team Alibaba / UCSD / CUHK | [arXiv 2605.05204](https://arxiv.org/abs/2605.05204) | **D-OPSD** — OPSD for continuously tuning *step-distilled* diffusion models; exploits that LLM/VLM-encoder T2I models inherit in-context ability, so feeding the encoder the target image is free privileged context |\n| [PW-OPSD](https://github.com/SaFo-Lab/PW-OPSD) | \u003cimg src=\"https://img.shields.io/github/stars/SaFo-Lab/PW-OPSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | SaFo Lab / UW–Madison | [arXiv 2605.21606](https://arxiv.org/abs/2605.21606) | **PW-OPSD — When Are Teacher Tokens Reliable?**; a branch-viability diagnostic shows *position* separates genuinely-uncertain from merely-diverse teacher tokens (AUROC 0.83) where every local uncertainty measure fails (≤0.57) |\n| [OPSD Predictive Law](https://github.com/Tufalabs/opsd-predictive-law) | \u003cimg src=\"https://img.shields.io/github/stars/Tufalabs/opsd-predictive-law?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | Tufa Labs, Zürich | [arXiv 2605.30070](https://arxiv.org/abs/2605.30070) | **A Predictive Law for OPSD From World Feedback** — the *pre-training* student–self-teacher accuracy gap linearly predicts final OPSD gain (R² 0.949 / 0.996), so privileged-context designs can be screened before training (ICML RLxF 2026) |\n| [SDSD Diversity](https://arxiv.org/abs/2606.26091) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.26091) | 2026.06 | Mila / Univ. de Montréal / FAIR at Meta | [arXiv 2606.26091](https://arxiv.org/abs/2606.26091) | **OPSD with Sampled Demonstrations Reduces Output Diversity** — proves the SDSD optimum tilts the base policy by *pointwise conditional mutual information* rather than reward, so it amplifies dominant modes; pass@1 up, pass@k flat |\n| [PHF](https://arxiv.org/abs/2606.29340) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2606.29340) | 2026.06 | HKUST(GZ) / NUAA / NUDT | [arXiv 2606.29340](https://arxiv.org/abs/2606.29340) | **PHF — Privileged Hidden Flow**; adds a residual-stream *transition-geometry* channel (direction + Gram-matrix CKA) on top of the standard OPSD output loss, with proven invariance to per-trajectory offsets and rescaling |\n| [Visual-OPSD](https://github.com/TiezMind/Visual-OPSD) | \u003cimg src=\"https://img.shields.io/github/stars/TiezMind/Visual-OPSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.06 | XJTU (MOE KLINNS) / SYSU | [arXiv 2606.18974](https://arxiv.org/abs/2606.18974) | **Visual-OPSD** — shows a unified multimodal model's rendered \"visual thoughts\" matter as a *generation pathway*, not as pixels, then distills that pathway into a text-only student: **+3.40 pp over its own teacher at 14.3× speedup** |\n| [Denser ≠ Better](https://github.com/Moenupa/SDPO-CL) | \u003cimg src=\"https://img.shields.io/github/stars/Moenupa/SDPO-CL?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.07 | HKISI CAS / CASIA / UCAS / NJUST | [arXiv 2607.01763](https://arxiv.org/abs/2607.01763) | **Denser ≠ Better: Limits of OPSD for Continual Post-Training** — SDPO specialises harder than GRPO but forgets far more (ToolUse −80% vs GRPO +18%); dense self-distillation is a specialisation accelerator, not a continual-learning stabiliser |\n| [Purified OPSD](https://arxiv.org/abs/2607.02234) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.02234) | 2026.07 | ZJU / Tongyi Lab Alibaba / HUST / Jilin | [arXiv 2607.02234](https://arxiv.org/abs/2607.02234) | **Purified OPSD — Without Losing How to Think**; decomposes the teacher update via a *reference-only* teacher and finds the reference-memorisation component dominates while the useful component is **actively opposed** (cos ≈ −0.95); re-anchors on a PMI target |\n| [DemoPSD](https://arxiv.org/abs/2607.02502) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.02502) | 2026.07 | CityU HK / Tsinghua / SIAT / CUHK-SZ | [arXiv 2607.02502](https://arxiv.org/abs/2607.02502) | **DemoPSD — Disagreement-Modulated Policy Self-Distillation**; attenuates *privileged-information leakage* by pulling toward a reverse-KL barycenter whose interpolation weight is driven by per-token teacher–student disagreement |\n| [Rethinking OPSD](https://github.com/princeton-pli/rethinking-opsd-for-thinking-models) | \u003cimg src=\"https://img.shields.io/github/stars/princeton-pli/rethinking-opsd-for-thinking-models?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.07 | Princeton PLI (Kaur, Ri, He, Fowl, Arora) | [arXiv 2607.05184](https://arxiv.org/abs/2607.05184) | **Rethinking OPSD for Thinking Models** — privileged self-distillation **degrades all five thinking models tested** (up to −17% rel.), while the same students *improve* under unprivileged OPD; mechanism is **fork suppression** of self-correction cues |\n| [GeoSD](https://arxiv.org/abs/2607.06855) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.06855) | 2026.07 | ILLC Amsterdam / ILCC Edinburgh (Jukić \u0026 Titov) | [arXiv 2607.06855](https://arxiv.org/abs/2607.06855) | **GeoSD — Geometric Self-Distillation**; replaces KL with squared **Hellinger** distance so the pull vanishes as the student withdraws support, plus a Fisher–Rao proximal term and K-FAC natural gradient. +5.7–8.6 OOD where FKL/RKL *lose* up to 8.1 / 7.5 points on the worst model family |\n| [ShortOPD](https://arxiv.org/abs/2607.13124) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.13124) | 2026.07 | ByteDance / ISCAS / UCAS | [arXiv 2607.13124](https://arxiv.org/abs/2607.13124) | **ShortOPD — Short-to-Long OPD for pruned LLMs** (self-teacher = the model's own *pre-compression* checkpoint, hence OPSD rather than an external teacher); pruning *demotes* rather than erases correct trajectories (pass@k recovers), so recovery is an OPD problem; adds a rollout-budget controller driven by suffix-repetition EMAs |\n| [AD-OPSD](https://arxiv.org/abs/2607.10805) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.10805) | 2026.07 | Beihang / HK PolyU / Alibaba | [arXiv 2607.10805](https://arxiv.org/abs/2607.10805) | **Diagnosing and Mitigating Thinking Collapse in OPSD**; OPSD cuts epistemic-token density 10.3→7.9 per 1k; entropy masking restores density but not accuracy (an \"optimization deadlock\"), fixed by an asymmetric teacher-unreliability interpolation |\n| [dOPSD](https://github.com/tuandattt/dOPSD) | \u003cimg src=\"https://img.shields.io/github/stars/tuandattt/dOPSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.07 | NUS | [arXiv 2607.04428](https://arxiv.org/abs/2607.04428) | **dOPSD — OPSD for diffusion LMs**; privilege comes from *later, more-decoded steps of the student's own denoising trajectory* (\"peek-ahead\"), needing no reference solution. Only method beating base on all four benchmarks where SFT and GRPO both degrade |\n| [PromptSD](https://arxiv.org/abs/2607.18293) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2607.18293) | 2026.07 | UW–Madison / JHU / NTU | [arXiv 2607.18293](https://arxiv.org/abs/2607.18293) | **One Student, Many Teachers** — formalises three teacher desiderata (thinking-pattern consistency, absorbable gap, **no post-hoc rationalization**) and shows every existing teacher family violates one; makes the teacher a *learned soft prompt* over the frozen student backbone |\n| [PAINT](https://arxiv.org/abs/2604.26573) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2604.26573) | 2026.04 | Tsinghua / Beihang | [arXiv 2604.26573](https://arxiv.org/abs/2604.26573) | **PAINT — Partial-Solution Adaptive Interpolated Training**; teacher sees a *suffix-masked* verified solution, with overlap-adaptive masking + entropy-gated logit interpolation. Qwen3-8B +2.1 over prior OPSD ⚠️ promised repo 404 at listing |\n| [PBSD](https://arxiv.org/abs/2605.05040) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.05040) | 2026.05 | Penn State / TikTok | [arXiv 2605.05040](https://arxiv.org/abs/2605.05040) | **Preference-Based Self-Distillation — Beyond KL Matching**; replaces KL matching with a reward-regularized Bradley–Terry preference loss between the privileged teacher’s positive and the student’s own on-policy negative |\n| [vOPD](https://github.com/holi-lab/vOPD) | \u003cimg src=\"https://img.shields.io/github/stars/holi-lab/vOPD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | Seoul National University | [arXiv 2605.07865](https://arxiv.org/abs/2605.07865) | **KL for a KL — OPD with Control Variate Baseline**; derives the OPD value function in closed form as the per-token negative reverse KL (already computed in the forward pass) and subtracts it as an **unbiased variance-reducing baseline**. −57.7% wall-clock vs full-vocab KL |\n| [OPHSD](https://github.com/zzy1127/OPHSD-On-Policy-Harness-Self-Distillation) | \u003cimg src=\"https://img.shields.io/github/stars/zzy1127/OPHSD-On-Policy-Harness-Self-Distillation?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | PKU | [arXiv 2605.08741](https://arxiv.org/abs/2605.08741) | **Training with Harnesses — On-Policy Harness Self-Distillation**; the privileged context is an *inference-time programmatic harness* (draft-verify, plan-solve with oracle) rather than a text hint. +10.83% over OPSD on HMMT25 |\n| [AntiSD](https://github.com/FloyedShen/AntiSD) | \u003cimg src=\"https://img.shields.io/github/stars/FloyedShen/AntiSD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | Xiaohongshu / CASIA | [arXiv 2605.11609](https://arxiv.org/abs/2605.11609) | **Anti-Self-Distillation via Pointwise Mutual Information**; deliberately **ascends** JSD (sign-reversed vs normal self-distillation) under an entropy gate. Reaches GRPO’s peak in 2–10× fewer steps, up to +11.5 avg |\n| [CREDIT](https://arxiv.org/abs/2605.11613) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.11613) | 2026.05 | Xiaohongshu / CASIA | [arXiv 2605.11613](https://arxiv.org/abs/2605.11613) | **From Generic Correlation to Input-Specific Credit in OPSD**; shows the OPSD token reward is a Bayesian-filtering PMI increment, then subtracts an input-*generic* baseline averaged over mismatched contrastive inputs |\n| [Adaptive Teacher Exposure](https://arxiv.org/abs/2605.11458) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.11458) | 2026.05 | ByteDance Douyin | [arXiv 2605.11458](https://arxiv.org/abs/2605.11458) | **Adaptive Teacher Exposure**; reveals only a *learned fraction* of the reference CoT to the teacher, with the exposure ratio sampled from a Beta-policy controller trained by REINFORCE on learning progress |\n| [OGLS-SD](https://arxiv.org/abs/2605.12400) | [![Paper](https://img.shields.io/badge/📄-paper-845C40?style=for-the-badge)](https://arxiv.org/abs/2605.12400) | 2026.05 | UNC Chapel Hill / NVIDIA | [arXiv 2605.12400](https://arxiv.org/abs/2605.12400) | **Outcome-Guided Logit Steering**; averages teacher logits separately over verified-correct and incorrect rollout pools and *contrasts* them into a steered target, applied only to incorrect rollouts |\n| [Multi-Rollout OPD](https://github.com/viviable/mopd_code) | \u003cimg src=\"https://img.shields.io/github/stars/viviable/mopd_code?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | Microsoft / CMU / Purdue | [arXiv 2605.12652](https://arxiv.org/abs/2605.12652) | **Multi-Rollout OPD via Peer Successes and Failures**; conditions the self-teacher on **peer rollouts from the same group** — successes as demonstrations, failures as contrastive evidence. AIME25 mean@8 **25.41 vs SDPO 7.81** |\n| [RESD](https://github.com/horizon-llm/RESD) | \u003cimg src=\"https://img.shields.io/github/stars/horizon-llm/RESD?style=for-the-badge\u0026logo=github\u0026logoColor=white\u0026labelColor=181717\u0026color=ffd700\" alt=\"Stars\"\u003e | 2026.05 | UC San Diego / Amazon / Georgia Tech | [arXiv 2605.12741](https://arxiv.org/abs/2605.12741) | **Reflection-Enhanced Self-Distillation**; enriches the teacher context with a r","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/thinkwee%2Fawesomeopd/projects"}