{"id":87984,"url":"https://github.com/fscdc/Awesome-Efficient-Reasoning-Models","name":"Awesome-Efficient-Reasoning-Models","description":"[TMLR 2025] Efficient Reasoning Models: A Survey","projects_count":372,"last_synced_at":"2026-07-17T08:00:36.345Z","repository":{"id":288026485,"uuid":"964653555","full_name":"fscdc/Awesome-Efficient-Reasoning-Models","owner":"fscdc","description":"[TMLR 2025] Efficient Reasoning Models: A Survey","archived":false,"fork":false,"pushed_at":"2026-06-26T02:29:32.000Z","size":43752,"stargazers_count":314,"open_issues_count":0,"forks_count":23,"subscribers_count":9,"default_branch":"master","last_synced_at":"2026-06-28T22:05:38.481Z","etag":null,"topics":["chain-of-thought","compression","efficient-reasoning"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2504.10903","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/fscdc.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-04-11T15:04:28.000Z","updated_at":"2026-06-26T08:05:43.000Z","dependencies_parsed_at":"2025-04-29T12:48:14.760Z","dependency_job_id":"1c218c04-4745-41ab-8a05-ac9c5339b6c8","html_url":"https://github.com/fscdc/Awesome-Efficient-Reasoning-Models","commit_stats":null,"previous_names":["fscdc/awesome-efficient-reasoning-models"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/fscdc/Awesome-Efficient-Reasoning-Models","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fscdc%2FAwesome-Efficient-Reasoning-Models","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fscdc%2FAwesome-Efficient-Reasoning-Models/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fscdc%2FAwesome-Efficient-Reasoning-Models/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fscdc%2FAwesome-Efficient-Reasoning-Models/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/fscdc","download_url":"https://codeload.github.com/fscdc/Awesome-Efficient-Reasoning-Models/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fscdc%2FAwesome-Efficient-Reasoning-Models/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35573737,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-07-17T02:00:06.162Z","response_time":116,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2025-04-24T05:03:30.728Z","updated_at":"2026-07-17T08:00:36.349Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Full list","Updates"],"sub_categories":["Evaluation and Benchmarks","Build SLM with Strong Reasoning Ability","Let Decoding More Efficient","Make Long CoT Short","Background Papers","Competition","Efficient Multimodal Reasoning","Efficient Agentic Reasoning"],"readme":"\u003cdiv align=\"center\"\u003e\n\n  \u003ch2\u003e\u003cb\u003e [TMLR 2025] Efficient Reasoning Models: A Survey \u003c/b\u003e\u003c/h2\u003e\n  \u003ch4\u003e An overview of research in efficient reasoning models\u003c/h4\u003e\n\n\u003c/div\u003e\n\n\n\u003cdiv align=\"center\"\u003e\n\n![](https://img.shields.io/github/stars/fscdc/Awesome-Efficient-Reasoning-Models?color=yellow)\n![](https://img.shields.io/github/forks/fscdc/Awesome-Efficient-Reasoning-Models?color=lightblue)\n![](https://img.shields.io/github/last-commit/fscdc/Awesome-Efficient-Reasoning-Models?color=green)\n![](https://img.shields.io/badge/PRs-Welcome-blue)\n\u003ca href=\"https://arxiv.org/abs/2504.10903\" target=\"_blank\"\u003e\u003cimg src=\"https://img.shields.io/badge/arXiv-2504.10903-009688.svg\" alt=\"arXiv\"\u003e\u003c/a\u003e\n\n\u003c/div\u003e\n\n\u003cdiv align=\"center\"\u003e\n\n**[\u003ca href=\"https://arxiv.org/abs/2504.10903\"\u003earXiv\u003c/a\u003e]** **[\u003ca href=\"https://x.com/si_feng32704/status/1912378179718901843\"\u003eTwitter\u003c/a\u003e]**\n\n\u003c/div\u003e\n\n\n\nThis repository is for our paper:\n\n\u003e **[Efficient Reasoning Models: A Survey](https://arxiv.org/abs/2504.10903)** \\\n\u003e [Sicheng Feng](https://fscdc.github.io/)\u003csup\u003e1,2\u003c/sup\u003e, [Gongfan Fang](https://fangggf.github.io/)\u003csup\u003e1\u003c/sup\u003e, [Xinyin Ma](https://horseee.github.io/)\u003csup\u003e1\u003c/sup\u003e, [Xinchao Wang](https://sites.google.com/site/sitexinchaowang/)\u003csup\u003e1,*\u003c/sup\u003e \\\n\u003e \u003csup\u003e1\u003c/sup\u003eNational University of Singapore, Singapore \\\n\u003e \u003csup\u003e2\u003c/sup\u003eNankai University, Tianjin, China \\\n\u003e \u003csup\u003e∗\u003c/sup\u003eCorresponding author: xinchao@nus.edu.sg\n\n---\n\u003e\n\u003e 🙋 Please let us know if you find out a mistake or have any suggestions!\n\u003e \n\u003e 🌟 If you find this resource helpful, please consider to star this repository and cite our [research](#citation)!\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"assets/figure2.svg\" width = \"95%\" alt=\"\" align=center /\u003e\n\u003c/p\u003e\n\n\n## Updates\n\n- 2025-10-30: 🚀 We open a new section about [efficient agentic reasoning](#efficient-agentic-reasoning)! Welcome merge your papers!\n- 2025-09-16: 🎉 Our survey has been accepted by TMLR!\n- 2025-05-24: 🚀 We present a fine-grained visual reasoning benchmark - [ReasonMap](https://fscdc.github.io/Reason-Map/)!\n- 2025-05-18: 🎉 We open a new section about [efficient multimodal reasoning methods](#efficient-multimodal-reasoning)!\n- 2025-05-16: 🎉 Two-month milestone! Special thanks to [VainF](https://github.com/VainF), [horseee](https://github.com/horseee), [CHEN1594](https://github.com/CHEN1594), [ZhenyuSun-Walker](https://github.com/ZhenyuSun-Walker), [xianzuwu](https://github.com/xianzuwu)!\n- 2025-04-16: 📝 The survey is now available on [arXiv](https://arxiv.org/abs/2504.10903)!\n- 2025-04-11: 📚 The full paper list is now available, and our survey is coming soon!\n- 2025-03-16: 🚀 Efficient Reasoning Repo launched!\n\n\n## Full list\n\n\n\u003e **Contributions**\n\u003e\n\u003e If you want to add your paper or update details like conference info or code URLs, please submit a pull request. You can generate the necessary markdown for each paper by filling out `generate_item.py` and running `python generate_item.py`. We greatly appreciate your contributions. Alternatively, you can email me ([Gmail](fscnkucs@gmail.com)) the links to your paper and code, and I will add your paper to the list as soon as possible.\n\n---\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"assets/taxonomy.png\" width = \"95%\" alt=\"\" align=center /\u003e\n\u003c/p\u003e\n\n### Quick Links\n  - [Make Long CoT Short](#Make-Long-CoT-Short)\n    - [SFT-based Methods](#SFT-based-Methods)\n    - [RL-based Methods](#RL-based-Methods)\n    - [Prompt-driven Methods](#Prompt-driven-Methods)\n    - [Latent Reasoning](#Latent-Reasoning)\n  - [Build SLM with Strong Reasoning Ability](#Build-SLM-with-Strong-Reasoning-Ability)\n    - [Distillation](#Distillation)\n    - [Quantization and Pruning](#Quantization-and-Pruning)\n    - [RL+SLM Methods](#rlslm-methods)\n  - [Let Decoding More Efficient](#Let-Decoding-More-Efficient)\n    - [Efficient TTS](#Efficient-TTS)\n    - [Other Optimal Methods](#Other-Optimal-Methods)\n  - [Efficient Multimodal Reasoning](#efficient-agentic-reasoning)\n  - [Efficient Agentic Reasoning](#Efficient-Agentic-Reasoning)\n  - [Evaluation and Benchmarks](#Evaluation-and-Benchmarks)\n  - [Background Papers](#Background-Papers)\n  - [Competition](#Competition)\n\n\n\n\n\n### Make Long CoT Short\n\n#### SFT-based Methods\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/horseee/CoT-Valve.svg?style=social\u0026label=Star)](https://github.com/horseee/CoT-Valve)\u003cbr\u003e[CoT-Valve: Length-Compressible Chain-of-Thought Tuning](https://arxiv.org/abs/2502.09601) \u003cbr\u003e Xinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang, Xinchao Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/cot_valve.png\"\u003e |[Github](https://github.com/horseee/CoT-Valve) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.09601)|[//]: #03/16\n|[Flexible Realignment of Language Models](https://arxiv.org/abs/2506.12704) \u003cbr\u003e Wenhong Zhu, Ruobing Xie, Weinan Zhang, Rui Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/inra.png\"\u003e |[Paper](https://arxiv.org/abs/2506.12704)| [//]: #10/30\n|[From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement](https://www.arxiv.org/abs/2509.22144) \u003cbr\u003e Jianzhi Yan, Le Liu, Youcheng Pan, Shiwei Chen, Zike Yuan, Yang Xiang, Buzhou Tang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.22144v1/x2.png\"\u003e |[Paper](https://www.arxiv.org/abs/2509.22144)| [//]: #10/19\n|[![Star](https://img.shields.io/github/stars/LWL-cpu/Question-Free-Fine-Tuning.svg?style=social\u0026label=Star)](https://github.com/LWL-cpu/Question-Free-Fine-Tuning)\u003cbr\u003e[QFFT, Question-Free Fine-Tuning for Adaptive Reasoning](https://arxiv.org/abs/2506.12860) \u003cbr\u003e Wanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin, Ke Ji, Wenyu Chen, Yan Xu, Yasheng Wang, Lifeng Shang, Benyou Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/qfft.png\"\u003e |[Github](https://github.com/LWL-cpu/Question-Free-Fine-Tuning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2506.12860)| [//]: #06/24\n|[OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation](https://arxiv.org/abs/2506.02397) \u003cbr\u003e Shengjia Zhang, Junjie Wu, Jiawei Chen, Changwang Zhang, Xingyu Lou, Wangchunshu Zhou, Sheng Zhou, Can Wang, Jun Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/othink.png\"\u003e |[Paper](https://arxiv.org/abs/2506.02397)| [//]: #06/13\n|[Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting](https://arxiv.org/abs/2505.19716) \u003cbr\u003e Yifan Wu, Jingze Shi, Bingheng Wu, Jiayi Zhang, Xiaotian Lin, Nan Tang, Yuyu Luo |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/dap.png\"\u003e |[Paper](https://arxiv.org/abs/2505.19716)| [//]: #06/13\n|[Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition](https://arxiv.org/abs/2505.19788) \u003cbr\u003e Zihao Zeng, Xuyao Huang, Boxiu Li, Hao Zhang, Zhijie Deng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.19788v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.19788)| [//]: #06/13\n|[Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN](https://arxiv.org/abs/2505.17153) \u003cbr\u003e Yao Xu, Mingyu Xu, Fangyu Lei, Wangtao Sun, Xiangrong Zeng, Bingning Wang, Guang Liu, Shizhu He, Jun Zhao, Kang Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.17153v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2505.17153)| [//]: #06/11\n|[![Star](https://img.shields.io/github/stars/YuxuanJiang1/DRP.svg?style=social\u0026label=Star)](https://github.com/YuxuanJiang1/DRP)\u003cbr\u003e[DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models](https://arxiv.org/abs/2505.13975) \u003cbr\u003e Yuxuan Jiang, Dawei Li, Frank Ferraro |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/YuxuanJiang1/DRP/blob/main/resources/overview.png\"\u003e |[Github](https://github.com/YuxuanJiang1/DRP) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.13975)| [//]: #05/26\n  |[![Star](https://img.shields.io/github/stars/czg1225/VeriThinker.svg?style=social\u0026label=Star)](https://github.com/czg1225/VeriThinker)\u003cbr\u003e[VeriThinker: Learning to Verify Makes Reasoning Model Efficient](https://arxiv.org/pdf/2505.17941) \u003cbr\u003e Zigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu, Xinchao Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/verithinker.png\"\u003e |[Github](https://github.com/czg1225/VeriThinker) \u003cbr\u003e [Paper](https://arxiv.org/pdf/2505.17941)|[//]: #05/23\n|[Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning](https://arxiv.org/abs/2505.14582) \u003cbr\u003e Shangziqi Zhao, Jiahao Yuan, Guisong Yang, Usman Naseem |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14582v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14582)| [//]: #05/26\n|[![Star](https://img.shields.io/github/stars/ZJU-REAL/Self-Braking-Tuning.svg?style=social\u0026label=Star)](https://github.com/ZJU-REAL/Self-Braking-Tuning)\u003cbr\u003e[Let LLMs Break Free from Overthinking via Self-Braking Tuning](https://arxiv.org/abs/2505.14604) \u003cbr\u003e Haoran Zhao, Yuchen Yan, Yongliang Shen, Haolei Xu, Wenqi Zhang, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, Yueting Zhuang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14604v2/x1.png\"\u003e |[Github](https://github.com/ZJU-REAL/Self-Braking-Tuning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.14604)| [//]: #05/24\n|[![Star](https://img.shields.io/github/stars/w-yibo/R1-Compress.svg?style=social\u0026label=Star)](https://github.com/w-yibo/R1-Compress)\u003cbr\u003e[R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search](https://arxiv.org/abs/2505.16838) \u003cbr\u003e Yibo Wang, Li Shen, Huanjin Yao, Tiansheng Huang, Rui Liu, Naiqiang Tan, Jiaxing Huang, Kai Zhang, Dacheng Tao |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/r1-compress.png\"\u003e |[Github](https://github.com/w-yibo/R1-Compress) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.16838)| [//]: #05/24\n|[Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought](https://arxiv.org/abs/2505.15431) \u003cbr\u003e Hunyuan team |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/hunyuan.png\"\u003e |[Paper](https://arxiv.org/abs/2505.15431)| [//]: #05/22\n|[![Star](https://img.shields.io/github/stars/ZJU-REAL/Self-Braking-Tuning.svg?style=social\u0026label=Star)](https://github.com/ZJU-REAL/Self-Braking-Tuning)\u003cbr\u003e[Let LLMs Break Free from Overthinking via Self-Braking Tuning](https://arxiv.org/abs/2505.14604) \u003cbr\u003e Haoran Zhao, Yuchen Yan, Yongliang Shen, Haolei Xu, Wenqi Zhang, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, Yueting Zhuang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14604v1/x2.png\"\u003e |[Github](https://github.com/ZJU-REAL/Self-Braking-Tuning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.14604)| [//]: #05/22\n|[![Star](https://img.shields.io/github/stars/ZGCA-AI4Edu/LS-Mixture.svg?style=social\u0026label=Star)](https://github.com/ZGCA-AI4Edu/LS-Mixture)\u003cbr\u003e[Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models](https://arxiv.org/abs/2505.03469) \u003cbr\u003e Bin Yu, Hang Yuan, Yuliang Wei, Bailing Wang, Weizhen Qi, Kai Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/mix-sft.png\"\u003e |[Github](https://github.com/ZGCA-AI4Edu/LS-Mixture) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.03469)| [//]: #05/17\n|[Llama-Nemotron: Efficient Reasoning Models](https://arxiv.org/abs/2505.00949) \u003cbr\u003e NVIDIA |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.00949v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.00949)| [//]: #05/05\n|[![Publish](https://img.shields.io/badge/Conference-AAAI_2025-blue)]()\u003cbr\u003e[C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness](https://arxiv.org/abs/2412.11664) \u003cbr\u003e Yu Kang, Xianghui Sun, Liangyu Chen, Wei Zou |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/co3t.png\"\u003e |[Paper](https://arxiv.org/abs/2412.11664)|[//]: #03/16\n|[![Star](https://img.shields.io/github/stars/tengxiaoliu/LM_skip.svg?style=social\u0026label=Star)](https://github.com/tengxiaoliu/LM_skip) [![Publish](https://img.shields.io/badge/Conference-NeurIPS_2024-blue)]()\u003cbr\u003e[Can Language Models Learn to Skip Steps?](https://arxiv.org/abs/2411.01855) \u003cbr\u003e Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang, Xipeng Qiu, Zheng Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/skip_step.png\"\u003e |[Github](https://github.com/tengxiaoliu/LM_skip) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.01855)|[//]: #03/16\n|[Distilling System 2 into System 1](https://arxiv.org/abs/2407.06023) \u003cbr\u003e Ping Yu, Jing Xu, Jason Weston, Ilia Kulikov |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/distill_sys1_sys2.png\"\u003e |[Paper](https://arxiv.org/abs/2407.06023)|[//]: #03/16\n|[![Star](https://img.shields.io/github/stars/hemingkx/TokenSkip.svg?style=social\u0026label=Star)](https://github.com/hemingkx/TokenSkip)\u003cbr\u003e[TokenSkip: Controllable Chain-of-Thought Compression in LLMs](https://arxiv.org/abs/2502.12067) \u003cbr\u003e Heming Xia, Yongqi Li, Chak Tou Leong, Wenjie Wang, Wenjie Li |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/TokenSkip.png\"\u003e |[Github](https://github.com/hemingkx/TokenSkip) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.12067)|[//]: #03/20\n|[Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models](https://arxiv.org/abs/2502.13260) \u003cbr\u003e Yingqian Cui, Pengfei He, Jingying Zeng, Hui Liu, Xianfeng Tang, Zhenwei Dai, Yan Han, Chen Luo, Jing Huang, Zhen Li, Suhang Wang, Yue Xing, Jiliang Tang, Qi He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.13260v1/extracted/6214965/pics/merge.png\"\u003e |[Paper](https://arxiv.org/abs/2502.13260)| [//]: #04/08\n|[Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning](https://arxiv.org/abs/2502.18080) \u003cbr\u003e Wenkai Yang, Shuming Ma, Yankai Lin, Furu Wei |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.18080v1/x10.png\"\u003e |[Paper](https://arxiv.org/abs/2502.18080)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/TergelMunkhbat/concise-reasoning.svg?style=social\u0026label=Star)](https://github.com/TergelMunkhbat/concise-reasoning)\u003cbr\u003e[Self-Training Elicits Concise Reasoning in Large Language Models](https://arxiv.org/abs/2502.20122) \u003cbr\u003e Tergel Munkhbat, Namgyu Ho, Seo Hyun Kim, Yongjin Yang, Yujin Kim, Se-Young Yun |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.20122v2/x1.png\"\u003e |[Github](https://github.com/TergelMunkhbat/concise-reasoning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.20122)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/GeniusHTX/TALE.svg?style=social\u0026label=Star)](https://github.com/GeniusHTX/TALE)\u003cbr\u003e[Token-Budget-Aware LLM Reasoning](https://arxiv.org/abs/2412.18547) \u003cbr\u003e Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.18547v4/x10.png\"\u003e |[Github](https://github.com/GeniusHTX/TALE) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.18547)| [//]: #04/08\n\n\n#### RL-based Methods\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/VainF/Thinkless.svg?style=social\u0026label=Star)](https://github.com/VainF/Thinkless)\u003cbr\u003e[Thinkless: LLM Learns When to Think](https://arxiv.org/abs/2505.13379) \u003cbr\u003e Gongfan Fang, Xinyin Ma, Xinchao Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.13379v1/x1.png\"\u003e |[Github](https://github.com/VainF/Thinkless) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.13379)| [//]: #05/20\n|[Rethinking Thinking Tokens: LLMs as Improvement Operators](https://arxiv.org/abs/2510.01123) \u003cbr\u003e Lovish Madaan, Aniket Didolkar, Suchin Gururangan, John Quan, Ruan Silva, Ruslan Salakhutdinov, Manzil Zaheer, Sanjeev Arora, Anirudh Goyal |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2510.01123v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2510.01123)| [//]: #10/30\n|[Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners](https://arxiv.org/abs/2509.26226) \u003cbr\u003e Xin Xu, Cliveb AI, Kai Yang, Tianhao Chen, Yang Wang, Saiyong Yang, Can Yang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.26226v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2509.26226)| [//]: #10/30\n|[SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression](https://arxiv.org/abs/2509.25176) \u003cbr\u003e Haoming Wen, Yushi Bai, Juanzi Li, Jie Tang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.25176v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2509.25176)| [//]: #10/30\n|[Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking](https://arxiv.org/abs/2509.23392) \u003cbr\u003e Jinyi Han, Ying Huang, Ying Liao, Zishang Jiang, Xikun Lu, Haiquan Zhao, Xinyi Wang, Guanghao Zhou, Sihang Jiang, Jiaqing Liang, Weikang Zhou, Zeye Sun, Fei Yu, Yanghua Xiao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.23392v2/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2509.23392)| [//]: #10/19\n|[Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models](https://arxiv.org/abs/2510.03805) \u003cbr\u003e Canhui Wu, Qiong Cao, Chang Li, Zhenfang Wang, Chao Xue, Yuwei Fan, Wei Xi, Xiaodong He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2510.03805v1/fig/intro.png\"\u003e |[Paper](https://arxiv.org/abs/2510.03805)| [//]: #10/11\n|[![Star](https://img.shields.io/github/stars/Optimization-AI/DRPO.svg?style=social\u0026label=Star)](https://github.com/Optimization-AI/DRPO)\u003cbr\u003e[DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization](https://arxiv.org/abs/2510.04474) \u003cbr\u003e Gang Li, Yan Chen, Ming Lin, Tianbao Yang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2510.04474v1/x1.png\"\u003e |[Github](https://github.com/Optimization-AI/DRPO) \u003cbr\u003e [Paper](https://arxiv.org/abs/2510.04474)| [//]: #10/11\n|[ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models](https://arxiv.org/abs/2508.18773) \u003cbr\u003e Qianyu He, Siyu Yuan, Xuefeng Li, Mingxuan Wang, Jiangjie Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2508.18773v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2508.18773)| [//]: #09/03\n|[Hawkeye:Efficient Reasoning with Model Collaboration](https://arxiv.org/abs/2504.00424) \u003cbr\u003e Jianshu She, Zhuohao Li, Zhemin Huang, Qi Li, Peiran Xu, Haonan Li, Qirong Ho |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.00424v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2504.00424)| [//]: #09/03\n|[![Star](https://img.shields.io/github/stars/hammoudhasan/curriculum_grpo.svg?style=social\u0026label=Star)](https://github.com/hammoudhasan/curriculum_grpo)\u003cbr\u003e[Train Long, Think Short: Curriculum Learning for Efficient Reasoning](https://arxiv.org/abs/2508.08940) \u003cbr\u003e Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud, Elie Bou-Zeid, Marzyeh Ghassemi, Bernard Ghanem |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2508.08940v1/x1.png\"\u003e |[Github](https://github.com/hammoudhasan/curriculum_grpo) \u003cbr\u003e [Paper](https://arxiv.org/abs/2508.08940)| [//]: #08/26\n|[Compressing Chain-of-Thought in LLMs via Step Entropy](https://arxiv.org/abs/2508.03346) \u003cbr\u003e Zeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen, Zhijian Xu, Yingying Cheng, Fan Zhang, Qiang Xu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2508.03346v1/1st_fig.png\"\u003e |[Paper](https://arxiv.org/abs/2508.03346)| [//]: #08/08\n|[![Star](https://img.shields.io/github/stars/analokmaus/kaggle-aimo2-fast-math-r1.svg?style=social\u0026label=Star)](https://github.com/analokmaus/kaggle-aimo2-fast-math-r1) [![Publish](https://img.shields.io/badge/Conference-ICML_Workshop-blue)]()\u003cbr\u003e[A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning](https://arxiv.org/abs/2507.08267) \u003cbr\u003e Hiroshi Yoshihara, Taiki Yamaguchi, Yuichi Inoue |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2507.08267v1/x1.png\"\u003e |[Github](https://github.com/analokmaus/kaggle-aimo2-fast-math-r1) \u003cbr\u003e [Paper](https://arxiv.org/abs/2507.08267)| [//]: #07/21\n|[![Star](https://img.shields.io/github/stars/RazvanDu/ConciseRL.svg?style=social\u0026label=Star)](https://github.com/RazvanDu/ConciseRL)\u003cbr\u003e[ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models](https://www.arxiv.org/abs/2505.17250) \u003cbr\u003e Razvan-Gabriel Dumitru, Darius Peteleaza, Vikas Yadav, Liangming Pan |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/conciserl.png\"\u003e |[Github](https://github.com/RazvanDu/ConciseRL) \u003cbr\u003e [Paper](https://www.arxiv.org/abs/2505.17250)| [//]: #07/20\n|[![Star](https://img.shields.io/github/stars/zxiangx/LC-R1.svg?style=social\u0026label=Star)](https://github.com/zxiangx/LC-R1)\u003cbr\u003e[Optimizing Length Compression in Large Reasoning Models](https://arxiv.org/abs/2506.14755) \u003cbr\u003e Zhengxiang Cheng, Dongping Chen, Mingyang Fu, Tianyi Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.14755v1/x3.png\"\u003e |[Github](https://github.com/zxiangx/LC-R1) \u003cbr\u003e [Paper](https://arxiv.org/abs/2506.14755)| [//]: #06/24\n|[PATS: Process-Level Adaptive Thinking Mode Switching](https://arxiv.org/abs/2505.19250) \u003cbr\u003e Yi Wang, Junxiao Liu, Shimao Zhang, Jiajun Chen, Shujian Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.19250v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.19250)| [//]: #06/13\n|[Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition](https://arxiv.org/abs/2505.19788) \u003cbr\u003e Zihao Zeng, Xuyao Huang, Boxiu Li, Hao Zhang, Zhijie Deng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.19788v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.19788)| [//]: #06/13\n|[AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting](https://arxiv.org/abs/2505.18822) \u003cbr\u003e Shijue Huang, Hongru Wang, Wanjun Zhong, Zhaochen Su, Jiazhan Feng, Bowen Cao, Yi R. Fung |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.18822v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.18822)| [//]: #06/12\n|[Bingo: Boosting Efficient Reasoning of LLMs via Dynamic and Significance-based Reinforcement Learning](https://arxiv.org/abs/2506.08125) \u003cbr\u003e Hanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou, Haoyu Dong, Xiaojun Ma, Shi Han, Dongmei Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.08125v1/x3.png\"\u003e | [Paper](https://arxiv.org/abs/2506.08125)| [//]: #06/09\n|[ARM: Adaptive Reasoning Model](https://arxiv.org/abs/2505.20258) \u003cbr\u003e Siye Wu, Jian Xie, Yikai Zhang, Aili Chen, Kai Zhang, Yu Su, Yanghua Xiao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.20258v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.20258)| [//]: #06/12\n|[When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning](https://arxiv.org/abs/2505.15400) \u003cbr\u003e Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Haodong Zhao, Hao Li, Jiansong Chen, Ke Zeng, Xunliang Cai |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.15400v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2505.15400)| [//]: #05/23\n|[![Star](https://img.shields.io/github/stars/hkust-nlp/Laser.svg?style=social\u0026label=Star)](https://github.com/hkust-nlp/Laser)\u003cbr\u003e[Learn to Reason Efficiently with Adaptive Length-based Reward Shaping](https://arxiv.org/abs/2505.15612) \u003cbr\u003e Wei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang, Junteng Liu, Yuntian Deng, Yizhe Zhang, Junxian He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.15612v1/x1.png\"\u003e |[Github](https://github.com/hkust-nlp/Laser) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.15612)| [//]: #05/23\n|[Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought](https://arxiv.org/abs/2505.15431) \u003cbr\u003e Hunyuan team |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/hunyuan.png\"\u003e |[Paper](https://arxiv.org/abs/2505.15431)| [//]: #05/22\n|[Think Only When You Need with Large Hybrid-Reasoning Models](https://arxiv.org/abs/2505.14631) \u003cbr\u003e Lingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong, Xingxing Zhang, Tengchao Lv, Lei Cui, Furu Wei |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14631v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14631)| [//]: #05/22\n|[Reward Reasoning Model](https://arxiv.org/abs/2505.14674) \u003cbr\u003e Jiaxin Guo, Zewen Chi, Li Dong, Qingxiu Dong, Xun Wu, Shaohan Huang, Furu Wei |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/rrm.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14674)| [//]: #05/22\n|[![Star](https://img.shields.io/github/stars/THU-KEG/AdaptThink.svg?style=social\u0026label=Star)](https://github.com/THU-KEG/AdaptThink)\u003cbr\u003e[AdaptThink: Reasoning Models Can Learn When to Think](https://arxiv.org/abs/2505.13417) \u003cbr\u003e Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, Juanzi Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.13417v1/x1.png\"\u003e |[Github](https://github.com/THU-KEG/AdaptThink) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.13417)| [//]: #05/20\n|[Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning](https://arxiv.org/abs/2505.11827) \u003cbr\u003e Yansong Ning, Wei Li, Jun Fang, Naiqiang Tan, Hao Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.11827v1/extracted/6447996/Figure/fig3.png\"\u003e |[Paper](https://arxiv.org/abs/2505.11827)| [//]: #05/20\n|[ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving](https://arxiv.org/abs/2505.12717) \u003cbr\u003e Haoyuan Wu, Xueyi Chen, Rui Ming, Jilong Gao, Shoubo Hu, Zhuolun He, Bei Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.12717v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.12717)| [//]: #05/20\n|[AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning](https://arxiv.org/abs/2505.11896) \u003cbr\u003e Chenwei Lou, Zewei Sun, Xinnian Liang, Meng Qu, Wei Shen, Wenqi Wang, Yuntao Li, Qingping Yang, Shuangzhi Wu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.11896v1/extracted/6446095/pareto_optimal_boundary_new_data.png\"\u003e |[Paper](https://arxiv.org/abs/2505.11896)| [//]: #05/20\n|[Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs](https://arxiv.org/abs/2505.10425) \u003cbr\u003e Jingyao Wang, Wenwen Qiang, Zeen Song, Changwen Zheng, Hui Xiong |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.10425v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.10425)| [//]: #05/18\n|[MilChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Remote Sensing](https://arxiv.org/abs/2505.07984) \u003cbr\u003e Aybora Koksal, A. Aydin Alatan |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.07984v1/extracted/6432707/figures/sample_sam.png\"\u003e |[Paper](https://arxiv.org/abs/2505.07984)| [//]: #05/17\n|[Scalable Chain of Thoughts via Elastic Reasoning](https://arxiv.org/abs/2505.05315) \u003cbr\u003e Yuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo, Junnan Li, Caiming Xiong |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.05315v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.05315)| [//]: #05/17\n|[![Star](https://img.shields.io/github/stars/CodeGoat24/UnifiedReward.svg?style=social\u0026label=Star)](https://github.com/CodeGoat24/UnifiedReward)\u003cbr\u003e[Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning](https://arxiv.org/abs/2505.03318) \u003cbr\u003e Yibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang, Qinglin Lu, Cheng Jin, Jiaqi Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/umrf.png\"\u003e |[Github](https://github.com/CodeGoat24/UnifiedReward) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.03318)| [//]: #05/17\n|[Llama-Nemotron: Efficient Reasoning Models](https://arxiv.org/abs/2505.00949) \u003cbr\u003e NVIDIA |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.00949v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.00949)| [//]: #05/05\n|[![Star](https://img.shields.io/github/stars/StarDewXXX/AdaR1.svg?style=social\u0026label=Star)](https://github.com/StarDewXXX/AdaR1)\u003cbr\u003e[AdaR1: From Long-CoT to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization](https://arxiv.org/abs/2504.21659) \u003cbr\u003e Haotian Luo, Haiying He, Yibo Wang, Jinluan Yang, Rui Liu, Naiqiang Tan, Xiaochun Cao, Dacheng Tao, Li Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/AdaR1.png\"\u003e |[Github](https://github.com/StarDewXXX/AdaR1) \u003cbr\u003e [Paper](https://arxiv.org/abs/2504.21659)|[//]: #05/02\n|[![Star](https://img.shields.io/github/stars/StarDewXXX/O1-Pruner.svg?style=social\u0026label=Star)](https://github.com/StarDewXXX/O1-Pruner)\u003cbr\u003e[O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning](https://arxiv.org/abs/2501.12570) \u003cbr\u003e Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, Dacheng Tao |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/o1_pruner.png\"\u003e |[Github](https://github.com/StarDewXXX/O1-Pruner) \u003cbr\u003e [Paper](https://arxiv.org/abs/2501.12570)|[//]: #03/16\n|[Kimi k1.5: Scaling Reinforcement Learning with LLMs](https://arxiv.org/abs/2501.12599) \u003cbr\u003e Kimi Team |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2501.12599v2/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2501.12599)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/eddycmu/demystify-long-cot.svg?style=social\u0026label=Star)](https://github.com/eddycmu/demystify-long-cot)\u003cbr\u003e[Demystifying Long Chain-of-Thought Reasoning in LLMs](https://arxiv.org/abs/2502.03373) \u003cbr\u003e Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, Xiang Yue |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.03373v1/x1.png\"\u003e |[Github](https://github.com/eddycmu/demystify-long-cot) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.03373)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Zanette-Labs/efficient-reasoning.svg?style=social\u0026label=Star)](https://github.com/Zanette-Labs/efficient-reasoning)\u003cbr\u003e[Training Language Models to Reason Efficiently](https://arxiv.org/abs/2502.04463) \u003cbr\u003e Daman Arora, Andrea Zanette |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.04463v2/x3.png\"\u003e |[Github](https://github.com/Zanette-Labs/efficient-reasoning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.04463)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/cmu-l3/l1.svg?style=social\u0026label=Star)](https://github.com/cmu-l3/l1)\u003cbr\u003e[L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning](https://www.arxiv.org/abs/2503.04697) \u003cbr\u003e Pranjal Aggarwal, Sean Welleck |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.04697v1/x2.png\"\u003e |[Github](https://github.com/cmu-l3/l1) \u003cbr\u003e [Paper](https://www.arxiv.org/abs/2503.04697)| [//]: #04/08\n|[DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models](https://arxiv.org/abs/2503.04472) \u003cbr\u003e Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wenjing Zhang, Jiangze Yan, Ning Wang, Kai Wang, Shiguo Lian |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.04472v1/extracted/6254851/DAST.png\"\u003e |[Paper](https://arxiv.org/abs/2503.04472)| [//]: #04/08\n|[Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning](https://arxiv.org/abs/2503.15952) \u003cbr\u003e Chen Li, Nazhou Liu, Kai Yang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/agpo.png\"\u003e |[Paper](https://arxiv.org/abs/2503.15952)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/UCSB-NLP-Chang/ThinkPrune.svg?style=social\u0026label=Star)](https://github.com/UCSB-NLP-Chang/ThinkPrune)\u003cbr\u003e[ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning](https://arxiv.org/abs/2504.01296) \u003cbr\u003e Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.01296v1/x1.png\"\u003e |[Github](https://github.com/UCSB-NLP-Chang/ThinkPrune) \u003cbr\u003e [Paper](https://arxiv.org/abs/2504.01296)| [//]: #04/08\n|[Think When You Need: Self-Adaptive Chain-of-Thought Learning](https://arxiv.org/abs/2504.03234) \u003cbr\u003e Junjie Yang, Ke Lin, Xing Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.03234v1/extracted/6335120/alg_illu.png\"\u003e |[Paper](https://arxiv.org/abs/2504.03234)| [//]: #04/08\n|[The Art of Efficient Reasoning: Data, Reward, and Optimization](https://arxiv.org/pdf/2602.20945) \u003cbr\u003e Taiqiang Wu, Zenan Xu, Bo Zhou, Ngai Wong |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2602.20945v2/x1.png\"\u003e |[Project](https://wutaiqiang.github.io/project/Art) [Weights](https://huggingface.co/collections/taki555/the-art-of-efficient-reasoning)| [//]: #04/08\n\n\n#### Prompt-driven Methods\n\n##### Prompt-guided Efficint Reasoning\n\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt](https://arxiv.org/abs/2505.23480) \u003cbr\u003e Keqin Peng, Liang Ding, Yuanxin Ouyang, Meng Fang, Dacheng Tao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.23480v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.23480)| [//]: #06/11\n|[![Star](https://img.shields.io/github/stars/ZJU-REAL/Self-Braking-Tuning.svg?style=social\u0026label=Star)](https://github.com/ZJU-REAL/Self-Braking-Tuning)\u003cbr\u003e[Let LLMs Break Free from Overthinking via Self-Braking Tuning](https://arxiv.org/abs/2505.14604) \u003cbr\u003e Haoran Zhao, Yuchen Yan, Yongliang Shen, Haolei Xu, Wenqi Zhang, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, Yueting Zhuang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14604v1/x2.png\"\u003e |[Github](https://github.com/ZJU-REAL/Self-Braking-Tuning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.14604)| [//]: #05/22\n|[Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation](https://arxiv.org/abs/2505.03320) \u003cbr\u003e Junyu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.03320v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.03320)| [//]: #05/17\n|[Time's Up! An Empirical Study of LLM Reasoning Ability Under Output Length Constraint](https://arxiv.org/abs/2504.14350) \u003cbr\u003e Yi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu, Xiangyu Li, Hao Wen, Huiwen Zheng, Yan Liang, Yuanchun Li, Yunxin Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/time_up.png\"\u003e |[Paper](https://arxiv.org/abs/2504.14350)| [//]: #04/23\n|[CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Models](https://arxiv.org/abs/2504.13534) \u003cbr\u003e Feiyang Li, Peng Fang, Zhan Shi, Arijit Khan, Fang Wang, Dan Feng, Weihao Wang, Xin Zhang, Yongjian Cui |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.13534v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2504.13534)| [//]: #04/21\n|[Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models](https://arxiv.org/abs/2504.13626) \u003cbr\u003e Yule Liu, Jingyi Zheng, Zhen Sun, Zifan Peng, Wenhan Dong, Zeyang Sha, Shiwen Cui, Weiqiang Wang, Xinlei He |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/thoughtmani.png\"\u003e |[Paper](https://arxiv.org/abs/2504.13626)| [//]: #04/21\n|[![Star](https://img.shields.io/github/stars/GeniusHTX/TALE.svg?style=social\u0026label=Star)](https://github.com/GeniusHTX/TALE)\u003cbr\u003e[Token-Budget-Aware LLM Reasoning](https://arxiv.org/abs/2412.18547) \u003cbr\u003e Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.18547v4/x10.png\"\u003e |[Github](https://github.com/GeniusHTX/TALE) \u003cbr\u003e [Paper](https://arxiv.org/abs/2412.18547)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/matthewrenze/jhu-concise-cot.svg?style=social\u0026label=Star)](https://github.com/matthewrenze/jhu-concise-cot) [![Publish](https://img.shields.io/badge/Conference-FLLM_2024-blue)]()\u003cbr\u003e[The Benefits of a Concise Chain of Thought on Problem-Solving in Large Language Models](https://arxiv.org/abs/2401.05618) \u003cbr\u003e Matthew Renze, Erhan Guven |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2401.05618v3/x1.png\"\u003e |[Github](https://github.com/matthewrenze/jhu-concise-cot) \u003cbr\u003e [Paper](https://arxiv.org/abs/2401.05618)| [//]: #04/08\n|[Break the Chain: Large Language Models Can be Shortcut Reasoners](https://arxiv.org/abs/2406.06580) \u003cbr\u003e Mengru Ding, Hanmeng Liu, Zhizhang Fu, Jian Song, Wenbo Xie, Yue Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2406.06580v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2406.06580)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/sileix/chain-of-draft.svg?style=social\u0026label=Star)](https://github.com/sileix/chain-of-draft)\u003cbr\u003e[Chain of Draft: Thinking Faster by Writing Less](https://arxiv.org/abs/2502.18600) \u003cbr\u003e Silei Xu, Wenhao Xie, Lingxiao Zhao, Pengcheng He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.18600v2/extracted/6244873/plot.png\"\u003e |[Github](https://github.com/sileix/chain-of-draft) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.18600)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/LightChen233/reasoning-boundary.svg?style=social\u0026label=Star)](https://github.com/LightChen233/reasoning-boundary) [![Publish](https://img.shields.io/badge/Conference-NeurIPS_2024-blue)]()\u003cbr\u003e[Unlocking the Capabilities of Thought: A Reasoning Boundary Framework to Quantify and Optimize Chain-of-Thought](https://arxiv.org/abs/2410.05695) \u003cbr\u003e Qiguang Chen, Libo Qin, Jiaqi Wang, Jinxuan Zhou, Wanxiang Che |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.05695v2/x1.png\"\u003e |[Github](https://github.com/LightChen233/reasoning-boundary) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.05695)| [//]: #04/08\n|[How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach](https://arxiv.org/abs/2503.01141) \u003cbr\u003e Ayeong Lee, Ethan Che, Tianyi Peng |\u003cimg src=\"https://arxiv.org/html/2503.01141v2/extracted/6325669/plot/mmlu-pro-legend.png\" width=\"45%\"\u003e \u003cimg src=\"https://arxiv.org/html/2503.01141v2/extracted/6325669/plot/Anthropic/claude-3-5-sonnet-20241022-mmlu-main.png\" width=\"45%\"\u003e |[Paper](https://arxiv.org/abs/2503.01141)| [//]: #04/08\n\n\n\n##### Prompt Attribute-Aware Reasoning Routing\n\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning](https://arxiv.org/abs/2505.15154) \u003cbr\u003e Jinghui Lu, Haiyang Yu, Siliang Xu, Shiwei Ran, Guozhi Tang, Siqi Wang, Bin Shan, Teng Fu, Hao Feng, Jingqun Tang, Han Wang, Can Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.15154v1/x10.png\"\u003e |[Paper](https://arxiv.org/abs/2505.15154)| [//]: #05/26\n|[Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers](https://arxiv.org/abs/2505.12601) \u003cbr\u003e Yang Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.12601v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.12601)| [//]: #05/20\n|[How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach](https://arxiv.org/abs/2503.01141) \u003cbr\u003e Ayeong Lee, Ethan Che, Tianyi Peng |\u003cimg src=\"https://arxiv.org/html/2503.01141v2/extracted/6325669/plot/mmlu-pro-legend.png\" width=\"45%\"\u003e \u003cimg src=\"https://arxiv.org/html/2503.01141v2/extracted/6325669/plot/Anthropic/claude-3-5-sonnet-20241022-mmlu-main.png\" width=\"45%\"\u003e |[Paper](https://arxiv.org/abs/2503.01141)| [//]: #04/08\n| [![Publish](https://img.shields.io/badge/Conference-ICLR_2025-blue)]()\u003cbr\u003e[RouteLLM: Learning to Route LLMs with Preference Data](https://arxiv.org/abs/2406.18665) \u003cbr\u003e Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, Ion Stoica |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2406.18665v4/extracted/6226172/Figs/gsm8k.png\"\u003e |[Paper](https://arxiv.org/abs/2406.18665)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/SimonAytes/SoT.svg?style=social\u0026label=Star)](https://github.com/SimonAytes/SoT)\u003cbr\u003e[Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching](https://arxiv.org/abs/2503.05179) \u003cbr\u003e Simon A. Aytes, Jinheon Baek, Sung Ju Hwang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.05179v1/x1.png\"\u003e |[Github](https://github.com/SimonAytes/SoT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2503.05179)| [//]: #04/08\n|[Learning to Route LLMs with Confidence Tokens](https://arxiv.org/abs/2410.13284) \u003cbr\u003e Yu-Neng Chuang, Helen Zhou, Prathusha Kameswara Sarma, Parikshit Gopalan, John Boccio, Sara Bolouki, Xia Hu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.13284v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2410.13284)| [//]: #04/08\n|[Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization](https://arxiv.org/abs/2502.04428) \u003cbr\u003e Yu-Neng Chuang, Leisheng Yu, Guanchu Wang, Lizhe Zhang, Zirui Liu, Xuanting Cai, Yang Sui, Vladimir Braverman, Xia Hu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.04428v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2502.04428)| [//]: #04/08\n\n###### Blog\n* [Claude 3.7 Sonnet](https://www.anthropic.com/news/claude-3-7-sonnet). Claude team. [[Paper]](https://www.anthropic.com/news/claude-3-7-sonnet)\n\n\n#### Latent Reasoning\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge](https://arxiv.org/abs/2601.08808) \u003cbr\u003e Yao Tang, Li Dong, Yaru Hao, Qingxiu Dong, Furu Wei, Jiatao Gu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://github.com/GMLR-Penn/Multiplex-Thinking/raw/main/figs/teaser.png\"\u003e |[Paper](https://arxiv.org/abs/2601.08808)| [//]: #3/9\n|[SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs](https://arxiv.org/abs/2510.05069) \u003cbr\u003e Dachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan, Leyan Pan, Wenke Lee, Wen Xiao |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/swireasoning.png\"\u003e |[Paper](https://arxiv.org/abs/2510.05069)| [//]: #10/19\n|[LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking](https://arxiv.org/abs/2508.03440) \u003cbr\u003e Chünhung Wu, Jinliang Lu, Zixuan Ren, Gangqiang Hu, Zhi Wu, Dai Dai, Hua Wu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/soft_thinking.png\"\u003e |[Paper](https://arxiv.org/abs/2508.03440)| [//]: #08/09\n|[CTRLS: Chain-of-Thought Reasoning via Latent State-Transition](https://arxiv.org/abs/2507.08182) \u003cbr\u003e Junda Wu, Yuxin Xiong, Xintong Li, Zhengmian Hu, Tong Yu, Rui Wang, Xiang Chen, Jingbo Shang, Julian McAuley |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2507.08182v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2507.08182)| [//]: #07/19\n|[![Star](https://img.shields.io/github/stars/wenquanlu/huginn-latent-cot.svg?style=social\u0026label=Star)](https://github.com/wenquanlu/huginn-latent-cot)\u003cbr\u003e[Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer](https://arxiv.org/abs/2507.02199) \u003cbr\u003e Wenquan Lu, Yuechuan Yang, Kyle Lee, Yanshu Li, Enqi Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2507.02199v1/extracted/6591393/img/lens_arch2.jpg\"\u003e |[Github](https://github.com/wenquanlu/huginn-latent-cot) \u003cbr\u003e [Paper](https://arxiv.org/abs/2507.02199)| [//]: #07/06\n|[![Star](https://img.shields.io/github/stars/UMass-Embodied-AGI/Mirage.svg?style=social\u0026label=Star)](https://github.com/UMass-Embodied-AGI/Mirage)\u003cbr\u003e[Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens](https://arxiv.org/abs/2506.17218) \u003cbr\u003e Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, Chuang Gan |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.17218v1/x1.png\"\u003e |[Github](https://github.com/UMass-Embodied-AGI/Mirage) \u003cbr\u003e [Paper](https://arxiv.org/abs/2506.17218)| [//]: #07/01\n|[![Star](https://img.shields.io/github/stars/whyNLP/PCCoT.svg?style=social\u0026label=Star)](https://github.com/whyNLP/PCCoT)\u003cbr\u003e[Parallel Continuous Chain-of-Thought with Jacobi Iteration](https://arxiv.org/abs/2506.18582) \u003cbr\u003e Haoyi Wu, Zhihao Teng, Kewei Tu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.18582v1/x2.png\"\u003e |[Github](https://github.com/whyNLP/PCCoT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2506.18582)| [//]: #07/01\n|[DART: Distilling Autoregressive Reasoning to Silent Thought](https://arxiv.org/abs/2506.11752) \u003cbr\u003e Nan Jiang, Ziming Wu, De-Chuan Zhan, Fuming Lai, Shaobing Lian |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.11752v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2506.11752)| [//]: #06/22\n|[System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts](https://arxiv.org/abs/2505.18962) \u003cbr\u003e Xiaoqiang Wang, Suyuchen Wang, Yun Zhu, Bang Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.18962v3/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.18962)| [//]: #06/13\n|[Hybrid Latent Reasoning via Reinforcement Learning](https://arxiv.org/abs/2505.18454) \u003cbr\u003e Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang, Zhen Qin, Jinsung Yoon, Lanyu Shang, Jiawei Han, Dong Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.18454v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.18454)| [//]: #06/12\n|[SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought](https://arxiv.org/abs/2505.24181) \u003cbr\u003e Guanghao Li,Wenhao Jiang,Mingfeng Chen,Yan Li,Hao Yu,Shuting Dong,Tao Ren,Ming Tang,Chun Yuan |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.24181v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2505.24181)| [//]: #06/11\n|[Continuous Chain of Thought Enables Parallel Exploration and Reasoning](https://arxiv.org/abs/2505.23648) \u003cbr\u003e Halil Alperen Gozeten,M. Emrullah Ildiz,Xuechen Zhang,Hrayr Harutyunyan,Ankit Singh Rawat,Samet Oymak |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.23648v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.23648)| [//]: #06/11\n|[![Star](https://img.shields.io/github/stars/Rohan-GRH/CoUT.svg?style=social\u0026label=Star)](https://github.com/Rohan-GRH/CoUT)\u003cbr\u003e[Efficient Reasoning via Chain of Unconscious Thought](https://arxiv.org/abs/2505.19756) \u003cbr\u003e Ruihan Gong, Yue Liu, Wenjie Qu, Mingzhe Du, Yufei He, Yingwei Ma, Yulin Chen, Xiang Liu, Yi Wen, Xinfeng Li, Ruidong Wang, Xinzhong Zhu, Bryan Hooi, Jiaheng Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/cout.png\"\u003e |[Github](https://github.com/Rohan-GRH/CoUT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.19756)| [//]: #06/11\n|[![Star](https://img.shields.io/github/stars/EIT-NLP/Awesome-Latent-CoT.svg?style=social\u0026label=Star)](https://github.com/EIT-NLP/Awesome-Latent-CoT)\u003cbr\u003e[Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning](https://arxiv.org/abs/2505.16782) \u003cbr\u003e Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, Xiaoyu Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.16782v1/x1.png\"\u003e |[Github](https://github.com/EIT-NLP/Awesome-Latent-CoT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.16782)| [//]: #05/24\n|[![Star](https://img.shields.io/github/stars/xiaomi-research/colar.svg?style=social\u0026label=Star)](https://github.com/xiaomi-research/colar)\u003cbr\u003e[Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains](https://arxiv.org/abs/2505.16552) \u003cbr\u003e Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, Ruihua Song |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.16552v1/x1.png\"\u003e |[Github](https://github.com/xiaomi-research/colar) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.16552)| [//]: #05/24\n|[![Star](https://img.shields.io/github/stars/eric-ai-lab/Soft-Thinking.svg?style=social\u0026label=Star)](https://github.com/eric-ai-lab/Soft-Thinking)\u003cbr\u003e[Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space](https://arxiv.org/abs/2505.15778) \u003cbr\u003e Zhen Zhang, Xuehai He, Weixiang Yan, Ao Shen, Chenyang Zhao, Shuohang Wang, Yelong Shen, Xin Eric Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.15778v1/x1.png\"\u003e |[Github](https://github.com/eric-ai-lab/Soft-Thinking) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.15778)| [//]: #05/22\n|[Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models](https://arxiv.org/abs/2505.15634) \u003cbr\u003e Zihao Li, Xu Wang, Yuzhe Yang, Ziyu Yao, Haoyi Xiong, Mengnan Du |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.15634v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.15634)| [//]: #05/22\n|[Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space](https://arxiv.org/abs/2505.13308) \u003cbr\u003e Hengli Li, Chenxi Li, Tong Wu, Xuekai Zhu, Yuxuan Wang, Zhaoxin Yu, Eric Hanchen Jiang, Song-Chun Zhu, Zixia Jia, Ying Nian Wu, Zilong Zheng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.13308v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.13308)| [//]: #05/20\n|[Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought](https://arxiv.org/abs/2505.12514) \u003cbr\u003e Hanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao, Stuart Russell, Yuandong Tian |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.12514v1/extracted/6451418/figs/input_format.png\"\u003e |[Paper](https://arxiv.org/abs/2505.12514)| [//]: #05/20\n|[![Star](https://img.shields.io/github/stars/xuyige/SoftCoT.svg?style=social\u0026label=Star)](https://github.com/xuyige/SoftCoT)\u003cbr\u003e[SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning](https://arxiv.org/abs/2505.11484) \u003cbr\u003e Yige Xu, Xu Guo, Zhiwei Zeng, Chunyan Miao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.11484v1/x1.png\"\u003e |[Github](https://github.com/xuyige/SoftCoT) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.11484)| [//]: #05/19\n|[Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models](https://arxiv.org/abs/2504.10615) \u003cbr\u003e Thilo Hagendorff, Sarah Fabi |\u003cimg width=\"1002\" alt=\"image\" src=\"./figures/BCoT.png\"\u003e |[Paper](https://arxiv.org/abs/2504.10615)|[//]: #04/17\n|[Distilling System 2 into System 1](https://arxiv.org/abs/2407.06023) \u003cbr\u003e Ping Yu, Jing Xu, Jason Weston, Ilia Kulikov |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/distill_sys1_sys2.png\"\u003e |[Paper](https://arxiv.org/abs/2407.06023)|[//]: #03/16\n|[![Star](https://img.shields.io/github/stars/da03/implicit_chain_of_thought.svg?style=social\u0026label=Star)](https://github.com/da03/implicit_chain_of_thought/)\u003cbr\u003e[Implicit Chain of Thought Reasoning via Knowledge Distillation](https://arxiv.org/abs/2311.01460) \u003cbr\u003e Yuntian Deng, Kiran Prasad, Roland Fernandez, Paul Smolensky, Vishrav Chaudhary, Stuart Shieber |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/explicit2implicit.png\"\u003e |[Github](https://github.com/da03/implicit_chain_of_thought/) \u003cbr\u003e [Paper](https://arxiv.org/abs/2311.01460)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/HKUNLP/diffusion-of-thoughts.svg?style=social\u0026label=Star)](https://github.com/HKUNLP/diffusion-of-thoughts) [![Publish](https://img.shields.io/badge/Conference-NeurIPS_2024-blue)]()\u003cbr\u003e[Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models](https://arxiv.org/abs/2402.07754) \u003cbr\u003e Jiacheng Ye, Shansan Gong, Liheng Chen, Lin Zheng, Jiahui Gao, Han Shi, Chuan Wu, Xin Jiang, Zhenguo Li, Wei Bi, Lingpeng Kong |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/diffusion_thought.png\"\u003e |[Github](https://github.com/HKUNLP/diffusion-of-thoughts) \u003cbr\u003e [Paper](https://arxiv.org/abs/2402.07754)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/da03/Internalize_CoT_Step_by_Step.svg?style=social\u0026label=Star)](https://github.com/da03/Internalize_CoT_Step_by_Step)\u003cbr\u003e[From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step](https://arxiv.org/abs/2405.14838) \u003cbr\u003e Yuntian Deng, Yejin Choi, Stuart Shieber |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2405.14838v1/extracted/2405.14838v1/training_illustration.png\"\u003e |[Github](https://github.com/da03/Internalize_CoT_Step_by_Step) \u003cbr\u003e [Paper](https://arxiv.org/abs/2405.14838)| [//]: #04/08\n|[Compressed Chain of Thought: Efficient Reasoning Through Dense Representations](https://arxiv.org/abs/2412.13171) \u003cbr\u003e Jeffrey Cheng, Benjamin Van Durme |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.13171v1/extracted/6074157/figures/fig1.png\"\u003e |[Paper](https://arxiv.org/abs/2412.13171)| [//]: #04/08\n|[SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs](https://arxiv.org/abs/2502.12134) \u003cbr\u003e Yige Xu, Xu Guo, Zhiwei Zeng, Chunyan Miao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.12134v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2502.12134)| [//]: #04/08\n| [![Publish](https://img.shields.io/badge/Conference-ICLR_2025-blue)]()\u003cbr\u003e[Reasoning with Latent Thoughts: On the Power of Looped Transformers](https://arxiv.org/abs/2502.17416) \u003cbr\u003e Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.17416v1/extracted/6229618/Media/looping_illustration2.png\"\u003e |[Paper](https://arxiv.org/abs/2502.17416)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/qifanyu/RELAY.svg?style=social\u0026label=Star)](https://github.com/qifanyu/RELAY)\u003cbr\u003e[Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning](https://arxiv.org/abs/2502.08482) \u003cbr\u003e Qifan Yu, Zhenyu He, Sijie Li, Xun Zhou, Jun Zhang, Jingjing Xu, Di He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.08482v1/x1.png\"\u003e |[Github](https://github.com/qifanyu/RELAY) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.08482)| [//]: #04/08\n|[CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation](https://arxiv.org/abs/2502.21074) \u003cbr\u003e Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, Yulan He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.21074v1/extracted/6241542/figures/codi_illustrate12.png\"\u003e |[Paper](https://arxiv.org/abs/2502.21074)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/zjunlp/LightThinker.svg?style=social\u0026label=Star)](https://github.com/zjunlp/LightThinker)\u003cbr\u003e[LightThinker: Thinking Step-by-Step Compression](https://arxiv.org/abs/2502.15589) \u003cbr\u003e Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, Ningyu Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.15589v1/x1.png\"\u003e |[Github](https://github.com/zjunlp/LightThinker) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.15589)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/WANGXinyiLinda/planning_tokens.svg?style=social\u0026label=Star)](https://github.com/WANGXinyiLinda/planning_tokens) [![Publish](https://img.shields.io/badge/Conference-COLM_2024-blue)]()\u003cbr\u003e[Guiding Language Model Reasoning with Planning Tokens](https://arxiv.org/abs/2310.05707) \u003cbr\u003e Xinyi Wang, Lucas Caccia, Oleksiy Ostapenko, Xingdi Yuan, William Yang Wang, Alessandro Sordoni |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2310.05707v4/extracted/5777851/img/overview.png\"\u003e |[Github](https://github.com/WANGXinyiLinda/planning_tokens) \u003cbr\u003e [Paper](https://arxiv.org/abs/2310.05707)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/JacobPfau/fillerTokens.svg?style=social\u0026label=Star)](https://github.com/JacobPfau/fillerTokens) [![Publish](https://img.shields.io/badge/Conference-COLM_2024-blue)]()\u003cbr\u003e[Let's Think Dot by Dot: Hidden Computation in Transformer Language Models](https://arxiv.org/abs/2404.15758) \u003cbr\u003e Jacob Pfau, William Merrill, Samuel R. Bowman |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2404.15758v1/extracted/2404.15758v1/figs/scale_len.png\"\u003e |[Github](https://github.com/JacobPfau/fillerTokens) \u003cbr\u003e [Paper](https://arxiv.org/abs/2404.15758)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/MingyuJ666/Disentangling-Memory-and-Reasoning.svg?style=social\u0026label=Star)](https://github.com/MingyuJ666/Disentangling-Memory-and-Reasoning)\u003cbr\u003e[Disentangling Memory and Reasoning Ability in Large Language Models](https://arxiv.org/abs/2411.13504) \u003cbr\u003e Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, Yongfeng Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.13504v2/x1.png\"\u003e |[Github](https://github.com/MingyuJ666/Disentangling-Memory-and-Reasoning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2411.13504)| [//]: #04/08\n|[Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning](https://arxiv.org/abs/2502.03275) \u003cbr\u003e DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, Qinqing Zheng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.03275v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2502.03275)| [//]: #04/08\n|[Training Large Language Models to Reason in a Continuous Latent Space](https://arxiv.org/abs/2412.06769) \u003cbr\u003e Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2412.06769v2/extracted/6060815/figures/figure_1_meta_3.png\"\u003e |[Paper](https://arxiv.org/abs/2412.06769)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/shawnricecake/Heima.svg?style=social\u0026label=Star)](https://github.com/shawnricecake/Heima)\u003cbr\u003e[Efficient Reasoning with Hidden Thinking](https://arxiv.org/abs/2501.19201) \u003cbr\u003e Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, Jiuxiang Gu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2501.19201v1/x1.png\"\u003e |[Github](https://github.com/shawnricecake/Heima) \u003cbr\u003e [Paper](https://arxiv.org/abs/2501.19201)| [//]: #04/08\n| [![Publish](https://img.shields.io/badge/Conference-ICLR_2024-blue)]()\u003cbr\u003e[Think before you speak: Training Language Models With Pause Tokens](https://arxiv.org/abs/2310.02226) \u003cbr\u003e Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, Vaishnavh Nagarajan |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/pause_token.png\"\u003e |[Paper](https://arxiv.org/abs/2310.02226)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/seal-rg/recurrent-pretraining.svg?style=social\u0026label=Star)](https://github.com/seal-rg/recurrent-pretraining)\u003cbr\u003e[Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach](https://arxiv.org/abs/2502.05171) \u003cbr\u003e Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.05171v2/x2.png\"\u003e |[Github](https://github.com/seal-rg/recurrent-pretraining) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.05171)| [//]: #04/08\n|[Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning](https://arxiv.org/abs/2504.10646) \u003cbr\u003e Saif Punjwani, Larry Heck |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.10646v1/extracted/6355099/cotvswot.png\"\u003e |[Paper](https://arxiv.org/abs/2504.10646)|[//]: #04/16\n|[![Star](https://img.shields.io/github/stars/MobiusDai/LRT.svg?style=social\u0026label=Star)](https://github.com/MobiusDai/LRT) [![Publish](https://img.shields.io/badge/Conference-ICLR-blue)]()\u003cbr\u003e[Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations](https://openreview.net/forum?id=CbK7lYbmv8) \u003cbr\u003e Cong Jiang, Xiaofeng Zhang, Fangzhi Zhu, XiaoWei Chen, Junxiong Zhu, Zheng Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/rellmr.png\"\u003e |[Github](https://github.com/MobiusDai/LRT) \u003cbr\u003e [Paper](https://openreview.net/forum?id=CbK7lYbmv8)| [//]: #06/26\n\n### Build SLM with Strong Reasoning Ability\n\n\n#### Distillation \n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning](https://arxiv.org/abs/2505.24850) \u003cbr\u003e Shuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu, Wei Chu, Yuan Qi |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/redi.png\"\u003e |[Paper](https://arxiv.org/abs/2505.24850)| [//]: #06/13\n|[Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster](https://arxiv.org/abs/2505.18642) \u003cbr\u003e Xiao Chen, Sihang Zhou, Ke Liang, Xiaoyu Sun, Xinwang Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.18642v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.18642)| [//]: #06/11\n|[Llama-Nemotron: Efficient Reasoning Models](https://arxiv.org/abs/2505.00949) \u003cbr\u003e NVIDIA |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.00949v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.00949)| [//]: #05/05\n|[Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math](https://arxiv.org/abs/2504.21233) \u003cbr\u003e Haoran Xu, Baolin Peng, Hany Awadalla, Dongdong Chen, Yen-Chun Chen, Mei Gao, Young Jin Kim, Yunsheng Li, Liliang Ren, Yelong Shen, Shuohang Wang, Weijian Xu, Jianfeng Gao, Weizhu Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/phi_4_mini_reasoning.png\"\u003e |[Paper](https://arxiv.org/abs/2504.21233)|[//]: #05/02\n|[Phi-4-reasoning Technical Report](https://arxiv.org/abs/2504.21318) \u003cbr\u003e Marah Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, Vibhav Vineet, Yue Wu, Safoora Yousefi, Guoqing Zheng |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/phi_4_reasoning.png\"\u003e |[Paper](https://arxiv.org/abs/2504.21318)|[//]: #05/02\n| [![Publish](https://img.shields.io/badge/Conference-ACL_2023-blue)]()\u003cbr\u003e[Teaching Small Language Models to Reason](https://arxiv.org/abs/2212.08410) \u003cbr\u003e Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, Aliaksei Severyn |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/slm_kd.png\"\u003e |[Paper](https://arxiv.org/abs/2212.08410)| [//]: #04/08\n| [![Publish](https://img.shields.io/badge/Conference-EMNLP_2024-blue)]()\u003cbr\u003e[Mixed Distillation Helps Smaller Language Model Better Reasoning](https://arxiv.org/abs/2312.10730) \u003cbr\u003e Chenglin Li, Qianglong Chen, Liangyue Li, Caiyu Wang, Yicheng Li, Zulong Chen, Yin Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/mix_distillation.png\"\u003e |[Paper](https://arxiv.org/abs/2312.10730)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Small-Model-Gap/Small-Model-Learnability-Gap.svg?style=social\u0026label=Star)](https://github.com/Small-Model-Gap/Small-Model-Learnability-Gap)\u003cbr\u003e[Small Models Struggle to Learn from Strong Reasoners](https://arxiv.org/abs/2502.12143) \u003cbr\u003e Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, Radha Poovendran |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.12143v2/x1.png\"\u003e |[Github](https://github.com/Small-Model-Gap/Small-Model-Learnability-Gap) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.12143)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Yiwei98/TDG.svg?style=social\u0026label=Star)](https://github.com/Yiwei98/TDG) [![Publish](https://img.shields.io/badge/Conference-AAAI_2024-blue)]()\u003cbr\u003e[Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data](https://arxiv.org/abs/2312.12832) \u003cbr\u003e Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Bin Sun, Xinglin Wang, Heda Wang, Kan Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2312.12832v1/x1.png\"\u003e |[Github](https://github.com/Yiwei98/TDG) \u003cbr\u003e [Paper](https://arxiv.org/abs/2312.12832)| [//]: #04/08\n| [![Publish](https://img.shields.io/badge/Conference-EMNLP_2024-blue)]()\u003cbr\u003e[Teaching Small Language Models Reasoning through Counterfactual Distillation](https://aclanthology.org/2024.emnlp-main.333/) \u003cbr\u003e Tao Feng, Yicheng Li, Li Chenglin, Hao Chen, Fei Yu, Yin Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/counterfactual_distillation.png\"\u003e |[Paper](https://aclanthology.org/2024.emnlp-main.333/)| [//]: #04/08\n|[Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation](https://arxiv.org/abs/2503.16385) \u003cbr\u003e Yijia Luo, Yulin Song, Xingyao Zhang, Jiaheng Liu, Weixun Wang, GengRu Chen, Wenbo Su, Bo Zheng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.16385v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2503.16385)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/yunx-z/SCORE.svg?style=social\u0026label=Star)](https://github.com/yunx-z/SCORE) [![Publish](https://img.shields.io/badge/Conference-ACL_Findings_2024-blue)]()\u003cbr\u003e[Small Language Models Need Strong Verifiers to Self-Correct Reasoning](https://arxiv.org/abs/2404.17140) \u003cbr\u003e Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, Lu Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2404.17140v2/x1.png\"\u003e |[Github](https://github.com/yunx-z/SCORE) \u003cbr\u003e [Paper](https://arxiv.org/abs/2404.17140)| [//]: #04/08\n|[Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation](https://arxiv.org/abs/2411.14698) \u003cbr\u003e Xunyu Zhu, Jian Li, Can Ma, Weiping Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2411.14698v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2411.14698)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Xnhyacinth/SKIntern.svg?style=social\u0026label=Star)](https://github.com/Xnhyacinth/SKIntern) [![Publish](https://img.shields.io/badge/Conference-COLING_2025-blue)]()\u003cbr\u003e[SKIntern : Internalizing Symbolic Knowledge for Distilling Better CoT Capabilities into Small Language Models](https://arxiv.org/abs/2409.13183) \u003cbr\u003e Huanxuan Liao, Shizhu He, Yupu Hao, Xiang Li, Yuanzhe Zhang, Jun Zhao, Kang Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2409.13183v2/x1.png\"\u003e |[Github](https://github.com/Xnhyacinth/SKIntern) \u003cbr\u003e [Paper](https://arxiv.org/abs/2409.13183)| [//]: #04/08\n| [![Publish](https://img.shields.io/badge/Conference-COLING_2024-blue)]()\u003cbr\u003e[Probe then Retrieve and Reason: Distilling Probing and Reasoning Capabilities into Smaller Language Models](https://aclanthology.org/2024.lrec-main.1140.pdf) \u003cbr\u003e Yichun Zhao, Shuheng Zhou, Huijia Zhu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/prr.png\"\u003e |[Paper](https://aclanthology.org/2024.lrec-main.1140.pdf)| [//]: #04/08\n|[Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners](https://arxiv.org/abs/2502.20339) \u003cbr\u003e Daniele Paliotta, Junxiong Wang, Matteo Pagliardini, Kevin Y. Li, Aviv Bick, J. Zico Kolter, Albert Gu, François Fleuret, Tri Dao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.20339v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2502.20339)| [//]: #04/08\n|[Distilling Reasoning Ability from Large Language Models with Adaptive Thinking](https://arxiv.org/abs/2404.09170) \u003cbr\u003e Xiaoshu Chen, Sihang Zhou, Ke Liang, Xinwang Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2404.09170v5/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2404.09170)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/EIT-NLP/Distilling-CoT-Reasoning.svg?style=social\u0026label=Star)](https://github.com/EIT-NLP/Distilling-CoT-Reasoning)\u003cbr\u003e[Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning](https://arxiv.org/abs/2502.18001) \u003cbr\u003e Xinghao Chen, Zhijing Sun, Wenjin Guo, Miaoran Zhang, Yanjun Chen, Yirong Sun, Hui Su, Yijie Pan, Dietrich Klakow, Wenjie Li, Xiaoyu Shen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.18001v1/x1.png\"\u003e |[Github](https://github.com/EIT-NLP/Distilling-CoT-Reasoning) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.18001)| [//]: #04/08\n\n#### Quantization and Pruning\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Towards Reasoning Ability of Small Language Models](https://arxiv.org/abs/2502.11569) \u003cbr\u003e Gaurav Srivastava, Shuxiang Cao, Xuan Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/slm_reasoning.png\"\u003e |[Paper](https://arxiv.org/abs/2502.11569)| [//]: #04/14\n|[![Star](https://img.shields.io/github/stars/ruikangliu/Quantized-Reasoning-Models.svg?style=social\u0026label=Star)](https://github.com/ruikangliu/Quantized-Reasoning-Models)\u003cbr\u003e[Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models](https://arxiv.org/abs/2504.04823) \u003cbr\u003e Ruikang Liu, Yuxuan Sun, Manyi Zhang, Haoli Bai, Xianzhi Yu, Tiezheng Yu, Chun Yuan, Lu Hou |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/quant_hurt.png\"\u003e |[Github](https://github.com/ruikangliu/Quantized-Reasoning-Models) \u003cbr\u003e [Paper](https://arxiv.org/abs/2504.04823)| [//]: #04/14\n|[When Reasoning Meets Compression: Benchmarking Compressed Large Reasoning Models on Complex Reasoning Tasks](https://arxiv.org/abs/2504.02010) \u003cbr\u003e Nan Zhang, Yusen Zhang, Prasenjit Mitra, Rui Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/when_compression.png\"\u003e |[Paper](https://arxiv.org/abs/2504.02010)| [//]: #04/14\n\n\n\n#### RL+SLM Methods\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[Replacing thinking with tool usage enables reasoning in small language models](https://arxiv.org/abs/2507.05065) \u003cbr\u003e Corrado Rainone, Tim Bakker, Roland Memisevic |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/tool_replace.png\"\u003e |[Paper](https://arxiv.org/abs/2507.05065)| [//]: #07/21\n|[Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning](https://arxiv.org/abs/2505.24850) \u003cbr\u003e Shuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu, Wei Chu, Yuan Qi |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/redi.png\"\u003e |[Paper](https://arxiv.org/abs/2505.24850)| [//]: #06/13\n|[Llama-Nemotron: Efficient Reasoning Models](https://arxiv.org/abs/2505.00949) \u003cbr\u003e NVIDIA |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.00949v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.00949)| [//]: #05/05\n|[Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math](https://arxiv.org/abs/2504.21233) \u003cbr\u003e Haoran Xu, Baolin Peng, Hany Awadalla, Dongdong Chen, Yen-Chun Chen, Mei Gao, Young Jin Kim, Yunsheng Li, Liliang Ren, Yelong Shen, Shuohang Wang, Weijian Xu, Jianfeng Gao, Weizhu Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/phi_4_mini_reasoning.png\"\u003e |[Paper](https://arxiv.org/abs/2504.21233)|[//]: #05/02\n|[Phi-4-reasoning Technical Report](https://arxiv.org/abs/2504.21318) \u003cbr\u003e Marah Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, Vibhav Vineet, Yue Wu, Safoora Yousefi, Guoqing Zheng |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/phi_4_reasoning.png\"\u003e |[Paper](https://arxiv.org/abs/2504.21318)|[//]: #05/02\n|[![Star](https://img.shields.io/github/stars/shangshang-wang/Tina.svg?style=social\u0026label=Star)](https://github.com/shangshang-wang/Tina)\u003cbr\u003e[Tina: Tiny Reasoning Models via LoRA](https://arxiv.org/abs/2504.15777) \u003cbr\u003e Shangshang Wang, Julian Asilis, Ömer Faruk Akgül, Enes Burak Bilgin, Ollie Liu, Willie Neiswanger |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.15777v1/x4.png\"\u003e |[Github](https://github.com/shangshang-wang/Tina) \u003cbr\u003e [Paper](https://arxiv.org/abs/2504.15777)| [//]: #04/25\n|[![Star](https://img.shields.io/github/stars/knoveleng/open-rs.svg?style=social\u0026label=Star)](https://github.com/knoveleng/open-rs)\u003cbr\u003e[Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't](https://arxiv.org/abs/2503.16219) \u003cbr\u003e Quy-Anh Dang, Chris Ngo |\u003cimg src=\"https://arxiv.org/html/2503.16219v1/extracted/6296504/images/pass1.png\" width=\"45%\"\u003e \u003cimg src=\"https://arxiv.org/html/2503.16219v1/extracted/6296504/images/costs.png\" width=\"45%\"\u003e |[Github](https://github.com/knoveleng/open-rs) \u003cbr\u003e [Paper](https://arxiv.org/abs/2503.16219)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/hkust-nlp/simpleRL-reason.svg?style=social\u0026label=Star)](https://github.com/hkust-nlp/simpleRL-reason)\u003cbr\u003e[SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild](https://arxiv.org/abs/2503.18892) \u003cbr\u003e Weihao Zeng, Yuzhen Huang, Qian Liu, Wei Liu, Keqing He, Zejun Ma, Junxian He |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/simplerl_zoo.png\"\u003e |[Github](https://github.com/hkust-nlp/simpleRL-reason) \u003cbr\u003e [Paper](https://arxiv.org/abs/2503.18892)| [//]: #04/08\n\n###### Repo\n\n* [DeepScaleR](https://github.com/agentica-project/deepscaler). DeepScaleR team. [Webpage](https://agentica-project.com/)\n\n\n\n### Let Decoding More Efficient\n\n\n#### Efficient TTS\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/kurt232/RLKV.svg?style=social\u0026label=Star)](https://github.com/kurt232/RLKV)\u003cbr\u003e[Which Heads Matter for Reasoning? RL-Guided KV Cache Compression](https://arxiv.org/abs/2510.08525) \u003cbr\u003e Wenjie Du, Li Jiang, Keda Tao, Xue Liu, Huan Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2510.08525v1/x3.png\"\u003e |[Github](https://github.com/kurt232/RLKV) \u003cbr\u003e[Paper](https://arxiv.org/abs/2510.08525)| [//]: #01/02\n|[Intra-request branch orchestration for efficient LLM reasoning](https://arxiv.org/abs/2509.24957) \u003cbr\u003e Weifan Jiang, Rana Shahout, Yilun Du, Michael Mitzenmacher, Minlan Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.24957v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2509.24957)| [//]: #10/30\n|[Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts](https://arxiv.org/abs/2509.21743) \u003cbr\u003e Ammar Ahmed, Azal Ahmad Khan, Ayaan Ahmad, Sheng Di, Zirui Liu, Ali Anwar |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.21743v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2509.21743)| [//]: #10/19\n|[A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning](https://arxiv.org/abs/2509.22044) \u003cbr\u003e Ziqi Wang, Boye Niu, Zhongli Li, Linghui Meng, Jing Liu, Zhi Zheng, Tong Xu, Hua Wu, Haifeng Wang, Enhong Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.22044v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2509.22044)| [//]: #10/19\n|[From Long to Short: LLMs Excel at Trimming Own Reasoning Chains](https://arxiv.org/abs/2509.06174) \u003cbr\u003e Wei Han, Geng Zhan, Sicheng Yu, Chenyu Wang, Bryan Hooi |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/l2s.png\"\u003e |[Paper](https://arxiv.org/abs/2509.06174)| [//]: #09/30\n|[Deep Think with Confidence](https://arxiv.org/abs/2508.15260) \u003cbr\u003e Yichao Fu, Xuewei Wang, Yuandong Tian, Jiawei Zhao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2508.15260v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2508.15260)| [//]: #08/26\n|[Large Reasoning Models Know How to Think Efficiently](https://openreview.net/forum?id=pLKDeGm2t1) ![](https://img.shields.io/badge/ESFoMoIII@ICML-2025-blue) \u003cbr\u003e Zeyu XING, Xing Li, Huiling Zhen, Xianzhi Yu, Mingxuan Yuan, Sinno Jialin Pan  |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/selfthink_esfomo3_icml25.png\"\u003e |[Paper](https://openreview.net/forum?id=pLKDeGm2t1)| [//]:\n|[TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling](https://arxiv.org/abs/2505.17155) ![](https://img.shields.io/badge/SCALR@COLM-2025-blue) \u003cbr\u003e Weizhe Lin, Xing Li, Zhiyuan Yang, Xiaojin Fu, Hui-Ling Zhen, Yaoyuan Wang, Xianzhi Yu, Wulong Liu, Xiaosong Li, Mingxuan Yuan  |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/trimr_scalr_colm25.png\"\u003e |[Paper](https://arxiv.org/abs/2505.17155)| [//]:\n|[Inference-Time Hyper-Scaling with KV Cache Compression](https://arxiv.org/abs/2506.05345) \u003cbr\u003e Adrian Łańcucki, Konrad Staniszewski, Piotr Nawrot, Edoardo M. Ponti |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.05345v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2506.05345)| [//]:\n|[Control-R: Towards controllable test-time scaling](https://arxiv.org/abs/2506.00189) \u003cbr\u003e Di Zhang, Weida Wang, Junxian Li, Xunzhi Wang, Jiatong Li, Jianbo Wu, Jingdi Lei, Haonan He, Peng Ye, Shufei Zhang, Wanli Ouyang, Yuqiang Li, Dongzhan Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.00189v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2506.00189)| [//]: #06/13\n|[Plan and Budget: Effective and Efficient Test-Time Scaling on Large Language Model Reasoning](https://arxiv.org/abs/2505.16122) \u003cbr\u003e Junhong Lin, Xinyue Zeng, Jie Zhu, Song Wang, Julian Shun, Jun Wu, Dawei Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/bbam.png\"\u003e |[Paper](https://arxiv.org/abs/2505.16122)| [//]: #06/13\n|[First Finish Search: Efficient Test-Time Scaling in Large Language Models](https://arxiv.org/abs/2505.18149) \u003cbr\u003e Aradhye Agarwal, Ayan Sengupta, Tanmoy Chakraborty |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/FFS.png\"\u003e |[Paper](https://arxiv.org/abs/2505.18149)| [//]: #06/13\n|[LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling](https://arxiv.org/abs/2505.19187) \u003cbr\u003e Yang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li, Pengfei Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.19187v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.19187)| [//]: #06/13\n|[Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence](https://arxiv.org/abs/2505.20325) \u003cbr\u003e Amirhosein Ghasemabadi, Keith G. Mills, Baochun Li, Di Niu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/etts.png\"\u003e |[Paper](https://arxiv.org/abs/2505.20325)| [//]: #06/13\n|[Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones](https://arxiv.org/abs/2505.21825) \u003cbr\u003e Parsa Mirtaheri, Ezra Edelman, Samy Jelassi, Eran Malach, Enric Boix-Adsera |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.21825v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.21825)| [//]: #06/11\n|[Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning](https://arxiv.org/abs/2505.17813) \u003cbr\u003e Michael Hassid, Gabriel Synnaeve, Yossi Adi, Roy Schwartz |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.17813v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.17813)| [//]: #06/11\n|[![Star](https://img.shields.io/github/stars/kaiwenw/value-guided-search.svg?style=social\u0026label=Star)](https://github.com/kaiwenw/value-guided-search)\u003cbr\u003e[Value-Guided Search for Efficient Chain-of-Thought Reasoning](https://arxiv.org/abs/2505.17373) \u003cbr\u003e Kaiwen Wang, Jin Peng Zhou, Jonathan Chang, Zhaolin Gao, Nathan Kallus, Kianté Brantley, Wen Sun |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.17373v1/x2.png\"\u003e |[Github](https://github.com/kaiwenw/value-guided-search) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.17373)| [//]: #06/11\n|[Accelerated Test-Time Scaling with Model-Free Speculative Sampling](https://arxiv.org/abs/2506.04708) \u003cbr\u003e Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/stand.png\"\u003e |[Paper](https://arxiv.org/abs/2506.04708)|[//]: #06/05\n|[Learning to Rank Chain-of-Thought: An Energy-Based Approach with Outcome Supervision](https://arxiv.org/abs/2505.14999) \u003cbr\u003e Eric Hanchen Jiang, Haozheng Luo, Shengyuan Pang, Xiaomin Li, Zhenting Qi, Hengli Li, Cheng-Fu Yang, Zongyu Lin, Xinfeng Li, Hao Xu, Kai-Wei Chang, Ying Nian Wu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/eorm.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14999)| [//]: #05/22\n|[Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling](https://arxiv.org/abs/2505.11730) \u003cbr\u003e Hao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo, Masato Motomura, Hongxiang Fan |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.11730v1/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2505.11730)| [//]: #05/22\n|[Reward Reasoning Model](https://arxiv.org/abs/2505.14674) \u003cbr\u003e Jiaxin Guo, Zewen Chi, Li Dong, Qingxiu Dong, Xun Wu, Shaohan Huang, Furu Wei |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/rrm.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14674)| [//]: #05/22\n|[Fractured Chain-of-Thought Reasoning](https://arxiv.org/abs/2505.12992) \u003cbr\u003e Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.12992v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.12992)| [//]: #05/20\n|[Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately](https://arxiv.org/abs/2505.13326) \u003cbr\u003e Yuhang Wang, Youhe Jiang, Bin Cui, Fangcheng Fu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.13326v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.13326)| [//]: #05/20\n|[Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers](https://arxiv.org/abs/2505.04842) \u003cbr\u003e Kusha Sareen, Morgane M Moss, Alessandro Sordoni, Rishabh Agarwal, Arian Hosseini |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/valueback.png\"\u003e |[Paper](https://arxiv.org/abs/2505.04842)| [//]: #05/19\n|[Think Deep, Think Fast: Investigating Efficiency of Verifier-free Inference-time-scaling Methods](https://arxiv.org/abs/2504.14047) \u003cbr\u003e Junlin Wang, Shang Zhu, Jon Saad-Falcon, Ben Athiwaratkun, Qingyang Wu, Jue Wang, Shuaiwen Leon Song, Ce Zhang, Bhuwan Dhingra, James Zou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.14047v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2504.14047)| [//]: #04/23\n|[![Star](https://img.shields.io/github/stars/IAAR-Shanghai/xVerify.svg?style=social\u0026label=Star)](https://github.com/IAAR-Shanghai/xVerify)\u003cbr\u003e[xVerify: Efficient Answer Verifier for Reasoning Model Evaluations](https://arxiv.org/abs/2504.10481) \u003cbr\u003e Ding Chen, Qingchen Yu, Pengyuan Wang, Wentao Zhang, Bo Tang, Feiyu Xiong, Xinchi Li, Minchuan Yang, Zhiyu Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2504.10481v1/x1.png\"\u003e |[Github](https://github.com/IAAR-Shanghai/xVerify) \u003cbr\u003e [Paper](https://arxiv.org/abs/2504.10481)| [//]: #04/17\n|[![Star](https://img.shields.io/github/stars/Pranjal2041/AdaptiveConsistency.svg?style=social\u0026label=Star)](https://github.com/Pranjal2041/AdaptiveConsistency)\u003cbr\u003e[Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs](https://arxiv.org/abs/2305.11860) \u003cbr\u003e Pranjal Aggarwal, Aman Madaan, Yiming Yang, Mausam |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/asc.png\"\u003e |[Github](https://github.com/Pranjal2041/AdaptiveConsistency) \u003cbr\u003e [Paper](https://arxiv.org/abs/2305.11860)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Yiwei98/ESC.svg?style=social\u0026label=Star)](https://github.com/Yiwei98/ESC) [![Publish](https://img.shields.io/badge/Conference-ICLR_2024-blue)]()\u003cbr\u003e[Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning](https://arxiv.org/abs/2401.10480) \u003cbr\u003e Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun, Heda Wang, Kan Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2401.10480v1/x1.png\"\u003e |[Github](https://github.com/Yiwei98/ESC) \u003cbr\u003e [Paper](https://arxiv.org/abs/2401.10480)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/WangXinglin/DSC.svg?style=social\u0026label=Star)](https://github.com/WangXinglin/DSC) [![Publish](https://img.shields.io/badge/Conference-NAACL_Findings_2025-blue)]()\u003cbr\u003e[Make Every Penny Count: Difficulty-Adaptive Self-Consistency for Cost-Efficient Reasoning](https://arxiv.org/abs/2408.13457) \u003cbr\u003e Xinglin Wang, Shaoxiong Feng, Yiwei Li, Peiwen Yuan, Yueqi Zhang, Chuyi Tan, Boyuan Pan, Yao Hu, Kan Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2408.13457v3/x3.png\"\u003e |[Github](https://github.com/WangXinglin/DSC) \u003cbr\u003e [Paper](https://arxiv.org/abs/2408.13457)| [//]: #04/08\n|[Path-Consistency: Prefix Enhancement for Efficient Inference in LLM](https://arxiv.org/abs/2409.01281) \u003cbr\u003e Jiace Zhu, Yingtao Shen, Jie Zhao, An Zou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2409.01281v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2409.01281)| [//]: #04/08\n|[Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning](https://arxiv.org/abs/2502.00511) \u003cbr\u003e Zhi Zhou, Tan Yuhao, Zenan Li, Yuan Yao, Lan-Zhe Guo, Xiaoxing Ma, Yu-Feng Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.00511v2/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2502.00511)| [//]: #04/08\n|[Confidence Improves Self-Consistency in LLMs](https://arxiv.org/abs/2502.06233) \u003cbr\u003e Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, Gal Yona |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/cisc.png\"\u003e |[Paper](https://arxiv.org/abs/2502.06233)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Chengsong-Huang/Self-Calibration.svg?style=social\u0026label=Star)](https://github.com/Chengsong-Huang/Self-Calibration)\u003cbr\u003e[Efficient Test-Time Scaling via Self-Calibration](https://arxiv.org/abs/2503.00031) \u003cbr\u003e Chengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu, Jiaxin Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.00031v1/x2.png\"\u003e |[Github](https://github.com/Chengsong-Huang/Self-Calibration) \u003cbr\u003e [Paper](https://arxiv.org/abs/2503.00031)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Zanette-Labs/SpeculativeRejection.svg?style=social\u0026label=Star)](https://github.com/Zanette-Labs/SpeculativeRejection) [![Publish](https://img.shields.io/badge/Conference-NeurIPS_2024-blue)]()\u003cbr\u003e[Fast Best-of-N Decoding via Speculative Rejection](https://arxiv.org/abs/2410.20290) \u003cbr\u003e Hanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang, Jiahao Qiu, Ming Yin, Mengdi Wang, Peter Bartlett, Andrea Zanette |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2410.20290v2/x1.png\"\u003e |[Github](https://github.com/Zanette-Labs/SpeculativeRejection) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.20290)| [//]: #04/08\n|[Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding](https://arxiv.org/abs/2503.01422) \u003cbr\u003e Yiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang, Zhuosheng Zhang, Fei Huang, Rui Wang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.01422v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2503.01422)| [//]: #04/08\n|[FastMCTS: A Simple Sampling Strategy for Data Synthesis](https://www.arxiv.org/abs/2502.11476) \u003cbr\u003e Peiji Li, Kai Lv, Yunfan Shao, Yichuan Ma, Linyang Li, Xiaoqing Zheng, Xipeng Qiu, Qipeng Guo |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.11476v1/x2.png\"\u003e |[Paper](https://www.arxiv.org/abs/2502.11476)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/chang-github-00/LLM-Predictive-Decoding.svg?style=social\u0026label=Star)](https://github.com/chang-github-00/LLM-Predictive-Decoding) [![Publish](https://img.shields.io/badge/Conference-ICLR_2025-blue)]()\u003cbr\u003e[Non-myopic Generation of Language Models for Reasoning and Planning](https://arxiv.org/abs/2410.17195) \u003cbr\u003e Chang Ma, Haiteng Zhao, Junlei Zhang, Junxian He, Lingpeng Kong |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/predictive_decoding.png\"\u003e |[Github](https://github.com/chang-github-00/LLM-Predictive-Decoding) \u003cbr\u003e [Paper](https://arxiv.org/abs/2410.17195)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/ethanm88/self-taught-lookahead.svg?style=social\u0026label=Star)](https://github.com/ethanm88/self-taught-lookahead)\u003cbr\u003e[Language Models can Self-Improve at State-Value Estimation for Better Search](https://arxiv.org/abs/2503.02878) \u003cbr\u003e Ethan Mendes, Alan Ritter |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.02878v1/x1.png\"\u003e |[Github](https://github.com/ethanm88/self-taught-lookahead) \u003cbr\u003e [Paper](https://arxiv.org/abs/2503.02878)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/xufangzhi/phi-Decoding.svg?style=social\u0026label=Star)](https://github.com/xufangzhi/phi-Decoding)\u003cbr\u003e[ϕ-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation](https://arxiv.org/abs/2503.13288) \u003cbr\u003e Fangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Jun Liu, Qika Lin, Zhiyong Wu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2503.13288v1/x2.png\"\u003e |[Github](https://github.com/xufangzhi/phi-Decoding) \u003cbr\u003e [Paper](https://arxiv.org/abs/2503.13288)| [//]: #04/08\n|[Dynamic Parallel Tree Search for Efficient LLM Reasoning](https://arxiv.org/abs/2502.16235) \u003cbr\u003e Yifu Ding, Wentao Jiang, Shunyu Liu, Yongcheng Jing, Jinyang Guo, Yingjie Wang, Jing Zhang, Zengmao Wang, Ziwei Liu, Bo Du, Xianglong Liu, Dacheng Tao |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.16235v2/x5.png\"\u003e |[Paper](https://arxiv.org/abs/2502.16235)| [//]: #04/08\n|[![Star](https://img.shields.io/github/stars/Soistesimmer/Fetch.svg?style=social\u0026label=Star)](https://github.com/Soistesimmer/Fetch)\u003cbr\u003e[Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls](https://arxiv.org/abs/2502.11183) \u003cbr\u003e Ante Wang, Linfeng Song, Ye Tian, Dian Yu, Haitao Mi, Xiangyu Duan, Zhaopeng Tu, Jinsong Su, Dong Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2502.11183v2/extracted/6301324/figures/method.png\"\u003e |[Github](https://github.com/Soistesimmer/Fetch) \u003cbr\u003e [Paper](https://arxiv.org/abs/2502.11183)| [//]: #04/08\n\n\n\n#### Other Optimal Methods\n| Title \u0026 Authors | Introduction | Links |\n|:--|  :----: | :---:|\n|[![Star](https://img.shields.io/github/stars/giovanni-vaccarino/PUMA.svg?style=social\u0026label=Star)](https://github.com/giovanni-vaccarino/PUMA)\u003cbr\u003e[Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models](https://arxiv.org/abs/2605.17672) \u003cbr\u003e Dehai Min, Giovanni Vaccarino, Huiyi Chen, Yongliang Wu, Gal Yona, Lu Cheng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2605.17672v1/x1.png\"\u003e |[Github](https://github.com/giovanni-vaccarino/PUMA) \u003cbr\u003e [Paper](https://arxiv.org/abs/2605.17672)| [//]: #05/24\n|[Entropy After ⟨/𝚃𝚑𝚒𝚗𝚔⟩ for reasoning model early exiting](https://arxiv.org/abs/2509.26522) \u003cbr\u003e Xi Wang, James McInerney, Lequn Wang, Nathan Kallus |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.26522v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2509.26522)| [//]: #10/30\n|[SpecExit: Accelerating Large Reasoning Model via Speculative Exit](https://arxiv.org/abs/2509.24248) \u003cbr\u003e Rubing Yang, Huajun Bai, Song Liu, Guanghua Yu, Runzhi Fan, Yanbin Dang, Jiejing Zhang, Kai Liu, Jianchen Zhu, Peng Chen |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.24248v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2509.24248)| [//]: #10/19\n|[SPEC-RL: Accelerating On-Policy Reinforcement Learning via Speculative Rollouts](https://arxiv.org/abs/2509.23232) \u003cbr\u003e Bingshuai Liu, Ante Wang, Zijun Min, Liang Yao, Haibo Zhang, Yang Liu, Anxiang Zeng, Jinsong Su |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.23232v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2509.23232)| [//]: #10/19\n|[FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning](https://arxiv.org/abs/2509.21792) \u003cbr\u003e Yizhou Zhang, Ning Lv, Teng Wang, Jisheng Dang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.21792v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2509.21792)| [//]: #10/19\n|[Reasoning with Sampling: Your Base Model is Smarter Than You Think](https://arxiv.org/abs/2510.14901) \u003cbr\u003e Aayush Karan, Yilun Du |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2510.14901v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2510.14901)| [//]: #10/19\n|[Early Stopping Chain-of-thoughts in Large Language Models](https://arxiv.org/abs/2509.14004) \u003cbr\u003e Minjia Mao, Bowen Yin, Yu Zhu, Xiao Fang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2509.14004v1/fig/ESCoT_framework.png\"\u003e |[Paper](https://arxiv.org/abs/2509.14004)| [//]: #09/30\n|[Stop Spinning Wheels: Mitigating LLM Overthinking via Mining Patterns for Early Reasoning Exit](https://arxiv.org/abs/2508.17627) \u003cbr\u003e Zihao Wei, Liang Pang, Jiahao Liu, Jingcheng Deng, Shicheng Xu, Zenghao Duan, Jingang Wang, Fei Sun, Xunliang Cai, Huawei Shen, Xueqi Cheng |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2508.17627v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2508.17627)| [//]: #09/03\n|[Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression](https://arxiv.org/abs/2508.05337) \u003cbr\u003e Jiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen, Di He, Lu Hou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2508.05337v1/imgs/case_study.png\"\u003e |[Paper](https://arxiv.org/abs/2508.05337)| [//]: #08/26\n|[![Star](https://img.shields.io/github/stars/ArminAzizi98/ASC.svg?style=social\u0026label=Star)](https://github.com/ArminAzizi98/ASC)\u003cbr\u003e[Activation Steering for Chain-of-Thought Compression](https://arxiv.org/abs/2507.04742) \u003cbr\u003e Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/asc_2.png\"\u003e |[Github](https://github.com/ArminAzizi98/ASC) \u003cbr\u003e [Paper](https://arxiv.org/abs/2507.04742)| [//]: #07/12\n|[Scaling Speculative Decoding with Lookahead Reasoning](https://arxiv.org/abs/2506.19830) \u003cbr\u003e Yichao Fu, Rui Ge, Zelei Shao, Zhijie Deng, Hao Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.19830v1/extracted/6567712/figure/LookaheadReasoningStep.jpg\"\u003e |[Paper](https://arxiv.org/abs/2506.19830)| [//]: #07/01\n|[Wait, We Don't Need to 'Wait'! Removing Thinking Tokens Improves Reasoning Efficiency](https://arxiv.org/abs/2506.08343) \u003cbr\u003e Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, Tianyi Zhou |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.08343v2/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2506.08343)| [//]: #06/24\n|[Steering LLM Thinking with Budget Guidance](https://arxiv.org/abs/2506.13752) \u003cbr\u003e Junyan Li, Wenshuo Zhao, Yang Zhang, Chuang Gan |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.13752v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2506.13752)| [//]: #06/23\n|[![Star](https://img.shields.io/github/stars/microsoft/SeerAttention.svg?style=social\u0026label=Star)](https://github.com/microsoft/SeerAttention)\u003cbr\u003e[SeerAttention-R: Sparse Attention Adaptation for Long Reasoning](http://gooarxiv.org/abs/2506.08889) \u003cbr\u003e Yizhao Gao, Shuming Guo, Shijie Cao, Yuqing Xia, Yu Cheng, Lei Wang, Lingxiao Ma, Yutao Sun, Tianzhu Ye, Li Dong, Hayden Kwok-Hay So, Yu Hua, Ting Cao, Fan Yang, Mao Yang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.08889v1/x1.png\"\u003e |[Github](https://github.com/microsoft/SeerAttention) \u003cbr\u003e [Paper](http://gooarxiv.org/abs/2506.08889)| [//]: #06/16\n|[Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs](https://arxiv.org/abs/2506.07240) \u003cbr\u003e Roy Eisenstadt, Itamar Zimerman, Lior Wolf |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.07240v1/extracted/6523627/figures/loading_Bar2.jpg\"\u003e |[Paper](https://arxiv.org/abs/2506.07240)| [//]: #06/16\n|[![Star](https://img.shields.io/github/stars/tsinghua-fib-lab/Token_Signature.svg?style=social\u0026label=Star)](https://github.com/tsinghua-fib-lab/Token_Signature)\u003cbr\u003e[Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models](https://arxiv.org/abs/2506.06008) \u003cbr\u003e Peijie Liu, Fengli Xu, Yong Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2506.06008v1/x2.png\"\u003e |[Github](https://github.com/tsinghua-fib-lab/Token_Signature) \u003cbr\u003e [Paper](https://arxiv.org/abs/2506.06008)| [//]: #06/16\n|[![Star](https://img.shields.io/github/stars/ASTRAL-Group/AlphaOne.svg?style=social\u0026label=Star)](https://github.com/ASTRAL-Group/AlphaOne)\u003cbr\u003e[AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time](https://arxiv.org/abs/2505.24863) \u003cbr\u003e Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, Huan Zhang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.24863v1/x27.png\"\u003e |[Github](https://github.com/ASTRAL-Group/AlphaOne) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.24863)| [//]: #06/13\n|[ProxyThinker: Test-Time Guidance through Small Visual Reasoners](https://arxiv.org/abs/2505.24872) \u003cbr\u003e Zilin Xiao,Jaywon Koo,Siru Ouyang,Jefferson Hernandez,Yu Meng,Vicente Ordonez |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.24872v1/x3.png\"\u003e |[Paper](https://arxiv.org/abs/2505.24872)| [//]: #06/11\n|[A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings](https://arxiv.org/abs/2505.24550) \u003cbr\u003e Xiaoang Xu,Shuo Wang,Xu Han,Zhenghao Liu,Huijia Wu,Peipei Li,Zhiyuan Liu,Maosong Sun,Zhaofeng He |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.24550v1/x2.png\"\u003e |[Paper](https://arxiv.org/abs/2505.24550)| [//]: #06/11\n|[Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models](https://arxiv.org/abs/2505.17697) \u003cbr\u003e Zekai Zhao, Qi Liu, Kun Zhou, Zihan Liu, Yifei Shao, Zhiting Hu, Biwei Huang |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.17697v1/x7.png\"\u003e |[Paper](https://arxiv.org/abs/2505.17697)| [//]: #06/11\n|[Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training](https://arxiv.org/abs/2505.14681) \u003cbr\u003e Mengru Wang, Xingyu Chen, Yue Wang, Zhiwei He, Jiahao Xu, Tian Liang, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang, Ruotian Ma, Haitao Mi, Ningyu Zhang, Zhaopeng Tu, Xiaolong Li, Dong Yu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14681v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14681)| [//]: #05/22\n|[![Star](https://img.shields.io/github/stars/jiwonsong-dev/ReasoningPathCompression.svg?style=social\u0026label=Star)](https://github.com/jiwonsong-dev/ReasoningPathCompression)\u003cbr\u003e[Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning](https://arxiv.org/abs/2505.13866) \u003cbr\u003e Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/rpc_new.png\"\u003e |[Github](https://github.com/jiwonsong-dev/ReasoningPathCompression) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.13866)| [//]: #05/22\n|[RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning](https://arxiv.org/abs/2505.14140) \u003cbr\u003e Qianyue Hao, Sibo Li, Jian Yuan, Yong Li |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.14140v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.14140)| [//]: #05/22\n|[Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity](https://arxiv.org/abs/2505.11107) \u003cbr\u003e Chan-Jan Hsu, Davide Buffelli, Jamie McGowan, Feng-Ting Liao, Yi-Chang Chen, Sattar Vakili, Da-shan Shiu |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.11107v1/extracted/6445446/figures/gt_main_new.png\"\u003e |[Paper](https://arxiv.org/abs/2505.11107)| [//]: #05/19\n|[![Star](https://img.shields.io/github/stars/LYC127/RPG.svg?style=social\u0026label=Star)](https://github.com/LYC127/RPG) [![Publish](https://img.shields.io/badge/Conference-ACL_main_2025-blue)]()\u003cbr\u003e[Rethinking Repetition Problems of LLMs in Code Generation](https://arxiv.org/abs/2505.10402) \u003cbr\u003e Yihong Dong, Yuchen Liu, Xue Jiang, Zhi Jin, Ge Li |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/code_repeat.png\"\u003e |[Github](https://github.com/LYC127/RPG) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.10402)| [//]: #05/18\n|[Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping](https://arxiv.org/abs/2505.08392) \u003cbr\u003e Ren Zhuang, Ben Wang, Shuifa Sun |\u003cimg width=\"1002\" alt=\"image\" src=\"https://arxiv.org/html/2505.08392v1/x1.png\"\u003e |[Paper](https://arxiv.org/abs/2505.08392)| [//]: #05/17\n|[![Star](https://img.shields.io/github/stars/zch65458525/L2T.svg?style=social\u0026label=Star)](https://github.com/zch65458525/L2T) [![Publish](https://img.shields.io/badge/Conference-IJCAI-blue)]()\u003cbr\u003e[Learn to Think: Bootstrapping LLM Reasoning Capability Through Graph Learning](https://arxiv.org/abs/2505.06321) \u003cbr\u003e Hang Gao, Chenhao Zhang, Tie Wang, Junsuo Zhao, Fengge Wu, Changwen Zheng, Huaping Liu |\u003cimg width=\"1002\" alt=\"image\" src=\"figures/learn2think.png\"\u003e |[Github](https://github.com/zch65458525/L2T) \u003cbr\u003e [Paper](https://arxiv.org/abs/2505.06321)| [//]: #05/17\n|[![Star](https://img.shields.io/github/stars/hammoudhasan/SubthoughtReasoner.svg?style=social\u0026label=Star)](https://github.","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/fscdc%2Fawesome-efficient-reasoning-models/projects"}