{"id":87334,"url":"https://github.com/hemingkx/Awesome-Efficient-Reasoning","name":"Awesome-Efficient-Reasoning","description":"Paper list for Efficient Reasoning.","projects_count":488,"last_synced_at":"2026-09-04T05:00:30.798Z","repository":{"id":282163217,"uuid":"947605037","full_name":"hemingkx/Awesome-Efficient-Reasoning","owner":"hemingkx","description":"Paper list for Efficient Reasoning.","archived":false,"fork":false,"pushed_at":"2026-05-29T03:59:53.000Z","size":242,"stargazers_count":901,"open_issues_count":0,"forks_count":47,"subscribers_count":8,"default_branch":"main","last_synced_at":"2026-08-15T12:06:43.950Z","etag":null,"topics":["chain-of-thought","efficient-reasoning"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hemingkx.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-03-13T00:36:14.000Z","updated_at":"2026-08-11T04:02:09.000Z","dependencies_parsed_at":"2025-04-18T12:27:33.863Z","dependency_job_id":"258b065c-ba31-46f6-b5fc-7677612aebae","html_url":"https://github.com/hemingkx/Awesome-Efficient-Reasoning","commit_stats":null,"previous_names":["hemingkx/awesome-efficient-reasoning"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/hemingkx/Awesome-Efficient-Reasoning","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hemingkx%2FAwesome-Efficient-Reasoning","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hemingkx%2FAwesome-Efficient-Reasoning/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hemingkx%2FAwesome-Efficient-Reasoning/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hemingkx%2FAwesome-Efficient-Reasoning/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hemingkx","download_url":"https://codeload.github.com/hemingkx/Awesome-Efficient-Reasoning/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hemingkx%2FAwesome-Efficient-Reasoning/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":37020126,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-04T02:00:06.169Z","response_time":115,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2025-04-07T13:01:40.679Z","updated_at":"2026-09-04T05:00:30.798Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Papers","Keywords Convention","Resources","Blog \u0026 Project","Talks"],"sub_categories":["Optimal Test-Time Scaling","Latent Chain-of-Thought","Long-to-Short Chain-of-Thought","Speculative Decoding for CoT Efficiency","Efficient Training","Survey","Applications","Reasoning Shortcuts","Reasoning Step Decomposition","Small Reasoning Models \u0026 CoT Distillation","Efficient Sampling","Efficient Self-Consistency","Long-Context Reasoning Efficiency","Parallel Thinking","Other Work","Benchmarks","Analysis","Small \u0026 Large Reasoning Model Collaboration","Adaptive Thinking","Sparse Attention \u0026 KV Cache","Multimodal Reasoning Efficiency","Efficient Sampling Methods"],"readme":"This repository contains a regularly updated paper list for **Efficient Reasoning**.\n\n[![Awesome](https://awesome.re/badge.svg)](https://awesome.re) [![License](https://img.shields.io/badge/License-Apache_2.0-green.svg)](./LICENSE) ![GitHub last commit (branch)](https://img.shields.io/github/last-commit/hemingkx/Awesome-Efficient-Reasoning/main?logo=github\u0026color=blue) ![Static Badge](https://img.shields.io/badge/Contributions-welcome-blue.svg?style=flat) \n\n## Content\n\n- [Content](#content)\n- [Keywords Convention](#keywords-convention)\n- [Papers](#papers)\n  - [Survey](#survey)\n  - [Efficient Training](#efficient-training)\n  - [Latent Chain-of-Thought](#latent-chain-of-thought)\n  - [Long-to-Short Chain-of-Thought](#long-to-short-chain-of-thought)\n  - [Adaptive Thinking](#adaptive-thinking)\n  - [Reasoning Shortcuts](#reasoning-shortcuts)\n  - [Reasoning Step Decomposition](#reasoning-step-decomposition)\n  - [Small Reasoning Models \\\u0026 CoT Distillation](#small-reasoning-models--cot-distillation)\n  - [Small \\\u0026 Large Reasoning Model Collaboration](#small--large-reasoning-model-collaboration)\n  - [Speculative Decoding for CoT Efficiency](#speculative-decoding-for-cot-efficiency)\n  - [Parallel Thinking](#parallel-thinking)\n  - [Sparse Attention \\\u0026 KV Cache](#sparse-attention--kv-cache)\n  - [Optimal Test-Time Scaling](#optimal-test-time-scaling)\n  - [Efficient Self-Consistency](#efficient-self-consistency)\n  - [Efficient Sampling Methods](#efficient-sampling-methods)\n  - [Long-Context Reasoning Efficiency](#long-context-reasoning-efficiency)\n  - [Multimodal Reasoning Efficiency](#multimodal-reasoning-efficiency)\n  - [Other Work](#other-work)\n  - [Benchmarks](#benchmarks)\n  - [Analysis](#analysis)\n  - [Applications](#applications)\n- [Blog \\\u0026 Project](#blog--project)\n- [Talks](#talks)\n- [Resources](#resources)\n- [Contributors](#contributors)\n- [Contributing to this paper list](#contributing-to-this-paper-list)\n\n\n## Keywords Convention\n\n![](https://img.shields.io/badge/COCONUT-blue) Abbreviation\n\n![](https://img.shields.io/badge/ACL2024-orange) Conference\n\n![](https://img.shields.io/badge/Multimodal-green) Main Features\n\n## Papers\n\n### Survey\n\n- **Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models**  \n  *Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen (Henry) Zhong, Hanjie Chen, Xia Hu*. [[pdf](https://arxiv.org/pdf/2503.16419)], [[paper list](https://github.com/Eclipsess/Awesome-Efficient-Reasoning-LLMs)], 2025.03. ![](https://img.shields.io/badge/TMLR2025-orange)\n- **A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond**  \n  *Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, Yu Cheng*. [[pdf](https://arxiv.org/pdf/2503.21614)], [[paper list](https://github.com/XiaoYee/Awesome_Efficient_LRM_Reasoning)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Efficient Inference for Large Reasoning Models: A Survey**  \n  *Yue Liu, Jiaying Wu, Yufei He, Hongcheng Gao, Hongyu Chen, Baolong Bi, Jiaheng Zhang, Zhiqi Huang, Bryan Hooi*. [[pdf](https://arxiv.org/pdf/2503.23077)], [[paper list](https://github.com/yueliu1999/Awesome-Efficient-Inference-for-LRMs)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models**  \n  *Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, Kam-Fai Wong*. [[pdf](https://arxiv.org/pdf/2503.24377)], [[paper list](https://github.com/DevoAllen/Awesome-Reasoning-Economy-Papers)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Efficient Reasoning Models: A Survey**  \n  *Sicheng Feng, Gongfan Fang, Xinyin Ma, Xinchao Wang*. [[pdf](https://arxiv.org/pdf/2504.10903)], [[paper list](https://github.com/fscdc/Awesome-Efficient-Reasoning-Models)], 2025.04. ![](https://img.shields.io/badge/TMLR2025-orange)\n- **Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning**  \n  *Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, Xiaoyu Shen*. [[pdf](https://arxiv.org/pdf/2505.16782)], [[paper list](https://github.com/EIT-NLP/Awesome-Latent-CoT)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs**  \n  *Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun, Soumyasundar Pal, Zhanguang Zhang, Yaochen Hu, Rohan Deepak Ajwani, Antonios Valkanas, Raika Karimi, Peng Cheng, Yunzhou Wang, Pengyi Liao, Hanrui Huang, Bin Wang, Jianye Hao, Mark Coates*. [[pdf](https://arxiv.org/pdf/2507.02076)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange)\n- **A Survey on Latent Reasoning**  \n  *Rui-Jie Zhu, Tianhao Peng, Tianhao Cheng, Xingwei Qu, Jinfa Huang, Dawei Zhu, Hao Wang, Kaiwen Xue, Xuanliang Zhang, Yong Shan, Tianle Cai, Taylor Kergan, Assel Kembay, Andrew Smith, Chenghua Lin, Binh Nguyen, Yuqi Pan, Yuhong Chou, Zefan Cai, Zhenhe Wu, Yongchi Zhao, Tianyu Liu, Jian Yang, Wangchunshu Zhou, Chujie Zheng, Chongxuan Li, Yuyin Zhou, Zhoujun Li, Zhaoxiang Zhang, Jiaheng Liu, Ge Zhang, Wenhao Huang, Jason Eshraghian*. [[pdf](https://arxiv.org/pdf/2507.06203)], [[paper list](https://github.com/multimodal-art-projection/LatentCoT-Horizon/)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey**  \n  *Jason Zhu, Hongyu Li*. [[pdf](https://arxiv.org/pdf/2507.09662)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models**  \n  *Linan Yue, Yichao Du, Yizhi Wang, Weibo Gao, Fangzhou Yao, Li Wang, Ye Liu, Ziyu Xu, Qi Liu, Shimin Di, Min-Ling Zhang*. [[pdf](https://arxiv.org/pdf/2508.02120)], [[paper list](https://github.com/yuelinan/Awesome-Efficient-R1-style-LRMs)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Implicit Reasoning in Large Language Models: A Comprehensive Survey**  \n  *Jindong Li, Yali Fu, Li Fan, Jiahong Liu, Yao Shu, Chengwei Qin, Menglin Yang, Irwin King, Rex Ying*. [[pdf](https://arxiv.org/pdf/2509.02350)], [[paper list](https://github.com/digailab/awesome-llm-implicit-reasoning)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange)\n- **A Survey on Parallel Reasoning**  \n  *Ziqi Wang, Boye Niu, Zipeng Gao, Zhi Zheng, Tong Xu, Linghui Meng, Zhongli Li, Jing Liu, Yilong Chen, Chen Zhu, Hua Wu, Haifeng Wang, Enhong Chen*. [[pdf](https://arxiv.org/pdf/2510.12164)], [[paper list](https://github.com/PPPP-kaqiu/Awesome-Parallel-Reasoning)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Models**  \n  *Chao Wu, Baoheng Li, Mingchen Gao, Zhenyi Wang*. [[pdf](https://arxiv.org/pdf/2511.10788)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook**  \n  *Xinlei Yu, Zhangquan Chen, Yongbo He, Tianyu Fu, Cheng Yang, Chengming Xu, Yue Ma, Xiaobin Hu, Zhe Cao, Jie Xu, Guibin Zhang, Jiale Tao, Jiayi Zhang, Siyuan Ma, Kaituo Feng, Haojie Huang, Youxing Li, Ronghao Chen, Huacan Wang, Chenglin Wu, Zikun Su, Xiaogang Xu, Kelu Yao, Kun Wang, Chen Gao, Yue Liao, Ruqi Huang, Tao Jin, Cheng Tan, Jiangning Zhang, Wenqi Ren, Yanwei Fu, Yong Liu, Yu Wang, Xiangyu Yue, Yu-Gang Jiang, Shuicheng Yan*. [[pdf](https://arxiv.org/pdf/2604.02029)], [[paper list](https://github.com/YU-deep/Awesome-Latent-Space)], 2026.04. ![](https://img.shields.io/badge/Arxiv-orange)\n\n### Efficient Training\n\n- **s1: Simple test-time scaling**  \n  *Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto*. [[pdf](https://arxiv.org/pdf/2501.19393)], [[code](https://github.com/simplescaling/s1)], 2025.01. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/s1-blue)\n- **LIMO: Less is More for Reasoning**  \n  *Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, Pengfei Liu*. [[pdf](https://arxiv.org/pdf/2502.03387)], [[code](https://github.com/GAIR-NLP/LIMO)], 2025.02. ![](https://img.shields.io/badge/COLM2025-orange) ![](https://img.shields.io/badge/LIMO-blue)\n- **TreeRL: LLM Reinforcement Learning with On-Policy Tree Search**  \n  *Zhenyu Hou, Ziniu Hu, Yujiang Li, Rui Lu, Jie Tang, Yuxiao Dong*. [[pdf](https://arxiv.org/pdf/2506.11902)], [[code](https://github.com/THUDM/TreeRL)], 2025.02. ![](https://img.shields.io/badge/ACL2025-orange) ![](https://img.shields.io/badge/QFFT-blue)\n- **Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond**  \n  *Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Lifu Tang, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, Xiangzheng Zhang*. [[pdf](https://aclanthology.org/2025.acl-industry.24/)], [[code](https://github.com/Qihoo360/Light-R1)], 2025.03. ![](https://img.shields.io/badge/ACL2025--industry-orange) ![](https://img.shields.io/badge/Light--R1-blue)\n- **DAPO: An Open-Source LLM Reinforcement Learning System at Scale**  \n  *Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Jinhua Zhu, Jiaze Chen, Jiangjie Chen, Chengyi Wang, Hongli Yu, Weinan Dai, Yuxuan Song, Xiangpeng Wei, Hao Zhou, Jingjing Liu, Wei-Ying Ma, Ya-Qin Zhang, Lin Yan, Mu Qiao, Yonghui Wu, Mingxuan Wang*. [[pdf](https://arxiv.org/pdf/2503.14476)], [[code](https://github.com/BytedTsinghua-SIA/DAPO)], [[homepage](https://dapo-sia.github.io/)], 2025.03. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/DAPO-blue)\n- **FastCuRL: Curriculum Reinforcement Learning with Progressive Context Extension for Efficient Training R1-like Reasoning Models**  \n  *Mingyang Song, Mao Zheng, Zheng Li, Wenjie Yang, Xuan Luo, Yue Pan, Feng Zhang*. [[pdf](https://arxiv.org/pdf/2503.17287)], [[code](https://github.com/nick7nlp/FastCuRL)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/FastCuRL-blue)\n- **Understanding R1-Zero-Like Training: A Critical Perspective**  \n  *Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin*. [[pdf](https://arxiv.org/pdf/2503.20783)], [[code](https://github.com/sail-sg/understand-r1-zero)], 2025.03. ![](https://img.shields.io/badge/COLM2025-orange) ![](https://img.shields.io/badge/Dr.GRPO-blue)\n- **Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training**  \n  *Brian R. Bartoldson, Siddarth Venkatraman, James Diffenderfer, Moksh Jain, Tal Ben-Nun, Seanie Lee, Minsu Kim, Johan Obando-Ceron, Yoshua Bengio, Bhavya Kailkhura*. [[pdf](https://arxiv.org/pdf/2503.18929)], 2025.03. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/TBA-blue)\n- **CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning Models**  \n  *Zhihang Lin, Mingbao Lin, Yuan Xie, Rongrong Ji*. [[pdf](https://arxiv.org/pdf/2503.22342)], [[code](https://github.com/lzhxmu/CPPO)], 2025.03. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/CPPO-blue)\n- **Efficient Reinforcement Finetuning via Adaptive Curriculum Learning**  \n  *Taiwei Shi, Yiyang Wu, Linxin Song, Tianyi Zhou, Jieyu Zhao*. [[pdf](https://arxiv.org/pdf/2504.05520)], [[code](https://github.com/uscnlp-lime/verl)], 2025.04. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ADARFT-blue)\n- **VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks**  \n  *Yu Yue, Yufeng Yuan, Qiying Yu, Xiaochen Zuo, Ruofei Zhu, Wenyuan Xu, Jiaze Chen, Chengyi Wang, TianTian Fan, Zhengyin Du, Xiangpeng Wei, Xiangyu Yu, Gaohong Liu, Juncai Liu, Lingjun Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Ru Zhang, Xin Liu, Mingxuan Wang, Yonghui Wu, Lin Yan*. [[pdf](https://arxiv.org/pdf/2504.05118)], 2025.04. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/VAPO-blue)\n- **Accelerating RL for LLM Reasoning with Optimal Advantage Regression**  \n  *Kianté Brantley, Mingyu Chen, Zhaolin Gao, Jason D. Lee, Wen Sun, Wenhao Zhan, Xuezhou Zhang*. [[pdf](https://arxiv.org/pdf/2505.20686)], [[code](https://github.com/ZhaolinGao/A-PO)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/A*--PO-blue)\n- **AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning**  \n  *Wei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu, Zhiyu Mei, Chuyi He, Shusheng Xu, Guo Wei, Jun Mei, Jiashu Wang, Tongkai Yang, Binhang Yuan, Yi Wu*. [[pdf](https://arxiv.org/pdf/2505.24298)], [[code](https://github.com/inclusionAI/AReaL/)], [[homepage](https://inclusionai.github.io/AReaL/intro.html)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/AReal-blue)\n- **Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning**  \n  *Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, Kai Dang, Xionghui Chen, Jianxin Yang, Zhenru Zhang, Yuqiong Liu, An Yang, Andrew Zhao, Yang Yue, Shiji Song, Bowen Yu, Gao Huang, Junyang Lin*. [[pdf](https://arxiv.org/pdf/2506.01939)], [[homepage](https://shenzhi-wang.github.io/high-entropy-minority-tokens-rlvr/)], 2025.06. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Training_with_High--Entropy_Tokens-green)\n- **Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts**  \n  *Haizhong Zheng, Yang Zhou, Brian R. Bartoldson, Bhavya Kailkhura, Fan Lai, Jiawei Zhao, Beidi Chen*. [[pdf](https://arxiv.org/pdf/2506.02177)], [[homepage](https://infini-ai-lab.github.io/GRESO/)], [[code](https://github.com/Infini-AI-Lab/GRESO)], 2025.06. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Selective_Rollouts-green)\n- **EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation**  \n  *Jinghan Jia, Hadi Reisizadeh, Chongyu Fan, Nathalie Baracaldo, Mingyi Hong, Sijia Liu*. [[pdf](https://arxiv.org/pdf/2506.04205)], [[code](https://github.com/OPTML-Group/EPiC)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/EPiC-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **SPEED-RL: Faster Training of Reasoning Models via Online Curriculum Learning**  \n  *Ruiqi Zhang, Daman Arora, Song Mei, Andrea Zanette*. [[pdf](https://arxiv.org/pdf/2506.09016)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SPEED--RL-blue)\n- **Truncated Proximal Policy Optimization**  \n  *Tiantian Fan, Lingjun Liu, Yu Yue, Jiaze Chen, Chengyi Wang, Qiying Yu, Chi Zhang, Zhiqi Lin, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Bole Ma, Mofan Zhang, Gaohong Liu, Ru Zhang, Haotian Zhou, Cong Xie, Ruidong Zhu, Zhi Zhang, Xin Liu, Mingxuan Wang, Lin Yan, Yonghui Wu*. [[pdf](https://arxiv.org/pdf/2506.15050)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/T--PPO-blue)\n- **QFFT, Question-Free Fine-Tuning for Adaptive Reasoning**  \n  *Wanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin, Ke Ji, Wenyu Chen, Yan Xu, Yasheng Wang, Lifeng Shang, Benyou Wang*. [[pdf](https://arxiv.org/pdf/2506.12860)], [[code](https://github.com/LWL-cpu/Question-Free-Fine-Tuning)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/QFFT-blue)\n- **TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling**  \n  *Yizhi Li, Qingshui Gu, Zhoufutu Wen, Ziniu Li, Tianshun Xing, Shuyue Guo, Tianyu Zheng, Xin Zhou, Xingwei Qu, Wangchunshu Zhou, Zheng Zhang, Wei Shen, Qian Liu, Chenghua Lin, Jian Yang, Ge Zhang, Wenhao Huang*. [[pdf](https://arxiv.org/pdf/2508.17445)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/TreePO-blue)\n- **History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL**  \n  *Jingkai He, Tianjian Li, Erhu Feng, Dong Du, Qian Liu, Tao Liu, Yubin Xia, Haibo Chen*. [[pdf](https://arxiv.org/pdf/2508.18588)], 2025.08. ![](https://img.shields.io/badge/ASPLOS2026-orange) ![](https://img.shields.io/badge/RhymeRL-blue)\n- **FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning**  \n  *Yizhou Zhang, Ning Lv, Teng Wang, Jisheng Dang*. [[pdf](https://arxiv.org/pdf/2509.21792)], [[code](https://github.com/yedaotian9/GRPO_speculative)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/FastGRPO-blue)\n- **Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward**  \n  *Xinyu Tang, Zhenduo Zhang, Yurou Liu, Wayne Xin Zhao, Zujie Wen, Zhiqiang Zhang, Jun Zhou*. [[pdf](https://arxiv.org/pdf/2509.01321)], 2025.09. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/DEPO-blue)\n- **SPEC-RL: Accelerating On-Policy Reinforcement Learning via Speculative Rollouts**  \n  *Bingshuai Liu, Ante Wang, Zijun Min, Liang Yao, Haibo Zhang, Yang Liu, Anxiang Zeng, Jinsong Su*. [[pdf](https://arxiv.org/pdf/2509.23232)], [[code](https://github.com/ShopeeLLM/Spec-RL)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SPEC--RL-blue)\n- **RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training**  \n  *Wei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng, Wei Wang*. [[pdf](https://arxiv.org/pdf/2509.21009)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/RollPacker-blue)\n- **Self-Aligned Reward: Towards Effective and Efficient Reasoners**  \n  *Peixuan Han, Adit Krishnan, Gerald Friedland, Jiaxuan You, Chris Kong*. [[pdf](https://arxiv.org/pdf/2509.05489)], 2025.09. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/SAR-blue)\n- **CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models**  \n  *Runpeng Dai, Linfeng Song, Haolin Liu, Zhenwen Liang, Dian Yu, Haitao Mi, Zhaopeng Tu, Rui Liu, Tong Zheng, Hongtu Zhu, Dong Yu*. [[pdf](https://arxiv.org/pdf/2509.09675)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CDE-blue)\n- **On Predictability of Reinforcement Learning Dynamics for Large Language Models**  \n  *Yuchen Cai, Ding Cao, Xin Xu, Zijun Yao, Yuqing Huang, Zhenyu Tan, Benyi Zhang, Guiquan Liu, Junfeng Fang*. [[pdf](https://arxiv.org/pdf/2510.00553)], [[code](https://github.com/caiyuchen-ustc/Alpha-RL)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/AlphaRL-blue)\n- **ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems**  \n  *Qiaoling Chen, Zijun Liu, Peng Sun, Shenggui Li, Guoteng Wang, Ziming Liu, Yonggang Wen, Siyuan Feng, Tianwei Zhang*. [[pdf](https://arxiv.org/pdf/2510.26475)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ReSpec-blue)\n- **CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs**  \n  *Yongcheng Zeng, Zexu Sun, Bokai Ji, Erxue Min, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Haifeng Zhang, Xu Chen, Jun Wang*. [[pdf](https://arxiv.org/pdf/2510.01037)], [[code](https://github.com/ZexuSun/CurES)], 2025.10. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/CurES-blue)\n- **Training Large Reasoning Models Efficiently via Progressive Thought Encoding**  \n  *Zeliang Zhang, Xiaodong Liu, Hao Cheng, Hao Sun, Chenliang Xu, Jianfeng Gao*. [[pdf](https://openreview.net/pdf?id=q4iJxp47CT)], 2025.10. ![](https://img.shields.io/badge/ICLR2026-orange)\n- **SRT: Accelerating Reinforcement Learning via Speculative Rollout with Tree-Structured Cache**  \n  *Chi-Chih Chang, Siqi Zhu, Zhichen Zeng, Haibin Lin, Xin Liu, Jiaxuan You, Mohamed S. Abdelfattah, Ziheng Jiang, Xuehai Qian*. [[pdf](https://openreview.net/attachment?id=UpEKKJStAY\u0026name=pdf)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning**  \n  *Ruoyu Qin, Weiran He, Weixiao Huang, Yangkun Zhang, Yikai Zhao, Bo Pang, Xinran Xu, Yingdi Shan, Yongwei Wu, Mingxing Zhang*. [[pdf](https://arxiv.org/pdf/2511.14617)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Seer-blue)\n- **Beat the long tail: Distribution-Aware Speculative Decoding for RL Training**  \n  *Zelei Shao, Vikranth Srivatsa, Sanjana Srivastava, Qingyang Wu, Alpay Ariyak, Xiaoxia Wu, Ameen Patel, Jue Wang, Percy Liang, Tri Dao, Ce Zhang, Yiying Zhang, Ben Athiwaratkun, Chenfeng Xu, Junxiong Wang*. [[pdf](https://arxiv.org/pdf/2511.13841)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter**  \n  *Qinghao Hu, Shang Yang, Junxian Guo, Xiaozhe Yao, Yujun Lin, Yuxian Gu, Han Cai, Chuang Gan, Ana Klimovic, Song Han*. [[pdf](https://arxiv.org/pdf/2511.16665)], [[code](https://github.com/mit-han-lab/fastrl)], 2025.11. ![](https://img.shields.io/badge/ASPLOS'26-orange) ![](https://img.shields.io/badge/TLT-blue)\n- **Fast LLM Post-training via Decoupled and Best-of-N Speculation**  \n  *Rongxin Cheng, Kai Zhou, Xingda Wei, Siyuan Liu, Mingcong Han, Mingjing Ai, Yeju Zhou, Baoquan Zhong, Wencong Xiao, Rong Chen, Haibo Chen*. [[pdf](https://arxiv.org/pdf/2511.16193)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting**  \n  *Siqi Wang, Hailong Yang, Junjie Zhu, Xuezhu Wang, Yufan Xu, Depei Qian*. [[pdf](https://arxiv.org/pdf/2512.04752)], 2025.12. ![](https://img.shields.io/badge/Arxiv-orange)\n\n### Latent Chain-of-Thought\n\n- **Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning**  \n  *Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, Xiaoyu Shen*. [[pdf](https://arxiv.org/pdf/2505.16782)], [[paper list](https://github.com/EIT-NLP/Awesome-Latent-CoT)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Survey-green)\n- **Think before you speak: Training Language Models With Pause Tokens**  \n  *Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, Vaishnavh Nagarajan*. [[pdf](https://openreview.net/pdf?id=ph04CRkPdC)], 2023.10. ![](https://img.shields.io/badge/ICLR2024-orange)\n- **Guiding Language Model Reasoning with Planning Tokens**  \n  *Xinyi Wang, Lucas Caccia, Oleksiy Ostapenko, Xingdi Yuan, William Yang Wang, Alessandro Sordoni*. [[pdf](https://openreview.net/pdf?id=wi9IffRhVM)], 2023.10. ![](https://img.shields.io/badge/COLM2024-orange)\n- **Implicit Chain of Thought Reasoning via Knowledge Distillation**   \n  *Yuntian Deng, Kiran Prasad, Roland Fernandez, Paul Smolensky, Vishrav Chaudhary, Stuart Shieber*. [[pdf](https://arxiv.org/pdf/2311.01460)], 2023.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models**  \n  *Jiacheng Ye, Shansan Gong, Liheng Chen, Lin Zheng, Jiahui Gao, Han Shi, Chuan Wu, Xin Jiang, Zhenguo Li, Wei Bi, Lingpeng Kong*. [[pdf](https://openreview.net/pdf?id=G0v0TxX01N)], 2024.02. ![](https://img.shields.io/badge/NIPS2024-orange) ![](https://img.shields.io/badge/DoT-blue)\n- **Let's Think Dot by Dot: Hidden Computation in Transformer Language Models**  \n  *Jacob Pfau, William Merrill, Samuel R. Bowman*. [[pdf](https://openreview.net/pdf?id=NikbrdtYvG)], 2024.04. ![](https://img.shields.io/badge/COLM2024-orange) ![](https://img.shields.io/badge/Filler-blue)\n- **From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step**  \n  *Yuntian Deng, Yejin Choi, Stuart Shieber*. [[pdf](https://arxiv.org/pdf/2405.14838)], 2024.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ImplicitCoT-blue)\n- **Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decoding**  \n  *Tianqiao Liu, Zui Chen, Zitao Liu, Mi Tian, Weiqi Luo*. [[pdf](https://arxiv.org/pdf/2409.08561)], 2024.09. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Do LLMs Really Think Step-by-step In Implicit Reasoning?**  \n  *Yijiong Yu*. [[pdf](https://arxiv.org/pdf/2411.15862v3)], 2024.11. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Analysis-green)\n- **Training Large Language Models to Reason in a Continuous Latent Space**  \n  *Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, Yuandong Tian*. [[pdf](https://openreview.net/pdf?id=Itxz7S4Ip3)], [[code](https://github.com/facebookresearch/coconut)], 2024.12. ![](https://img.shields.io/badge/COLM2025-orange) ![](https://img.shields.io/badge/COCONUT-blue)\n- **Compressed Chain of Thought: Efficient Reasoning Through Dense Representations**  \n  *Jeffrey Cheng, Benjamin Van Durme*. [[pdf](https://arxiv.org/pdf/2412.13171)], 2024.12. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CCoT-blue)\n- **Efficient Reasoning with Hidden Thinking**  \n  *Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, Jiuxiang Gu*. [[pdf](https://arxiv.org/pdf/2501.19201)], 2025.01. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Heima-blue) ![](https://img.shields.io/badge/Multimodal-green)\n- **Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking**  \n  *Yilong Chen, Junyuan Shang, Zhenyu Zhang, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang*. [[pdf](https://arxiv.org/pdf/2502.13842)], 2025.02. ![](https://img.shields.io/badge/ACL2025-orange) ![](https://img.shields.io/badge/ITT-blue)\n- **LightThinker: Thinking Step-by-Step Compression**  \n  *Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, Ningyu Zhang*. [[pdf](https://arxiv.org/pdf/2502.15589)], [[code](https://github.com/zjunlp/LightThinker)], 2025.02. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/LightThinker-blue)\n- **Reasoning with Latent Thoughts: On the Power of Looped Transformers**  \n  *Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi*. [[pdf](https://openreview.net/pdf?id=din0lGfZFd)], 2025.02. ![](https://img.shields.io/badge/ICLR2025-orange)\n- **CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation**  \n  *Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, Yulan He*. [[pdf](https://arxiv.org/pdf/2502.21074)], 2025.02. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/CODI-blue)\n- **Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach**  \n  *Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, Tom Goldstein*. [[pdf](https://arxiv.org/pdf/2502.05171)], [[code](https://github.com/seal-rg/recurrent-pretraining)], 2025.02. ![](https://img.shields.io/badge/NeurIPS2025-orange)\n- **LLM Pretraining with Continuous Concepts**  \n  *Jihoon Tack, Jack Lanchantin, Jane Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason Weston, Xian Li*. [[pdf](https://arxiv.org/pdf/2502.08524)], [[code](https://github.com/facebookresearch/RAM/tree/main/projects/cocomix)], 2025.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CoCoMix-blue) ![](https://img.shields.io/badge/Pretrain-green)\n- **Latent Thought Models with Variational Bayes Inference-Time Computation**  \n  *Deqian Kong, Minglu Zhao, Dehong Xu, Bo Pang, Shu Wang, Edouardo Honig, Zhangzhang Si, Chuan Li, Jianwen Xie, Sirui Xie, Ying Nian Wu*. [[pdf](https://arxiv.org/pdf/2502.01567)], 2025.02. ![](https://img.shields.io/badge/ICML2025-orange) ![](https://img.shields.io/badge/LTM-blue)\n- **Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning**  \n  *Qifan Yu, Zhenyu He, Sijie Li, Xun Zhou, Jun Zhang, Jingjing Xu, Di He*. [[pdf](https://arxiv.org/pdf/2502.08482)], [[code](https://github.com/qifanyu/RELAY)], 2025.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/RELAY-blue)\n- **Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning**  \n  *DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, Qinqing Zheng*. [[pdf](https://arxiv.org/pdf/2502.03275)], 2025.02. ![](https://img.shields.io/badge/ICML2025-orange)\n- **Implicit Reasoning in Transformers is Reasoning through Shortcuts**  \n  *Tianhe Lin, Jian Xie, Siyu Yuan, Deqing Yang*. [[pdf](https://arxiv.org/pdf/2503.07604)], 2025.03. ![](https://img.shields.io/badge/ACL2025Findings-orange) ![](https://img.shields.io/badge/Analysis-green)\n- **Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendation**  \n  *Jiakai Tang, Sunhao Dai, Teng Shi, Jun Xu, Xu Chen, Wen Chen, Wu Jian, Yuning Jiang*. [[pdf](https://arxiv.org/pdf/2503.22675)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ReaRec-blue) ![](https://img.shields.io/badge/Sequential_Recommendation-green)\n- **Efficient Pretraining Length Scaling**  \n  *Bohong Wu, Shen Yan, Sijun Zhang, Jianqiao Lu, Yutao Zeng, Ya Wang, Xun Zhou*. [[pdf](https://arxiv.org/pdf/2504.14992)], 2025.04. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/PHD--Transformer-blue)\n- **Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space**  \n  *Zhen Zhang, Xuehai He, Weixiang Yan, Ao Shen, Chenyang Zhao, Shuohang Wang, Yelong Shen, Xin Eric Wang*. [[pdf](https://arxiv.org/pdf/2505.15778)], [[code](https://github.com/eric-ai-lab/Soft-Thinking)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Soft--Thinking-blue)\n- **Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains**  \n  *Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, Ruihua Song*. [[pdf](https://arxiv.org/pdf/2505.16552)], [[homepage](https://colar-latent-reasoning.github.io/)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/CoLaR-blue)\n- **Hybrid Latent Reasoning via Reinforcement Learning**  \n  *Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang, Zhen Qin, Jinsung Yoon, Lanyu Shang, Jiawei Han, Dong Wang*. [[pdf](https://arxiv.org/pdf/2505.18454)], [[code](https://github.com/Yueeeeeeee/HRPO)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/HRPO-blue)\n- **Efficient Post-Training Refinement of Latent Reasoning in Large Language Models**  \n  *Xinyuan Wang, Dongjie Wang, Wangyang Ying, Haoyue Bai, Nanxu Gong, Sixun Dong, Kunpeng Liu, Yanjie Fu*. [[pdf](https://arxiv.org/pdf/2506.08552)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens**  \n  *Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen, Chuang Gan*. [[pdf](https://arxiv.org/pdf/2506.17218)], [[code](https://github.com/UMass-Embodied-AGI/Mirage)], [[homepage](https://vlm-mirage.github.io/)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Mirage-blue) ![](https://img.shields.io/badge/Multimodal-green)\n- **DART: Distilling Autoregressive Reasoning to Silent Thought**  \n  *Nan Jiang, Ziming Wu, De-Chuan Zhan, Fuming Lai, Shaobing Lian*. [[pdf](https://arxiv.org/pdf/2506.11752)], 2025.06. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/DART-blue)\n- **Parallel Continuous Chain-of-Thought with Jacobi Iteration**  \n  *Haoyi Wu, Zhihao Teng, Kewei Tu*. [[pdf](https://arxiv.org/pdf/2506.18582)], [[code](https://github.com/whyNLP/PCCoT)], 2025.06. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/PCCoT-blue)\n- **Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models**  \n  *Tan-Hanh Pham, Chris Ngo*. [[pdf](https://arxiv.org/pdf/2508.12587)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/MCOUT-blue)\n- **LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking**  \n  *Chünhung Wu, Jinliang Lu, Zixuan Ren, Gangqiang Hu, Zhi Wu, Dai Dai, Hua Wu*. [[pdf](https://arxiv.org/pdf/2508.03440)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Soft Tokens, Hard Truths**  \n  *Natasha Butt, Ariel Kwiatkowski, Ismail Labiad, Julia Kempe, Yann Ollivier*. [[pdf](https://arxiv.org/pdf/2509.19170)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange)\n- **SIM-CoT: Supervised Implicit Chain-of-Thought**  \n  *Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Jiaqi Wang, Xipeng Qiu, Dahua Lin*. [[pdf](https://arxiv.org/pdf/2509.20317)], [[code](https://github.com/InternLM/SIM-CoT)], 2025.09. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/SIM--CoT-blue)\n- **KaVa: Latent Reasoning via Compressed KV-Cache Distillation**  \n  *Anna Kuzina, Maciej Pioro, Paul N. Whatmough, Babak Ehteshami Bejnordi*. [[pdf](https://arxiv.org/pdf/2510.02312)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/KaVa-blue)\n- **SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs**  \n  *Dachuan Shi, Abedelkadir Asi, Keying Li, Xiangchi Yuan, Leyan Pan, Wenke Lee, Wen Xiao*. [[pdf](https://arxiv.org/pdf/2510.05069)], [[code](https://github.com/sdc17/SwiReasoning)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SwiReasoning-blue)\n- **Parallel Test-Time Scaling for Latent Reasoning Models**\n  *Runyang You, Yongqi Li, Meng Liu, Wenjie Wang, Liqiang Nie, Wenjie Li*. [[pdf](https://arxiv.org/pdf/2510.07745),[code](https://github.com/ModalityDance/LatentTTS/)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Towards Inference-time Scaling for Continuous Space Reasoning**  \n  *Minghan Wang, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari*. [[pdf](https://arxiv.org/pdf/2510.12167)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Analysis-green)\n- **Latent Reasoning in LLMs as a Vocabulary-Space Superposition**  \n  *Jingcheng Deng, Liang Pang, Zihao Wei, Shichen Xu, Zenghao Duan, Kun Xu, Yang Song, Huawei Shen, Xueqi Cheng*. [[pdf](https://arxiv.org/pdf/2510.15522)], [[code](https://github.com/DJC-GO-SOLO/Latent-SFT)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space**  \n  *Chao Chen, Zhixin Ma, Yongqi Li, Yupeng Hu, Yinwei Wei, Wenjie Li, Liqiang Nie*. [[pdf](https://arxiv.org/pdf/2510.12603)], [[code](https://github.com/FYYDCC/IVT-LR)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/IVT--LR-blue) ![](https://img.shields.io/badge/Multimodal-green)\n- **Continuous Autoregressive Language Models**  \n  *Chenze Shao, Darren Li, Fandong Meng, Jie Zhou*. [[pdf](https://arxiv.org/pdf/2510.27688)], [[code](https://github.com/shaochenze/calm)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CALM-blue)\n- **Rethinking LLM Reasoning: From Explicit Trajectories to Latent Representations**  \n  *Cong Jiang, Xiaofeng Zhang, Fangzhi Zhu, XiaoWei Chen, Junxiong Zhu, Zheng Zhang*. [[pdf](https://openreview.net/pdf?id=CbK7lYbmv8)], [[code](https://github.com/MobiusDai/LRT)], 2025.09. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/LRT-blue)\n- **Learning to Reason over Continuous Tokens with Reinforcement Learning**  \n  *Yiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong, Junnan Li*. [[pdf](https://openreview.net/pdf?id=lebJ6wz1vj)], [[code](https://github.com/zhaoyiran924/HyRea)], 2025.09. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/HyRea-blue)\n- **Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models**  \n  *Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, Yu Wang*. [[pdf](https://arxiv.org/abs/2511.08577)], [[code](https://github.com/thu-nics/TaH)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/TaH-blue)\n- **Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought**  \n  *Zhikang Chen, Sen Cui, Deheng Ye, Yu Zhang, Yatao Bian, Tingting Zhu*. [[pdf](https://arxiv.org/pdf/2511.07124)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/EBM--CoT-blue)\n- **SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization**  \n  *Zhi Zheng, Yu Gu, Wei Liu, Yee Whye Teh, Wee Sun Lee*. [[pdf](https://arxiv.org/pdf/2511.06411)], [[code](https://github.com/zz1358m/SofT-GRPO-master)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SofT--GRPO-blue)\n- **Mull-Tokens: Modality-Agnostic Latent Thinking**  \n*Arijit Ray, Ahmed Abdelkader, Chengzhi Mao, Bryan A. Plummer, Kate Saenko, Ranjay Krishna, Leonidas Guibas, Wen-Sheng Chu*. [[pdf](https://arxiv.org/pdf/2512.10941)], [[homepage](https://arijitray.com/multimodal_thinking/)], 2025.12. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Mull--Tokens-blue) ![](https://img.shields.io/badge/Multimodal-green)\n- **Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space**  \n*Chengzhi Liu, Yuzhe Yang, Yue Fan, Qingyue Wei, Sheng Liu, Xin Eric Wang*. [[pdf](https://arxiv.org/pdf/2512.12623)], [[homepage](https://mllm-dmlr.github.io/)], [[code](https://github.com/eric-ai-lab/DMLR)], 2025.12. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DMLR-blue)\n- **Beyond Imitation: Reinforcement Learning for Active Latent Planning**  \n*Zhi Zheng, Wee Sun Lee*. [[pdf](https://arxiv.org/pdf/2601.21598)], [[code](https://github.com/zz1358m/ATP-Latent-master)], 2026.01. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ATP--Latent-blue)\n- **Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning**  \n*Chi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen, Jan Kautz, Yu-Chiang Frank Wang, Fu-En Yang*. [[pdf](https://arxiv.org/pdf/2601.09708)], [[homepage](https://jasper0314-huang.github.io/fast-thinkact/)], 2025.12. ![](https://img.shields.io/badge/CVPR2026-orange) ![](https://img.shields.io/badge/Multimodal-green)\n- **VaLR: Vision-aligned Latent Reasoning for Multi-modal Large Language Model**  \n*Byungwoo Jeon, Yoonwoo Jeong, Hyunseok Lee, Minsu Cho, Jinwoo Shin*. [[pdf](https://arxiv.org/pdf/2602.04476)], [[homepage](https://rootyjeon.github.io/valr/)], 2026.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/VaLR-blue) ![](https://img.shields.io/badge/Multimodal-green)\n\n### Long-to-Short Chain-of-Thought\n\n- **Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models**  \n  *Hanxu Hu, Hongyuan Lu, Huajian Zhang, Yun-Ze Song, Wai Lam, Yue Zhang*. [[pdf](https://arxiv.org/pdf/2305.10276)], [[code](https://github.com/hanxuhu/chain-of-symbol-planning)], 2023.05. ![](https://img.shields.io/badge/COLM2024-orange)\n- **The Benefits of a Concise Chain of Thought on Problem-Solving in Large Language Models**  \n  *Matthew Renze, Erhan Guven*. [[pdf](https://arxiv.org/pdf/2401.05618)], [[code](https://github.com/matthewrenze/jhu-concise-cot)], 2024.01. ![](https://img.shields.io/badge/FLLM2024-orange) ![](https://img.shields.io/badge/CCoT-blue)\n- **Efficiently Serving LLM Reasoning Programs with Certaindex**  \n  *Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Aurick Qiao, Hao Zhang*. [[pdf](https://arxiv.org/pdf/2412.20993)], 2024.12. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Dynasor-blue)\n- **C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness**  \n  *Yu Kang, Xianghui Sun, Liangyu Chen, Wei Zou*. [[pdf](https://arxiv.org/pdf/2412.11664)], 2024.12. ![](https://img.shields.io/badge/AAAI2025-orange) ![](https://img.shields.io/badge/C3oT-blue)\n- **Token-Budget-Aware LLM Reasoning**  \n  *Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen*. [[pdf](https://arxiv.org/pdf/2412.18547)], [[code](https://github.com/GeniusHTX/TALE)], 2024.12. ![](https://img.shields.io/badge/ACL2025--findings-orange) ![](https://img.shields.io/badge/TALE-blue) ![](https://img.shields.io/badge/Prompt-green)\n- **O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning**  \n  *Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, Dacheng Tao*. [[pdf](https://arxiv.org/pdf/2501.12570)], [[code](https://github.com/StarDewXXX/O1-Pruner)], 2025.01. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/O1--Pruner-blue)\n- **Kimi k1.5: Scaling Reinforcement Learning with LLMs**  \n  *Kimi Team*. [[pdf](https://arxiv.org/pdf/2501.12599)], 2025.01. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Kimi_k1.5-blue)\n- **Training Language Models to Reason Efficiently**  \n  *Daman Arora, Andrea Zanette*. [[pdf](https://arxiv.org/pdf/2502.04463)], [[code](https://github.com/Zanette-Labs/efficient-reasoning)], [[homepage](https://zanette-labs.github.io/efficient-reasoning/)], 2025.02. ![](https://img.shields.io/badge/NeurIPS2025-orange)\n- **Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models**  \n  *Yuan Sui, Yufei He, Tri Cao, Simeng Han, Bryan Hooi*. [[pdf](https://arxiv.org/pdf/2502.19918)], 2025.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Meta--Reasoner-blue)\n- **CoT-Valve: Length-Compressible Chain-of-Thought Tuning**  \n  *Xinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang, Xinchao Wang*. [[pdf](https://arxiv.org/pdf/2502.09601)], [[code](https://github.com/horseee/CoT-Valve)], 2025.02. ![](https://img.shields.io/badge/ACL2025-orange) ![](https://img.shields.io/badge/CoT--Valve-blue)\n- **TokenSkip: Controllable Chain-of-Thought Compression in LLMs**  \n  *Heming Xia, Yongqi Li, Chak Tou Leong, Wenjie Wang, Wenjie Li*. [[pdf](https://arxiv.org/pdf/2502.12067)], [[code](https://github.com/hemingkx/TokenSkip)], 2025.02. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/TokenSkip-blue) ![](https://img.shields.io/badge/Controllable_Compression-green)\n- **Self-Training Elicits Concise Reasoning in Large Language Models**  \n  *Tergel Munkhbat, Namgyu Ho, Seo Hyun Kim, Yongjin Yang, Yujin Kim, Se-Young Yun*. [[pdf](https://arxiv.org/pdf/2502.20122)], [[code](https://github.com/TergelMunkhbat/concise-reasoning)], 2025.02. ![](https://img.shields.io/badge/ACL2025--findings-orange)\n- **Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning**  \n  *Wenkai Yang, Shuming Ma, Yankai Lin, Furu Wei*. [[pdf](https://arxiv.org/pdf/2502.18080)], 2025.02. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Over--thinking-green)\n- **Chain of Draft: Thinking Faster by Writing Less**  \n  *Silei Xu, Wenhao Xie, Lingxiao Zhao, Pengcheng He*. [[pdf](https://arxiv.org/pdf/2502.18600)], [[code](https://github.com/sileix/chain-of-draft)], 2025.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CoD-blue) ![](https://img.shields.io/badge/Prompt-green)\n- **L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning**  \n  *Pranjal Aggarwal, Sean Welleck*. [[pdf](https://www.arxiv.org/pdf/2503.04697)], [[code](https://github.com/cmu-l3/l1)], [[homepage](https://cmu-l3.github.io/l1/)], 2025.03. ![](https://img.shields.io/badge/COLM2025-orange) ![](https://img.shields.io/badge/L1-blue) ![](https://img.shields.io/badge/Length_Control-green)\n- **DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models**  \n  *Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wenjing Zhang, Jiangze Yan, Ning Wang, Kai Wang, Shiguo Lian*. [[pdf](https://arxiv.org/pdf/2503.04472)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DAST-blue)\n- **How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach**  \n  *Ayeong Lee, Ethan Che, Tianyi Peng*. [[pdf](https://arxiv.org/pdf/2503.01141)], [[code](https://github.com/Compressed-CoT/compressed-cot)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Prompt-green)\n- **Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching**  \n  *Simon A. Aytes, Jinheon Baek, Sung Ju Hwang*. [[pdf](https://arxiv.org/pdf/2503.05179)], [[code](https://github.com/SimonAytes/SoT)], 2025.03. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/SoT-blue) ![](https://img.shields.io/badge/Route--and--Prompt-green)\n- **Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning**  \n  *Chen Li, Nazhou Liu, Kai Yang*. [[pdf](https://arxiv.org/pdf/2503.15952)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/AGPO-blue) ![](https://img.shields.io/badge/length--based_reward-green)\n- **Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging**  \n  *Han Wu, Yuxuan Yao, Shuqi Liu, Zehua Liu, Xiaojin Fu, Xiongwei Han, Xing Li, Hui-Ling Zhen, Tao Zhong, Mingxuan Yuan*. [[pdf](https://arxiv.org/pdf/2503.20641)], [[code](https://github.com/hahahawu/Long-to-Short-via-Model-Merging)], 2025.03. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Model_Merging-green)\n- **Think When You Need: Self-Adaptive Chain-of-Thought Learning**  \n  *Junjie Yang, Ke Lin, Xing Yu*. [[pdf](https://arxiv.org/pdf/2504.03234)], 2025.04. ![](https://img.shields.io/badge/Arxiv-orange)\n- **ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning**  \n  *Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang*. [[pdf](https://arxiv.org/pdf/2504.01296)], [[code](https://github.com/UCSB-NLP-Chang/ThinkPrune)], 2025.04. ![](https://img.shields.io/badge/TMLR-orange) ![](https://img.shields.io/badge/ThinkPrune-blue)\n- **Reasoning Models Can Be Effective Without Thinking**  \n  *Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, Matei Zaharia*. [[pdf](https://arxiv.org/pdf/2504.09858)], 2025.04. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/NoThinking-blue) ![](https://img.shields.io/badge/Prompt-green)\n- **ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning**  \n  *Jingyang Yi, Jiazheng Wang*. [[pdf](https://arxiv.org/pdf/2504.21370)], 2025.04. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/ShorterBetter-blue) ![](https://img.shields.io/badge/Group--Relative_Length_Reward-green)\n- **Dynamic Early Exit in Reasoning Models**  \n  *Chenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu, Chenyu Zhu, Zheng Lin, Li Cao, Weiping Wang*. [[pdf](https://arxiv.org/pdf/2504.15895)], 2025.04. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/DEER-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **AdaR1: From Long-CoT to Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization**  \n  *Haotian Luo, Haiying He, Yibo Wang, Jinluan Yang, Rui Liu, Naiqiang Tan, Xiaochun Cao, Dacheng Tao, Li Shen*. [[pdf](https://arxiv.org/pdf/2504.21659)], [[code](https://github.com/StarDewXXX/AdaR1)], 2025.04. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/AdaR1-blue)\n- **Concise Reasoning via Reinforcement Learning**  \n  *Mehdi Fatemi, Banafsheh Rafiee, Mingjie Tang, Kartik Talamadupula*. [[pdf](https://arxiv.org/pdf/2504.05185)], 2025.04. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models**  \n  *Bin Yu, Hang Yuan, Yuliang Wei, Bailing Wang, Weizhen Qi, Kai Chen*. [[pdf](https://arxiv.org/pdf/2505.03469)], [[code](https://github.com/ZGCA-AI4Edu/LS-Mixture)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/LS--Mixture_SFT-blue)\n- **ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning**  \n  *Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang*. [[pdf](https://arxiv.org/pdf/2505.04881)], 2025.05. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/ConCISE-blue)\n- **Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens**  \n  *Xixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li, Yefeng Zheng, Xian Wu*. [[pdf](https://arxiv.org/pdf/2505.18237)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange)\n- **Activation-Guided Consensus Merging for Large Language Models**  \n  *Yuxuan Yao, Shuqi Liu, Zehua Liu, Qintong Li, Mingyang Liu, Xiongwei Han, Zhijiang Guo, Han Wu, Linqi Song*. [[pdf](https://arxiv.org/pdf/2505.14009)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange)\n- **Scalable Chain of Thoughts via Elastic Reasoning**  \n  *Yuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo, Junnan Li, Caiming Xiong*. [[pdf](https://arxiv.org/pdf/2505.05315)], 2025.05. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/Elastic_Reasoning-blue) ![](https://img.shields.io/badge/Length_Control-green)\n- **Let LRMs Break Free from Overthinking via Self-Braking Tuning**  \n  *Haoran Zhao, Yuchen Yan, Yongliang Shen, Haolei Xu, Wenqi Zhang, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, Yueting Zhuang*. [[pdf](https://arxiv.org/pdf/2505.14604)], [[code](https://github.com/ZJU-REAL/Self-Braking-Tuning)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Self--Braking--Tuning-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time**  \n  *Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, Huan Zhang*. [[pdf](https://arxiv.org/pdf/2505.24863)], [[homepage](https://alphaone-project.github.io/)], [[code](https://github.com/ASTRAL-Group/AlphaOne)], 2025.05. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/AlphaOne-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models**  \n  *Muzhi Dai, Chenxu Yang, Qingyi Si*. [[pdf](https://arxiv.org/pdf/2505.07686)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/S--GRPO-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement**  \n  *Xuechen Zhang, Zijian Huang, Chenchun Ni, Ziyang Xiong, Jiasi Chen, Samet Oymak*. [[pdf](https://arxiv.org/pdf/2505.07961)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Length_Penalty-green)\n- **Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping**  \n  *Ren Zhuang, Ben Wang, Shuifa Sun*. [[pdf](https://arxiv.org/pdf/2505.08392)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Token_Skipping-green)\n- **SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning**  \n  *Zheng Li, Qingxiu Dong, Jingyuan Ma, Di Zhang, Zhifang Sui*. [[pdf](https://arxiv.org/pdf/2505.11274)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SelfBudgeter-blue) ![](https://img.shields.io/badge/Adaptive_Token_Budget-green)\n- **Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning**  \n  *Yansong Ning, Wei Li, Jun Fang, Naiqiang Tan, Hao Liu*. [[pdf](https://arxiv.org/pdf/2505.11827)], [[code](https://github.com/yasNing/Long-otimes-Short/)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Fractured Chain-of-Thought Reasoning**  \n  *Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong*. [[pdf](https://arxiv.org/pdf/2505.12992)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Efficient RL Training for Reasoning Models via Length-Aware Optimization**  \n  *Danlong Yuan, Tian Xie, Shaohan Huang, Zhuocheng Gong, Huishuai Zhang, Chong Luo, Furu Wei, Dongyan Zhao*. [[pdf](https://arxiv.org/pdf/2505.12284)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning**  \n  *Shangziqi Zhao, Jiahao Yuan, Guisong Yang, Usman Naseem*. [[pdf](https://arxiv.org/pdf/2505.14582)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Prune--on--Logic-blue)\n- **DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models**  \n  *Yuxuan Jiang, Dawei Li, Frank Ferraro*. [[pdf](https://arxiv.org/pdf/2505.13975)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DRP-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **SEAL: Steerable Reasoning Calibration of Large Language Models for Free**  \n  *Runjin Chen, Zhenyu Zhang, Junyuan Hong, Souvik Kundu, Zhangyang Wang*. [[pdf](https://arxiv.org/pdf/2504.07986)], [[code](https://github.com/VITA-Group/SEAL)], 2025.05. ![](https://img.shields.io/badge/COLM2025-orange) ![](https://img.shields.io/badge/SEAL-blue) ![](https://img.shields.io/badge/Steering-green)\n- **FlashThink: An Early Exit Method For Efficient Reasoning**  \n  *Guochao Jiang, Guofeng Quan, Zepeng Ding, Ziqin Luo, Dixuan Wang, Zheng Hu*. [[pdf](https://arxiv.org/pdf/2505.13949)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/FlashThink-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **Optimizing Anytime Reasoning via Budget Relative Policy Optimization**  \n  *Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin*. [[pdf](https://arxiv.org/pdf/2505.13438)], [[code](https://github.com/sail-sg/AnytimeReasoner)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/BRPO-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **VeriThinker: Learning to Verify Makes Reasoning Model Efficient**  \n  *Zigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu, Xinchao Wang*. [[pdf](https://arxiv.org/pdf/2505.17941)], [[code](https://github.com/czg1225/VeriThinker)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/VeriThinker-blue) ![](https://img.shields.io/badge/CoTs_by_InstructLM-green)\n- **Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning**  \n  *Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim*. [[pdf](https://arxiv.org/pdf/2505.13866)], [[code](https://github.com/jiwonsong-dev/ReasoningPathCompression)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/KV_Cache_Pruning-green)\n- **ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy**  \n  *Gengyang Li, Yifeng Gao, Yuming Li, Yunfang Wu*. [[pdf](https://arxiv.org/pdf/2505.15684)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ThinkLess-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n- **Learn to Reason Efficiently with Adaptive Length-based Reward Shaping**  \n  *Wei Liu, Ruochen Zhou, Yiyun Deng, Yuzhen Huang, Junteng Liu, Yuntian Deng, Yizhe Zhang, Junxian He*. [[pdf](https://arxiv.org/pdf/2505.15612)], [[code](https://github.com/hkust-nlp/Laser)], 2025.05. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/Laser-blue)\n- **R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search**  \n  *Yibo Wang, Li Shen, Huanjin Yao, Tiansheng Huang, Rui Liu, Naiqiang Tan, Jiaxing Huang, Kai Zhang, Dacheng Tao*. [[pdf](https://arxiv.org/pdf/2505.16838)], [[code](https://github.com/w-yibo/R1-Compress)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/R1--Compress-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning**  \n  *Xiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang, Wayne Xin Zhao, Xinyu Kong, Zhiqiang Zhang*. [[pdf](https://arxiv.org/pdf/2505.16315)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/ACPO-blue)\n- **Plan and Budget: Effective and Efficient Test-Time Scaling on Large Language Model Reasoning**  \n  *Junhong Lin, Xinyue Zeng, Jie Zhu, Song Wang, Julian Shun, Jun Wu, Dawei Zhou*. [[pdf](https://www.arxiv.org/pdf/2505.16122)], [[code](https://github.com/junhongmit/P-and-B)], 2025.05. ![](https://img.shields.io/badge/ICLR2026-orange)\n- **ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models**  \n  *Razvan-Gabriel Dumitru, Darius Peteleaza, Vikas Yadav, Liangming Pan*. [[pdf](https://www.arxiv.org/pdf/2505.17250)], [[code](https://github.com/RazvanDu/ConciseRL)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ConciseRL-blue)\n- **TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling**  \n  *Weizhe Lin, Xing Li, Zhiyuan Yang, Xiaojin Fu, Hui-Ling Zhen, Yaoyuan Wang, Xianzhi Yu, Wulong Liu, Xiaosong Li, Mingxuan Yuan*. [[pdf](https://arxiv.org/pdf/2505.17155)], 2025.05. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/TrimR-blue)  \n- **Not All Tokens Are What You Need In Thinking**  \n  *Hang Yuan, Bin Yu, Haotian Li, Shijun Yang, Christina Dan Wang, Zhou Yu, Xueyin Xu, Weizhen Qi, Kai Chen*. [[pdf](https://arxiv.org/pdf/2505.17827)], [[code](https://github.com/Faustrazor/Not-All-Thinking-Tokens)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Token_Skipping-green)\n- **LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling**  \n  *Yang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li, Pengfei Liu*. [[pdf](https://arxiv.org/pdf/2505.19187)], [[code](https://github.com/GAIR-NLP/LIMOPro)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/LIMOPro-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Interleaved Reasoning for Large Language Models via Reinforcement Learning**  \n  *Roy Xie, David Qiu, Deepak Gopinath, Dong Lin, Yanchao Sun, Chong Wang, Saloni Potdar, Bhuwan Dhingra*. [[pdf](https://arxiv.org/pdf/2505.19640)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning**  \n  *Mingyang Song, Mao Zheng*. [[pdf](https://arxiv.org/pdf/2505.21178)], [[code](https://github.com/nick7nlp/ConciseR)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Length_Penalty-green)\n- **AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting**  \n  *Shijue Huang, Hongru Wang, Wanjun Zhong, Zhaochen Su, Jiazhan Feng, Bowen Cao, Yi R. Fung*. [[pdf](https://arxiv.org/pdf/2505.18822)], [[code](https://github.com/JoeYing1019/AdaCtrl)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Length_Penalty-green)\n- **CoThink: Token-Efficient Reasoning via Instruct Models Guiding Reasoning Models**  \n  *Siqi Fan, Peng Han, Shuo Shang, Yequan Wang, Aixin Sun*. [[pdf](https://arxiv.org/pdf/2505.22017)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Outline_by_InstructLM-green)\n- **Stable Reinforcement Learning for Efficient Reasoning**  \n  *Muzhi Dai, Shixuan Liu, Qingyi Si*. [[pdf](https://arxiv.org/pdf/2505.18086)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Length_Penalty-green)\n- **Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models**  \n  *Sohyun An, Ruochen Wang, Tianyi Zhou, Cho-Jui Hsieh*. [[pdf](https://arxiv.org/pdf/2505.21765)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/DTO-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **LLMs Can Reason Faster Only If We Let Them**  \n  *Bilgehan Sel, Lifu Huang, Naren Ramakrishnan, Ruoxi Jia, Ming Jin*. [[pdf](https://openreview.net/pdf?id=uTv5rOPZr4)], 2025.05. ![](https://img.shields.io/badge/ICML2025-orange)\n- **DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning**  \n  *Shih-Yang Liu, Xin Dong, Ximing Lu, Shizhe Diao, Mingjie Liu, Min-Hung Chen, Hongxu Yin, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Yejin Choi, Jan Kautz, Pavlo Molchanov*. [[pdf](https://arxiv.org/pdf/2510.15110)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DLER-blue)\n- **A\\*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings**  \n  *Xiaoang Xu, Shuo Wang, Xu Han, Zhenghao Liu, Huijia Wu, Peipei Li, Zhiyuan Liu, Maosong Sun, Zhaofeng He*. [[pdf](https://arxiv.org/pdf/2505.24550)], [[code](https://github.com/AI9Stars/AStar-Thought)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/A*--Thought-blue)\n- **TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression**  \n  *Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Ying Nian Wu, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu*. [[pdf](https://arxiv.org/pdf/2506.02678)], [[code](https://github.com/zzli2022/TLDR)], 2025.06. ![](https://img.shields.io/badge/ACL2026-orange) ![](https://img.shields.io/badge/TL;DR-blue)\n- **Answer Convergence as a Signal for Early Stopping in Reasoning**  \n  *Xin Liu, Lu Wang*. [[pdf](https://arxiv.org/pdf/2506.02536)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Early_Exit-green)\n- **How Far Are We from Optimal Reasoning Efficiency?**  \n  *Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang, Shusheng Xu, Wei Fu, Zhiyu Mei, Kaifeng Lyu, Yi Wu*. [[pdf](https://arxiv.org/pdf/2506.07104)], [[code](https://github.com/samjia2000/Optimal-Reasoning-Efficiency)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/REO--RL-blue)\n- **Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs**  \n  *Roy Eisenstadt, Itamar Zimerman, Lior Wolf*. [[pdf](https://arxiv.org/pdf/2506.07240)], [[homepage](https://royeisen.github.io/OverclockingLLMReasoning-paper/)], [[code](https://github.com/royeisen/reasoning_loading_bar)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Bingo: Boosting Efficient Reasoning of LLMs via Dynamic and Significance-based Reinforcement Learning**  \n  *Hanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou, Haoyu Dong, Xiaojun Ma, Shi Han, Dongmei Zhang*. [[pdf](https://arxiv.org/pdf/2506.08125)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Significance--aware_Length_Penalty-green)\n- **Brevity is the soul of sustainability: Characterizing LLM response lengths**  \n  *Soham Poddar, Paramita Koley, Janardan Misra, Sanjay Podder, Navveen Balani, Niloy Ganguly, Saptarshi Ghosh*. [[pdf](https://arxiv.org/pdf/2506.08686)], 2025.06. ![](https://img.shields.io/badge/ACL2025--findings-orange) ![](https://img.shields.io/badge/Prompting-green)\n- **Wait, We Don't Need to \"Wait\"! Removing Thinking Tokens Improves Reasoning Efficiency**  \n  *Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, Tianyi Zhou*. [[pdf](https://arxiv.org/pdf/2506.08343)], 2025.06. ![](https://img.shields.io/badge/EMNLP2025--findings-orange) ![](https://img.shields.io/badge/NoWait-blue)\n- **Causal Sufficiency and Necessity Improves Chain-of-Thought Reasoning**  \n  *Xiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li, Anjie Liu, Xiao Xue, Jun Wang, Mengyue Yang*. [[pdf](https://arxiv.org/pdf/2506.09853)], 2025.06. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/In--context_Learning\u0026SFT-green)\n- **PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models**  \n  *Ye Yu, Yaoning Yu, Haohan Wang*. [[pdf](https://arxiv.org/pdf/2506.10716)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Prompting-green)\n- **Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning**  \n  *Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, Nick Haber*. [[pdf](https://arxiv.org/pdf/2506.05256)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ALP-blue) ![](https://img.shields.io/badge/Length_Control-green)\n- **ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization**  \n  *Zhensheng Jin, Xinze Li, Yifan Ji, Chunyi Peng, Zhenghao Liu, Qi Shi, Yukun Yan, Shuo Wang, Furong Peng, Ge Yu*. [[pdf](https://arxiv.org/pdf/2506.10822)], [[code](https://github.com/NEUIR/ReCUT)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty**  \n  *Zehui Ling, Deshu Chen, Hongwei Zhang, Yifeng Jiao, Xin Guo, Yuan Cheng*. [[pdf](https://arxiv.org/pdf/2506.10446)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models**  \n  *Kaiyuan Liu, Chen Shen, Zhanwei Zhang, Junjie Liu, Xiaosong Yuan, Jieping ye*. [[pdf](https://arxiv.org/pdf/2506.12353)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Steering LLM Thinking with Budget Guidance**  \n  *Junyan Li, Wenshuo Zhao, Yang Zhang, Chuang Gan*. [[pdf](https://arxiv.org/pdf/2506.13752)], [[code](https://github.com/UMass-Embodied-AGI/BudgetGuidance)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Optimizing Length Compression in Large Reasoning Models**  \n  *Zhengxiang Cheng, Dongping Chen, Mingyang Fu, Tianyi Zhou*. [[pdf](https://arxiv.org/pdf/2506.14755)], [[code](https://github.com/zxiangx/LC-R1)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement**  \n  *Weixiang Zhao, Jiahe Guo, Yang Deng, Xingyu Sui, Yulin Hu, Yanyan Zhao, Wanxiang Che, Bing Qin, Tat-Seng Chua, Ting Liu*. [[pdf](https://arxiv.org/pdf/2506.15647)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Steering-green)\n- **ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation**  \n  *Siao Tang, Xinyin Ma, Gongfan Fang, Xinchao Wang*. [[pdf](https://arxiv.org/pdf/2506.18810)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **AdapThink: Adaptive Thinking Preferences for Reasoning Language Model**  \n  *Xu Wan, Wei Wang, Wenyue Xu, Wotao Yin, Jie Song, Mingyang Sun*. [[pdf](https://arxiv.org/pdf/2506.18237)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs**  \n  *Kang Chen, Mengdi Zhang, Yixin Cao*. [[pdf](https://arxiv.org/pdf/2506.18341)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Multilingual-green)\n- **AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control**  \n  *Ruosen Li, Ziming Luo, Quan Zhang, Ruochen Li, Ben Zhou, Ali Payani, Xinya Du*. [[pdf](https://arxiv.org/pdf/2506.20160)], [[code](https://github.com/du-nlp-lab/LengthReward)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model**  \n  *Bowen Ding, Yuhan Chen, Futing Wang, Lingfeng Ming, Tao Lin*. [[pdf](https://arxiv.org/pdf/2506.23840)], [[code](https://github.com/Danield21/Dual-Policy-Preference-Optimization)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange)\n- **EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning**  \n  *Sanchit Ahuja, Praneetha Vaddamanu, Barun Patra*. [[pdf](https://arxiv.org/pdf/2507.00246)], [[code](https://github.com/microsoft/EfficientXLang)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Multilingual-green)\n- **Activation Steering for Chain-of-Thought Compression**  \n  *Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram*. [[pdf](https://arxiv.org/pdf/2507.04742)], [[code](https://github.com/ArminAzizi98/ASC)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Steering-green)\n- **SmartThinker: Learning to Compress and Preserve Reasoning by Step-Level Length Control**  \n  *Xingyang He, Xiao Ling, Jie Liu*. [[pdf](https://arxiv.org/pdf/2507.04348)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Controlling Thinking Speed in Reasoning Models**  \n  *Zhengkai Lin, Zhihang Fu, Ze Chen, Chao Chen, Liang Xie, Wenxiao Wang, Deng Cai, Zheng Wang, Jieping Ye*. [[pdf](https://arxiv.org/pdf/2507.03704v1)], 2025.07. ![](https://img.shields.io/badge/NeurIPS2025_Spotlight-orange) ![](https://img.shields.io/badge/Steering-green)\n- **Verbosity-Aware Rationale Reduction: Sentence-Level Rationale Reduction for Efficient and Effective Reasoning**  \n  *Joonwon Jang, Jaehee Kim, Wonbin Kweon, Seonghyeon Lee, Hwanjo Yu*. [[pdf](https://aclanthology.org/2025.findings-acl.1068.pdf)], 2025.07. ![](https://img.shields.io/badge/ACL2025--findings-orange)\n- **Test-time Prompt Intervention**  \n  *Chenxu Yang, Qingyi Si, Muzhi Dai, Dingyu Yao, Mingyu Zheng, Minghui Chen, Zheng Lin, Weiping Wang*. [[pdf](https://arxiv.org/pdf/2508.02511)], 2025.08. ![](https://img.shields.io/badge/AAAI2026-orange) ![](https://img.shields.io/badge/PI-blue)\n- **Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning**  \n  *Jialiang Hong, Taihang Zhen, Kai Chen, Jiaheng Liu, Wenpeng Zhu, Jing Huo, Yang Gao, Depeng Wang, Haitao Wan, Xi Yang, Boyan Wang, Fanyu Meng*. [[pdf](https://www.arxiv.org/pdf/2508.02178)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Compressing Chain-of-Thought in LLMs via Step Entropy**  \n  *Zeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen, Zhijian Xu, Yingying Cheng, Fan Zhang, Qiang Xu*. [[pdf](https://arxiv.org/pdf/2508.03346)], [[code](https://github.com/staymylove/COT_Compresstion_via_Step_entropy)], 2025.08. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Efficient Reasoning for Large Reasoning Language Models via Certainty-Guided Reflection Suppression**  \n  *Jiameng Huang, Baijiong Lin, Guhao Feng, Jierun Chen, Di He, Lu Hou*. [[pdf](https://arxiv.org/pdf/2508.05337)], 2025.08. ![](https://img.shields.io/badge/AAAI2026-orange) ![](https://img.shields.io/badge/CGRS-blue)\n- **Train Long, Think Short: Curriculum Learning for Efficient Reasoning**  \n  *Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud, Elie Bou-Zeid, Marzyeh Ghassemi, Bernard Ghanem*. [[pdf](https://arxiv.org/pdf/2508.08940)], [[code](https://github.com/hammoudhasan/curriculum_grpo)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning**  \n  *Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran, Shivam Garg, Harkirat Behl, Dimitris Papailiopoulos*. [[pdf](https://arxiv.org/pdf/2508.09726)], 2025.08. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/GFPO-blue)\n- **SABER: Switchable and Balanced Training for Efficient LLM Reasoning**  \n  *Kai Zhao, Yanjun Zhao, Jiaming Song, Shien He, Lusheng Zhang, Qiang Zhang, Tianjiao Li*. [[pdf](https://arxiv.org/pdf/2508.10026)], 2025.08. ![](https://img.shields.io/badge/AAAI2026-orange) ![](https://img.shields.io/badge/SABER-blue)\n- **Promoting Efficient Reasoning with Verifiable Stepwise Reward**  \n  *Chuhuai Yue, Chengqi Dong, Yinan Gao, Hang He, Jiajun Chai, Guojun Yin, Wei Lin*. [[pdf](https://arxiv.org/pdf/2508.10293)], 2025.08. ![](https://img.shields.io/badge/AAAI2026-orange) ![](https://img.shields.io/badge/VSRM-blue)\n- **Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models**  \n  *Qiguang Chen, Dengyun Peng, Jinhao Liu, HuiKang Su, Jiannan Guan, Libo Qin, Wanxiang Che*. [[pdf](https://arxiv.org/pdf/2508.11582)], [[code](https://github.com/sfasfaffa/DR_SAF)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DR.SAF-blue)\n- **Stop Spinning Wheels: Mitigating LLM Overthinking via Mining Patterns for Early Reasoning Exit**  \n  *Zihao Wei, Liang Pang, Jiahao Liu, Jingcheng Deng, Shicheng Xu, Zenghao Duan, Jingang Wang, Fei Sun, Xunliang Cai, Huawei Shen, Xueqi Cheng*. [[pdf](https://arxiv.org/pdf/2508.17627)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Early_Exit-green)\n- **BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens**  \n  *Hao Wen, Xinrui Wu, Yi Sun, Feifei Zhang, Liye Chen, Jie Wang, Yunxin Liu, Ya-Qin Zhang, Yuanchun Li*. [[pdf](https://arxiv.org/pdf/2508.17196)], [[code](https://github.com/MobileLLM/BudgetThinker)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/BudgetThinker-blue)\n- **DRQA: Dynamic Reasoning Quota Allocation for Controlling Overthinking in Reasoning Large Language Models**  \n  *Kaiwen Yan, Xuanqing Shi, Hongcheng Guo, Wenxuan Wang, Zhuosheng Zhang, Chengwei Qin*. [[pdf](https://arxiv.org/pdf/2508.17803)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DRQA-blue)\n- **CAC-CoT: Connector-Aware Compact Chain-of-Thought for Efficient Reasoning Data Synthesis Across Dual-System Cognitive Tasks**  \n  *Sunguk Choi, Yonghoon Kwon, Heondeuk Lee*. [[pdf](https://arxiv.org/pdf/2508.18743)], 2025.08. ![](https://img.shields.io/badge/EMNLP2025--findings-orange) ![](https://img.shields.io/badge/CAC--CoT-blue)\n- **ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models**  \n  *Qianyu He, Siyu Yuan, Xuefeng Li, Mingxuan Wang, Jiangjie Chen*. [[pdf](https://arxiv.org/pdf/2508.18773)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ThinkDial-blue) ![](https://img.shields.io/badge/gpt--oss--style-green)\n- **Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation**  \n  *Abdul Waheed, Chancharik Mitra, Laurie Z. Wang, Deva Ramanan, Bhiksha Raj*. [[pdf](https://arxiv.org/pdf/2509.05226)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange)\n- **From Long to Short: LLMs Excel at Trimming Own Reasoning Chains**  \n  *Wei Han, Geng Zhan, Sicheng Yu, Chenyu Wang, Bryan Hooi*. [[pdf](https://arxiv.org/pdf/2509.06174)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/EDIT-blue)\n- **Hierarchical Budget Policy Optimization for Adaptive Reasoning**  \n  *Shangke Lyu, Linjuan Wu, Yuchen Yan, Xingyu Wu, Hao Li, Yongliang Shen, Peisheng Jiang, Weiming Lu, Jun Xiao, Yueting Zhuang*. [[pdf](https://arxiv.org/pdf/2507.15844v2)], [[code](https://github.com/zju-real/hbpo)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/HBPO-blue)\n- **Early Stopping Chain-of-thoughts in Large Language Models**  \n  *Minjia Mao, Bowen Yin, Yu Zhu, Xiao Fang*. [[pdf](https://arxiv.org/pdf/2509.14004)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/ES--COT-blue)\n- **Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors**  \n  *Aniket Didolkar, Nicolas Ballas, Sanjeev Arora, Anirudh Goyal*. [[pdf](https://arxiv.org/pdf/2509.13237)], 2025.09. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Revisiting Model Interpolation for Efficient Reasoning**  \n  *Taiqiang Wu, Runming Yang, Tao Liu, Jiahao Wang, Ngai Wong*. [[pdf](https://arxiv.org/pdf/2510.10977)], [[code](https://github.com/wutaiqiang/MI)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/MI-blue)\n- **From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement**  \n  *Jianzhi Yan, Le Liu, Youcheng Pan, Shiwei Chen, Zike Yuan, Yang Xiang, Buzhou Tang*. [[pdf](https://arxiv.org/pdf/2509.22144)], [[code](https://github.com/Leon221220/MACC)], 2025.09. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/MACC-blue)\n- **Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking**  \n  *Jinyi Han, Ying Huang, Ying Liao, Zishang Jiang, Xikun Lu, Haiquan Zhao, Xinyi Wang, Guanghao Zhou, Sihang Jiang, Jiaqing Liang, Weikang Zhou, Zeye Sun, Fei Yu, Yanghua Xiao*. [[pdf](https://arxiv.org/pdf/2509.23392)], [[code](https://github.com/JinyiHan99/Just-Enough-Think)], 2025.09. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/JET-blue)\n- **Entropy After ⟨/𝚃𝚑𝚒𝚗𝚔⟩ for reasoning model early exiting**  \n  *Xi Wang, James McInerney, Lequn Wang, Nathan Kallus*. [[pdf](https://arxiv.org/pdf/2509.26522)], [[code](https://github.com/xidulu/EAT)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Early_Exit-green)\n- **SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression**  \n  *Haoming Wen, Yushi Bai, Juanzi Li, Jie Tang*. [[pdf](https://arxiv.org/pdf/2509.25176)], [[huggingface](https://huggingface.co/collections/THU-KEG/siri)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SIRI-blue)\n- **Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models**  \n  *Canhui Wu, Qiong Cao, Chang Li, Zhenfang Wang, Chao Xue, Yuwei Fan, Wei Xi, Xiaodong He*. [[pdf](https://arxiv.org/pdf/2510.03805)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation**  \n  *Tianyi Jiang, Yi Bin, Yujuan Ding, Kainian Zhu, Fei Ma, Jingkuan Song, Heng Tao Shen*. [[pdf](https://arxiv.org/pdf/2510.02249)], [[code](https://github.com/AusertDream/CumulativeEntropyRegulation)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression**  \n  *Joykirat Singh, Justin Chih-Yao Chen, Archiki Prasad, Elias Stengel-Eskin, Akshay Nambi, Mohit Bansal*. [[pdf](https://arxiv.org/pdf/2510.01581)], [[code](https://github.com/joykirat18/TRAAC)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Step_Shortcut-green) ![](https://img.shields.io/badge/TRAAC-blue)\n- **DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization**  \n  *Gang Li, Yan Chen, Ming Lin, Tianbao Yang*. [[pdf](https://arxiv.org/pdf/2510.04474)], [[code](https://github.com/Optimization-AI/DRPO)], 2025.10. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/DRPO-blue)\n- **Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression**  \n  *Chengzhengxu Li, Xiaoming Liu, Zhaohan Zhang, Shaochu Zhang, Shengchao Liu, Guoxin Ma, Yu Lan, Chao Shen*. [[pdf](https://arxiv.org/pdf/2510.04474)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/UCoT-blue)\n- **Mitigating Overthinking through Reasoning Shaping**  \n  *Feifan Song, Shaohang Wei, Bofei Gao, Yejie Wang, Wen Luo, Wei Li, Linli Yao, Weimin Xiong, Liang Chen, Tianyu Liu, Houfeng Wang*. [[pdf](https://arxiv.org/pdf/2510.09535)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/GRSP-blue)\n- **PAC Reasoning: Controlling the Performance Loss for Efficient Reasoning**  \n  *Hao Zeng, Jianguo Huang, Bingyi Jing, Hongxin Wei, Bo An*. [[pdf](https://arxiv.org/pdf/2510.09133)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning**  \n  *Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen, Wei Wang*. [[pdf](https://arxiv.org/pdf/2510.10103)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Merlin's Whisper: Enabling Efficient Reasoning in LLMs via Black-box Persuasive Prompting**  \n  *Heming Xia, Cunxiao Du, Rui Li, Chak Tou Leong, Yongqi Li, Wenjie Li*. [[pdf](https://arxiv.org/pdf/2510.10528)], [[code](https://github.com/hemingkx/Whisper)], 2025.10. ![](https://img.shields.io/badge/ACL2026-orange) ![](https://img.shields.io/badge/Prompting-green) ![](https://img.shields.io/badge/Whisper-blue)\n- **Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning**  \n  *Yujian Zhang, Keyu Chen, Zhifeng Shen, Ruizhi Qiao, Xing Sun*. [[pdf](https://arxiv.org/pdf/2510.10207)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling**  \n  *Shuyang Jiang, Yusheng Liao, Ya Zhang, Yanfeng Wang, Yu Wang*. [[pdf](https://openreview.net/pdf?id=kdeiRledV6)], [[code](https://github.com/pixas/DECS)], 2025.10. ![](https://img.shields.io/badge/ICLR2026--oral-orange)\n- **Concise Reasoning in the Lens of Lagrangian Optimization**  \n  *Chengqian Gao, Haonan Li, Taylor W. Killian, Jianshu She, Renxi Wang, Liqun Ma, Zhoujun Cheng, Shibo Hao, Zhiqiang Xu*. [[pdf](https://arxiv.org/pdf/2510.10168)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Towards Flash Thinking via Decoupled Advantage Policy Optimization**  \n  *Zezhong Tan, Hang Gao, Xinhong Ma, Feng Zhang, Ziqiang Dong*. [[pdf](https://arxiv.org/pdf/2510.15374)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **DART: Difficulty-Adaptive Reasoning Truncation for Efficient Large Language Models**  \n  *Ruofan Zhang, Bin Xia, Zhen Cheng, Cairen Jian, Minglun Yang, Ngai Wong, Yuan Cheng*. [[pdf](https://arxiv.org/pdf/2511.01170)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DART-blue)\n- **e1: Learning Adaptive Control of Reasoning Effort**  \n  *Michael Kleinman, Matthew Trager, Alessandro Achille, Wei Xia, Stefano Soatto*. [[pdf](https://arxiv.org/pdf/2510.27042)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Efficient Reasoning via Reward Model**  \n  *Yuhao Wang, Xiaopeng Li, Cheng Gong, Ziru Liu, Suiyun Zhang, Rui Liu, Xiangyu Zhao*. [[pdf](https://arxiv.org/pdf/2511.09158)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **ORION: Teaching Language Models to Reason Efficiently in the Language of Thought**  \n  *Kumar Tanmay, Kriti Aggarwal, Paul Pu Liang, Subhabrata Mukherjee*. [[pdf](https://arxiv.org/pdf/2511.22891)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Efficient Reasoning via Thought-Training and Thought-Free Inference**  \n  *Canhui Wu, Qiong Cao, Chao Xue, Wei Xi, Xiaodong He*. [[pdf](https://arxiv.org/pdf/2511.03408)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Dual-Density Inference for Efficient Language Model Reasoning**  \n  *Zhengyi Zhao, Shubo Zhang, Yuxi Zhang, Huimin Wang, Binyang Li, Kam-Fai Wong*. [[pdf](https://arxiv.org/pdf/2512.15358)], 2025.12. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Beyond Model Scaling: Test-Time Intervention for Efficient Deep Reasoning**  \n  *Qianyue Wang, Jinwu Hu, Yufeng Wang, Huanxiang Lin, Bolin Chen, Zhiquan Wen, Yaofo Chen, Mingkui Tan*. [[pdf](https://arxiv.org/pdf/2601.11252)], 2026.01. ![](https://img.shields.io/badge/Arxiv-orange)\n- **FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning**  \n  *Haozheng Luo, Zhuolin Jiang, Md Zahid Hasan, Yan Chen, Soumalya Sarkar*. [[pdf](https://arxiv.org/pdf/2601.19001)], [[code](https://github.com/robinzixuan/FROST)], 2026.01. ![](https://img.shields.io/badge/ICLR2026-orange) ![](https://img.shields.io/badge/FROST-blue)\n- **Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models**  \n  *Wei Wu, Liyi Chen, Congxi Xiao, Tianfu Wang, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Hui Xiong*. [[pdf](https://arxiv.org/pdf/2601.03969)], 2026.01. ![](https://img.shields.io/badge/Arxiv-orange)\n- **ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning**  \n  *Minda Hu, Zexuan Qiu, Zenan Xu, Kun Li, Bo Zhou, Irwin King*. [[pdf](https://arxiv.org/pdf/2601.04973)], 2026.01. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression**  \n  *Yuntian Tang, Bohan Jia, Wenxuan Huang, Lianyue Zhang, Jiao Xie, Wenxi Li, Rongrong Ji, Shaohui Lin*. [[pdf](https://arxiv.org/pdf/2602.08324)], [[code](https://github.com/Mwie1024/Extra-CoT)], 2026.02. ![](https://img.shields.io/badge/Arxiv-orange)\n- **On-Policy Supervised Fine-Tuning for Efficient Reasoning**  \n  *Anhao Zhao, Ziyang Chen, Junlong Tong, Yingqi Fan, Fanghua Ye, Shuhao Li, Yunpu Ma, Wenjie Li, Xiaoyu Shen*. [[pdf](https://arxiv.org/pdf/2602.13407)], 2026.02. ![](https://img.shields.io/badge/Arxiv-orange)\n- **The Art of Efficient Reasoning: Data, Reward, and Optimization**  \n  *Taiqiang Wu, Zenan Xu, Bo Zhou, Ngai Wong*. [[pdf](https://arxiv.org/pdf/2602.20945)], [[homepage](https://wutaiqiang.github.io/project/Art)], 2026.02. ![](https://img.shields.io/badge/Arxiv-orange)\n- **Constraint-Rectified Training for Efficient Chain-of-Thought**  \n  *Qinhang Wu, Sen Lin, Ming Zhang, Yingbin Liang, Ness B. Shroff*. [[pdf](https://arxiv.org/pdf/2602.12526)], 2026.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CRT-blue) ![](https://img.shields.io/badge/Accuracy_Guard-green)\n- **ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning**  \n  *Shuaiyi Nie, Siyu Ding, Wenyuan Zhang, Linhao Yu, Tianmeng Yang, Yao Chen, Tingwen Liu, Weichong Yin, Yu Sun, Hua Wu*. [[pdf](https://arxiv.org/pdf/2602.09953)], 2026.02. ![](https://img.shields.io/badge/ACL2026-orange)\n- **Early Stopping for Large Reasoning Models via Confidence Dynamics**   \n  *Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn, Soheil Feizi*. [[pdf](https://arxiv.org/pdf/2604.04930)], [[code](https://github.com/sudoparsa/CoDE-Stop)], 2026.04. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/CoDE--Stop-blue) ![](https://img.shields.io/badge/Early_Exit-green)\n\n### Adaptive Thinking\n\n- **Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL**  \n  *Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, Dongbin Zhao*. [[pdf](https://arxiv.org/pdf/2505.10832)], [[code](https://github.com/TU2021/AutoThink)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange)\n- **AdaptThink: Reasoning Models Can Learn When to Think**  \n  *Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, Juanzi Li*. [[pdf](https://arxiv.org/pdf/2505.13417)], [[code](https://github.com/THU-KEG/AdaptThink)], 2025.05. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/AdaptThink-blue)\n- **Thinkless: LLM Learns When to Think**  \n  *Gongfan Fang, Xinyin Ma, Xinchao Wang*. [[pdf](https://arxiv.org/pdf/2505.13379)], [[code](https://github.com/VainF/Thinkless)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/Thinkless-blue)\n- **Think Only When You Need with Large Hybrid-Reasoning Models**  \n  *Lingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong, Xingxing Zhang, Tengchao Lv, Lei Cui, Furu Wei*. [[pdf](https://arxiv.org/pdf/2505.14631)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange)\n- **ThinkSwitcher: When to Think Hard, When to Think Fast**  \n  *Guosheng Liang, Longguang Zhong, Ziyi Yang, Xiaojun Quan*. [[pdf](https://arxiv.org/pdf/2505.14183)], 2025.05. ![](https://img.shields.io/badge/EMNLP2025--findings-orange) ![](https://img.shields.io/badge/ThinkSwitcher-blue)\n- **ARM: Adaptive Reasoning Model**  \n  *Siye Wu, Jian Xie, Yikai Zhang, Aili Chen, Kai Zhang, Yu Su, Yanghua Xiao*. [[pdf](https://arxiv.org/pdf/2505.20258)], [[homepage](https://team-arm.github.io/arm/)], [[code](https://github.com/TEAM-ARM/ARM)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/ARM-blue)\n- **When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning**  \n  *Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Haodong Zhao, Hao Li, Jiansong Chen, Ke Zeng, Xunliang Cai*. [[pdf](https://arxiv.org/pdf/2505.15400)], 2025.05. ![](https://img.shields.io/badge/EMNLP2025--findings-orange) ![](https://img.shields.io/badge/ASRR-blue)\n- **AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning**  \n  *Chenwei Lou, Zewei Sun, Xinnian Liang, Meng Qu, Wei Shen, Wenqi Wang, Yuntao Li, Qingping Yang, Shuangzhi Wu*. [[pdf](https://arxiv.org/pdf/2505.11896)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/AdaCoT-blue)\n- **AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models**  \n  *Feng Luo, Yu-Neng Chuang, Guanchu Wang, Hoang Anh Duy Le, Shaochen Zhong, Hongyi Liu, Jiayi Yuan, Yang Sui, Vladimir Braverman, Vipin Chaudhary, Xia Hu*. [[pdf](https://arxiv.org/pdf/2505.22662)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/AutoL2S-blue)\n- **OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation**  \n  *Shengjia Zhang, Junjie Wu, Jiawei Chen, Changwang Zhang, Xingyu Lou, Wangchunshu Zhou, Sheng Zhou, Can Wang, Jun Wang*. [[pdf](https://arxiv.org/pdf/2506.02397)], [[code](https://github.com/AgenticIR-Lab/OThink-R1)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/OThink--R1-blue)\n- **Long or short CoT? Investigating Instance-level Switch of Large Reasoning Models**  \n  *Ruiqi Zhang, Changyi Xiao, Yixin Cao*. [[pdf](https://arxiv.org/pdf/2506.04182)], 2025.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SwitchCoT-blue)\n- **Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models**  \n  *Peijie Liu, Fengli Xu, Yong Li*. [[pdf](https://arxiv.org/pdf/2506.06008)], [[code](https://github.com/tsinghua-fib-lab/Token_Signature)], 2025.06. ![](https://img.shields.io/badge/ICML2025-orange)\n- **SynapseRoute: An Auto-Route Switching Framework on Dual-State Large Language Model**  \n  *Wencheng Zhang, Shiqin Qiao, Lingjie Luo, Yinfeng Li, Chuanyang Zheng, Qian Xu, Meng Li, Yong Gui, Yijun He, Jianing Qiu, Jindong Hong, Jiankai Sun*. [[pdf](https://arxiv.org/pdf/2507.02822)], 2025.07. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/SynapseRoute-blue)\n- **Large Reasoning Models Know How to Think Efficiently**  \n  *Zeyu XING, Xing Li, Huiling Zhen, Xianzhi Yu, Mingxuan Yuan, Sinno Jialin Pan*. [[pdf](https://openreview.net/forum?id=pLKDeGm2t1)], 2025.07. ![](https://img.shields.io/badge/ESFoMoIII@ICML2025-orange) ![](https://img.shields.io/badge/SelfThink-blue)\n- **Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs**  \n  *Jaeseong Lee, Dayoung Kwon, seung-won hwang*. [[pdf](https://arxiv.org/pdf/2510.06750)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **MixReasoning: Switching Modes to Think**  \n  *Haiquan Lu, Gongfan Fang, Xinyin Ma, Qi Li, Xinchao Wang*. [[pdf](https://arxiv.org/pdf/2510.06052)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **When to Reason: Semantic Router for vLLM**  \n  *Chen Wang, Xunzhuo Liu, Yuhan Liu, Yue Zhu, Xiangxi Mo, Junchen Jiang, Huamin Chen*. [[pdf](https://arxiv.org/pdf/2510.08731)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference**  \n  *Xiang Liu, Xuming Hu, Xiaowen Chu, Eunsol Choi*. [[pdf](https://arxiv.org/pdf/2510.19669)], 2025.10. ![](https://img.shields.io/badge/Arxiv-orange)\n- **DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains**  \n  *Tian Liang, Wenxiang Jiao, Zhiwei He, Jiahao Xu, Haitao Mi, Dong Yu*. [[pdf](https://arxiv.org/pdf/2510.27419)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n- **MuTIS: Enhancing Reasoning Efficiency through Multi-Turn Intervention Sampling in Reinforcement Learning**  \n  *Wenshuo Zhao, Haoxing Zhai, Xinyu Qiu, Zhenting Qi, Shuhe Li, Linchao Zhu*. [[pdf](https://arxiv.org/pdf/2510.27419)], 2025.11. ![](https://img.shields.io/badge/EMNLP2025-orange)\n- **Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning**  \n  *Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto*. [[pdf](https://arxiv.org/pdf/2511.02130)], 2025.11. ![](https://img.shields.io/badge/Arxiv-orange)\n\n### Reasoning Shortcuts\n\n- **Not All Neuro-Symbolic Concepts Are Created Equal: Analysis and Mitigation of Reasoning Shortcuts**  \n  *Emanuele Marconato, Stefano Teso, Antonio Vergari, Andrea Passerini*. [[pdf](https://arxiv.org/pdf/2305.19951)], 2023.05. ![](https://img.shields.io/badge/NIPS2023-orange)\n- **Break the Chain: Large Language Models Can be Shortcut Reasoners**  \n  *Mengru Ding, Hanmeng Liu, Zhizhang Fu, Jian Song, Wenbo Xie, Yue Zhang*. [[pdf](https://arxiv.org/pdf/2406.06580)], 2024.06. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Break_the_Chain-blue)\n- **Can Language Models Learn to Skip Steps?**  \n  *Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang, Xipeng Qiu, Zheng Zhang*. [[pdf](https://arxiv.org/pdf/2411.01855)], [[code](https://github.com/tengxiaoliu/LM_skip)], 2024.11. ![](https://img.shields.io/badge/NIPS2024-orange) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **TokenSkip: Controllable Chain-of-Thought Compression in LLMs**  \n  *Heming Xia, Yongqi Li, Chak Tou Leong, Wenjie Wang, Wenjie Li*. [[pdf](https://arxiv.org/pdf/2502.12067)], [[code](https://github.com/hemingkx/TokenSkip)], 2025.02. ![](https://img.shields.io/badge/EMNLP2025-orange) ![](https://img.shields.io/badge/TokenSkip-blue) ![](https://img.shields.io/badge/Token_Shortcut-green)\n- **Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models**  \n  *Yingqian Cui, Pengfei He, Jingying Zeng, Hui Liu, Xianfeng Tang, Zhenwei Dai, Yan Han, Chen Luo, Jing Huang, Zhen Li, Suhang Wang, Yue Xing, Jiliang Tang, Qi He*. [[pdf](https://arxiv.org/pdf/2502.13260)], 2025.02. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping**  \n  *Ren Zhuang, Ben Wang, Shuifa Sun*. [[pdf](https://arxiv.org/pdf/2505.08392)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Token_Shortcut-green)\n- **DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models**  \n  *Yuxuan Jiang, Dawei Li, Frank Ferraro*. [[pdf](https://arxiv.org/pdf/2505.13975)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DRP-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search**  \n  *Yibo Wang, Li Shen, Huanjin Yao, Tiansheng Huang, Rui Liu, Naiqiang Tan, Jiaxing Huang, Kai Zhang, Dacheng Tao*. [[pdf](https://arxiv.org/pdf/2505.16838)], [[code](https://github.com/w-yibo/R1-Compress)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/R1--Compress-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Not All Tokens Are What You Need In Thinking**  \n  *Hang Yuan, Bin Yu, Haotian Li, Shijun Yang, Christina Dan Wang, Zhou Yu, Xueyin Xu, Weizhen Qi, Kai Chen*. [[pdf](https://arxiv.org/pdf/2505.17827)], [[code](https://github.com/Faustrazor/Not-All-Thinking-Tokens)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Token_Shortcut-green)\n- **LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling**  \n  *Yang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li, Pengfei Liu*. [[pdf](https://arxiv.org/pdf/2505.19187)], [[code](https://github.com/GAIR-NLP/LIMOPro)], 2025.05. ![](https://img.shields.io/badge/NeurIPS2025-orange) ![](https://img.shields.io/badge/LIMOPro-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models**  \n  *Sohyun An, Ruochen Wang, Tianyi Zhou, Cho-Jui Hsieh*. [[pdf](https://arxiv.org/pdf/2505.21765)], 2025.05. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/DTO-blue) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Compressing Chain-of-Thought in LLMs via Step Entropy**  \n  *Zeju Li, Jianyuan Zhong, Ziyang Zheng, Xiangyu Wen, Zhijian Xu, Yingying Cheng, Fan Zhang, Qiang Xu*. [[pdf](https://arxiv.org/pdf/2508.03346)], [[code](https://github.com/staymylove/COT_Compresstion_via_Step_entropy)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Pruning the Unsurprising: Efficient Code Reasoning via First-Token Surprisal**  \n  *Wenhao Zeng, Yaoning Wang, Chao Hu, Yuling Shi, Chengcheng Wan, Hongyu Zhang, Xiaodong Gu*. [[pdf](https://arxiv.org/pdf/2508.05988)], [[code](https://github.com/Zengwh02/ASAP)], 2025.08. ![](https://img.shields.io/badge/Arxiv-orange) ![](https://img.shields.io/badge/Step_Shortcut-green)\n- **Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models**  \n  *Canhui Wu, Qiong Cao, Chang Li, Zhenfang Wang, Chao Xue, Yuwei Fan, Wei Xi, X","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/hemingkx%2Fawesome-efficient-reasoning/projects"}