{"id":14848857,"url":"https://github.com/Furyton/awesome-language-model-analysis","last_synced_at":"2025-09-18T06:30:46.896Z","repository":{"id":238154332,"uuid":"791883093","full_name":"Furyton/awesome-language-model-analysis","owner":"Furyton","description":"This paper list focuses on the theoretical and empirical analysis of language models, especially large language models (LLMs). The papers in this list investigate the learning behavior, generalization ability, and other properties of language models through theoretical analysis, empirical analysis, or a combination of both.","archived":false,"fork":false,"pushed_at":"2024-09-18T11:12:08.000Z","size":622,"stargazers_count":31,"open_issues_count":1,"forks_count":0,"subscribers_count":3,"default_branch":"main","last_synced_at":"2024-09-18T17:16:57.352Z","etag":null,"topics":["ai","analysis","analytics","awesome","chatgpt","deep-learning","generative-ai","large-language-models","llm","nlp","theory","transformers"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc0-1.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Furyton.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-04-25T14:55:57.000Z","updated_at":"2024-09-17T14:48:25.000Z","dependencies_parsed_at":"2024-06-02T18:51:27.442Z","dependency_job_id":"6352bef3-a509-4b08-a036-4ff6c75561fc","html_url":"https://github.com/Furyton/awesome-language-model-analysis","commit_stats":null,"previous_names":["furyton/awesome-transformers-lm-analytics"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Furyton%2Fawesome-language-model-analysis","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Furyton%2Fawesome-language-model-analysis/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Furyton%2Fawesome-language-model-analysis/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Furyton%2Fawesome-language-model-analysis/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Furyton","download_url":"https://codeload.github.com/Furyton/awesome-language-model-analysis/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":219856522,"owners_count":16553571,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","analysis","analytics","awesome","chatgpt","deep-learning","generative-ai","large-language-models","llm","nlp","theory","transformers"],"created_at":"2024-09-19T14:00:42.159Z","updated_at":"2025-09-18T06:30:41.496Z","avatar_url":"https://github.com/Furyton.png","language":"Python","funding_links":[],"categories":["Other Papers","Other Lists","其他相关论文"],"sub_categories":["TeX Lists"],"readme":"# Awesome Language Model Analysis [![Awesome](https://awesome.re/badge-flat2.svg)](https://awesome.re)\n\nThis paper list focuses on the **theoretical and empirical analysis** of language models, especially **large language models** (LLMs).\nThe papers in this list investigate the learning behavior, generalization ability, and other properties of language models through theoretical analysis, empirical analysis, or a combination of both.\n\nScope of this list:\n- Currently, this list focuses on **transformer-based** models.\n- We hope to collect papers that only focus on the theoretical and empirical analysis of language models, instead of papers that aim to improve the performance of language models.\n\nLimitations of this list:\n- This list is not exhaustive, and we may miss some very important papers.\n- This list is not well-organized yet, and we may need to reorganize the list in the future.\n- Some popular topics are not well-covered yet, such as mechanistic engineering, probing, and interpretability.\n\nStatistics of This paper list:\n- Total number of different papers: **571**\n- For more detailed statistics, please refer to the end of this page.\n\nIf you have any suggestions or want to contribute, please feel free to open an issue or a pull request.\n\nFor details on how to contribute, please refer to the [contribution guidelines](CONTRIBUTING.md).\n\nYou can also share your thoughts and discuss with others in the [Discussions](https://github.com/Furyton/awesome-language-model-analysis/discussions).\n\n\u003e [!NOTE]  \n\u003e For uncategorized version, please refer to [here](README.uncategorized.md).\n\nTable of Content\n====================\n\u003c!--ts--\u003e\n- [Awesome Language Model Analysis](#awesome-language-model-analysis-)\n- [Table of Content](#table-of-content)\n  - [**Phenomena of Interest**](#phenomena-of-interest)\n    - [**In-Context Learning**](#in-context-learning)\n    - [**Chain-of-Thought**](#chain-of-thought)\n    - [**Hallucination**](#hallucination)\n    - [**Reversal Curse**](#reversal-curse)\n    - [**Scaling Laws / Emergent Abilities / Grokking / etc.**](#scaling-laws--emergent-abilities--grokking--etc)\n    - [**Knowledge / Memory Mechanisms**](#knowledge--memory-mechanisms)\n    - [**Training Dynamics / Landscape / Optimization / Fine-tuning / etc.**](#training-dynamics--landscape--optimization--fine-tuning--etc)\n    - [**Learning / Generalization / Reasoning / Weak to Strong Generalization**](#learning--generalization--reasoning--weak-to-strong-generalization)\n    - [**Other Phenomena / Discoveries**](#other-phenomena--discoveries)\n  - [**Representational Capacity**](#representational-capacity)\n    - [**What Can Transformer Do? / Properties of Transformer**](#what-can-transformer-do--properties-of-transformer)\n    - [**What Can Transformer Not Do? / Limitation of Transformer**](#what-can-transformer-not-do--limitation-of-transformer)\n  - [**Architectural Effectivity**](#architectural-effectivity)\n    - [**Layer-normalization**](#layer-normalization)\n    - [**Tokenization / Embedding**](#tokenization--embedding)\n    - [**Linear Attention / State Space Models / Recurrent Language Models / etc.**](#linear-attention--state-space-models--recurrent-language-models--etc)\n  - [**Training Paradigms**](#training-paradigms)\n  - [**Mechanistic Engineering / Probing / Interpretability**](#mechanistic-engineering--probing--interpretability)\n  - [**Miscellanea**](#miscellanea)\n\u003c!--te--\u003e\n\n\n## **Phenomena of Interest**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nCategories focusing on different phenomena, properties, and behaviors observed in large language models (LLMs) and transformer-based models.\n\n### **In-Context Learning**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers focusing on the theoretical and empirical analysis of in-context learning in large language models.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Transformers are Deep Optimizers: Provable In-Context Learning for Deep Model Training** [[paper link]](http://arxiv.org/abs/2411.16549) 2024-11-25  \nWeimin Wu; Maojiang Su; Jerry Yao-Chieh Hu; Zhao Song; Han Liu\n\n\n\n- **Can a Large Language Model Learn Matrix Functions In Context?** [[paper link]](http://arxiv.org/abs/2411.15675) 2024-11-24  \nPaimon Goulart; Evangelos E. Papalexakis\n\n\n\n- **Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models** [[paper link]](http://arxiv.org/abs/2410.09701) 2024-11-13  \nChengshuai Shi; Kun Yang; Jing Yang; Cong Shen\n\n\n\n- **Adversarial Robustness of In-Context Learning in Transformers for Linear Regression** [[paper link]](http://arxiv.org/abs/2411.05189) 2024-11-07  \nUsman Anwar; Johannes Von Oswald; Louis Kirsch; David Krueger; Spencer Frei\n\n\n\n- **Provable In-Context Learning with Transformers: A Case Study on Linear Regression** [[paper link]](http://arxiv.org/abs/2411.02199) 2024-11-04  \nDake Bu; Wei Huang; Andi Han; Atsushi Nitanda; Taiji Suzuki; Qingfu Zhang; Hau-San Wong\n\n\n\n- **Pretrained transformer efficiently learns low-dimensional target functions in-context** [[paper link]](http://arxiv.org/abs/2411.02544) 2024-11-04  \nKazusato Oko; Yujin Song; Taiji Suzuki; Denny Wu\n\n\n\n- **Toward Understanding In-context vs. In-weight Learning** [[paper link]](http://arxiv.org/abs/2410.23042) 2024-10-30  \nBryan Chan; Xinyi Chen; András György; Dale Schuurmans\n\n\n\n- **On the Role of Depth and Looping for In-Context Learning with Task Diversity** [[paper link]](http://arxiv.org/abs/2410.21698) 2024-10-29  \nKhashayar Gatmiry; Nikunj Saunshi; Sashank J. Reddi; Stefanie Jegelka; Sanjiv Kumar\n\n\n\n- **Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks** [[paper link]](http://arxiv.org/abs/2410.17498) 2024-10-23  \nPaul Smolensky; Roland Fernandez; Zhenghao Herbert Zhou; Mattia Opper; Jianfeng Gao\n\n\n\n- **Can Transformers In-Context Learn Behavior of a Linear Dynamical System?** [[paper link]](http://arxiv.org/abs/2410.16546) 2024-10-21  \nUsman Akram; Haris Vikalo\n\n\n\n- **Bayesian scaling laws for in-context learning** [[paper link]](http://arxiv.org/abs/2410.16531) 2024-10-21  \nAryaman Arora; Dan Jurafsky; Christopher Potts; Noah D. Goodman\n\n\n\n- **Provable In-context Learning for Mixture of Linear Regressions using Transformers** [[paper link]](http://arxiv.org/abs/2410.14183) 2024-10-18  \nYanhao Jin; Krishnakumar Balasubramanian; Lifeng Lai\n\n\n\n- **In-context learning and Occam's razor** [[paper link]](http://arxiv.org/abs/2410.14086) 2024-10-17  \nEric Elmoznino; Tom Marty; Tejas Kasetty; Leo Gagnon; Sarthak Mittal; Mahan Fathi; Dhanya Sridhar; Guillaume Lajoie\n\n\n\n- **Context-Scaling versus Task-Scaling in In-Context Learning** [[paper link]](http://arxiv.org/abs/2410.12783) 2024-10-16  \nAmirhesam Abedsoltan; Adityanarayanan Radhakrishnan; Jingfeng Wu; Mikhail Belkin\n\n\n\n- **Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent** [[paper link]](http://arxiv.org/abs/2410.11268) 2024-10-15  \nBo Chen; Xiaoyu Li; Yingyu Liang; Zhenmei Shi; Zhao Song\n\n\n\n- **How Transformers Implement Induction Heads: Approximation and Optimization Analysis** [[paper link]](http://arxiv.org/abs/2410.11474) 2024-10-15  \nMingze Wang; Ruoxi Yu; Weinan E; Lei Wu\n\n\n\n- **On the Training Convergence of Transformers for In-Context Classification** [[paper link]](http://arxiv.org/abs/2410.11475) 2024-10-15  \nWei Shen; Ruida Zhou; Jing Yang; Cong Shen\n\n\n\n- **Transformers learn variable-order Markov chains in-context** [[paper link]](http://arxiv.org/abs/2410.05493) 2024-10-07  \nRuida Zhou; Chao Tian; Suhas Diggavi\n\n\n\n- **Revisiting In-context Learning Inference Circuit in Large Language Models** [[paper link]](http://arxiv.org/abs/2410.04468) 2024-10-06  \nHakaze Cho; Mariko Kato; Yoshihiro Sakai; Naoya Inoue\n\n\n\n- **Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context** [[paper link]](http://arxiv.org/abs/2410.01774) 2024-10-02  \nSpencer Frei; Gal Vardi\n\n\n\n- **Transformers Handle Endogeneity in In-Context Linear Regression** [[paper link]](http://arxiv.org/abs/2410.01265) 2024-10-02  \nHaodong Liang; Krishnakumar Balasubramanian; Lifeng Lai\n\n\n\n- **Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers** [[paper link]](http://arxiv.org/abs/2409.10559) 2024-09-10  \nSiyu Chen; Heejune Sheen; Tianhao Wang; Zhuoran Yang\n\n\n\n- **Learning vs Retrieval: The Role of In-Context Examples in Regression with LLMs** [[paper link]](http://arxiv.org/abs/2409.04318) 2024-09-06  \nAliakbar Nafar; Kristen Brent Venable; Parisa Kordjamshidi\n\n\n\n- **Transformers are Minimax Optimal Nonparametric In-Context Learners** [[paper link]](http://arxiv.org/abs/2408.12186) 2024-08-22  \nJuno Kim; Tai Nakamaki; Taiji Suzuki\n\n\n\n- **Memorisation In In-Context Learning** [[paper link]](http://arxiv.org/abs/2408.11546) 2024-08-21  \nShahriar Golchin; Mihai Surdeanu; Steven Bethard; Eduardo Blanco; Ellen Riloff\n\n\n\n- **In-Context Learning with Representations: Contextual Generalization of Trained Transformers** [[paper link]](http://arxiv.org/abs/2408.10147) 2024-08-19  \nTong Yang; Yu Huang; Yingbin Liang; Yuejie Chi\n\n\n\n- **Fast Training Dataset Attribution via In-Context Learning** [[paper link]](http://arxiv.org/abs/2408.11852) 2024-08-14  \nMilad Fotouhi; Mohammad Taha Bahadori; Oluwaseyi Feyisetan; Payman Arabshahi; David Heckerman\n\n\n\n- **How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression** [[paper link]](http://arxiv.org/abs/2408.04532) 2024-08-08  \nXingwu Chen; Lei Zhao; Difan Zou\n\n\n\n- **Transformers are Universal In-context Learners** [[paper link]](http://arxiv.org/abs/2408.01367) 2024-08-02  \nTakashi Furuya; Maarten V. de Hoop; Gabriel Peyré\n\n\n\n- **Polynomial Regression as a Task for Understanding In-context Learning Through Finetuning and Alignment** [[paper link]](http://arxiv.org/abs/2407.19346) 2024-07-27  \nMax Wilcoxson; Morten Svendgård; Ria Doshi; Dylan Davis; Reya Vir; Anant Sahai\n\n\n\n- **Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism** [[paper link]](http://arxiv.org/abs/2407.17011) 2024-07-24  \nAnhao Zhao; Fanghua Ye; Jinlan Fu; Xiaoyu Shen\n\n\n\n- **One-Layer Transformer Provably Learns One-Nearest Neighbor In Context** [[paper link]](https://klusowski.princeton.edu/sites/g/files/toruqf5901/files/documents/li2024one.pdf) 2024-07-24  \nZihao Li; Yuan Cao; Cheng Gao; Yihan He; Han Liu; Jason M. Klusowski; Jianqing Fan; Mengdi Wang\n\n\n\n- **When can transformers compositionally generalize in-context?** [[paper link]](http://arxiv.org/abs/2407.12275) 2024-07-17  \nSeijin Kobayashi; Simon Schug; Yassir Akram; Florian Redhardt; Johannes von Oswald; Razvan Pascanu; Guillaume Lajoie; João Sacramento\n\n\n\n- **In-Context In-Context Learning with Transformer Neural Processes** [[paper link]](http://arxiv.org/abs/2406.13493) 2024-06-19  \nMatthew Ashman; Cristiana Diaconu; Adrian Weller; Richard E. Turner\n\n\n\n- **Probing the Decision Boundaries of In-context Learning in Large Language Models** [[paper link]](http://arxiv.org/abs/2406.11233) 2024-06-17  \nSiyan Zhao; Tung Nguyen; Aditya Grover\n\n\n\n- **State Soup: In-Context Skill Learning, Retrieval and Mixing** [[paper link]](http://arxiv.org/abs/2406.08423) 2024-06-12  \nMaciej Pióro; Maciej Wołczyk; Razvan Pascanu; Johannes von Oswald; João Sacramento\n\n\n\n- **Estimating the Hallucination Rate of Generative AI** [[paper link]](http://arxiv.org/abs/2406.07457) 2024-06-11  \nAndrew Jesson; Nicolas Beltran-Velez; Quentin Chu; Sweta Karlekar; Jannik Kossen; Yarin Gal; John P. Cunningham; David Blei\n\n\n\n- **BERTs are Generative In-Context Learners** [[paper link]](http://arxiv.org/abs/2406.04823) 2024-06-07  \nDavid Samuel\n\n\n\n- **Enhancing In-Context Learning Performance with just SVD-Based Weight Pruning: A Theoretical Perspective** [[paper link]](http://arxiv.org/abs/2406.03768) 2024-06-06  \nXinhao Yao; Xiaolin Hu; Shenzhi Yang; Yong Liu\n\n\n\n- **What Do Language Models Learn in Context? The Structured Task Hypothesis** [[paper link]](http://arxiv.org/abs/2406.04216) 2024-06-06  \nJiaoda Li; Yifan Hou; Mrinmaya Sachan; Ryan Cotterell\n\n\n\n- **Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers** [[paper link]](http://arxiv.org/abs/2406.02847) 2024-06-05  \nBrian K Chen; Tianyang Hu; Hui Jin; Hwee Kuan Lee; Kenji Kawaguchi\n\n\n\n- **Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks** [[paper link]](http://arxiv.org/abs/2406.02550) 2024-06-04  \nTianyu He; Darshil Doshi; Aritra Das; Andrey Gromov\n\n\n\n- **Why Larger Language Models Do In-context Learning Differently?** [[paper link]](http://arxiv.org/abs/2405.19592) 2024-05-30  \nZhenmei Shi; Junyi Wei; Zhuoyan Xu; Yingyu Liang\n\n\n\n- **Is In-Context Learning Sufficient for Instruction Following in LLMs?** [[paper link]](http://arxiv.org/abs/2405.19874) 2024-05-30  \nHao Zhao; Maksym Andriushchenko; Francesco Croce; Nicolas Flammarion\n\n\n\n- **Does learning the right latent variables necessarily improve in-context learning?** [[paper link]](http://arxiv.org/abs/2405.19162) 2024-05-29  \nSarthak Mittal; Eric Elmoznino; Leo Gagnon; Sangnie Bhardwaj; Dhanya Sridhar; Guillaume Lajoie\n\n\n\n- **A Theory of In-Context Learning in Transformers** [[paper link]](http://arxiv.org/abs/2405.18634) 2024-05-29  \nYifei Wang; Yuyang Wu; Zeming Wei; Stefanie Jegelka; Yisen Wang\n\n\n\n- **On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability** [[paper link]](http://arxiv.org/abs/2405.16845) 2024-05-27  \nChenyu Zheng; Wei Huang; Rongzhen Wang; Guoqiang Wu; Jun Zhu; Chongxuan Li\n\n\n\n- **Transformer In-Context Learning for Categorical Data** [[paper link]](http://arxiv.org/abs/2405.17248) 2024-05-27  \nAaron T. Wang; Ricardo Henao; Lawrence Carin\n\n\n\n- **Automatic Domain Adaptation by Transformers in In-Context Learning** [[paper link]](http://arxiv.org/abs/2405.16819) 2024-05-27  \nRyuichiro Hataya; Kota Matsui; Masaaki Imaizumi\n\n\n\n- **Unifying Demonstration Selection and Compression for In-Context Learning** [[paper link]](http://arxiv.org/abs/2405.17062) 2024-05-27  \nJun Gao\n\n\n\n- **On the Noise Robustness of In-Context Learning for Text Generation** [[paper link]](http://arxiv.org/abs/2405.17264) 2024-05-27  \nHongfu Gao; Feipeng Zhang; Wenyu Jiang; Jun Shu; Feng Zheng; Hongxin Wei\n\n\n\n- **MLPs Learn In-Context** [[paper link]](http://arxiv.org/abs/2405.15618) 2024-05-24  \nWilliam L. Tong; Cengiz Pehlevan\n\n\n\n- **Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification** [[paper link]](http://arxiv.org/abs/2405.15115) 2024-05-24  \nShang Liu; Zhongze Cai; Guanting Chen; Xiaocheng Li\n\n\n\n- **Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?** [[paper link]](https://openreview.net/pdf?id=o8AaRKbP9K) 2024-05-02  \nKhashayar Gatmiry; Nikunj Saunshi; Sashank J. Reddi; Stefanie Jegelka; Sanjiv Kumar\n\n\n\n- **In-context Learning on Function Classes Unveiled for Transformers** [[paper link]](https://openreview.net/pdf?id=rJkGOARXns) 2024-05-02  \nZhijie Wang; Bo Jiang; Shuai Li\n\n\n\n- **In-Context Learning with Long-Context Models: An In-Depth Exploration** [[paper link]](http://arxiv.org/abs/2405.00200) 2024-04-30  \nAmanda Bertsch; Maor Ivgi; Uri Alon; Jonathan Berant; Matthew R. Gormley; Graham Neubig\n\n\n\n- **What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation** [[paper link]](http://arxiv.org/abs/2404.07129) 2024-04-10  \nAaditya K. Singh; Ted Moskovitz; Felix Hill; Stephanie C. Y. Chan; Andrew M. Saxe\n\n\n\n- **Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability** [[paper link]](http://arxiv.org/abs/2310.08049) 2024-04-01  \nIvan Lee; Nan Jiang; Taylor Berg-Kirkpatrick\n\n\n\n- **Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality** [[paper link]](http://arxiv.org/abs/2402.19442) 2024-02-29  \nSiyu Chen; Heejune Sheen; Tianhao Wang; Zhuoran Yang\n\n\n\n- **How Transformers Learn Causal Structure with Gradient Descent** [[paper link]](http://arxiv.org/abs/2402.14735) 2024-02-22  \nEshaan Nichani; Alex Damian; Jason D. Lee\n\n\n\n- **In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization** [[paper link]](http://arxiv.org/abs/2402.14951) 2024-02-22  \nRuiqi Zhang; Jingfeng Wu; Peter L. Bartlett\n\n\n\n- **Identifying Semantic Induction Heads to Understand In-Context Learning** [[paper link]](http://arxiv.org/abs/2402.13055) 2024-02-20  \nJie Ren; Qipeng Guo; Hang Yan; Dongrui Liu; Xipeng Qiu; Dahua Lin\n\n\n\n- **How do Transformers perform In-Context Autoregressive Learning?** [[paper link]](http://arxiv.org/abs/2402.05787) 2024-02-08  \nMichael E. Sander; Raja Giryes; Taiji Suzuki; Mathieu Blondel; Gabriel Peyré\n\n\n\n- **Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks** [[paper link]](http://arxiv.org/abs/2402.04248) 2024-02-06  \nJongho Park; Jaeseung Park; Zheyang Xiong; Nayoung Lee; Jaewoong Cho; Samet Oymak; Kangwook Lee; Dimitris Papailiopoulos\n\n\n\n- **An Information-Theoretic Analysis of In-Context Learning** [[paper link]](http://arxiv.org/abs/2401.15530) 2024-01-28  \nHong Jun Jeon; Jason D. Lee; Qi Lei; Benjamin Van Roy\n\n\n\n- **The Transient Nature of Emergent In-Context Learning in Transformers** [[paper link]](http://arxiv.org/abs/2311.08360) 2023-12-11  \nAaditya K. Singh; Stephanie C. Y. Chan; Ted Moskovitz; Erin Grant; Andrew M. Saxe; Felix Hill\n\n\n\n- **In-Context Learning Functions with Varying Number of Minima** [[paper link]](http://arxiv.org/abs/2311.12538) 2023-11-21  \nDavid Oniani; Yanshan Wang\n\n\n\n- **Exploring the Relationship between In-Context Learning and Instruction Tuning** [[paper link]](http://arxiv.org/abs/2311.10367) 2023-11-17  \nHanyu Duan; Yixuan Tang; Yi Yang; Ahmed Abbasi; Kar Yan Tam\n\n\n\n- **When does In-context Learning Fall Short and Why? A Study on Specification-Heavy Tasks** [[paper link]](http://arxiv.org/abs/2311.08993) 2023-11-15  \nHao Peng; Xiaozhi Wang; Jianhui Chen; Weikai Li; Yunjia Qi; Zimu Wang; Zhili Wu; Kaisheng Zeng; Bin Xu; Lei Hou; Juanzi Li\n\n\n\n- **In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax** [[paper link]](http://arxiv.org/abs/2311.07811) 2023-11-13  \nAaron Mueller; Albert Webson; Jackson Petty; Tal Linzen\n\n\n\n- **Transformers learn to implement preconditioned gradient descent for in-context learning** [[paper link]](http://arxiv.org/abs/2306.00297) 2023-11-09  \nKwangjun Ahn; Xiang Cheng; Hadi Daneshmand; Suvrit Sra\n\n\n\n- **Transformers Learn Higher-Order Optimization Methods for In-Context Learning: A Study with Linear Models** [[paper link]](https://arxiv.org/abs/2310.17086v1) 2023-10-26  \nDeqing Fu; Tian-Qi Chen; Robin Jia; Vatsal Sharan\n\n\n\n- **In-Context Learning Creates Task Vectors** [[paper link]](http://arxiv.org/abs/2310.15916) 2023-10-24  \nRoee Hendel; Mor Geva; Amir Globerson\n\n\n\n- **Function Vectors in Large Language Models** [[paper link]](http://arxiv.org/abs/2310.15213) 2023-10-23  \nEric Todd; Millicent L. Li; Arnab Sen Sharma; Aaron Mueller; Byron C. Wallace; David Bau\n\n\n\n- **In-context Learning with Transformer Is Really Equivalent to a Contrastive Learning Pattern** [[paper link]](http://arxiv.org/abs/2310.13220) 2023-10-19  \nRuifeng Ren; Yong Liu\n\n\n\n- **Trained Transformers Learn Linear Models In-Context** [[paper link]](http://arxiv.org/abs/2306.09927) 2023-10-19  \nRuiqi Zhang; Spencer Frei; Peter L. Bartlett\n\n\n\n- **How Do Transformers Learn In-Context Beyond Simple Functions? A Case Study on Learning with Representations** [[paper link]](http://arxiv.org/abs/2310.10616) 2023-10-16  \nTianyu Guo; Wei Hu; Song Mei; Huan Wang; Caiming Xiong; Silvio Savarese; Yu Bai\n\n\n\n- **Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions** [[paper link]](https://openreview.net/forum?id=ekeyCgeRfC) 2023-10-13  \nSatwik Bhattamishra; Arkil Patel; Phil Blunsom; Varun Kanade\n\n\n\n- **How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?** [[paper link]](https://openreview.net/forum?id=vSh5ePa0ph) 2023-10-13  \nJingfeng Wu; Difan Zou; Zixiang Chen; Vladimir Braverman; Quanquan Gu; Peter Bartlett\n\n\n\n- **In-Context Learning Learns Label Relationships but Is Not Conventional Learning** [[paper link]](https://openreview.net/forum?id=YPIA7bgd5y) 2023-10-13  \nJannik Kossen; Yarin Gal; Tom Rainforth\n\n\n\n- **In-context Convergence of Transformers** [[paper link]](https://openreview.net/forum?id=kxpswbhr1r) 2023-10-13  \nYu Huang; Yuan Cheng; Yingbin Liang\n\n\n\n- **In-Context Learning through the Bayesian Prism** [[paper link]](https://openreview.net/forum?id=HX5ujdsSon) 2023-10-13  \nMadhur Panwar; Kabir Ahuja; Navin Goyal\n\n\n\n- **Do pretrained Transformers Really Learn In-context by Gradient Descent?** [[paper link]](http://arxiv.org/abs/2310.08540) 2023-10-12  \nLingfeng Shen; Aayush Mishra; Daniel Khashabi\n\n\n\n- **What and How does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and Generalization** [[paper link]](http://arxiv.org/abs/2305.19420) 2023-10-10  \nYufeng Zhang; Fengzhuo Zhang; Zhuoran Yang; Zhaoran Wang\n\n\n\n- **Explaining Emergent In-Context Learning as Kernel Regression** [[paper link]](http://arxiv.org/abs/2305.12766) 2023-10-05  \nChi Han; Ziqi Wang; Han Zhao; Heng Ji\n\n\n\n- **CausalLM is not optimal for in-context learning** [[paper link]](http://arxiv.org/abs/2308.06912) 2023-09-02  \nNan Ding; Tomer Levinboim; Jialin Wu; Sebastian Goodman; Radu Soricut\n\n\n\n- **One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention** [[paper link]](http://arxiv.org/abs/2307.03576) 2023-07-07  \nArvind Mahankali; Tatsunori B. Hashimoto; Tengyu Ma\n\n\n\n- **Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection** [[paper link]](http://arxiv.org/abs/2306.04637) 2023-07-06  \nYu Bai; Fan Chen; Huan Wang; Caiming Xiong; Song Mei\n\n\n\n- **Transformers Learn In-Context by Gradient Descent** [[paper link]](https://openreview.net/forum?id=tHvXrFQma5) 2023-06-15  \nJohannes Von Oswald; Eyvind Niklasson; Ettore Randazzo; Joao Sacramento; Alexander Mordvintsev; Andrey Zhmoginov; Max Vladymyrov\n\n\n\n- **The Closeness of In-Context Learning and Weight Shifting for Softmax Regression** [[paper link]](http://arxiv.org/abs/2304.13276) 2023-04-26  \nShuai Li; Zhao Song; Yu Xia; Tong Yu; Tianyi Zhou\n\n\n\n- **A Theory of Emergent In-Context Learning as Implicit Structure Induction** [[paper link]](http://arxiv.org/abs/2303.07971) 2023-03-14  \nMichael Hahn; Navin Goyal\n\n\n\n- **The Learnability of In-Context Learning** [[paper link]](http://arxiv.org/abs/2303.07895) 2023-03-14  \nNoam Wies; Yoav Levine; Amnon Shashua\n\n\n\n- **What Can Transformers Learn In-Context? A Case Study of Simple Function Classes** [[paper link]](http://arxiv.org/abs/2208.01066) 2023-01-14  \nShivam Garg; Dimitris Tsipras; Percy Liang; Gregory Valiant\n\n\n\n- **Transformers generalize differently from information stored in context vs in weights** [[paper link]](http://arxiv.org/abs/2210.05675) 2022-10-13  \nStephanie C. Y. Chan; Ishita Dasgupta; Junkyung Kim; Dharshan Kumaran; Andrew K. Lampinen; Felix Hill\n\n\n\n- **In-Context Learning and Induction Heads** [[paper link]](http://arxiv.org/abs/2209.11895) 2022-09-24  \nCatherine Olsson; Nelson Elhage; Neel Nanda; Nicholas Joseph; Nova DasSarma; Tom Henighan; Ben Mann; Amanda Askell; Yuntao Bai; Anna Chen; Tom Conerly; Dawn Drain; Deep Ganguli; Zac Hatfield-Dodds; Danny Hernandez; Scott Johnston; Andy Jones; Jackson Kernion; Liane Lovitt; Kamal Ndousse; Dario Amodei; Tom Brown; Jack Clark; Jared Kaplan; Sam McCandlish; Chris Olah\n\n\u003c/details\u003e\n\n\n### **Chain-of-Thought**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers analyzing the chain-of-thought phenomenon in large language models, exploring theoretical and empirical perspectives.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Rethinking Thinking Tokens: Understanding Why They Underperform in Practice** [[paper link]](http://arxiv.org/abs/2411.11371) 2024-11-18  \nSreeram Vennam; David Valente; David Herel; Ponnurangam Kumaraguru\n\n\n\n- **What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective** [[paper link]](http://arxiv.org/abs/2410.23743) 2024-10-31  \nMing Li; Yanhong Li; Tianyi Zhou\n\n\n\n- **A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration** [[paper link]](https://arxiv.org/abs/2410.16540v1) 2024-10-21  \nYingqian Cui; Pengfei He; Xianfeng Tang; Qi He; Chen Luo; Jiliang Tang; Yue Xing\n\n\n\n- **Transformers Provably Solve Parity Efficiently with Chain of Thought** [[paper link]](http://arxiv.org/abs/2410.08633) 2024-10-11  \nJuno Kim; Taiji Suzuki\n\n\n\n- **From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency** [[paper link]](http://arxiv.org/abs/2410.05459) 2024-10-07  \nKaiyue Wen; Huaqing Zhang; Hongzhou Lin; Jingzhao Zhang\n\n\n\n- **Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis** [[paper link]](http://arxiv.org/abs/2410.02167) 2024-10-03  \nHongkang Li; Meng Wang; Songtao Lu; Xiaodong Cui; Pin-Yu Chen\n\n\n\n- **Autoregressive + Chain of Thought (CoT) ≃ Recurrent: Recurrence's Role in Language Models and a Revist of Recurrent Transformer** [[paper link]](http://arxiv.org/abs/2409.09239) 2024-09-14  \nXiang Zhang; Muhammad Abdul-Mageed; Laks V.S. Lakshmanan\n\n\n\n- **Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods** [[paper link]](http://arxiv.org/abs/2408.14511) 2024-08-25  \nXinyang Hu; Fengzhuo Zhang; Siyu Chen; Zhuoran Yang\n\n\n\n- **Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning** [[paper link]](http://arxiv.org/abs/2407.01687) 2024-07-01  \nAkshara Prabhakar; Thomas L. Griffiths; R. Thomas McCoy\n\n\n\n- **On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning** [[paper link]](http://arxiv.org/abs/2406.14197) 2024-06-20  \nFranz Nowak; Anej Svete; Alexandra Butoi; Ryan Cotterell\n\n\n\n- **Iteration Head: A Mechanistic Study of Chain-of-Thought** [[paper link]](http://arxiv.org/abs/2406.02128) 2024-06-04  \nVivien Cabannes; Charles Arnal; Wassim Bouaziz; Alice Yang; Francois Charton; Julia Kempe\n\n\n\n- **Let's Think Dot by Dot: Hidden Computation in Transformer Language Models** [[paper link]](http://arxiv.org/abs/2404.15758) 2024-04-24  \nJacob Pfau; William Merrill; Samuel R. Bowman\n\n\n\n- **Chain of Thought Empowers Transformers to Solve Inherently Serial Problems** [[paper link]](http://arxiv.org/abs/2402.12875) 2024-02-20  \nZhiyuan Li; Hong Liu; Denny Zhou; Tengyu Ma\n\n\n\n- **Towards Revealing the Mystery behind Chain of Thought: A Theoretical Perspective** [[paper link]](http://arxiv.org/abs/2305.15408) 2023-12-22  \nGuhao Feng; Bohang Zhang; Yuntian Gu; Haotian Ye; Di He; Liwei Wang\n\n\n\n- **Why Can Large Language Models Generate Correct Chain-of-Thoughts?** [[paper link]](http://arxiv.org/abs/2310.13571) 2023-10-20  \nRasul Tutunov; Antoine Grosnit; Juliusz Ziomek; Jun Wang; Haitham Bou-Ammar\n\n\n\n- **How Large Language Models Implement Chain-of-Thought?** [[paper link]](https://openreview.net/forum?id=b2XfOm3RJa) 2023-10-13  \nYiqun Wang; Sile Hu; Yonggang Zhang; Xiang Tian; Xuesong Liu; Yaowu Chen; Xu Shen; Jieping Ye\n\n\n\n- **The Expressive Power of Transformers with Chain of Thought** [[paper link]](https://openreview.net/forum?id=NjNGlPh8Wh) 2023-10-13  \nWilliam Merrill; Ashish Sabharwal\n\n\u003c/details\u003e\n\n\n### **Hallucination**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers examining the hallucination phenomenon in language models, including both theoretical and empirical analysis.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse** [[paper link]](http://arxiv.org/abs/2411.09642) 2024-11-14  \nAlkis Kalavasis; Anay Mehrotra; Grigoris Velegkas\n\n\n\n- **No Free Lunch: Fundamental Limits of Learning Non-Hallucinating Generative Models** [[paper link]](http://arxiv.org/abs/2410.19217) 2024-10-24  \nChanglong Wu; Ananth Grama; Wojciech Szpankowski\n\n\n\n- **Shared Imagination: LLMs Hallucinate Alike** [[paper link]](http://arxiv.org/abs/2407.16604) 2024-07-23  \nYilun Zhou; Caiming Xiong; Silvio Savarese; Chien-Sheng Wu\n\n\n\n- **Estimating the Hallucination Rate of Generative AI** [[paper link]](http://arxiv.org/abs/2406.07457) 2024-06-11  \nAndrew Jesson; Nicolas Beltran-Velez; Quentin Chu; Sweta Karlekar; Jannik Kossen; Yarin Gal; John P. Cunningham; David Blei\n\n\n\n- **Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?** [[paper link]](http://arxiv.org/abs/2405.05904) 2024-05-09  \nZorik Gekhman; Gal Yona; Roee Aharoni; Matan Eyal; Amir Feder; Roi Reichart; Jonathan Herzig\n\n\n\n- **Mechanisms of non-factual hallucinations in language models** [[paper link]](http://arxiv.org/abs/2403.18167) 2024-03-26  \nLei Yu; Meng Cao; Jackie Chi Kit Cheung; Yue Dong\n\n\n\n- **Unfamiliar Finetuning Examples Control How Language Models Hallucinate** [[paper link]](http://arxiv.org/abs/2403.05612) 2024-03-08  \nKatie Kang; Eric Wallace; Claire Tomlin; Aviral Kumar; Sergey Levine\n\n\n\n- **In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation** [[paper link]](http://arxiv.org/abs/2403.01548) 2024-03-05  \nShiqi Chen; Miao Xiong; Junteng Liu; Zhengxuan Wu; Teng Xiao; Siyang Gao; Junxian He\n\n\n\n- **Calibrated Language Models Must Hallucinate** [[paper link]](http://arxiv.org/abs/2311.14648) 2023-11-24  \nAdam Tauman Kalai; Santosh S. Vempala\n\n\n\n- **The Curious Case of Hallucinatory Unanswerablity: Finding Truths in the Hidden States of Over-Confident Large Language Models** [[paper link]](http://arxiv.org/abs/2310.11877) 2023-10-18  \nAviv Slobodkin; Omer Goldman; Avi Caciularu; Ido Dagan; Shauli Ravfogel\n\n\u003c/details\u003e\n\n\n### **Reversal Curse**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers that analyze the reversal curse phenomenon in large language models.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics** [[paper link]](http://arxiv.org/abs/2405.04669) 2024-05-07  \nHanlin Zhu; Baihe Huang; Shaolun Zhang; Michael Jordan; Jiantao Jiao; Yuandong Tian; Stuart Russell\n\n\n\n- **The Reversal Curse: LLMs trained on \"A is B\" fail to learn \"B is A\"** [[paper link]](http://arxiv.org/abs/2309.12288) 2024-04-04  \nLukas Berglund; Meg Tong; Max Kaufmann; Mikita Balesni; Asa Cooper Stickland; Tomasz Korbak; Owain Evans\n\n\n\n- **An Investigation of LLMs' Inefficacy in Understanding Converse Relations** [[paper link]](https://aclanthology.org/2023.emnlp-main.429) 2023-12-01  \nChengwen Qi; Bowen Li; Binyuan Hui; Bailin Wang; Jinyang Li; Jinwang Wu; Yuanjun Laili\n\n\n\n- **Physics of Language Models: Part 3.2, Knowledge Manipulation** [[paper link]](http://arxiv.org/abs/2309.14402) 2023-09-25  \nZeyuan Allen-Zhu; Yuanzhi Li\n\n\n\n- **The Reversal Curse: Which Tokens You Predict Underlie the Factorization Curse and More** [[paper link]](http://arxiv.org/abs/2306.05183) 2023-06-07  \nOuail Kitouni; Niklas Nolte; Diane Bouchacourt; Adina Williams; Mike Rabbat; Mark Ibrahim\n\n\u003c/details\u003e\n\n\n### **Scaling Laws / Emergent Abilities / Grokking / etc.**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers exploring how model performance scales with model size, data size, or computational resources, and the emergence of unexpected abilities.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data** [[paper link]](http://arxiv.org/abs/2411.06646) 2024-11-11  \nAlex Havrilla; Wenjing Liao\n\n\n\n- **Scaling Laws for Precision** [[paper link]](http://arxiv.org/abs/2411.04330) 2024-11-07  \nTanishq Kumar; Zachary Ankner; Benjamin F. Spector; Blake Bordelon; Niklas Muennighoff; Mansheej Paul; Cengiz Pehlevan; Christopher Ré; Aditi Raghunathan\n\n\n\n- **Unlocking the Theory Behind Scaling 1-Bit Neural Networks** [[paper link]](http://arxiv.org/abs/2411.01663) 2024-11-03  \nMajid Daliri; Zhao Song; Chiwun Yang\n\n\n\n- **How Does Critical Batch Size Scale in Pre-training?** [[paper link]](http://arxiv.org/abs/2410.21676) 2024-10-29  \nHanlin Zhang; Depen Morwani; Nikhil Vyas; Jingfeng Wu; Difan Zou; Udaya Ghai; Dean Foster; Sham Kakade\n\n\n\n- **An Information Theory of Compute-Optimal Size Scaling, Emergence, and Plateaus in Language Models** [[paper link]](https://arxiv.org/abs/2410.01243) 2024-10-15  \nAnuj K. Nayak; Lav R. Varshney\n\n\n\n- **A Hitchhiker's Guide to Scaling Law Estimation** [[paper link]](http://arxiv.org/abs/2410.11840) 2024-10-15  \nLeshem Choshen; Yang Zhang; Jacob Andreas\n\n\n\n- **Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models** [[paper link]](https://arxiv.org/abs/2410.05661) 2024-10-08  \nSiqi Wang; Zhengyu Chen; Bei Li; Keqing He; Min Zhang; Jingang Wang\n\n\n\n- **Grokking at the Edge of Linear Separability** [[paper link]](https://arxiv.org/abs/2410.04489) 2024-10-06  \nAlon Beck; Noam Levi; Yohai Bar-Sinai\n\n\n\n- **An Empirical Study of Scaling Laws for Transfer** [[paper link]](https://arxiv.org/abs/2408.16947) 2024-08-30  \nMatthew Barnett\n\n\n\n- **A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language** [[paper link]](http://arxiv.org/abs/2408.12578) 2024-08-22  \nEkdeep Singh Lubana; Kyogo Kawaguchi; Robert P. Dick; Hidenori Tanaka\n\n\n\n- **Scaling Law with Learning Rate Annealing** [[paper link]](http://arxiv.org/abs/2408.11029) 2024-08-20  \nHowe Tissue; Venus Wang; Lu Wang\n\n\n\n- **Performance Law of Large Language Models** [[paper link]](http://arxiv.org/abs/2408.09895) 2024-08-19  \nChuhan Wu; Ruiming Tang\n\n\n\n- **Information-Theoretic Progress Measures reveal Grokking is an Emergent Phase Transition** [[paper link]](http://arxiv.org/abs/2408.08944) 2024-08-16  \nKenzo Clauw; Sebastiano Stramaglia; Daniele Marinazzo\n\n\n\n- **Large Language Monkeys: Scaling Inference Compute with Repeated Sampling** [[paper link]](http://arxiv.org/abs/2407.21787) 2024-07-31  \nBradley Brown; Jordan Juravsky; Ryan Ehrlich; Ronald Clark; Quoc V. Le; Christopher Ré; Azalia Mirhoseini\n\n\n\n- **Emergence in non-neural models: grokking modular arithmetic via average gradient outer product** [[paper link]](http://arxiv.org/abs/2407.20199) 2024-07-29  \nNeil Mallinar; Daniel Beaglehole; Libin Zhu; Adityanarayanan Radhakrishnan; Parthe Pandit; Mikhail Belkin\n\n\n\n- **Exploring Scaling Trends in LLM Robustness** [[paper link]](http://arxiv.org/abs/2407.18213) 2024-07-25  \nNikolaus Howe; Michał Zajac; Ian McKenzie; Oskar Hollinsworth; Tom Tseng; Pierre-Luc Bacon; Adam Gleave\n\n\n\n- **Understanding the Interplay of Scale, Data, and Bias in Language Models: A Case Study with BERT** [[paper link]](http://arxiv.org/abs/2407.21058) 2024-07-25  \nMuhammad Ali; Swetasudha Panda; Qinlan Shen; Michael Wick; Ari Kobren\n\n\n\n- **Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies** [[paper link]](http://arxiv.org/abs/2407.13623) 2024-07-18  \nChaofan Tao; Qian Liu; Longxu Dou; Niklas Muennighoff; Zhongwei Wan; Ping Luo; Min Lin; Ngai Wong\n\n\n\n- **Why Do You Grok? A Theoretical Analysis of Grokking Modular Addition** [[paper link]](http://arxiv.org/abs/2407.12332) 2024-07-17  \nMohamad Amin Mohamadi; Zhiyuan Li; Lei Wu; Danica J. Sutherland\n\n\n\n- **Predicting Emergent Capabilities by Finetuning** [[paper link]](https://openreview.net/pdf?id=vL8BIGuFTF) 2024-07-10  \nCharlie Victor Snell; Eric Wallace; Dan Klein; Sergey Levine\n\n\n\n- **Resolving Discrepancies in Compute-Optimal Scaling of Language Models** [[paper link]](http://arxiv.org/abs/2406.19146) 2024-06-25  \nTomer Porian; Mitchell Wortsman; Jenia Jitsev; Ludwig Schmidt; Yair Carmon\n\n\n\n- **Scaling Laws for Linear Complexity Language Models** [[paper link]](http://arxiv.org/abs/2406.16690) 2024-06-24  \nXuyang Shen; Dong Li; Ruitao Leng; Zhen Qin; Weigao Sun; Yiran Zhong\n\n\n\n- **Scaling Laws for Fact Memorization of Large Language Models** [[paper link]](http://arxiv.org/abs/2406.15720) 2024-06-22  \nXingyu Lu; Xiaonan Li; Qinyuan Cheng; Kai Ding; Xuanjing Huang; Xipeng Qiu\n\n\n\n- **Reconciling Kaplan and Chinchilla Scaling Laws** [[paper link]](http://arxiv.org/abs/2406.12907) 2024-06-12  \nTim Pearce; Jinyeop Song\n\n\n\n- **Deep Grokking: Would Deep Neural Networks Generalize Better?** [[paper link]](http://arxiv.org/abs/2405.19454) 2024-05-29  \nSimin Fan; Razvan Pascanu; Martin Jaggi\n\n\n\n- **Linguistic Collapse: Neural Collapse in (Large) Language Models** [[paper link]](https://arxiv.org/abs/2405.17767) 2024-05-28  \nRobert Wu; Vardan Papyan\n\n\n\n- **Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations** [[paper link]](http://arxiv.org/abs/2405.18392) 2024-05-28  \nAlexander Hägele; Elie Bakouch; Atli Kosson; Loubna Ben Allal; Leandro Von Werra; Martin Jaggi\n\n\n\n- **gzip Predicts Data-dependent Scaling Laws** [[paper link]](http://arxiv.org/abs/2405.16684) 2024-05-26  \nRohan Pandey\n\n\n\n- **Emergence of a High-Dimensional Abstraction Phase in Language Transformers** [[paper link]](http://arxiv.org/abs/2405.15471) 2024-05-24  \nEmily Cheng; Diego Doimo; Corentin Kervadec; Iuri Macocco; Jade Yu; Alessandro Laio; Marco Baroni\n\n\n\n- **A rationale from frequency perspective for grokking in training neural network** [[paper link]](http://arxiv.org/abs/2405.17479) 2024-05-24  \nZhangchen Zhou; Yaoyu Zhang; Zhi-Qin John Xu\n\n\n\n- **Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization** [[paper link]](http://arxiv.org/abs/2405.15071) 2024-05-23  \nBoshi Wang; Xiang Yue; Yu Su; Huan Sun\n\n\n\n- **Data Mixing Made Efficient: A Bivariate Scaling Law for Language Model Pretraining** [[paper link]](http://arxiv.org/abs/2405.14908) 2024-05-23  \nCe Ge; Zhijian Ma; Daoyuan Chen; Yaliang Li; Bolin Ding\n\n\n\n- **4+3 Phases of Compute-Optimal Neural Scaling Laws** [[paper link]](http://arxiv.org/abs/2405.15074) 2024-05-23  \nElliot Paquette; Courtney Paquette; Lechao Xiao; Jeffrey Pennington\n\n\n\n- **Slaves to the Law of Large Numbers: An Asymptotic Equipartition Property for Perplexity in Generative Language Models** [[paper link]](http://arxiv.org/abs/2405.13798) 2024-05-22  \nRaghu Mudumbai; Tyler Bell\n\n\n\n- **Quantifying Emergence in Large Language Models** [[paper link]](http://arxiv.org/abs/2405.12617) 2024-05-21  \nHang Chen; Xinyu Yang; Jiaying Zhu; Wenya Wang\n\n\n\n- **Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory** [[paper link]](http://arxiv.org/abs/2405.08707) 2024-05-14  \nXueyan Niu; Bo Bai; Lei Deng; Wei Han\n\n\n\n- **More Compute Is What You Need** [[paper link]](http://arxiv.org/abs/2404.19484) 2024-04-30  \nZhen Guo\n\n\n\n- **An exactly solvable model for emergence and scaling laws** [[paper link]](http://arxiv.org/abs/2404.17563) 2024-04-26  \nYoonsoo Nam; Nayara Fonseca; Seok Hyeong Lee; Ard Louis\n\n\n\n- **Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck** [[paper link]](http://arxiv.org/abs/2404.07647) 2024-04-11  \nNathan Godey; Éric de la Clergerie; Benoît Sagot\n\n\n\n- **A Large-Scale Exploration of $\\mu$-Transfer** [[paper link]](http://arxiv.org/abs/2404.05728) 2024-04-08  \nLucas Lingle\n\n\n\n- **Emergent Abilities in Reduced-Scale Generative Language Models** [[paper link]](http://arxiv.org/abs/2404.02204) 2024-04-02  \nSherin Muckatira; Vijeta Deshpande; Vladislav Lialin; Anna Rumshisky\n\n\n\n- **Understanding Emergent Abilities of Language Models from the Loss Perspective** [[paper link]](http://arxiv.org/abs/2403.15796) 2024-03-23  \nZhengxiao Du; Aohan Zeng; Yuxiao Dong; Jie Tang\n\n\n\n- **Unraveling the Mystery of Scaling Laws: Part I** [[paper link]](http://arxiv.org/abs/2403.06563) 2024-03-21  \nHui Su; Zhi Tian; Xiaoyu Shen; Xunliang Cai\n\n\n\n- **Language models scale reliably with over-training and on downstream tasks** [[paper link]](http://arxiv.org/abs/2403.08540) 2024-03-13  \nSamir Yitzhak Gadre; Georgios Smyrnis; Vaishaal Shankar; Suchin Gururangan; Mitchell Wortsman; Rulin Shao; Jean Mercat; Alex Fang; Jeffrey Li; Sedrick Keh; Rui Xin; Marianna Nezhurina; Igor Vasiljevic; Jenia Jitsev; Alexandros G. Dimakis; Gabriel Ilharco; Shuran Song; Thomas Kollar; Yair Carmon; Achal Dave; Reinhard Heckel; Niklas Muennighoff; Ludwig Schmidt\n\n\n\n- **When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method** [[paper link]](http://arxiv.org/abs/2402.17193) 2024-02-26  \nBiao Zhang; Zhongtao Liu; Colin Cherry; Orhan Firat\n\n\n\n- **Interpreting Grokked Transformers in Complex Modular Arithmetic** [[paper link]](https://arxiv.org/abs/2402.16726v2) 2024-02-26  \nHiroki Furuta; Gouki Minegishi; Yusuke Iwasawa; Yutaka Matsuo\n\n\n\n- **A Tale of Tails: Model Collapse as a Change of Scaling Laws** [[paper link]](https://arxiv.org/abs/2402.07043) 2024-02-10  \nElvis Dohmatob; Yunzhen Feng; Pu Yang; Francois Charton; Julia Kempe\n\n\n\n- **Scaling Data-Constrained Language Models** [[paper link]](http://arxiv.org/abs/2305.16264) 2023-10-25  \nNiklas Muennighoff; Alexander M. Rush; Boaz Barak; Teven Le Scao; Aleksandra Piktus; Nouamane Tazi; Sampo Pyysalo; Thomas Wolf; Colin Raffel\n\n\n\n- **The Cost of Down-Scaling Language Models: Fact Recall Deteriorates before In-Context Learning** [[paper link]](http://arxiv.org/abs/2310.04680) 2023-10-06  \nTian Jin; Nolan Clement; Xin Dong; Vaishnavh Nagarajan; Michael Carbin; Jonathan Ragan-Kelley; Gintare Karolina Dziugaite\n\n\n\n- **Are Emergent Abilities of Large Language Models a Mirage?** [[paper link]](https://arxiv.org/abs/2304.15004v2) 2023-04-28  \nRylan Schaeffer; Brando Miranda; Sanmi Koyejo\n\n\n\n- **Training Compute-Optimal Large Language Models** [[paper link]](http://arxiv.org/abs/2203.15556) 2022-03-29  \nJordan Hoffmann; Sebastian Borgeaud; Arthur Mensch; Elena Buchatskaya; Trevor Cai; Eliza Rutherford; Diego de Las Casas; Lisa Anne Hendricks; Johannes Welbl; Aidan Clark; Tom Hennigan; Eric Noland; Katie Millican; George van den Driessche; Bogdan Damoc; Aurelia Guy; Simon Osindero; Karen Simonyan; Erich Elsen; Jack W. Rae; Oriol Vinyals; Laurent Sifre\n\n\n\n- **Scaling Laws for Neural Language Models** [[paper link]](http://arxiv.org/abs/2001.08361) 2020-01-22  \nJared Kaplan; Sam McCandlish; Tom Henighan; Tom B. Brown; Benjamin Chess; Rewon Child; Scott Gray; Alec Radford; Jeffrey Wu; Dario Amodei\n\n\u003c/details\u003e\n\n\n### **Knowledge / Memory Mechanisms**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers focusing on how large language models store, retrieve, and utilize knowledge, analyzing the memory mechanisms involved.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **A Geometric Framework for Understanding Memorization in Generative Models** [[paper link]](http://arxiv.org/abs/2411.00113) 2024-10-31  \nBrendan Leigh Ross; Hamidreza Kamkari; Tongzi Wu; Rasa Hosseinzadeh; Zhaoyan Liu; George Stein; Jesse C. Cresswell; Gabriel Loaiza-Ganem\n\n\n\n- **Optimal Memorization Capacity of Transformers** [[paper link]](http://arxiv.org/abs/2409.17677) 2024-09-26  \nTokio Kajitsuka; Issei Sato\n\n\n\n- **Schrodingers Memory: Large Language Models** [[paper link]](https://arxiv.org/pdf/2409.10482) 2024-09-16  \nWei Wang; Qing Li\n\n\n\n- **Self-Attention Limits Working Memory Capacity of Transformer-Based Models** [[paper link]](http://arxiv.org/abs/2409.10715) 2024-09-16  \nDongyu Gong; Hantao Zhang\n\n\n\n- **Great Memory, Shallow Reasoning: Limits of kNN-LMs** [[paper link]](http://arxiv.org/abs/2408.11815) 2024-08-21  \nShangyi Geng; Wenting Zhao; Alexander M Rush\n\n\n\n- **Memorisation In In-Context Learning** [[paper link]](http://arxiv.org/abs/2408.11546) 2024-08-21  \nShahriar Golchin; Mihai Surdeanu; Steven Bethard; Eduardo Blanco; Ellen Riloff\n\n\n\n- **Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks** [[paper link]](http://arxiv.org/abs/2408.04965) 2024-08-09  \nVerna Dankers; Ivan Titov\n\n\n\n- **Understanding Memorisation in LLMs: Dynamics, Influencing Factors, and Implications** [[paper link]](http://arxiv.org/abs/2407.19262) 2024-07-27  \nTill Speicher; Mohammad Aflah Khan; Qinyuan Wu; Vedant Nanda; Soumi Das; Bishwamittra Ghosh; Krishna P. Gummadi; Evimaria Terzi\n\n\n\n- **Demystifying Verbatim Memorization in Large Language Models** [[paper link]](http://arxiv.org/abs/2407.17817) 2024-07-25  \nJing Huang; Diyi Yang; Christopher Potts\n\n\n\n- **From Internal Conflict to Contextual Adaptation of Language Models** [[paper link]](http://arxiv.org/abs/2407.17023) 2024-07-24  \nSara Vera Marjanović; Haeun Yu; Pepa Atanasova; Maria Maistro; Christina Lioma; Isabelle Augenstein\n\n\n\n- **Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data** [[paper link]](http://arxiv.org/abs/2407.14985) 2024-07-20  \nAntonis Antoniades; Xinyi Wang; Yanai Elazar; Alfonso Amayuelas; Alon Albalak; Kexun Zhang; William Yang Wang\n\n\n\n- **Physics of Language Models: Part 3.1, Knowledge Storage and Extraction** [[paper link]](http://arxiv.org/abs/2309.14316) 2024-07-16  \nZeyuan Allen-Zhu; Yuanzhi Li\n\n\n\n- **Induction Heads as an Essential Mechanism for Pattern Matching in In-context Learning** [[paper link]](http://arxiv.org/abs/2407.07011) 2024-07-09  \nJ. Crosbie; E. Shutova\n\n\n\n- **Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers** [[paper link]](http://arxiv.org/abs/2406.18400) 2024-06-26  \nYibo Jiang; Goutham Rajendran; Pradeep Ravikumar; Bryon Aragam\n\n\n\n- **Scaling Laws for Fact Memorization of Large Language Models** [[paper link]](http://arxiv.org/abs/2406.15720) 2024-06-22  \nXingyu Lu; Xiaonan Li; Qinyuan Cheng; Kai Ding; Xuanjing Huang; Xipeng Qiu\n\n\n\n- **Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data** [[paper link]](http://arxiv.org/abs/2406.14546) 2024-06-20  \nJohannes Treutlein; Dami Choi; Jan Betley; Cem Anil; Samuel Marks; Roger Baker Grosse; Owain Evans\n\n\n\n- **Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Large Language Models** [[paper link]](http://arxiv.org/abs/2406.14549) 2024-06-20  \nSunny Duan; Mikail Khona; Abhiram Iyer; Rylan Schaeffer; Ila R Fiete\n\n\n\n- **Understanding Finetuning for Factual Knowledge Extraction** [[paper link]](http://arxiv.org/abs/2406.14785) 2024-06-20  \nGaurav Ghosal; Tatsunori Hashimoto; Aditi Raghunathan\n\n\n\n- **Estimating Knowledge in Large Language Models Without Generating a Single Token** [[paper link]](http://arxiv.org/abs/2406.12673) 2024-06-18  \nDaniela Gottesman; Mor Geva\n\n\n\n- **How Do Large Language Models Acquire Factual Knowledge During Pretraining?** [[paper link]](http://arxiv.org/abs/2406.11813) 2024-06-17  \nHoyeon Chang; Jinho Park; Seonghyeon Ye; Sohee Yang; Youngkyung Seo; Du-Seong Chang; Minjoon Seo\n\n\n\n- **Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs** [[paper link]](http://arxiv.org/abs/2406.10209) 2024-06-14  \nAbhimanyu Hans; Yuxin Wen; Neel Jain; John Kirchenbauer; Hamid Kazemi; Prajwal Singhania; Siddharth Singh; Gowthami Somepalli; Jonas Geiping; Abhinav Bhatele; Tom Goldstein\n\n\n\n- **Knowledge Circuits in Pretrained Transformers** [[paper link]](http://arxiv.org/abs/2405.17969) 2024-05-28  \nYunzhi Yao; Ningyu Zhang; Zekun Xi; Mengru Wang; Ziwen Xu; Shumin Deng; Huajun Chen\n\n\n\n- **Upper and lower memory capacity bounds of transformers for next-token prediction** [[paper link]](http://arxiv.org/abs/2405.13718) 2024-05-22  \nLiam Madden; Curtis Fox; Christos Thrampoulidis\n\n\n\n- **A Multi-Perspective Analysis of Memorization in Large Language Models** [[paper link]](http://arxiv.org/abs/2405.11577) 2024-05-19  \nBowen Chen; Namgi Han; Yusuke Miyao\n\n\n\n- **Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws** [[paper link]](http://arxiv.org/abs/2404.05405) 2024-04-08  \nZeyuan Allen-Zhu; Yuanzhi Li\n\n\n\n- **Memorization Capacity of Multi-Head Attention in Transformers** [[paper link]](http://arxiv.org/abs/2306.02010) 2024-03-02  \nSadegh Mahdavi; Renjie Liao; Christos Thrampoulidis\n\n\n\n- **Birth of a Transformer: A Memory Viewpoint** [[paper link]](http://arxiv.org/abs/2306.00802) 2023-11-06  \nAlberto Bietti; Vivien Cabannes; Diane Bouchacourt; Herve Jegou; Leon Bottou\n\n\n\n- **Physics of Language Models: Part 3.2, Knowledge Manipulation** [[paper link]](http://arxiv.org/abs/2309.14402) 2023-09-25  \nZeyuan Allen-Zhu; Yuanzhi Li\n\n\n\n- **Can Neural Network Memorization Be Localized?** [[paper link]](http://arxiv.org/abs/2307.09542) 2023-07-18  \nPratyush Maini; Michael C. Mozer; Hanie Sedghi; Zachary C. Lipton; J. Zico Kolter; Chiyuan Zhang\n\n\n\n- **Quantifying Memorization Across Neural Language Models** [[paper link]](http://arxiv.org/abs/2202.07646) 2022-02-15  \nNicholas Carlini; Daphne Ippolito; Matthew Jagielski; Katherine Lee; Florian Tramer; Chiyuan Zhang\n\n\u003c/details\u003e\n\n\n### **Training Dynamics / Landscape / Optimization / Fine-tuning / etc.**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers discussing various aspects of the training process, including optimization, fine-tuning, and the training landscape of large language models.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Gradient dynamics for low-rank fine-tuning beyond kernels** [[paper link]](http://arxiv.org/abs/2411.15385) 2024-11-23  \nArif Kerem Dayi; Sitan Chen\n\n\n\n- **Unraveling the Gradient Descent Dynamics of Transformers** [[paper link]](http://arxiv.org/abs/2411.07538) 2024-11-12  \nBingqing Song; Boran Han; Shuai Zhang; Jie Ding; Mingyi Hong\n\n\n\n- **What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?** [[paper link]](http://arxiv.org/abs/2411.07681) 2024-11-12  \nKatie Kang; Amrith Setlur; Dibya Ghosh; Jacob Steinhardt; Claire Tomlin; Sergey Levine; Aviral Kumar\n\n\n\n- **Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis** [[paper link]](http://arxiv.org/abs/2410.09605) 2024-11-12  \nHongru Yang; Bhavya Kailkhura; Zhangyang Wang; Yingbin Liang\n\n\n\n- **Global Convergence in Training Large-Scale Transformers** [[paper link]](http://arxiv.org/abs/2410.23610) 2024-10-31  \nCheng Gao; Yuan Cao; Zihao Li; Yihan He; Mengdi Wang; Han Liu; Jason Matthew Klusowski; Jianqing Fan\n\n\n\n- **What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective** [[paper link]](http://arxiv.org/abs/2410.23743) 2024-10-31  \nMing Li; Yanhong Li; Tianyi Zhou\n\n\n\n- **Learning and Transferring Sparse Contextual Bigrams with Linear Transformers** [[paper link]](http://arxiv.org/abs/2410.23438) 2024-10-30  \nYunwei Ren; Zixuan Wang; Jason D. Lee\n\n\n\n- **Abrupt Learning in Transformers: A Case Study on Matrix Completion** [[paper link]](http://arxiv.org/abs/2410.22244) 2024-10-29  \nPulkit Gopalani; Ekdeep Singh Lubana; Wei Hu\n\n\n\n- **LoRA vs Full Fine-tuning: An Illusion of Equivalence** [[paper link]](http://arxiv.org/abs/2410.21228) 2024-10-28  \nReece Shuttleworth; Jacob Andreas; Antonio Torralba; Pratyusha Sharma\n\n\n\n- **A distributional simplicity bias in the learning dynamics of transformers** [[paper link]](http://arxiv.org/abs/2410.19637) 2024-10-25  \nRiccardo Rende; Federica Gerace; Alessandro Laio; Sebastian Goldt\n\n\n\n- **Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs** [[paper link]](http://arxiv.org/abs/2410.13835) 2024-10-17  \nTianyu Guo; Druv Pai; Yu Bai; Jiantao Jiao; Michael I. Jordan; Song Mei\n\n\n\n- **How Transformers Implement Induction Heads: Approximation and Optimization Analysis** [[paper link]](http://arxiv.org/abs/2410.11474) 2024-10-15  \nMingze Wang; Ruoxi Yu; Weinan E; Lei Wu\n\n\n\n- **What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis** [[paper link]](http://arxiv.org/abs/2410.10986) 2024-10-14  \nWeronika Ormaniec; Felix Dangel; Sidak Pal Singh\n\n\n\n- **Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?** [[paper link]](http://arxiv.org/abs/2410.05581) 2024-10-08  \nFırat Öncel; Matthias Bethge; Beyza Ermis; Mirco Ravanelli; Cem Subakan; Çağatay Yıldız\n\n\n\n- **On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent** [[paper link]](http://arxiv.org/abs/2410.04870) 2024-10-07  \nBingrui Li; Wei Huang; Andi Han; Zhanpeng Zhou; Taiji Suzuki; Jun Zhu; Jianfei Chen\n\n\n\n- **Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective** [[paper link]](http://arxiv.org/abs/2410.05192) 2024-10-07  \nKaiyue Wen; Zhiyuan Li; Jason Wang; David Hall; Percy Liang; Tengyu Ma\n\n\n\n- **Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis** [[paper link]](http://arxiv.org/abs/2410.02167) 2024-10-03  \nHongkang Li; Meng Wang; Songtao Lu; Xiaodong Cui; Pin-Yu Chen\n\n\n\n- **Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization** [[paper link]](http://arxiv.org/abs/2410.02247) 2024-10-03  \nXinhao Yao; Hongjin Qian; Xiaolin Hu; Gengze Xu; Yong Liu\n\n\n\n- **Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context** [[paper link]](http://arxiv.org/abs/2410.01774) 2024-10-02  \nSpencer Frei; Gal Vardi\n\n\n\n- **Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective** [[paper link]](http://arxiv.org/abs/2410.01720) 2024-10-02  \nZeyu Gan; Yong Liu\n\n\n\n- **Investigating the Impact of Model Complexity in Large Language Models** [[paper link]](http://arxiv.org/abs/2410.00699) 2024-10-01  \nJing Luo; Huiyuan Wang; Weiran Huang\n\n\n\n- **Benigh or Not-Benign Overfitting in Token Selection of Attention Mechanism** [[paper link]](http://arxiv.org/abs/2409.17625) 2024-09-26  \nKeitaro Sakamoto; Issei Sato\n\n\n\n- **Non-asymptotic Convergence of Training Transformers for Next-token Prediction** [[paper link]](http://arxiv.org/abs/2409.17335) 2024-09-25  \nRuiquan Huang; Yingbin Liang; Jing Yang\n\n\n\n- **Optimization Hyper-parameter Laws for Large Language Models** [[paper link]](http://arxiv.org/abs/2409.04777) 2024-09-07  \nXingyu Xie; Kuangyu Ding; Shuicheng Yan; Kim-Chuan Toh; Tianwen Wei\n\n\n\n- **The AdEMAMix Optimizer: Better, Faster, Older** [[paper link]](http://arxiv.org/abs/2409.03137) 2024-09-05  \nMatteo Pagliardini; Pierre Ablin; David Grangier\n\n\n\n- **Clustering and Alignment: Understanding the Training Dynamics in Modular Addition** [[paper link]](http://arxiv.org/abs/2408.09414) 2024-08-18  \nTiberiu Musat\n\n\n\n- **Global Convergence in Training Large-Scale Transformers** [[paper link]](https://klusowski.princeton.edu/sites/g/files/toruqf5901/files/documents/gao2024global.pdf) 2024-08  \nCheng Gao; Yuan Cao; Zihao Li; Yihan He; Mengdi Wang; Han Liu; Jason M. Klusowski; Jianqing Fan\n\n\n\n- **On the Convergence of Encoder-only Shallow Transformers** [[paper link]](https://proceedings.neurips.cc/paper_files/paper/2023/file/a3cf318fbeec1126da21e9185ae9908c-Paper-Conference.pdf) 2024-08  \nYongtao Wu; Fanghui Liu; Grigorios G Chrysos; Volkan Cevher\n\n\n\n- **Parameter-Efficient Fine-Tuning for Continual Learning: A Neural Tangent Kernel Perspective** [[paper link]](http://arxiv.org/abs/2407.17120) 2024-07-24  \nJingren Liu; Zhong Ji; YunLong Yu; Jiale Cao; Yanwei Pang; Jungong Han; Xuelong Li\n\n\n\n- **Learning Dynamics of LLM Finetuning** [[paper link]](http://arxiv.org/abs/2407.10490) 2024-07-15  \nYi Ren; Danica J. Sutherland\n\n\n\n- **Deconstructing What Makes a Good Optimizer for Language Models** [[paper link]](http://arxiv.org/abs/2407.07972) 2024-07-10  \nRosie Zhao; Depen Morwani; David Brandfonbrener; Nikhil Vyas; Sham Kakade\n\n\n\n- **Zero-Shot Generalization during Instruction Tuning: Insights from Similarity and Granularity** [[paper link]](http://arxiv.org/abs/2406.11721) 2024-06-17  \nBingxiang He; Ning Ding; Cheng Qian; Jia Deng; Ganqu Cui; Lifan Yuan; Huan-ang Gao; Huimin Chen; Zhiyuan Liu; Maosong Sun\n\n\n\n- **Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective** [[paper link]](http://arxiv.org/abs/2405.16747) 2024-05-27  \nAkiyoshi Tomihari; Issei Sato\n\n\n\n- **Infinite Limits of Multi-head Transformer Dynamics** [[paper link]](http://arxiv.org/abs/2405.15712) 2024-05-24  \nBlake Bordelon; Hamza Tahir Chaudhry; Cengiz Pehlevan\n\n\n\n- **Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics** [[paper link]](http://arxiv.org/abs/2405.04669) 2024-05-07  \nHanlin Zhu; Baihe Huang; Shaolun Zhang; Michael Jordan; Jiantao Jiao; Yuandong Tian; Stuart Russell\n\n\n\n- **Control Theoretic Approach to Fine-Tuning and Transfer Learning** [[paper link]](http://arxiv.org/abs/2404.11013) 2024-04-16  \nErkan Bayram; Shenyu Liu; Mohamed-Ali Belabbas; Tamer Başar\n\n\n\n- **Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think** [[paper link]](http://arxiv.org/abs/2404.08382) 2024-04-12  \nXinpeng Wang; Chengzhi Hu; Bolei Ma; Paul Röttger; Barbara Plank\n\n\n\n- **On Training Data Influence of GPT Models** [[paper link]](http://arxiv.org/abs/2404.07840) 2024-04-11  \nQingyi Liu; Yekun Chai; Shuohuan Wang; Yu Sun; Keze Wang; Hua Wu\n\n\n\n- **Best Practices and Lessons Learned on Synthetic Data for Language Models** [[paper link]](http://arxiv.org/abs/2404.07503) 2024-04-11  \nRuibo Liu; Jerry Wei; Fangyu Liu; Chenglei Si; Yanzhe Zhang; Jinmeng Rao; Steven Zheng; Daiyi Peng; Diyi Yang; Denny Zhou; Andrew M. Dai\n\n\n\n- **How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse** [[paper link]](http://arxiv.org/abs/2404.05090) 2024-04-07  \nMohamed El Amine Seddik; Suei-Wen Chen; Soufiane Hayou; Pierre Youssef; Merouane Debbah\n\n\n\n- **Unveiling the Generalization Power of Fine-Tuned Large Language Models** [[paper link]](http://arxiv.org/abs/2403.09162) 2024-03-14  \nHaoran Yang; Yumeng Zhang; Jiaqi Xu; Hongyuan Lu; Pheng Ann Heng; Wai Lam\n\n\n\n- **Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models** [[paper link]](http://arxiv.org/abs/2403.09635) 2024-03-14  \nAkhil Kedia; Mohd Abbas Zaidi; Sushil Khyalia; Jungho Jung; Harshith Goka; Haejun Lee\n\n\n\n- **Linear Attention is (Maybe) All You Need (to Understand Transformer Optimization)** [[paper link]](http://arxiv.org/abs/2310.01082) 2024-03-13  \nKwangjun Ahn; Xiang Cheng; Minhak Song; Chulhee Yun; Ali Jadbabaie; Suvrit Sra\n\n\n\n- **Hallmarks of Optimization Trajectories in Neural Networks and LLMs: The Lengths, Bends, and Dead Ends** [[paper link]](http://arxiv.org/abs/2403.07379) 2024-03-12  \nSidak Pal Singh; Bobby He; Thomas Hofmann; Bernhard Schölkopf\n\n\n\n- **The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models** [[paper link]](http://arxiv.org/abs/2403.03942) 2024-03-06  \nAdithya Bhaskar; Dan Friedman; Danqi Chen\n\n\n\n- **Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality** [[paper link]](http://arxiv.org/abs/2402.19442) 2024-02-29  \nSiyu Chen; Heejune Sheen; Tianhao Wang; Zhuoran Yang\n\n\n\n- **How Transformers Learn Causal Structure with Gradient Descent** [[paper link]](http://arxiv.org/abs/2402.14735) 2024-02-22  \nEshaan Nichani; Alex Damian; Jason D. Lee\n\n\n\n- **LoRA Training in the NTK Regime has No Spurious Local Minima** [[paper link]](http://arxiv.org/abs/2402.11867) 2024-02-19  \nUijeong Jang; Jason D. Lee; Ernest K. Ryu\n\n\n\n- **On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm** [[paper link]](http://arxiv.org/abs/2402.03660) 2024-02-06  \nZhanpeng Zhou; Zijun Chen; Yilan Chen; Bo Zhang; Junchi Yan\n\n\n\n- **Transformers learn through gradual rank increase** [[paper link]](http://arxiv.org/abs/2306.07042) 2023-12-10  \nEnric Boix-Adsera; Etai Littwin; Emmanuel Abbe; Samy Bengio; Joshua Susskind\n\n\n\n- **Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks** [[paper link]](http://arxiv.org/abs/2311.12786) 2023-11-21  \nSamyak Jain; Robert Kirk; Ekdeep Singh Lubana; Robert P. Dick; Hidenori Tanaka; Edward Grefenstette; Tim Rocktäschel; David Scott Krueger\n\n\n\n- **Connecting Pre-trained Language Model and Downstream Task via Properties of Representation** [[paper link]](https://openreview.net/forum?id=YLOJ4aKAka) 2023-11-02  \nChenwei Wu; Holden Lee; Rong Ge\n\n\n\n- **Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer** [[paper link]](http://arxiv.org/abs/2305.16380) 2023-07-02  \nYuandong Tian; Yiping Wang; Beidi Chen; Simon Du\n\n\n\n- **A Kernel-Based View of Language Model Fine-Tuning** [[paper link]](https://openreview.net/forum?id=49dTFIGdx8) 2023-06-15  \nSadhika Malladi; Alexander Wettig; Dingli Yu; Danqi Chen; Sanjeev Arora\n\n\n\n- **A Stability Analysis of Fine-Tuning a Pre-Trained Model** [[paper link]](https://arxiv.org/abs/2301.09820v2) 2023-01-24  \nZihao Fu; Anthony Man-Cho So; Nigel Collier\n\n\u003c/details\u003e\n\n\n### **Learning / Generalization / Reasoning / Weak to Strong Generalization**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers analyzing the learning capabilities and generalization performance of language models, from weak to strong generalization.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models** [[paper link]](http://arxiv.org/abs/2411.17182) 2024-11-26  \nYunzhe Hu; Difan Zou; Dong Xu\n\n\n\n- **What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?** [[paper link]](http://arxiv.org/abs/2411.07681) 2024-11-12  \nKatie Kang; Amrith Setlur; Dibya Ghosh; Jacob Steinhardt; Claire Tomlin; Sergey Levine; Aviral Kumar\n\n\n\n- **Generalization and Risk Bounds for Recurrent Neural Networks** [[paper link]](http://arxiv.org/abs/2411.02784) 2024-11-05  \nXuewei Cheng; Ke Huang; Shujie Ma\n\n\n\n- **Provable Length Generalization in Sequence Prediction via Spectral Filtering** [[paper link]](http://arxiv.org/abs/2411.01035) 2024-11-01  \nAnnie Marsden; Evan Dogariu; Naman Agarwal; Xinyi Chen; Daniel Suo; Elad Hazan\n\n\n\n- **RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner** [[paper link]](http://arxiv.org/abs/2410.23912) 2024-10-31  \nFu-Chieh Chang; Yu-Ting Lee; Hui-Ying Shih; Pei-Yuan Wu\n\n\n\n- **Mixture of Parrots: Experts improve memorization more than reasoning** [[paper link]](http://arxiv.org/abs/2410.19034) 2024-10-24  \nSamy Jelassi; Clara Mohri; David Brandfonbrener; Alex Gu; Nikhil Vyas; Nikhil Anand; David Alvarez-Melis; Yuanzhi Li; Sham M. Kakade; Eran Malach\n\n\n\n- **How Numerical Precision Affects Mathematical Reasoning Capabilities of LLMs** [[paper link]](http://arxiv.org/abs/2410.13857) 2024-10-17  \nGuhao Feng; Kai Yang; Yuntian Gu; Xinyue Ai; Shengjie Luo; Jiacheng Sun; Di He; Zhenguo Li; Liwei Wang\n\n\n\n- **On Rank-Dependent Generalisation Error Bounds for Transformers** [[paper link]](http://arxiv.org/abs/2410.11500) 2024-10-15  \nLan V. Truong\n\n\n\n- **Benign Overfitting in Single-Head Attention** [[paper link]](http://arxiv.org/abs/2410.07746) 2024-10-10  \nRoey Magen; Shuning Shang; Zhiwei Xu; Spencer Frei; Wei Hu; Gal Vardi\n\n\n\n- **Dynamics of Concept Learning and Compositional Generalization** [[paper link]](http://arxiv.org/abs/2410.08309) 2024-10-10  \nYongyi Yang; Core Francisco Park; Ekdeep Singh Lubana; Maya Okawa; Wei Hu; Hidenori Tanaka\n\n\n\n- **Benign Overfitting for Regression with Trained Two-Layer ReLU Networks** [[paper link]](http://arxiv.org/abs/2410.06191) 2024-10-08  \nJunhyung Park; Patrick Bloebaum; Shiva Prasad Kasiviswanathan\n\n\n\n- **Provable Weak-to-Strong Generalization via Benign Overfitting** [[paper link]](http://arxiv.org/abs/2410.04638) 2024-10-06  \nDavid X. Wu; Anant Sahai\n\n\n\n- **A Formal Framework for Understanding Length Generalization in Transformers** [[paper link]](http://arxiv.org/abs/2410.02140) 2024-10-03  \nXinting Huang; Andy Yang; Satwik Bhattamishra; Yash Sarrof; Andreas Krebs; Hattie Zhou; Preetum Nakkiran; Michael Hahn\n\n\n\n- **Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context** [[paper link]](http://arxiv.org/abs/2410.01774) 2024-10-02  \nSpencer Frei; Gal Vardi\n\n\n\n- **Lines of Thought in Large Language Models** [[paper link]](http://arxiv.org/abs/2410.01545) 2024-10-02  \nRaphaël Sarfati; Toni J. B. Liu; Nicolas Boullé; Christopher J. Earls\n\n\n\n- **Investigating the Impact of Model Complexity in Large Language Models** [[paper link]](http://arxiv.org/abs/2410.00699) 2024-10-01  \nJing Luo; Huiyuan Wang; Weiran Huang\n\n\n\n- **Benign or Not-Benign Overfitting in Token Selection of Attention Mechanism** [[paper link]](http://arxiv.org/abs/2409.17625) 2024-09-26  \nKeitaro Sakamoto; Issei Sato\n\n\n\n- **Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics** [[paper link]](http://arxiv.org/abs/2409.09626) 2024-09-15  \nYi Ren; Danica J. Sutherland\n\n\n\n- **Unforgettable Generalization in Language Models** [[paper link]](http://arxiv.org/abs/2409.02228) 2024-09-03  \nEric Zhang; Leshem Chosen; Jacob Andreas\n\n\n\n- **The Many Faces of Optimal Weak-to-Strong Learning** [[paper link]](http://arxiv.org/abs/2408.17148) 2024-08-30  \nMikael Møller Høgsgaard; Kasper Green Larsen; Markus Engelund Mathiasen\n\n\n\n- **Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems** [[paper link]](http://arxiv.org/abs/2408.16293) 2024-08-29  \nTian Ye; Zicheng Xu; Yuanzhi Li; Zeyuan Allen-Zhu\n\n\n\n- **Out-of-distribution generalization via composition: a lens through induction heads in Transformers** [[paper link]](http://arxiv.org/abs/2408.09503) 2024-08-18  \nJiajun Song; Zhuoyan Xu; Yiqiao Zhong\n\n\n\n- **On the Generalization of Preference Learning with DPO** [[paper link]](http://arxiv.org/abs/2408.03459) 2024-08-06  \nShawn Im; Yixuan Li\n\n\n\n- **Inductive or Deductive? Rethinking the Fundamental Reasoning Abilities of LLMs** [[paper link]](http://arxiv.org/abs/2408.00114) 2024-07-31  \nKewei Cheng; Jingfeng Yang; Haoming Jiang; Zhengyang Wang; Binxuan Huang; Ruirui Li; Shiyang Li; Zheng Li; Yifan Gao; Xian Li; Bing Yin; Yizhou Sun\n\n\n\n- **Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process** [[paper link]](http://arxiv.org/abs/2407.13123) 2024-07-29  \nTian Ye; Zicheng Xu; Yuanzhi Li; Zeyuan Allen-Zhu\n\n\n\n- **Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models** [[paper link]](http://arxiv.org/abs/2407.18158) 2024-07-25  \nSanae Lotfi; Yilun Kuang; Brandon Amos; Micah Goldblum; Marc Finzi; Andrew Gordon Wilson\n\n\n\n- **On Initialization of Transformers with Pre-trained Embeddings** [[paper link]](http://arxiv.org/abs/2407.12514) 2024-07-17  \nHa Young Kim; Niranjan Balasubramanian; Byungkon Kang\n\n\n\n- **When can transformers compositionally generalize in-context?** [[paper link]](http://arxiv.org/abs/2407.12275) 2024-07-17  \nSeijin Kobayashi; Simon Schug; Yassir Akram; Florian Redhardt; Johannes von Oswald; Razvan Pascanu; Guillaume Lajoie; João Sacramento\n\n\n\n- **Reasoning in Large Language Models: A Geometric Perspective** [[paper link]](http://arxiv.org/abs/2407.02678) 2024-07-02  \nRomain Cosentino; Sarath Shekkizhar\n\n\n\n- **Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis** [[paper link]](http://arxiv.org/abs/2406.17167) 2024-06-24  \nHongkang Li; Meng Wang; Shuai Zhang; Sijia Liu; Pin-Yu Chen\n\n\n\n- **How Truncating Weights Improves Reasoning in Language Models** [[paper link]](http://arxiv.org/abs/2406.03068) 2024-06-05  \nLei Chen; Joan Bruna; Alberto Bietti\n\n\n\n- **Understanding Transformer Reasoning Capabilities via Graph Algorithms** [[paper link]](http://arxiv.org/abs/2405.18512) 2024-05-28  \nClayton Sanford; Bahare Fatemi; Ethan Hall; Anton Tsitsulin; Mehran Kazemi; Jonathan Halcrow; Bryan Perozzi; Vahab Mirrokni\n\n\n\n- **Linguistic Collapse: Neural Collapse in (Large) Language Models** [[paper link]](https://arxiv.org/abs/2405.17767) 2024-05-28  \nRobert Wu; Vardan Papyan\n\n\n\n- **Reality Only Happens Once: Single-Path Generalization Bounds for Transformers** [[paper link]](http://arxiv.org/abs/2405.16563) 2024-05-26  \nYannick Limmer; Anastasis Kratsios; Xuwei Yang; Raeid Saqur; Blanka Horvath\n\n\n\n- **A statistical framework for weak-to-strong generalization** [[paper link]](http://arxiv.org/abs/2405.16236) 2024-05-25  \nSeamus Somerstep; Felipe Maia Polo; Moulinath Banerjee; Ya'acov Ritov; Mikhail Yurochkin; Yuekai Sun\n\n\n\n- **Theoretical Analysis of Weak-to-Strong Generalization** [[paper link]](http://arxiv.org/abs/2405.16043) 2024-05-25  \nHunter Lang; David Sontag; Aravindan Vijayaraghavan\n\n\n\n- **Quantifying the Gain in Weak-to-Strong Generalization** [[paper link]](http://arxiv.org/abs/2405.15116) 2024-05-24  \nMoses Charikar; Chirag Pabbaraju; Kirankumar Shiragur\n\n\n\n- **Towards Understanding How Transformer Perform Multi-step Reasoning with Matching Operation** [[paper link]](http://arxiv.org/abs/2405.15302) 2024-05-24  \nZhiwei Wang; Yunji Wang; Zhongwang Zhang; Zhangchen Zhou; Hui Jin; Tianyang Hu; Jiacheng Sun; Zhenguo Li; Yaoyu Zhang; Zhi-Qin John Xu\n\n\n\n- **Initialization is Critical to Whether Transformers Fit Composite Functions by Inference or Memorizing** [[paper link]](http://arxiv.org/abs/2405.05409) 2024-05-08  \nZhongwang Zhang; Pengxiao Lin; Zhiwei Wang; Yaoyu Zhang; Zhi-Qin John Xu\n\n\n\n- **On the Empirical Complexity of Reasoning and Planning in LLMs** [[paper link]](http://arxiv.org/abs/2404.11041) 2024-04-17  \nLiwei Kang; Zirui Zhao; David Hsu; Wee Sun Lee\n\n\n\n- **When can transformers reason with abstract symbols?** [[paper link]](http://arxiv.org/abs/2310.09753) 2024-04-16  \nEnric Boix-Adsera; Omid Saremi; Emmanuel Abbe; Samy Bengio; Etai Littwin; Joshua Susskind\n\n\n\n- **A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task** [[paper link]](http://arxiv.org/abs/2402.11917) 2024-02-19  \nJannik Brinkmann; Abhay Sheshadri; Victor Levoso; Paul Swoboda; Christian Bartelt\n\n\n\n- **Provably learning a multi-head attention layer** [[paper link]](http://arxiv.org/abs/2402.04084) 2024-02-06  \nSitan Chen; Yuanzhi Li\n\n\n\n- **Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks** [[paper link]](http://arxiv.org/abs/2311.12786) 2023-11-21  \nSamyak Jain; Robert Kirk; Ekdeep Singh Lubana; Robert P. Dick; Hidenori Tanaka; Edward Grefenstette; Tim Rocktäschel; David Scott Krueger\n\n\n\n- **The Impact of Depth and Width on Transformer Language Model Generalization** [[paper link]](http://arxiv.org/abs/2310.19956) 2023-10-30  \nJackson Petty; Sjoerd van Steenkiste; Ishita Dasgupta; Fei Sha; Dan Garrette; Tal Linzen\n\n\n\n- **Implicit meta-learning may lead language models to trust more reliable sources** [[paper link]](http://arxiv.org/abs/2310.15047) 2023-10-23  \nDmitrii Krasheninnikov; Egor Krasheninnikov; Bruno Mlodozeniec; Tegan Maharaj; David Krueger\n\n\n\n- **On the Optimization and Generalization of Multi-head Attention** [[paper link]](http://arxiv.org/abs/2310.12680) 2023-10-19  \nPuneesh Deora; Rouzbeh Ghaderi; Hossein Taheri; Christos Thrampoulidis\n\n\n\n- **Large Language Models Cannot Self-Correct Reasoning Yet** [[paper link]](https://openreview.net/forum?id=IkmD3fKBPQ) 2023-10-13  \nJie Huang; Xinyun Chen; Swaroop Mishra; Huaixiu Steven Zheng; Adams Wei Yu; Xinying Song; Denny Zhou\n\n\n\n- **How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition** [[paper link]](http://arxiv.org/abs/2310.05492) 2023-10-09  \nGuanting Dong; Hongyi Yuan; Keming Lu; Chengpeng Li; Mingfeng Xue; Dayiheng Liu; Wei Wang; Zheng Yuan; Chang Zhou; Jingren Zhou\n\n\n\n- **A Theory for Emergence of Complex Skills in Language Models** [[paper link]](http://arxiv.org/abs/2307.15936) 2023-07-29  \nSanjeev Arora; Anirudh Goyal\n\n\n\n- **On the Power of Foundation Models** [[paper link]](https://proceedings.mlr.press/v202/yuan23b.html) 2023-07-03  \nYang Yuan\n\n\n\n- **Task-Specific Skill Localization in Fine-tuned Language Models** [[paper link]](https://openreview.net/forum?id=Rgnaj43Pk0) 2023-06-15  \nAbhishek Panigrahi; Nikunj Saunshi; Haoyu Zhao; Sanjeev Arora\n\n\n\n- **Towards Understanding Why Mask-Reconstruction Pretraining Helps in Downstream Tasks** [[paper link]](http://arxiv.org/abs/2206.03826) 2023-02-11  \nJiachun Pan; Pan Zhou; Shuicheng Yan\n\n\n\n- **Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models** [[paper link]](http://arxiv.org/abs/2210.14199) 2022-10-25  \nHong Liu; Sang Michael Xie; Zhiyuan Li; Tengyu Ma\n\n\n\n- **Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt Tuning** [[paper link]](http://arxiv.org/abs/2106.09226) 2022-04-20  \nColin Wei; Sang Michael Xie; Tengyu Ma\n\n\n\n- **A Mathematical Exploration of Why Language Models Help Solve Downstream Tasks** [[paper link]](http://arxiv.org/abs/2010.03648) 2021-04-14  \nNikunj Saunshi; Sadhika Malladi; Sanjeev Arora\n\n\n\n- **Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning** [[paper link]](http://arxiv.org/abs/2012.13255) 2020-12-22  \nArmen Aghajanyan; Luke Zettlemoyer; Sonal Gupta\n\n\n\n- **How fine can fine-tuning be? Learning efficient language models** [[paper link]](https://proceedings.mlr.press/v108/radiya-dixit20a.html) 2020-06-03  \nEvani Radiya-Dixit; Xin Wang\n\n\u003c/details\u003e\n\n\n### **Other Phenomena / Discoveries**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers discussing other interesting phenomena or discoveries related to the behavior and properties of language models.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **On the loss of context-awareness in general instruction fine-tuning** [[paper link]](http://arxiv.org/abs/2411.02688) 2024-11-05  \nYihan Wang; Andrew Bai; Nanyun Peng; Cho-Jui Hsieh\n\n\n\n- **Weight decay induces low-rank attention layers** [[paper link]](http://arxiv.org/abs/2410.23819) 2024-10-31  \nSeijin Kobayashi; Yassir Akram; Johannes Von Oswald\n\n\n\n- **All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling** [[paper link]](http://arxiv.org/abs/2410.23501) 2024-10-30  \nEmanuele Marconato; Sébastien Lachapelle; Sebastian Weichwald; Luigi Gresele\n\n\n\n- **Looking Beyond The Top-1: Transformers Determine Top Tokens In Order** [[paper link]](http://arxiv.org/abs/2410.20210) 2024-10-26  \nDaria Lioubashevski; Tomer Schlank; Gabriel Stanovsky; Ariel Goldstein\n\n\n\n- **Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs** [[paper link]](http://arxiv.org/abs/2410.13835) 2024-10-17  \nTianyu Guo; Druv Pai; Yu Bai; Jiantao Jiao; Michael I. Jordan; Song Mei\n\n\n\n- **Emergent properties with repeated examples** [[paper link]](http://arxiv.org/abs/2410.07041) 2024-10-09  \nFrançois Charton; Julia Kempe\n\n\n\n- **Masked Mixers for Language Generation and Retrieval** [[paper link]](http://arxiv.org/abs/2409.01482) 2024-09-02  \nBenjamin L. Badger\n\n\n\n- **Monotonic Representation of Numeric Properties in Language Models** [[paper link]](http://arxiv.org/abs/2408.10381) 2024-08-15  \nBenjamin Heinzerling; Kentaro Inui\n\n\n\n- **Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models** [[paper link]](http://arxiv.org/abs/2408.06518) 2024-08-12  \nHila Gonen; Terra Blevins; Alisa Liu; Luke Zettlemoyer; Noah A. Smith\n\n\n\n- **Large Language Monkeys: Scaling Inference Compute with Repeated Sampling** [[paper link]](http://arxiv.org/abs/2407.21787) 2024-07-31  \nBradley Brown; Jordan Juravsky; Ryan Ehrlich; Ronald Clark; Quoc V. Le; Christopher Ré; Azalia Mirhoseini\n\n\n\n- **Transformers on Markov Data: Constant Depth Suffices** [[paper link]](http://arxiv.org/abs/2407.17686) 2024-07-25  \nNived Rajaraman; Marco Bondaschi; Kannan Ramchandran; Michael Gastpar; Ashok Vardhan Makkuva\n\n\n\n- **On the Benefits of Rank in Attention Layers** [[paper link]](http://arxiv.org/abs/2407.16153) 2024-07-23  \nNoah Amsel; Gilad Yehudai; Joan Bruna\n\n\n\n- **Transformer Alignment in Large Language Models** [[paper link]](http://arxiv.org/abs/2407.07810) 2024-07-10  \nMurdock Aubry; Haoming Meng; Anton Sugolov; Vardan Papyan\n\n\n\n- **Understanding Transformers via N-gram Statistics** [[paper link]](http://arxiv.org/abs/2407.12034) 2024-06-30  \nTimothy Nguyen\n\n\n\n- **Large Vocabulary Size Improves Large Language Models** [[paper link]](http://arxiv.org/abs/2406.16508) 2024-06-24  \nSho Takase; Ryokan Ri; Shun Kiyono; Takuya Kato\n\n\n\n- **Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data** [[paper link]](http://arxiv.org/abs/2406.14546) 2024-06-20  \nJohannes Treutlein; Dami Choi; Jan Betley; Cem Anil; Samuel Marks; Roger Baker Grosse; Owain Evans\n\n\n\n- **Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning** [[paper link]](http://arxiv.org/abs/2406.13858) 2024-06-19  \nYuval Shalev; Amir Feder; Ariel Goldstein\n\n\n\n- **Transcendence: Generative Models Can Outperform The Experts That Train Them** [[paper link]](http://arxiv.org/abs/2406.11741) 2024-06-17  \nEdwin Zhang; Vincent Zhu; Naomi Saphra; Anat Kleiman; Benjamin L. Edelman; Milind Tambe; Sham M. Kakade; Eran Malach\n\n\n\n- **Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens** [[paper link]](http://arxiv.org/abs/2406.10985) 2024-06-16  \nWeiyao Luo; Suncong Zheng; Heming Xia; Weikang Wang; Yan Lei; Tianyu Liu; Shuang Chen; Zhifang Sui\n\n\n\n- **Anisotropy is Not Inherent to Transformers** [[paper link]](https://aclanthology.org/2024.naacl-long.274) 2024-06  \nAnemily Machina; Robert Mercer\n\n\n\n- **Linguistic Collapse: Neural Collapse in (Large) Language Models** [[paper link]](https://arxiv.org/abs/2405.17767) 2024-05-28  \nRobert Wu; Vardan Papyan\n\n\n\n- **Exploring Activation Patterns of Parameters in Language Models** [[paper link]](https://arxiv.org/abs/2405.17799) 2024-05-28  \nYudong Wang; Damai Dai; Zhifang Sui\n\n\n\n- **Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs** [[paper link]](https://arxiv.org/abs/2405.16700) 2024-05-26  \nMustafa Shukor; Matthieu Cord\n\n\n\n- **Your Transformer is Secretly Linear** [[paper link]](https://arxiv.org/abs/2405.12250) 2024-05-19  \nAnton Razzhigaev; Matvey Mikhalchuk; Elizaveta Goncharova; Nikolai Gerasimenko; Ivan Oseledets; Denis Dimitrov; Andrey Kuznetsov\n\n\n\n- **The Platonic Representation Hypothesis** [[paper link]](https://arxiv.org/abs/2405.07987v1) 2024-05-13  \nMinyoung Huh; Brian Cheung; Tongzhou Wang; Phillip Isola\n\n\n\n- **By Tying Embeddings You Are Assuming the Distributional Hypothesis** [[paper link]](https://openreview.net/pdf?id=yyYMAprcAR) 2024-05-02  \nFrancesco Bertolotti; Walter Cazzola\n\n\n\n- **Emergent Representations of Program Semantics in Language Models Trained on Programs** [[paper link]](https://openreview.net/pdf?id=8PTx4CpNoT) 2024-05-02  \nCharles Jin; Martin Rinard\n\n\n\n- **Algorithmic progress in language models** [[paper link]](http://arxiv.org/abs/2403.05812) 2024-03-09  \nAnson Ho; Tamay Besiroglu; Ege Erdil; David Owen; Robi Rahman; Zifan Carl Guo; David Atkinson; Neil Thompson; Jaime Sevilla\n\n\n\n- **Massive Activations in Large Language Models** [[paper link]](http://arxiv.org/abs/2402.17762) 2024-02-27  \nMingjie Sun; Xinlei Chen; J. Zico Kolter; Zhuang Liu\n\n\n\n- **On the Emergence of Cross-Task Linearity in the Pretraining-Finetuning Paradigm** [[paper link]](http://arxiv.org/abs/2402.03660) 2024-02-06  \nZhanpeng Zhou; Zijun Chen; Yilan Chen; Bo Zhang; Junchi Yan\n\n\n\n- **Anisotropy Is Inherent to Self-Attention in Transformers** [[paper link]](http://arxiv.org/abs/2406.12143) 2024-01-24  \nNathan Godey; Éric de la Clergerie; Benoît Sagot\n\n\n\n- **The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers** [[paper link]](https://openreview.net/forum?id=TJ2nxciYCk-) 2023-02-01  \nZonglin Li; Chong You; Srinadh Bhojanapalli; Daliang Li; Ankit Singh Rawat; Sashank J. Reddi; Ke Ye; Felix Chern; Felix Yu; Ruiqi Guo; Sanjiv Kumar\n\n\u003c/details\u003e\n\n\n## **Representational Capacity**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nCategories focused on the representational capacities and limitations of transformers and language models.\n\n### **What Can Transformer Do? / Properties of Transformer**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers providing positive results into the capabilities and properties of transformer-based models, e.g., expressiveness and learning abilities.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency** [[paper link]](http://arxiv.org/abs/2411.16525) 2024-11-25  \nJerry Yao-Chieh Hu; Wei-Po Wang; Ammar Gilani; Chenyang Li; Zhao Song; Han Liu\n\n\n\n- **Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers** [[paper link]](http://arxiv.org/abs/2411.12118) 2024-11-18  \nTiberiu Musat\n\n\n\n- **Measure-to-measure interpolation using Transformers** [[paper link]](http://arxiv.org/abs/2411.04551) 2024-11-07  \nBorjan Geshkovski; Philippe Rigollet; Domènec Ruiz-Balet\n\n\n\n- **Ask, and it shall be given: Turing completeness of prompting** [[paper link]](http://arxiv.org/abs/2411.01992) 2024-11-04  \nRuizhong Qiu; Zhe Xu; Wenxuan Bao; Hanghang Tong\n\n\n\n- **Provable Optimal Transport with Transformers: The Essence of Depth and Prompt Engineering** [[paper link]](http://arxiv.org/abs/2410.19931) 2024-10-25  \nHadi Daneshmand\n\n\n\n- **On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery** [[paper link]](http://arxiv.org/abs/2410.13981) 2024-10-17  \nRenpu Liu; Ruida Zhou; Cong Shen; Jing Yang\n\n\n\n- **Theoretical Analysis of Hierarchical Language Recognition and Generation by Transformers without Positional Encoding** [[paper link]](http://arxiv.org/abs/2410.12413) 2024-10-16  \nDaichi Hayakawa; Issei Sato\n\n\n\n- **Memory-augmented Transformers can implement Linear First-Order Optimization Methods** [[paper link]](http://arxiv.org/abs/2410.07263) 2024-10-08  \nSanchayan Dutta; Suvrit Sra\n\n\n\n- **Transformers are Efficient Compilers, Provably** [[paper link]](http://arxiv.org/abs/2410.14706) 2024-10-07  \nXiyu Zhai; Runlong Zhou; Liao Zhang; Simon Shaolei Du\n\n\n\n- **Fundamental Limitations on Subquadratic Alternatives to Transformers** [[paper link]](http://arxiv.org/abs/2410.04271) 2024-10-05  \nJosh Alman; Hantao Yu\n\n\n\n- **Autoregressive Large Language Models are Computationally Universal** [[paper link]](http://arxiv.org/abs/2410.03170) 2024-10-04  \nDale Schuurmans; Hanjun Dai; Francesco Zanini\n\n\n\n- **Can Transformers Learn n-gram Language Models?** [[paper link]](http://arxiv.org/abs/2410.03001) 2024-10-03  \nAnej Svete; Nadav Borenstein; Mike Zhou; Isabelle Augenstein; Ryan Cotterell\n\n\n\n- **Towards Understanding the Universality of Transformers for Next-Token Prediction** [[paper link]](http://arxiv.org/abs/2410.03011) 2024-10-03  \nMichael E. Sander; Gabriel Peyré\n\n\n\n- **Large Language Models as Markov Chains** [[paper link]](http://arxiv.org/abs/2410.02724) 2024-10-03  \nOussama Zekri; Ambroise Odonnat; Abdelhakim Benechehab; Linus Bleistein; Nicolas Boullé; Ievgen Redko\n\n\n\n- **On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding** [[paper link]](http://arxiv.org/abs/2410.01405) 2024-10-02  \nKevin Xu; Issei Sato\n\n\n\n- **Attention layers provably solve single-location regression** [[paper link]](http://arxiv.org/abs/2410.01537) 2024-10-02  \nPierre Marion; Raphaël Berthier; Gérard Biau; Claire Boyer\n\n\n\n- **Transformers in Uniform TC0** [[paper link]](http://arxiv.org/abs/2409.13629) 2024-09-20  \nDavid Chiang\n\n\n\n- **How Transformers Learn Structured Data: Insights from Hierarchical Filtering** [[paper link]](http://arxiv.org/abs/2408.15138) 2024-08-27  \nJerome Garnier-Brun; Marc Mézard; Emanuele Moscato; Luca Saglietti\n\n\n\n- **Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations** [[paper link]](http://arxiv.org/abs/2408.15417) 2024-08-27  \nYize Zhao; Tina Behnia; Vala Vakilian; Christos Thrampoulidis\n\n\n\n- **A Law of Next-Token Prediction in Large Language Models** [[paper link]](http://arxiv.org/abs/2408.13442) 2024-08-24  \nHangfeng He; Weijie J. Su\n\n\n\n- **Transformers As Approximations of Solomonoff Induction** [[paper link]](http://arxiv.org/abs/2408.12065) 2024-08-22  \nNathan Young; Michael Witbrock\n\n\n\n- **Learning Randomized Algorithms with Transformers** [[paper link]](http://arxiv.org/abs/2408.10818) 2024-08-20  \nJohannes von Oswald; Seijin Kobayashi; Yassir Akram; Angelika Steger\n\n\n\n- **Attention is a smoothed cubic spline** [[paper link]](http://arxiv.org/abs/2408.09624) 2024-08-19  \nZehua Lai; Lek-Heng Lim; Yucong Liu\n\n\n\n- **Why Transformers are Obviously Good Models of Language** [[paper link]](http://arxiv.org/abs/2408.03855) 2024-08-07  \nFelix Hill\n\n\n\n- **Can LLMs predict the convergence of Stochastic Gradient Descent?** [[paper link]](http://arxiv.org/abs/2408.01736) 2024-08-03  \nOussama Zekri; Abdelhakim Benechehab; Ievgen Redko\n\n\n\n- **Transformers on Markov Data: Constant Depth Suffices** [[paper link]](http://arxiv.org/abs/2407.17686) 2024-07-25  \nNived Rajaraman; Marco Bondaschi; Kannan Ramchandran; Michael Gastpar; Ashok Vardhan Makkuva\n\n\n\n- **Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability** [[paper link]](http://arxiv.org/abs/2407.15720) 2024-07-22  \nZhuoyan Xu; Zhenmei Shi; Yingyu Liang\n\n\n\n- **Universal Approximation Theory: The basic theory for large language models** [[paper link]](http://arxiv.org/abs/2407.00958) 2024-07-01  \nWei Wang; Qing Li\n\n\n\n- **Seperations in the Representational Capabilities of Transformers and Recurrent Architectures** [[paper link]](http://arxiv.org/abs/2406.09347) 2024-06-13  \nSatwik Bhattamishra; Michael Hahn; Phil Blunsom; Varun Kanade\n\n\n\n- **Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot** [[paper link]](http://arxiv.org/abs/2406.06893) 2024-06-11  \nZixuan Wang; Stanley Wei; Daniel Hsu; Jason D. Lee\n\n\n\n- **What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages** [[paper link]](http://arxiv.org/abs/2406.04289) 2024-06-07  \nNadav Borenstein; Anej Svete; Robin Chan; Josef Valvoda; Franz Nowak; Isabelle Augenstein; Eleanor Chodroff; Ryan Cotterell\n\n\n\n- **Physics of Language Models: Part 1, Learning Hierarchical Language Structures** [[paper link]](http://arxiv.org/abs/2305.13673) 2024-06-02  \nZeyuan Allen-Zhu; Yuanzhi Li\n\n\n\n- **Transformers Can Do Arithmetic with the Right Embeddings** [[paper link]](http://arxiv.org/abs/2405.17399) 2024-05-27  \nSean McLeish; Arpit Bansal; Alex Stein; Neel Jain; John Kirchenbauer; Brian R. Bartoldson; Bhavya Kailkhura; Abhinav Bhatele; Jonas Geiping; Avi Schwarzschild; Tom Goldstein\n\n\n\n- **A One-Layer Decoder-Only Transformer is a Two-Layer RNN: With an Application to Certified Robustness** [[paper link]](http://arxiv.org/abs/2405.17361) 2024-05-27  \nYuhao Zhang; Aws Albarghouthi; Loris D'Antoni\n\n\n\n- **The Power of Hard Attention Transformers on Data Sequences: A Formal Language Theoretic Perspective** [[paper link]](http://arxiv.org/abs/2405.16166) 2024-05-25  \nPascal Bergsträßer; Chris Köcher; Anthony Widjaja Lin; Georg Zetzsche\n\n\n\n- **Transformers represent belief state geometry in their residual stream** [[paper link]](http://arxiv.org/abs/2405.15943) 2024-05-24  \nAdam S. Shai; Sarah E. Marzen; Lucas Teixeira; Alexander Gietelink Oldenziel; Paul M. Riechers\n\n\n\n- **ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models** [[paper link]](http://arxiv.org/abs/2405.09220) 2024-05-15  \n Siwei Wang; Yifei Shen; Shi Feng; Haoran Sun; Shang-Hua Teng; Wei Chen\n\n\n\n- **What Formal Languages Can Transformers Express? A Survey** [[paper link]](http://arxiv.org/abs/2311.00208) 2024-05-06  \nLena Strobl; William Merrill; Gail Weiss; David Chiang; Dana Angluin\n\n\n\n- **Transformers Can Represent $n$-gram Language Models** [[paper link]](http://arxiv.org/abs/2404.14994) 2024-04-23  \nAnej Svete; Ryan Cotterell\n\n\n\n- **Mechanics of Next Token Prediction with Self-Attention** [[paper link]](https://proceedings.mlr.press/v238/li24f.html) 2024-04-18  \nYingcong Li; Yixiao Huang; Muhammed E. Ildiz; Ankit Singh Rawat; Samet Oymak\n\n\n\n- **When can transformers reason with abstract symbols?** [[paper link]](http://arxiv.org/abs/2310.09753) 2024-04-16  \nEnric Boix-Adsera; Omid Saremi; Emmanuel Abbe; Samy Bengio; Etai Littwin; Joshua Susskind\n\n\n\n- **The Illusion of State in State-Space Models** [[paper link]](http://arxiv.org/abs/2404.08819) 2024-04-12  \nWilliam Merrill; Jackson Petty; Ashish Sabharwal\n\n\n\n- **Language Generation in the Limit** [[paper link]](http://arxiv.org/abs/2404.06757) 2024-04-10  \nJon Kleinberg; Sendhil Mullainathan\n\n\n\n- **Attention is Naturally Sparse with Gaussian Distributed Input** [[paper link]](http://arxiv.org/abs/2404.02690) 2024-04-03  \nYichuan Deng; Zhao Song; Chiwun Yang\n\n\n\n- **What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks** [[paper link]](http://arxiv.org/abs/2404.01601) 2024-04-01  \nXingwu Chen; Difan Zou\n\n\n\n- **The Topos of Transformer Networks** [[paper link]](http://arxiv.org/abs/2403.18415) 2024-03-27  \nMattia Jacopo Villani; Peter McBurney\n\n\n\n- **Simulating Weighted Automata over Sequences and Trees with Transformers** [[paper link]](http://arxiv.org/abs/2403.09728) 2024-03-12  \nMichael Rizvi; Maude Lizaire; Clara Lacroce; Guillaume Rabusseau\n\n\n\n- **Simplicity Bias of Transformers to Learn Low Sensitivity Functions** [[paper link]](http://arxiv.org/abs/2403.06925) 2024-03-11  \nBhavya Vasudeva; Deqing Fu; Tianyi Zhou; Elliott Kau; Youqi Huang; Vatsal Sharan\n\n\n\n- **On the Origins of Linear Representations in Large Language Models** [[paper link]](http://arxiv.org/abs/2403.03867) 2024-03-06  \nYibo Jiang; Goutham Rajendran; Pradeep Ravikumar; Bryon Aragam; Victor Veitch\n\n\n\n- **How Well Can Transformers Emulate In-context Newton's Method?** [[paper link]](http://arxiv.org/abs/2403.03183) 2024-03-05  \nAngeliki Giannou; Liu Yang; Tianhao Wang; Dimitris Papailiopoulos; Jason D. Lee\n\n\n\n- **RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval** [[paper link]](http://arxiv.org/abs/2402.18510) 2024-02-29  \nKaiyue Wen; Xingyu Dang; Kaifeng Lyu\n\n\n\n- **Implicit Bias of Next-Token Prediction** [[paper link]](http://arxiv.org/abs/2402.18551) 2024-02-28  \nChristos Thrampoulidis\n\n\n\n- **On the Expressive Power of a Variant of the Looped Transformer** [[paper link]](http://arxiv.org/abs/2402.13572) 2024-02-21  \nYihang Gao; Chuanyang Zheng; Enze Xie; Han Shi; Tianyang Hu; Yu Li; Michael K. Ng; Zhenguo Li; Zhaoqiang Liu\n\n\n\n- **From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers** [[paper link]](http://arxiv.org/abs/2402.13512) 2024-02-20  \nM. Emrullah Ildiz; Yixiao Huang; Yingcong Li; Ankit Singh Rawat; Samet Oymak\n\n\n\n- **Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context** [[paper link]](http://arxiv.org/abs/2312.06528) 2024-02-15  \nXiang Cheng; Yuxin Chen; Suvrit Sra\n\n\n\n- **Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks** [[paper link]](http://arxiv.org/abs/2311.12997) 2024-02-05  \nRahul Ramesh; Ekdeep Singh Lubana; Mikail Khona; Robert P. Dick; Hidenori Tanaka\n\n\n\n- **Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?** [[paper link]](http://arxiv.org/abs/2307.14023) 2024-01-29  \nTokio Kajitsuka; Issei Sato\n\n\n\n- **Transformers are Multi-State RNNs** [[paper link]](http://arxiv.org/abs/2401.06104) 2024-01-11  \nMatanel Oren; Michael Hassid; Yossi Adi; Roy Schwartz\n\n\n\n- **How Capable Can a Transformer Become? A Study on Synthetic, Interpretable Tasks** [[paper link]](https://openreview.net/forum?id=KIhFggzePM) 2023-12-12  \nRahul Ramesh; Mikail Khona; Robert P. Dick; Hidenori Tanaka; Ekdeep Singh Lubana\n\n\n\n- **Transformers can optimally learn regression mixture models** [[paper link]](http://arxiv.org/abs/2311.08362) 2023-11-14  \nReese Pathak; Rajat Sen; Weihao Kong; Abhimanyu Das\n\n\n\n- **The Expressive Power of Low-Rank Adaptation** [[paper link]](http://arxiv.org/abs/2310.17513) 2023-10-26  \nYuchen Zeng; Kangwook Lee\n\n\n\n- **What Algorithms can Transformers Learn? A Study in Length Generalization** [[paper link]](http://arxiv.org/abs/2310.16028) 2023-10-24  \nHattie Zhou; Arwen Bradley; Etai Littwin; Noam Razin; Omid Saremi; Josh Susskind; Samy Bengio; Preetum Nakkiran\n\n\n\n- **Transformers as Support Vector Machines** [[paper link]](http://arxiv.org/abs/2308.16898) 2023-09-07  \nDavoud Ataee Tarzanagh; Yingcong Li; Christos Thrampoulidis; Samet Oymak\n\n\n\n- **How Do Transformers Learn Topic Structure: Towards a Mechanistic Understanding** [[paper link]](https://openreview.net/forum?id=LMXgU4zrq6) 2023-06-15  \nYuchen Li; Yuanzhi Li; Andrej Risteski\n\n\n\n- **Tighter Bounds on the Expressivity of Transformer Encoders** [[paper link]](https://openreview.net/forum?id=XKcogevHj8) 2023-06-15  \nDavid Chiang; Peter Cholak; Anand Pillay\n\n\n\n- **Fast Attention Requires Bounded Entries** [[paper link]](https://arxiv.org/abs/2302.13214v2) 2023-02-26  \nJosh Alman; Zhao Song\n\n\n\n- **Transformers Learn Shortcuts to Automata** [[paper link]](https://openreview.net/forum?id=De4FYqjFueZ) 2023-02-01  \nBingbin Liu; Jordan T. Ash; Surbhi Goel; Akshay Krishnamurthy; Cyril Zhang\n\n\n\n- **Transformer Vs. MLP-Mixer: Exponential Expressive Gap For NLP Problems** [[paper link]](http://arxiv.org/abs/2208.08191) 2022-11-17  \nDan Navon; Alex M. Bronstein\n\n\n\n- **Small Transformers Compute Universal Metric Embeddings** [[paper link]](http://arxiv.org/abs/2209.06788) 2022-10-18  \nAnastasis Kratsios; Valentin Debarnot; Ivan Dokmanić\n\n\n\n- **The Lipschitz Constant of Self-Attention** [[paper link]](http://arxiv.org/abs/2006.04710) 2021-06-09  \nHyunjik Kim; George Papamakarios; Andriy Mnih\n\n\n\n- **On Identifiability in Transformers** [[paper link]](http://arxiv.org/abs/1908.04211) 2020-02-07  \nGino Brunner; Yang Liu; Damián Pascual; Oliver Richter; Massimiliano Ciaramita; Roger Wattenhofer\n\n\u003c/details\u003e\n\n\n### **What Can Transformer Not Do? / Limitation of Transformer**\n\n**[`^        back to top        ^`](#awesome-language-model-analysis-)**\n\nPapers investigating the limitations of transformer-based models, including expressiveness and learning constraints, e.g., limitations in reasoning.\n\n\n\u003cdetails open\u003e\n\u003csummary\u003e\u003cem\u003epaper list (click to fold / unfold)\u003c/em\u003e\u003c/summary\u003e\n\u003cbr\u003e\n\n- **Circuit Complexity Bounds for RoPE-based Transformer Architecture** [[paper link]](http://arxiv.org/abs/2411.07602) 2024-11-12  \nBo Chen; Xiaoyu Li; Yingyu Liang; Jiangxuan Long; Zhenmei Shi; Zhao Song\n\n\n\n- **Consistent Bidirectional Language Modelling: Expressive Power and Representational Conciseness** [[paper link]](http://aclanthology.org/2024.emnlp-main.328) 2024-11  \nGeorgi Shopov; Stefan Gerdjikov\n\n\n\n- **How Numerical Precision Affects Mathematical Reasoning Capabilities of LLMs** [[paper link]](http://arxiv.org/abs/2410.13857) 2024-10-17  \nGuhao Feng; Kai Yang; Yuntian Gu; Xinyue Ai; Shengjie Luo; Jiacheng Sun; Di He; Zhenguo Li; Liwei Wang\n\n\n\n- **Self-Attention Limits Working Memory Capacity of Transformer-Based Models** [[paper link]](http://arxiv.org/abs/2409.10715) 2024-09-16  \nDongyu Gong; Hantao Zhang\n\n\n\n- **One-layer transformers fail to solve the induction heads task** [[paper link]](http://arxiv.org/abs/2408.14332) 2024-08-26  \nClayton Sanford; Daniel Hsu; Matus Telgarsky\n\n\n\n- **Your Context Is Not an Array: Unveiling Random Access Limitations in Transformers** [[paper link]](http://arxiv.org/abs/2408.05506) 2024-08-10  \nMohammadReza Ebrahimi; Sunny Panchal; Roland Memisevic\n\n\n\n- **When Can Transformers Count to n?** [[paper link]](http://arxiv.org/abs/2407.15160) 2024-07-21  \nGilad Yehudai; Haim Kaplan; Asma Ghandeharioun; Mor Geva; Amir Globerson\n\n\n\n- **When can transformers compositionally generalize in-context?** [[paper link]](http://arxiv.org/abs/2407.12275) 2024-07-17  \nSeijin Kobayashi; Simon Schug; Yassir Akram; Florian Redhardt; Johannes von Oswald; Razvan Pascanu; Guillaume Lajoie; João Sacramento\n\n\n\n- **Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries** [[paper link]](http://arxiv.org/abs/2406.12775) 2024-06-18  \nEden Biran; Daniela Gottesman; Sohee Yang\n\n\n\n- **How Far Can Transformers Reason? The Locality Barrier and Inductive Scratchpad** [[paper link]](http://arxiv.org/abs/2406.06467) 2024-06-10  \nEmmanuel Abbe; Samy Bengio; Aryo Lotfi; Colin Sandon; Omid Saremi\n\n\n\n- **Transformers Need Glasses! Information Over-squashing in Language Tasks** [[paper link]](http://arxiv.org/abs/2406.04267) 2024-06-06  \nFederico Barbero; Andrea Banino; Steven Kapturowski; Dharshan Kumaran; João G.M. Araújo; Alex Vitvitskyi; Razvan Pascanu; Petar Veličković\n\n\n\n- **On Limitation of Transformer for Learning HMMs** [[paper link]](http://arxiv.org/abs/2406.04089) 2024-06-06  \nJiachen Hu; Qinghua Liu; Chi Jin\n\n\n\n- **Language Models Need Inductive Biases to Count Inductively** [[paper link]](http://arxiv.org/abs/2405.20131) 2024-05-30  \nYingshan Chang; Yonatan Bisk\n\n\n\n- **Limits of Deep Learning: Sequence Modeling through the Lens of Complexity Theory** [[paper link]](http://arxiv.org/abs/2405.16674) 2024-05-26  \nNikola Zubić; Federico Soldá; Aurelio Sulser; Davide Scaramuzza\n\n\n\n- **Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers** [[paper link]](http://arxiv.org/abs/2405.13536) 2024-05-22  \nTobias Leemann; Alina Fastowski; Felix Pfeiffer; Gjergji Kasneci\n\n\n\n- **Collapse of Self-trained Language Models** [[paper link]](http://arxiv.org/abs/2404.02305) 2024-04-02  \nDavid Herel; Tomas Mikolov\n\n\n\n- **The pitfalls of next-token prediction** [[paper link]](http://arxiv.org/abs/2403.06963) 2024-03-11  \nGregor Bachmann; Vaishnavh Nagarajan\n\n\n\n- **Why are Sensitive Functions Hard for Transformers?** [[paper link]](http://arxiv.org/abs/2402.09963) 2024-03-03  \nMichael Hahn; Mark Rofin\n\n\n\n- **Transformers are Expressive, But Are They Expressive Enough for Regression?** [[paper link]](http://arxiv.org/abs/2402.15478) 2024-02-23  \nSwaroop Nath; Harshad Khadilkar; Pushpak Bhattacharyya\n\n\n\n- **Limits of Transformer Language Models on Learning Algorithmic Compositions** [[paper link]](http://arxiv.org/abs/2402.05785) 2024-02-13  \nJonathan Thomm; Aleksandar Terzic; Geethan Karunaratne; Giacomo Camposampiero; Bernhard Schölkopf; Abbas Rahimi\n\n\n\n- **Representational Strengths and Limitations of Transformers** [[paper link]](http://arxiv.org/abs/2306.02896) 2023-11-16  \nClayton Sanford; Daniel Hsu; Matus Telgarsky\n\n\n\n- **Large Language Models Cannot Self-Correct Reasoning Yet** [[paper link]](https://openreview.net/forum?id=IkmD3fKBPQ) 2023-10-13  \nJie Huang; Xinyun Chen; Swaroop Mishra; Huaixiu Steven Zheng; Adams Wei Yu; Xinying Song; Denny Zhou\n\n\n\n- **Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth** [[paper link]](http://arxiv.org/abs/2103.03404) 2023-08-01  \nYihe Dong; Jean-Baptiste Cordonnier; Andreas Lo","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FFuryton%2Fawesome-language-model-analysis","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FFuryton%2Fawesome-language-model-analysis","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FFuryton%2Fawesome-language-model-analysis/lists"}