{"id":83303,"url":"https://github.com/HKUSTDial/NL2SQL_Handbook","name":"NL2SQL_Handbook","description":"This is a continuously updated handbook for readers to easily track the latest Text-to-SQL techniques in the literature and provide practical guidance for researchers and practitioners. ","projects_count":134,"last_synced_at":"2026-07-29T00:00:31.088Z","repository":{"id":252723526,"uuid":"803643412","full_name":"HKUSTDial/NL2SQL_Handbook","owner":"HKUSTDial","description":"This is a continuously updated handbook for readers to easily track the latest Text-to-SQL techniques in the literature and provide practical guidance for researchers and practitioners. ","archived":false,"fork":false,"pushed_at":"2026-05-06T16:30:13.000Z","size":230925,"stargazers_count":1526,"open_issues_count":0,"forks_count":97,"subscribers_count":28,"default_branch":"main","last_synced_at":"2026-07-10T00:03:56.517Z","etag":null,"topics":["ai4db","awesome","awesome-agents","awesome-nl2sql","awesome-text-to-sql","awesome-text2sql","db","llms","nl-to-code","nl-to-sql","nl2sql","nlp","text-to-code","text-to-sql","text2sql"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2408.05109","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/HKUSTDial.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2024-05-21T05:53:25.000Z","updated_at":"2026-07-09T09:45:16.000Z","dependencies_parsed_at":"2025-01-15T11:56:13.656Z","dependency_job_id":"217f9cbc-c2fa-4537-a58d-9736be0ec710","html_url":"https://github.com/HKUSTDial/NL2SQL_Handbook","commit_stats":null,"previous_names":["hkustdial/nl2sql_handbook"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/HKUSTDial/NL2SQL_Handbook","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/HKUSTDial%2FNL2SQL_Handbook","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/HKUSTDial%2FNL2SQL_Handbook/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/HKUSTDial%2FNL2SQL_Handbook/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/HKUSTDial%2FNL2SQL_Handbook/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/HKUSTDial","download_url":"https://codeload.github.com/HKUSTDial/NL2SQL_Handbook/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/HKUSTDial%2FNL2SQL_Handbook/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":36011657,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-28T02:00:06.341Z","response_time":109,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2025-01-17T00:00:58.490Z","updated_at":"2026-07-29T00:00:31.089Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["💾 Practical Guide for Novice","📰 Text-to-SQL Paper List","📚 Text-to-SQL Survey \u0026 Tutorial","📱 Text-to-SQL Related Applications:","📱 NL2SQL Related Applications:","📰 NL2SQL Paper List"],"sub_categories":["🔎How to evaluate your model:","🛠️ How to build an LLM-based Text-to-SQL model:","🗺️ Roadmap and Decision Flow"],"readme":" \u003ch1 align=\"center\"\u003eText-to-SQL Handbook\u003c/h1\u003e\n\n \u003ch3 align=\"center\"\u003eNL2SQL Handbook\u003c/h3\u003e\n\nFrom this repository, you can view the 📚[latest advancements](#-text-to-sql-survey--tutorial) in Text-to-SQL (a.k.a NL2SQL). This handbook corresponds to our survey paper[TKDE'2025]: [📖A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?](https://arxiv.org/abs/2408.05109). We also provide [tutorial slides](./slides/NL2SQL-VLDB2025.pptx) for VLDB'2025 Tutorial and summarize the key points of this survey. Based on language model trends, we've created a river diagram of Text-to-SQL methods to trace the field's evolution. \n\nIf you are a novice, don't worry—we have prepared a practical guide for you, covering a wide range of [foundational materials](#-practical-guide-for-novice) and related [applications](#-text-to-sql-related-applications). 📧If we missed any interesting work, [connect with us](#connect-with-us).\n\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"800\" src=\"./assets/river.svg\"/\u003e\n\u003c/p\u003e\n\n```bibtex\n@article{liu2025survey,\n  title={A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?},\n  author={Liu, Xinyu and Shen, Shuyu and Li, Boyan and Ma, Peixian and Jiang, Runzhi and Zhang, Yuxin and Fan, Ju and Li, Guoliang and Tang, Nan and Luo, Yuyu},\n  journal={IEEE Transactions on Knowledge and Data Engineering},\n  year={2025},\n  publisher={IEEE}\n}\n```\n\n## 🧭 Text-to-SQL Introduction \nTranslating users' natural language queries (NL) into SQL queries can significantly reduce barriers to accessing relational databases and support various commercial applications. The performance of Text-to-SQL has been greatly improved with the emergence of language models (LMs). In this context, it is crucial to assess our current position, determine the Text-to-SQL solutions that should be adopted for specific scenarios by practitioners, and identify the research topics that researchers should explore next.\n\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"600\" src=\"./assets/NL2SQL.jpg\"/\u003e\n\u003c/p\u003e\n\n## 📈 Text-to-SQL Lifecycle\n\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"800\" src=\"./assets/nl2sql_lifecycle.svg\"/\u003e\n\u003c/p\u003e\n\n+ Model: Text-to-SQL translation techniques that tackle not only NL ambiguity and under-specification, but also properly map NL with database schema and instances;\n\n+ Data: From the collection of training data, data synthesis due to training data scarcity, to Text-to-SQL benchmarks;\n\n+ Evaluation: Evaluating Text-to-SQL methods from multiple angles using different metrics and granularities;\n\n+ Error Analysis: analyzing Text-to-SQL errors to find the root cause and guiding Text-to-SQL models to evolve.\n\n## 🤔 Where Are We?\nWe categorize the challenges of Text-to-SQL into five levels, each addressing specific hurdles. The first three levels cover challenges that have been or are currently being addressed, reflecting the progressive development of Text-to-SQL. The fourth level represents the challenges we aim to tackle in the LLMs stage, while the fifth level outlines our vision for Text-to-SQL system in the next five years. \n\nWe describe the evolution of Text-to-SQL solutions from the perspective of language models, categorizing it into four stages.\nFor each stage of Text-to-SQL, we analyze the changes in target users and the extent to which challenges are addressed.\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"800\" src=\"./assets/The Evolution of NL2SQL Solutions from the Perspective of Language Models.svg\"/\u003e\n\u003c/p\u003e\n\n\n## 🧩 Module-based Text-to-SQL Methods\nWe summarize the key modules of Text-to-SQL solutions\nutilizing the language model. \n+ **Pre-processing** serves as an enhancement to the model’s inputs in the Text-to-SQL parsing process. You can get more details from this chapter: [Pre-Processing](chapter/Pre_Processing.md)\n+ **Text-to-SQL translation methods** constitute the core of the Text-to-SQL solution, responsible for converting input natural language queries into SQL queries. You can get more details from this chapter: [Text-to-SQL Translation Methods](chapter/Translation_method.md)\n+ **Post-processing** is a crucial step to refine the generated SQL queries, ensuring they meet user expectations more accurately. You can get more details from this chapter: [Post-Processing](chapter/Post_Processing.md)\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"600\" src=\"./assets/An Overview of NL2SQL Method in the LLM Era.svg\"/\u003e\n\u003c/p\u003e\n\n## 📚 Text-to-SQL Survey \u0026 Tutorial\n\n1. A Survey of Text-to-SQL in the Era of LLMs:\nWhere are we, and where are we going?\n\u003cimg src=\"https://img.shields.io/badge/TKDE'2025-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2408.05109) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/HKUSTDial/NL2SQL_Handbook)\n1. Natural Language to SQL: State of the Art and Open Problems. \u003cimg src=\"https://img.shields.io/badge/VLDB'2025-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dbgroup.cs.tsinghua.edu.cn/ligl/papers/VLDB25-NL2SQL.pdf)\n1. A Survey on Employing Large Language Models for Text-to-SQL Tasks.\n\u003cimg src=\"https://img.shields.io/badge/CSUR'2024-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2407.15186)\n1. Next-generation database interfaces: A survey of LLM-based Text-to-SQL.\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2406.08426)\n1. Large Language Model Enhanced Text-to-SQL Generation: A Survey.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2410.06011)\n1. From Natural Language to SQL: Review of LLM-based Text-to-SQL Systems.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2410.01066)\n1. Natural language interfaces for tabular data querying and visualization: A survey.\n\u003cimg src=\"https://img.shields.io/badge/TKDE'2024-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2310.17894)\n1. Natural Language Interfaces for Databases with Deep Learning.\u003cimg src=\"https://img.shields.io/badge/VLDB'2023-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/10.14778/3611540.3611575)\n1. A survey on deep learning approaches for text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/VLDBJ'2023-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/10.1007/s00778-022-00776-8)\n1. Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect.\n\u003cimg src=\"https://img.shields.io/badge/COLING'2022-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2022.coling-1.190/)\n1. A Deep Dive into Deep Learning Approaches for Text-to-SQL Systems.\n\u003cimg src=\"https://img.shields.io/badge/SIGMOD'2021-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/10.1145/3448016.3457543)\n1. State of the Art and Open Challenges in Natural Language Interfaces to Data.\n\u003cimg src=\"https://img.shields.io/badge/SIGMOD'2020-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/10.1145/3318464.3383128)\n1. Natural language to SQL: Where are we today? \u003cimg src=\"https://img.shields.io/badge/VLDB'2020-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.vldb.org/pvldb/vol13/p1737-kim.pdf)\n\n## 📰 Text-to-SQL Paper List\n1. Alpha-SQL: Zero-Shot Text-to-SQL using Monte Carlo Tree Search\n\u003cimg src=\"https://img.shields.io/badge/ICML'2025-brightgreen\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2502.17248) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://alpha-sql-hkust.github.io/)\n1. NL2SQL-BUGs: A Benchmark for Detecting Semantic Errors in NL2SQL Translation.\u003cimg src=\"https://img.shields.io/badge/SIGKDD'2025-B6FFBB\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2503.11984) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://nl2sql-bugs.github.io/)\n1. EllieSQL: Cost-Efficient Text-to-SQL with Complexity-Aware Routing. \u003cimg src=\"https://img.shields.io/badge/COLM'2025-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2503.22402) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://elliesql.github.io/)\n1. Structure-Guided Large Language Models for Text-to-SQL Generation. \u003cimg src=\"https://img.shields.io/badge/ICML'2025-brightgreen\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://icml.cc/virtual/2025/poster/44477)\n1. Sphinteract: Resolving Ambiguities in NL2SQL Through User Interaction.\n\u003cimg src=\"https://img.shields.io/badge/VLDB'2025-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.vldb.org/pvldb/vol18/p1145-zhao.pdf) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/ZhaoFuheng/Sphinteract)\n1. OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale. \u003cimg src=\"https://img.shields.io/badge/VLDB'2025-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e]([https://www.vldb.org/pvldb/vol18/p1145-zhao.pdf](https://arxiv.org/abs/2503.02240)) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/RUCKBReasoning/OmniSQL)\n1. EVOSCHEMA: TOWARDS TEXT-TO-SQL ROBUSTNESS AGAINST SCHEMA EVOLUTION. \u003cimg src=\"https://img.shields.io/badge/VLDB'2025-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://openreview.net/pdf?id=NfUHBaZdLw) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/zhangtianshu/EvoSchema)\n1. Is Long Context All You Need? Leveraging LLM's Extended Context for NL2SQL. \u003cimg src=\"https://img.shields.io/badge/VLDB'2025-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2501.12372) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/yeounoh/lc_nl2sql)\n1. The Power of Constraints in Natural Language to SQL Translation. \u003cimg src=\"https://img.shields.io/badge/VLDB'2025-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.vldb.org/pvldb/vol18/p2097-ren.pdf)\n[\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/httdty/REDSQL_VLDB)\n1. OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency Alignment. \u003cimg src=\"https://img.shields.io/badge/SIGMOD'2025-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2502.14913) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/OpenSearch-AI/OpenSearch-SQL)\n1. Reliable Text-to-SQL with Adaptive Abstention.\u003cimg src=\"https://img.shields.io/badge/SIGMOD'2025-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2501.10858) \n1. SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference.\u003cimg src=\"https://img.shields.io/badge/SIGMOD'2025-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/10.1145/3709727)\n1. Automated Validating and Fixing of Text-to-SQL Translation with Execution Consistency. \u003cimg src=\"https://img.shields.io/badge/SIGMOD'2025-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://ipads.se.sjtu.edu.cn/zh/publications/SQLDriller.pdf)\n1. Grounding Natural Language to SQL Translation with Data-Based Self-Explanations.\u003cimg src=\"https://img.shields.io/badge/ICDE'2025-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2411.02948) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/Kaimary/CycleSQL)\n1. AID-SQL: Adaptive In-Context Learning of Text-to-SQL with Difficulty-Aware Instruction and Retrieval-Augmented Generation. \u003cimg src=\"https://img.shields.io/badge/ICDE'2025-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.computer.org/csdl/proceedings-article/icde/2025/360300d945/26FZCc99mg0) \n1. CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL.\n\u003cimg src=\"https://img.shields.io/badge/ICDE'2025-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.computer.org/csdl/proceedings-article/icde/2025/360300d302/26FZBD2hBJe) \n1. CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/ICLR'2025-brightgreen\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2410.01943v1) \n1. Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows.\n\u003cimg src=\"https://img.shields.io/badge/ICLR'2025-brightgreen\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2411.07763) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/xlang-ai/Spider2)\n1. ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/ICLR'2025-brightgreen\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2412.10138)\n1. SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL.\u003cimg src=\"https://img.shields.io/badge/ACL'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2506.00391)\n1. DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph. \u003cimg src=\"https://img.shields.io/badge/ACL'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2505.19956)\n1. Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL.\u003cimg src=\"https://img.shields.io/badge/ACL'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2502.11656)\n1. STaR-SQL: Self-Taught Reasoner for Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/ACL'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2502.13550)\n1. SQLGenie: A Practical LLM based System for Reliable and Efficient SQL Generation \u003cimg src=\"https://img.shields.io/badge/ACL'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e]([https://arxiv.org/abs/2502.13550](https://aclanthology.org/2025.acl-industry.71/))\n1. Confidence Estimation for Error Detection in Text-to-SQL Systems. \u003cimg src=\"https://img.shields.io/badge/AAAI'2025-cyan\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2501.09527)\n1. SQLord: A Robust Enterprise Text-to-SQL Solution via Reverse Data Generation and Workflow Decomposition. \u003cimg src=\"https://img.shields.io/badge/WWW'2025-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/pdf/10.1145/3701716.3715541)\n1. DBCopilot: Scaling Natural Language Querying to Massive Databases.\u003cimg src=\"https://img.shields.io/badge/EDBT/ICDT'2025-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2312.03463) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/tshu-w/DBCopilot)\n1. Boosting Text-to-SQL through Multi-grained Error Identification.\n\u003cimg src=\"https://img.shields.io/badge/COLING'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2025.coling-main.289.pdf)\n1. Gen-SQL: Efficient Text-to-SQL By Bridging Natural Language Question And Database Schema With Pseudo-Schema.\n\u003cimg src=\"https://img.shields.io/badge/COLING'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2025.coling-main.256/)\n1. Utilising Large Language Models for Adversarial Attacks in Text-to-SQL: A Perpetrator and Victim Approach.\n\u003cimg src=\"https://img.shields.io/badge/BTW'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2502.20657) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/XGenerationLab/XiYan-DBDescGen)\n1. You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/NAACL'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2409.12172) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://sig4kg.github.io/archer-bench/) \n1. PARSQL: Enhancing Text-to-SQL through SQL Parsing and Reasoning. \u003cimg src=\"https://img.shields.io/badge/ACL(Findings)'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2024.findings-acl.120/)\n1. UCS-SQL: Uniting Content and Structure for Enhanced Semantic Bridging In Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/ACL(Findings)'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://openreview.net/forum?id=xnTouV7wyr)\n1. SQLForge: Synthesizing Reliable and Diverse Data to Enhance Text-to-SQL Reasoning in LLMs. \u003cimg src=\"https://img.shields.io/badge/ACL(Findings)'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2505.13725)\n1. Optimizing Reasoning for Text-to-SQL with Execution Feedback. \u003cimg src=\"https://img.shields.io/badge/ACL(Findings)'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2503.19988)\n1. Knowledge Base Construction for Knowledge-Augmented Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/ACL(Findings)'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2505.22096)\n1. SQLong: Enhanced NL2SQL for Longer Contexts with LLMs.\n\u003cimg src=\"https://img.shields.io/badge/ACL(Workshop)'2025-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2502.16747)\n1. Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/COLM'2025-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2503.23157)\n1. Automatic Metadata Extraction for Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2505.19988) \u003cimg src=\"https://img.shields.io/badge/BIRD Top2-blue\"\u003e\n1. CSC-SQL: Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2505.13271) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/CycloneBoy/csc_sql/)\n1. Cheaper, Better, Faster, Stronger: Robust Text-to-SQL without Chain-of-Thought or Fine-Tuning. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2505.14174) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/genaasia/N-rep)\n1. Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2505.04671) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/ruc-datalab/RewardSQL)\n1. SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2504.08600) \n1. Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2505.20315) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/snowflakedb/ArcticTraining)\n1. Corrective Self-Consistency in Text-to-SQL via Reinforcement Learning. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2505.13271) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/CycloneBoy/csc_sql)\n1. SQLForge: Synthesizing Reliable and Diverse Data to Enhance\nText-to-SQL Reasoning in LLMs.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2505.13725)\n1. Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2504.15077)\n1. Distill-C: Enhanced NL2SQL via Distilled Customization with LLMs. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2504.00048)\n1. OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2503.02240) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/RUCKBReasoning/OmniSQL)\n1. SQL-Factory: A Multi-Agent Framework for High-Quality and Large-Scale SQL Generation. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2504.14837)\n1. Text2SQL is Not Enough: Unifying AI and Databases with TAG. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2408.14717) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/TAG-Research/TAG-Bench) \n1. Automatic database description generation for Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2502.20657)\n[\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/XGenerationLab/XiYan-DBDescGen)\n1. MCTS-SQL: An Effective Framework for Text-to-SQL with Monte Carlo Tree Search.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2501.16607)\n1. SQL-o1: A Self-Reward Heuristic Dynamic Search Method for Text-to-SQL. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2502.11741)\n1. FEATHER-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2503.17811)\n1. FI-NL2PY2SQL: Financial Industry NL2SQL Innovation Model Based on Python and Large Language Model.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.mdpi.com/1999-5903/17/1/12)\n1. FGCSQL: A Three-Stage Pipeline for Large Language Model-Driven Chinese Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.mdpi.com/2079-9292/14/6/1214)\n1. Transforming Medical Data Access: The Role and Challenges of Recent Language Models in SQL Query Automation. \u003cimg src=\"https://img.shields.io/badge/arXiv'2025-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.mdpi.com/1999-4893/18/3/124)\n1. The Dawn of Natural Language to SQL: Are We Fully Ready?\n\u003cimg src=\"https://img.shields.io/badge/VLDB'2024-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2406.01265) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/HKUSTDial/NL2SQL360)\n1. Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation. \n\u003cimg src=\"https://img.shields.io/badge/VLDB'2024-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2308.15363) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/BeachWang/DAIL-SQL) \n1. Interleaving Pre-Trained Language Models and Large Language Models for Zero-Shot NL2SQL Generation. \n\u003cimg src=\"https://img.shields.io/badge/VLDB'2024-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2306.08891) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/ruc-datalab/ZeroNL2SQL)\n1. Generating Succinct Descriptions of Database Schemata for Cost-Efficient Prompting of Large Language Models. \n\u003cimg src=\"https://img.shields.io/badge/VLDB'2024-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/abs/10.14778/3681954.3682017) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/itrummer/schemacompression)\n1. ScienceBenchmark: A Complex Real-World Benchmark for Evaluating Natural Language to SQL Systems.\u003cimg src=\"https://img.shields.io/badge/VLDB'2024-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2306.04743) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://sciencebenchmark.cloudlab.zhaw.ch/)\n1. CodeS: Towards Building Open-source Language Models for Text-to-SQL. \n\u003cimg src=\"https://img.shields.io/badge/SIGMOD'2024-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2402.16347) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/RUCKBReasoning/codes)\n1. FinSQL: Model-Agnostic LLMs-based Text-to-SQL Framework for Financial Analysis. \n\u003cimg src=\"https://img.shields.io/badge/SIGMOD'2024-red\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2401.10506) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/bigbigwatermalon/FinSQL)\n1. PURPLE: Making a Large Language Model a Better SQL Writer. \n\u003cimg src=\"https://img.shields.io/badge/ICDE'2024-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2403.20014) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/httdty/purple)\n1. METASQL: A Generate-then-Rank Framework for Natural Language to SQL Translation. \n\u003cimg src=\"https://img.shields.io/badge/ICDE'2024-green\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2402.17144) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/Kaimary/MetaSQL)\n1. Archer: A Human-Labeled Text-to-SQL Dataset with Arithmetic, Commonsense and Hypothetical Reasoning.\n\u003cimg src=\"https://img.shields.io/badge/ACL'2024-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2024.eacl-long.6/) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://sig4kg.github.io/archer-bench/)\n1. Synthesizing Text-to-SQL Data from Weak and Strong LLMs.\n\u003cimg src=\"https://img.shields.io/badge/ACL'2024-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2408.03256) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/Yangjiaxi/Sense)\n1. Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark.\n\u003cimg src=\"https://img.shields.io/badge/ACL'2024-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2402.12243) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/niklaswretblad/the-effects-of-noise-in-text-to-SQL)\n1. I Need Help! Evaluating LLM’s Ability to Ask for Users’ Support: A Case Study on Text-to-SQL Generation.\n\u003cimg src=\"https://img.shields.io/badge/EMNLP'2024-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2407.14767) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/appier-research/i-need-help)\n1. PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/EMNLP'2024-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2409.14082) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/lrlbbzl/PTD-SQL)\n1. Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning.\n\u003cimg src=\"https://img.shields.io/badge/EMNLP'2024-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2407.03227)\n1. Data-Centric Text-to-SQL with Large Language Models. \n\u003cimg src=\"https://img.shields.io/badge/NeurIPS(workshop)'2024-yellow\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://openreview.net/pdf?id=gDKIjZcg93)\n1. Research and Practice on Database Interaction Based on Natural Language Processing\n\u003cimg src=\"https://img.shields.io/badge/AIAC'2024-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2310.17894)\n1. XiYan-SQL: A Multi-Generator Ensemble Framework for Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2411.08599)\n1. Structure Guided Large Language Model for SQL Generation. \n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2402.13284) \n1. A Plug-and-Play Natural Language Rewriter for Natural Language to SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2412.17068) \n1. RSL-SQL: Robust Schema Linking in Text-to-SQL Generation.   \n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2403.15879) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/glee4810/TrustSQL)\n1. In-Context Reinforcement Learning based Retrieval-Augmented Generation for Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://assets.amazon.science/09/f4/493c574346f895bbb0303282a501/in-context-reinforcement-learning-based-retrieval-augmented-generation-for-text-to-sql.pdf) \n1. TrustSQL: Benchmarking Text-to-SQL Reliability with Penalty-Based Scoring.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2411.00073) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/Laqcce-cao/RSL-SQL)\n1. LAIA-SQL: Enhancing Natural Language to SQL Generation in Multi-Table QA via Task Decomposition and Keyword Extraction\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://openreview.net/pdf?id=WYdpjwKQma)\n1. Research on Large Model Text-to-SQL Optimization Method for Intelligent Interaction in the Field of Construction Safety.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://ieeexplore.ieee.org/abstract/document/10810146)\n1. SQLh-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging.\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2408.12733v2)\n1. Grounding Natural Language to SQL Translation with Data-Based Self-Explanations.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2411.02948) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/Kaimary/CycleSQL)\n1. Towards Optimizing SQL Generation via LLM Routing.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2411.04319)\n1. E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2409.16751) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/HasanAlpCaferoglu/E-SQL)\n1. DB-GPT: Empowering Database Interactions with Private Large Language Models.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2312.17449) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/eosphoros-ai/DB-GPT)\n1. The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2408.07702)  \n1. CHESS: Contextual Harnessing for Efficient SQL Synthesis.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2405.16755) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/ShayanTalaei/CHESS)\n1. PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-consistency.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2403.09732) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/ruc-datalab/ZeroNL2SQL)\n1. CoE-SQL: In-Context Learning for Multi-Turn Text-to-SQL with Chain-of-Editions.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2405.02712) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/X-LANCE/text2sql-multiturn-GPT)\n1. AMBROSIA: A Benchmark for Parsing Ambiguous Questions into Database Queries.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2406.19073) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://ambrosia-benchmark.github.io/)\n1. Text-to-SQL Calibration: No Need to Ask—Just Rescale Model Probabilities.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2024-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/pdf/2411.16742) \n1. Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning.\n\u003cimg src=\"https://img.shields.io/badge/VLDB'2023-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://dl.acm.org/doi/abs/10.1145/3589292) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/ruc-datalab/SC-prompt)\n1. CatSQL: Towards Real World Natural Language to SQL Applications.\n\u003cimg src=\"https://img.shields.io/badge/VLDB'2023-blue\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://www.vldb.org/pvldb/vol16/p1534-fu.pdf) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/asfuhan/CatSQL)\n1. DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction. \n\u003cimg src=\"https://img.shields.io/badge/NeurIPS'2023-yellow\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2304.11015) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/MohammadrezaPourreza/Few-shot-NL2SQL-with-prompting/tree/main)\n1. Data Ambiguity Strikes Back: How Documentation Improves GPT's Text-to-SQL. \n\u003cimg src=\"https://img.shields.io/badge/NeurIPS(workshop)'2023-yellow\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://openreview.net/pdf?id=FflKTuIRTD) \n1. ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought.\n\u003cimg src=\"https://img.shields.io/badge/EMNLP'2023-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2310.17342) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/X-LANCE/text2sql-GPT)\n1. Selective Demonstrations for Cross-domain Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/EMNLP'2023-orange\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2310.06302) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/shuaichenchang/ODIS-Text-to-SQL)\n1. RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQL. \n\u003cimg src=\"https://img.shields.io/badge/AAAI'2023-cyan\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2302.05965) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/RUCKBReasoning/RESDSQL)\n1. Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing. \n\u003cimg src=\"https://img.shields.io/badge/AAAI'2023-cyan\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2301.07507) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/graphix)\n1. Improving Generalization in Language Model-based Text-to-SQL Semantic Parsing: Two Simple Semantic Boundary-based Techniques.\n\u003cimg src=\"https://img.shields.io/badge/ACL'2023-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://virtual2023.aclweb.org/paper_P4350.html) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/Dakingrai/ood-generalization-semantic-boundary-techniques)\n1. G\u003csup\u003e3\u003c/sup\u003eR: A Graph-Guided Generate-and-Rerank Framework for Complex and Cross-domain Text-to-SQL Generation.\n\u003cimg src=\"https://img.shields.io/badge/ACL(findings)'2023-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2023.findings-acl.23/) \n1. Importance of Synthesizing High-quality Data for Text-to-SQL Parsing.\n\u003cimg src=\"https://img.shields.io/badge/ACL(findings)'2023-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2023.findings-acl.86.pdf) \n1. Know What I don’t Know: Handling Ambiguous and Unknown Questions for Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/ACL(findings)'2023-9cf\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://aclanthology.org/2023.findings-acl.352/) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/wbbeyourself/DTE)\n1. C3: Zero-shot Text-to-SQL with ChatGPT \n\u003cimg src=\"https://img.shields.io/badge/arXiv'2023-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2307.07306) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/bigbigwatermalon/C3SQL)\n1. MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2023-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2312.11242) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/wbbeyourself/MAC-SQL)\n1. SQLformer: Deep Auto-Regressive Query Graph Generation for Text-to-SQL Translation.\n\u003cimg src=\"https://img.shields.io/badge/arXiv'2023-purple\"\u003e [\u003cimg src=\"https://img.shields.io/badge/Paper-grey\"\u003e](https://arxiv.org/abs/2310.18376) [\u003cimg src=\"https://img.shields.io/badge/Code-grey\"\u003e](https://github.com/AdrianBZG/SQLformer)\n\n\n## 📊 Text-to-SQL Benchmark\nWe create a timeline of the benchmark's development and mark relevant milestones. You can get more details from this chapter: [📊 Benchmark](chapter/Benchmark.md)\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"800\" src=\"./assets/Dataset_timeline.svg\"/\u003e\n\u003c/p\u003e\n\n## 🎯 Where Are We Going?\n\n* 🎯Solve Open Text-to-SQL Problem\n* 🎯Develop Cost-effective Text-to-SQL Methods\n* 🎯Make Text-to-SQL Solutions Trustworthy\n* 🎯Text-to-SQL with Ambiguous and Unspecified NL Queries\n* 🎯Adaptive Training Data Synthesis\n\n## 📖 Catalog for Our Survey\nYou can get more information from our subsection. We introduce representative papers on related concepts:\n* [Pre-Processing](chapter/Pre_Processing.md)\n* [Text-to-SQL Translation Methods](chapter/Translation_method.md)\n* [Post-Processing](chapter/Post_Processing.md)\n* [Benchmark](chapter/Benchmark.md)\n* [Evaluation](chapter/Evaluation.md)\n* [Error Analysis](chapter/Error_Analysis.md)\n\n## 💾 Practical Guide for Novice\n\n### 📊 How to get data:\n* We collect Text-to-SQL benchmark features and download links for you. You can get more details from this chapter: [Benchmark](chapter/Benchmark.md)\n* The analysis code for benchmarks is available in the `src/dataset_analysis` directory. Benchmark analysis reports can be found in the `report/` directory.\n\n### 🛠️ How to build an LLM-based Text-to-SQL model:\n\n* Litgpt [Repository Link](https://github.com/Lightning-AI/litgpt)\n\n    This repository offers access to over 20 high-performance large language models (LLMs) with comprehensive guides for pretraining, fine-tuning, and deploying at scale. It is designed to be beginner-friendly with from-scratch implementations and no complex abstractions.\n\n* LLaMA-Factory [Repository Link](https://github.com/hiyouga/LLaMA-Factory)\n    Unified Efficient Fine-Tuning of 100+ LLMs. Integrating various models with scalable training resources, advanced algorithms, practical tricks, and comprehensive experiment monitoring tools, this setup enables efficient and faster inference through optimized APIs and UIs.\n\n* Fine-tuning and In-Context learning for BIRD-SQL benchmark [Repository Link](https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/bird#fine-tuning-ft)\n    \n    A tutorial for both Fine-tuning and In-Context Learning is provided by the BIRD-SQL benchmark. \n\n### 🔎How to evaluate your model:\n\nWe collect NL2SQL evaluation metrics for you. You can get more details from this chapter: [Evaluation](chapter/Evaluation.md)\n\n* NLSQL360 [Repository Link](https://github.com/HKUSTDial/NL2SQL360) \n\n     NL2SQL360 is a testbed for fine-grained evaluation of NL2SQL solutions. Our testbed integrates existing NL2SQL benchmarks, a repository of NL2SQL models, and various evaluation metrics, which aims to provide an intuitive and user-friendly platform to enable both standard and customized performance evaluations. \u003cimg src=\"https://img.shields.io/badge/EX-red\"\u003e \u003cimg src=\"https://img.shields.io/badge/EM-green\"\u003e \u003cimg src=\"https://img.shields.io/badge/VES-blue\"\u003e \u003cimg src=\"https://img.shields.io/badge/QVT-orange\"\u003e\n\n* Test-suite-sql-eval [Repository Link](https://github.com/taoyds/test-suite-sql-eval)\n\n    This repo contains a test suite evaluation metric for 11 text-to-SQL tasks. It is now the official metric of [Spider](https://yale-lily.github.io/spider), [SParC](https://yale-lily.github.io/sparc), and [CoSQL](https://yale-lily.github.io/cosql), and is also now available for Academic, ATIS, Advising, Geography, IMDB, Restaurants, Scholar, and Yelp (building on the amazing work by [Catherine and Jonathan](https://github.com/jkkummerfeld/text2sql-data)).  \u003cimg src=\"https://img.shields.io/badge/EX-red\"\u003e \u003cimg src=\"https://img.shields.io/badge/EM-green\"\u003e\n\n* BIRD-SQL-Official [Repository Link](https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/bird#evaluation)\n\n    It is now the official tool of [BIRD-SQL](https://bird-bench.github.io/). It is the first tool to propose VES and give an official test suite. \u003cimg src=\"https://img.shields.io/badge/EX-red\"\u003e \u003cimg src=\"https://img.shields.io/badge/VES-blue\"\u003e\n\n\n### 🗺️ Roadmap and Decision Flow\n\nYou can get some inspiration from the Roadmap and Decision Flow.\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"800\" src=\"./assets/NL2SQL_Guidance.svg\"/\u003e\n\u003c/p\u003e\n\n## 📱 Text-to-SQL Related Applications:\n\n* Chat2DB: AI-driven database tool and SQL client, The hottest GUI client, supporting MySQL, Oracle, PostgreSQL, DB2, SQL Server, DB2, SQLite, H2, ClickHouse, and more. [\u003cimg src=\"https://img.shields.io/badge/Repositor Link-grey\"\u003e](https://github.com/codePhiliaX/Chat2DB) [\u003cimg src=\"https://img.shields.io/badge/Web Link-98f\"\u003e](https://chat2db-ai.com/zh-CN)\n* DB-GPT: AI Native Data App Development framework with AWEL(Agentic Workflow Expression Language) and Agents. [\u003cimg src=\"https://img.shields.io/badge/Repositor Link-grey\"\u003e](https://github.com/eosphoros-ai/DB-GPT) \n* Postgres.new: In-browser Postgres sandbox with AI assistance.  [\u003cimg src=\"https://img.shields.io/badge/Repositor Link-grey\"\u003e](https://github.com/supabase-community/postgres-new/tree/main) [\u003cimg src=\"https://img.shields.io/badge/Web Link-98f\"\u003e](https://postgres.new/)\n* QueryGPT – Natural Language to SQL Using Generative AI. [[\u003cimg src=\"https://img.shields.io/badge/Web Link-98f\"\u003e](https://www.uber.com/en-JP/blog/query-gpt/)\n\n## 📮Connect with Us\nPlease feel free to contact us if we missed any interesting work.\n\n📧 xliu371[at]connect.hkust-gz.edu.cn\n\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/hkustdial%2Fnl2sql_handbook/projects"}