Text-to-SQL on BIRD (Execution Accuracy EX)
73Execution AccuracyCHASE-SQL + Gemini 1.5
Evaluation Results
| Method | Links | |
|---|---|---|
| CHASE-SQL + Gemini 1.5Model Size=Closed-Source (very large), Use specific training set?=YES, Support Multi-Turn Text-to-SQL?=NO2025.10 | 73 | |
| PV-SQLInput tokens=3805, Output tokens=248, Average task completion time (seconds)=3.42026.04 | 65.1 | |
| LearNATLLMs=Qwen2.5-Coder-32B, Level=Ours2025.04 | 65 | |
| PV-SQL (w/o Probe)Input tokens=1068, Output tokens=67, Average task completion time (seconds)=1.02026.04 | 62.1 | |
| PV-SQL (w/o Repair)Input tokens=3521, Output tokens=223, Average task completion time (seconds)=3.12026.04 | 61.8 | |
| LearNATLLMs=Qwen2.5-Coder-14B, Level=Ours2025.04 | 61.2 | |
| TA-SQLInput tokens=6305, Output tokens=242, Average task completion time (seconds)=5.32026.04 | 60.43 | |
| MAC-SQLLLMs=GPT-4, Level=System-Level2025.04 | 59.4 | |
| E-SQLInput tokens=28177, Output tokens=642, Average task completion time (seconds)=21.42026.04 | 59.1 | |
| PV-SQL (LLM verify)Input tokens=4887, Output tokens=478, Average task completion time (seconds)=6.62026.04 | 59.1 | |
| MAC-SQLInput tokens=6067, Output tokens=387, Average task completion time (seconds)=4.72026.04 | 58.7 | |
| SuperSQLLLMs=GPT-4, Level=System-Level2025.04 | 58.5 | |
| CodeSLLMs=CodeS-15B, Level=Model-Level2025.04 | 58.5 | |
| LearNATLLMs=Qwen2.5-Coder-7B, Level=Ours2025.04 | 58.1 | |
| XiYan-SQLInput tokens=1689, Output tokens=128, Average task completion time (seconds)=2.42026.04 | 57.6 | |
| MAG-SQLLLMs=GPT-4, Level=System-Level2025.04 | 57.6 | |
| TS-SQLInput tokens=1803, Output tokens=310, Average task completion time (seconds)=3.62026.04 | 57.4 | |
| SFT CodeS-7BModel Size=7B, Use specific training set?=YES, Support Multi-Turn Text-to-SQL?=NO2025.10 | 57.1 | |
| CodeSLLMs=CodeS-7B, Level=Model-Level2025.04 | 57 | |
| Qwen3-4B + Long-Horizon Training (Ours)Model Size=4B, Use specific training set?=NO, Support Multi-Turn Text-to-SQL?=YES2025.10 | 56.9 | |
| DeepSeekModel Size=236B, Use specific training set?=NO, Support Multi-Turn Text-to-SQL?=YES2025.10 | 56.1 | |
| DIN-SQLInput tokens=26629, Output tokens=629, Average task completion time (seconds)=25.02026.04 | 54.8 | |
| DAIL-SQL + GPT-4Model Size=Closed-Source (very large), Use specific training set?=YES, Support Multi-Turn Text-to-SQL?=NO2025.10 | 54.76 | |
| DAIL-SQLLLMs=GPT-4, Level=Model-Level2025.04 | 54.3 | |
| DAIL-SQLInput tokens=2446, Output tokens=50, Average task completion time (seconds)=0.82026.04 | 54.1 | |
| ACT-SQLLLMs=GPT-4, Level=Model-Level2025.04 | 52.4 | |
| ZS-CoTInput tokens=787, Output tokens=191, Average task completion time (seconds)=1.942026.04 | 52.2 | |
| DIN-SQLLLMs=GPT-4, Level=System-Level2025.04 | 50.7 | |
| C3-SQLLLMs=GPT-4, Level=System-Level2025.04 | 50.2 | |
| MetaSQLLLMs=GPT-4, Level=System-Level2025.04 | 47.6 | |
| Qwen3-14BModel Size=14B, Use specific training set?=NO, Support Multi-Turn Text-to-SQL?=YES2025.10 | 32.6 | |
| Qwen3-4BModel Size=4B, Use specific training set?=NO, Support Multi-Turn Text-to-SQL?=YES2025.10 | 25.8 | |
| T5-3BModel Size=3B, Use specific training set?=YES, Support Multi-Turn Text-to-SQL?=NO2025.10 | 23.3 |