Multi-turn Text-to-SQL on SparC (dev)
68.5QM EMgpt-oss-20b + Rose-SQL
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| gpt-oss-20b + Rose-SQLLearning Strategy=In-Context Learning2026.05 | 68.5 | 80.25 | 75.3 | 52.4 | 65.4 | 59.5 | |
| RASAT + PICARDApproach=Fine-tuned Model, Framework=PICARD2026.03 | 67.7 | 73.3 | — | 49.1 | 54 | — | |
| RASAT + PICARDLearning Strategy=Fine-tuned2026.05 | 67.7 | 73.3 | — | 49.1 | 54 | — | |
| Qwen3-14B + Rose-SQLLearning Strategy=In-Context Learning2026.05 | 66.45 | 79.1 | 74.44 | 50.25 | 63.82 | 58.04 | |
| Qwen3-8B + Rose-SQLLearning Strategy=In-Context Learning2026.05 | 65.75 | 78.72 | 72.73 | 50 | 62.56 | 55.21 | |
| SFT Mistral 7B + Track-SQLApproach=Fine-tuned, Backbone=Mistral 7B, Framework=Track-SQL2026.03 | 65.41 | 73.23 | 67.83 | 46.91 | 54.73 | 48.57 | |
| Track-SQLLearning Strategy=Fine-tuned2026.05 | 65.41 | 75.39 | 69.16 | 46.91 | 57.81 | 50.71 | |
| SFT Deepseek 7B + Track-SQLApproach=Fine-tuned, Backbone=Deepseek 7B, Framework=Track-SQL2026.03 | 65.17 | 75.39 | 69.16 | 46.44 | 57.81 | 50.71 | |
| HIE-SQL + GraPPAApproach=Fine-tuned Model, Backbone=GraPPA2026.03 | 64.7 | — | — | 45 | — | — | |
| HIE-SQL + GraPPaLearning Strategy=Fine-tuned2026.05 | 64.7 | — | — | 45 | — | — | |
| SFT Deepseek 7BApproach=Fine-tuned, Backbone=Deepseek 7B2026.03 | 64.33 | 71.4 | 65.08 | 43.36 | 50.71 | 43.36 | |
| SFT Mistral 7BApproach=Fine-tuned, Backbone=Mistral 7B2026.03 | 64.17 | 70.82 | 65.58 | 43.6 | 52.13 | 45.49 | |
| Qwen3-4B + Rose-SQLLearning Strategy=In-Context Learning2026.05 | 61.51 | 74.23 | 68.74 | 43.83 | 55.98 | 48.95 | |
| QDA-SQLApproach=Fine-tuned Model, Framework=QDA-SQL2026.03 | 61.3 | — | — | 44.1 | — | — | |
| QDA-SQLLearning Strategy=Fine-tuned2026.05 | 61.3 | — | — | 44.1 | — | — | |
| SFT Codellama 7B + Track-SQLApproach=Fine-tuned, Backbone=Codellama 7B, Framework=Track-SQL2026.03 | 61.26 | 70.07 | 63.75 | 43.36 | 52.13 | 45.73 | |
| gpt-oss-20bLearning Strategy=In-Context Learning2026.05 | 60.1 | 73.5 | 66.8 | 33.2 | 50.8 | 36.5 | |
| SFT Codellama 7BApproach=Fine-tuned, Backbone=Codellama 7B2026.03 | 59.18 | 67.99 | 61.51 | 36.96 | 46.68 | 39.33 | |
| CoE-SQLApproach=In-Context Learning, Backbone=GPT 3.52026.03 | 56 | 70.3 | 63.3 | 36.5 | 50.5 | 41.9 | |
| CoE-SQLLearning Strategy=In-Context Learning2026.05 | 56 | 70.3 | 63.3 | 36.5 | 50.5 | 41.9 | |
| Qwen3-14BLearning Strategy=In-Context Learning2026.05 | 55.19 | 71.98 | 65.25 | 29.38 | 49.05 | 40.52 | |
| Qwen3-4BLearning Strategy=In-Context Learning2026.05 | 54.11 | 66 | 58.27 | 28.67 | 40.28 | 31.04 | |
| Qwen3-8BLearning Strategy=In-Context Learning2026.05 | 53.12 | 66.99 | 61.68 | 26.78 | 41.94 | 35.31 | |
| ACT-SQLApproach=In-Context Learning, Backbone=GPT 3.52026.03 | 51 | 63.8 | 56.9 | 24.4 | 38.9 | 29.6 | |
| ACT-SQLLearning Strategy=In-Context Learning2026.05 | 51 | 63.8 | 56.9 | 24.4 | 38.9 | 29.6 | |
| DeepSeek-R1-Distill-Llama-8B + Rose-SQLLearning Strategy=In-Context Learning2026.05 | 42.5 | 61.25 | 54.3 | 18.5 | 32.8 | 25.4 | |
| DeepSeek-R1-Distill-Llama-8BLearning Strategy=In-Context Learning2026.05 | 31.58 | 49.13 | 43.89 | 10.43 | 21.09 | 16.82 |