Text-to-SQL on Ambrosia (out-of-domain)
84.4Single Interpretation CoverageOurs
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OursApproach=Disambiguate-and-Parse, Method=Interpretation Infilling, Backbone=Llama-3.1 8B2025.02 | 84.4 | 1,880 | 51.9 | 24.2 | |
| Interp. PromptApproach=Disambiguate-and-Parse, Method=LLM Prompting, Backbone=Llama-3.1 8B2025.02 | 81.9 | 1,690 | 49 | 29.1 | |
| w. Self-CorrectionApproach=Disambiguate-and-Parse, Self-Correction=true, Backbone=Llama-3.1 8B2025.02 | 65.7 | 590 | 34.5 | 29.4 | |
| Gold Interp. SFTApproach=Disambiguate-and-Parse, Training=SFT on interpretations, Backbone=Llama-3.1 8B2025.02 | 62.6 | 30 | 32.1 | 49.5 | |
| SFTApproach=End-to-End Text-to-SQL, Training=SFT with LoRA, Backbone=Llama-3.1 8B2025.02 | 38 | 40 | 20 | 29.4 | |
| 3-shot PromptApproach=End-to-End Text-to-SQL, Shots=3, Backbone=Llama-3.1 8B2025.02 | 35.7 | 130 | 17.5 | 21.3 | |
| 0-shot PromptApproach=End-to-End Text-to-SQL, Shots=0, Backbone=Llama-3.1 8B2025.02 | 29.4 | 90 | 15 | 21.9 |