Text-to-SQL on EHR-Bench
45.5Execution AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4oDecoding=Maj@82025.05 | 45.5 | |
| GPT-4oDecoding=Greedy2025.05 | 44.9 | |
| GPT-4-TurboDecoding=Maj@82025.05 | 44.8 | |
| DeepSeek-V3 (671B MoE)Decoding=Maj@82025.05 | 43.5 | |
| DeepSeek-V3 (671B MoE)Decoding=Greedy2025.05 | 43.2 | |
| GPT-4-TurboDecoding=Greedy2025.05 | 43.1 | |
| Qwen2.5-72B-InstructDecoding=Maj@82025.05 | 41.2 | |
| OmniSQL + Qwen2.5-7BDecoding=Maj@82025.05 | 39.5 | |
| Arctic-SQL + Qwen2.5-7BDecoding=Maj@82025.05 | 38.9 | |
| SQL-R1 + Qwen2.5-7BDecoding=Maj@82025.05 | 38.8 | |
| Qwen2.5-Coder-7B-InstructDecoding=Maj@82025.05 | 36.9 | |
| Arctic-SQL + Qwen2.5-7BDecoding=Greedy2025.05 | 36.3 | |
| SQL-R1 + Qwen2.5-7BDecoding=Greedy2025.05 | 36 | |
| Qwen2.5-72B-InstructDecoding=Greedy2025.05 | 35 | |
| Reward-SQL + Qwen3-8BDecoding=Maj@82025.05 | 34.4 | |
| OmniSQL + Qwen2.5-7BDecoding=Greedy2025.05 | 34.3 | |
| Meta-Llama-3.1-8B-InstructDecoding=Maj@82025.05 | 33.7 | |
| Reward-SQL + Qwen3-8BDecoding=Greedy2025.05 | 29.6 | |
| Meta-Llama-3.1-8B-InstructDecoding=Greedy2025.05 | 24.6 | |
| Qwen2.5-Coder-7B-InstructDecoding=Greedy2025.05 | 24.3 |