Text-to-SQL on 5-benchmark aggregate Multi-turn
76.5AccuracyRefGRPO
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| RefGRPOBackbone=Qwen2.5-Coder-7B-Instruct, Turns=Multi-turn, Decoding Strategy=Greedy2026.06 | 76.5 | 77.8 | 22.4 | 7.7 | 76.5 | |
| GRPO+Backbone=Qwen2.5-Coder-7B-Instruct, Turns=Multi-turn, Decoding Strategy=Greedy2026.06 | 75.1 | 76.2 | 22.8 | 44.4 | 73 | |
| RefGRPOBackbone=Qwen2.5-Coder-3B-Instruct, Turns=Multi-turn, Decoding Strategy=Greedy2026.06 | 72.9 | 76.3 | 24.8 | 2.2 | 73.1 | |
| GRPO+Backbone=Qwen2.5-Coder-3B-Instruct, Turns=Multi-turn, Decoding Strategy=Greedy2026.06 | 72.5 | 75.8 | 24 | 36.5 | 70.8 | |
| BaseBackbone=Qwen2.5-Coder-7B-Instruct, Turns=Multi-turn, Decoding Strategy=Greedy2026.06 | 68.7 | 73.1 | 25 | 54.3 | 66 | |
| BaseBackbone=Qwen2.5-Coder-3B-Instruct, Turns=Multi-turn, Decoding Strategy=Greedy2026.06 | 57.2 | 70.7 | 32.8 | 18.7 | 55.8 |