Reasoning on Multi-Step Arithmetic
99.2AccuracyRoT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RoTBackbone=GPT-4, Shots=2-shot2024.10 | 99.2 | — | |
| RoTBackbone=GPT-4, Shots=1-shot2024.10 | 98.4 | — | |
| BOTBackbone=GPT-4, Shots=2-shot2024.10 | 96.8 | — | |
| BOTBackbone=GPT-4, Shots=1-shot2024.10 | 94.8 | — | |
| Meta-promptingBackbone=GPT-4, Shots=2-shot2024.10 | 90.7 | — | |
| TOTBackbone=GPT-4, Shots=2-shot2024.10 | 90.3 | — | |
| Meta-promptingBackbone=GPT-4, Shots=1-shot2024.10 | 89.6 | — | |
| RoTBackbone=GPT-3.5-turbo, Shots=2-shot2024.10 | 89.5 | — | |
| RoTBackbone=GPT-3.5-turbo, Shots=1-shot2024.10 | 89.2 | — | |
| TOTBackbone=GPT-4, Shots=1-shot2024.10 | 88.8 | — | |
| GOTBackbone=GPT-4, Shots=2-shot2024.10 | 88.7 | — | |
| BOTBackbone=GPT-3.5-turbo, Shots=2-shot2024.10 | 88.3 | — | |
| BOTBackbone=GPT-3.5-turbo, Shots=1-shot2024.10 | 87.6 | — | |
| GOTBackbone=GPT-4, Shots=1-shot2024.10 | 87.6 | — | |
| COTBackbone=GPT-4, Shots=2-shot2024.10 | 85.5 | — | |
| Meta-promptingBackbone=GPT-3.5-turbo, Shots=2-shot2024.10 | 85.1 | — | |
| Meta-promptingBackbone=GPT-3.5-turbo, Shots=1-shot2024.10 | 83.1 | — | |
| COTBackbone=GPT-4, Shots=1-shot2024.10 | 83.1 | — | |
| TOTBackbone=GPT-3.5-turbo, Shots=2-shot2024.10 | 82.3 | — | |
| TOTBackbone=GPT-3.5-turbo, Shots=1-shot2024.10 | 81.9 | — | |
| GOTBackbone=GPT-3.5-turbo, Shots=1-shot2024.10 | 79.1 | — | |
| GOTBackbone=GPT-3.5-turbo, Shots=2-shot2024.10 | 79 | — | |
| COTBackbone=GPT-3.5-turbo, Shots=2-shot2024.10 | 76.2 | — | |
| COTBackbone=GPT-3.5-turbo, Shots=1-shot2024.10 | 73.9 | — | |
| TurboConnModel=Qwen3-1.7B, Group Size=4, Finetuning=true2026.02 | 45.81 | 5.259 | |
| BaselineModel=Qwen3-8B, Finetuning=true2026.02 | 45.61 | 4.5 | |
| BaselineModel=Qwen3-4B, Finetuning=true2026.02 | 43.68 | 4.609 | |
| BaselineModel=Qwen3-1.7B, Finetuning=true2026.02 | 36.1 | 2.516 |