Mathematical Reasoning on ASDiv (test)
97.24AccuracySGE
Evaluation Results
| Method | Links | |
|---|---|---|
| SGEModel=GPT-4, Tool=Code Interpreter2024.05 | 97.24 | |
| Agent-GWOBackbone=GPT-4o-mini2026.04 | 95.5 | |
| GoTBackbone=GPT-4o-mini2026.04 | 94.2 | |
| Agent-GWOBackbone=Gemma-3-12b-it2026.04 | 94.2 | |
| AoTBackbone=GPT-4o-mini2026.04 | 94.1 | |
| AoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 94.1 | |
| Decomp PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 94.08 | |
| CoT + Skill-BasedPrompting=Chain-of-Thought with Skill-Based exemplar selection, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 94.03 | |
| GoTBackbone=Gemma-3-12b-it2026.04 | 93.9 | |
| ToTBackbone=GPT-4o-mini2026.04 | 93.6 | |
| CoT + PALPrompting=Hybrid CoT + PAL, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 93.5 | |
| AoTBackbone=Gemma-3-12b-it2026.04 | 93.5 | |
| AFlowBackbone=GPT-4o-mini2026.04 | 93.4 | |
| CoT PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 93.32 | |
| CoT-SC/n=5Backbone=GPT-4o-mini2026.04 | 93.3 | |
| Agent-GWOBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 93.3 | |
| CoT-SC/n=5Backbone=Gemma-3-12b-it2026.04 | 93.2 | |
| Refine PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 93.1 | |
| CoTPrompting=Chain-of-Thought, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 92.7 | |
| GoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 92.5 | |
| CoTBackbone=GPT-4o-mini2026.04 | 92.3 | |
| Self-RefineBackbone=GPT-4o-mini2026.04 | 92.2 | |
| AFlowBackbone=Gemma-3-12b-it2026.04 | 92.2 | |
| ToTBackbone=Gemma-3-12b-it2026.04 | 91.8 | |
| CoTBackbone=Gemma-3-12b-it2026.04 | 91 | |
| PALPrompting=Program-Aided Language Models, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 90.2 | |
| IO PromptingModel=GPT-4, Tool=Code Interpreter2024.05 | 90.1 | |
| AFlowBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 90.1 | |
| CoT-SC/n=5Backbone=Qwen2.5-Coder-7B-Instruct2026.04 | 89.8 | |
| ToTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 89.8 | |
| Self-RefineBackbone=Gemma-3-12b-it2026.04 | 89.2 | |
| Self-RefineBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 87.6 | |
| CoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 87.3 | |
| MMOS-CODEModel Size=13B2024.02 | 81.9 | |
| TORA-CODEModel Size=13B2024.02 | 81.4 | |
| MMOSModel Size=13B2024.02 | 80 | |
| TORA-CODEModel Size=7B2024.02 | 78.7 | |
| MMOS-CODEModel Size=7B2024.02 | 78.6 | |
| TORAModel Size=13B2024.02 | 77.2 | |
| MMOSModel Size=7B2024.02 | 76.8 | |
| MMOS-Min-CODEModel Size=7B2024.02 | 76.7 | |
| TORAModel Size=7B2024.02 | 73.9 | |
| CRESCENT (Llama3-8B-Instruct)Training=CRESCENT, Prompting=0-shot2025.02 | 65.9 | |
| WizardMathModel Size=13B2024.02 | 65.8 | |
| CRESCENT (Llama3-8B-Instruct)Training=CRESCENT, Prompting=5-shot2025.02 | 63.8 | |
| Llama3-8B-InstructTraining=Original, Prompting=5-shot2025.02 | 62.3 | |
| WizardMathModel Size=7B2024.02 | 59.1 | |
| LLAMA-2 SFTModel Size=13B2024.02 | 58.6 | |
| LLaMA-2Model Size=13B2024.02 | 56.3 | |
| LLaMA-2Model Size=7B2024.02 | 50.7 | |
| LLAMA-2 SFTModel Size=7B2024.02 | 47.4 | |
| CRESCENT (Llama2-7B-Chat)Training=CRESCENT, Prompting=0-shot2025.02 | 46 | |
| Llama2-7B-ChatTraining=Original, Prompting=5-shot2025.02 | 45.9 | |
| CRESCENT (Llama2-7B-Chat)Training=CRESCENT, Prompting=5-shot2025.02 | 45.2 | |
| Llama3-8B-InstructTraining=Original, Prompting=0-shot2025.02 | 43.6 | |
| Llama2-7B-ChatTraining=Original, Prompting=0-shot2025.02 | 41.7 | |
| ToolformerEvaluation Protocol=zero-shot, Tool Use=enabled2023.02 | 40.4 | |
| ToolformerEvaluation Protocol=zero-shot, Tool Use=disabled2023.02 | 14.8 | |
| GPT-3Model Parameters=175B, Evaluation Protocol=zero-shot2023.02 | 14 | |
| GPT-J + CCEvaluation Protocol=zero-shot2023.02 | 9.6 | |
| GPT-JEvaluation Protocol=zero-shot2023.02 | 7.5 | |
| OPTModel Parameters=66B, Evaluation Protocol=zero-shot2023.02 | 6 |