Mathematical Reasoning on SVAMP (test)
94AccuracySelf-Contrast
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Self-ContrastLLM=GPT4, #Call Avg.=7.82024.01 | 94 | — | — | — | |
| CoT + PALPrompting=Hybrid CoT + PAL, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 93.7 | — | — | — | |
| Math-PromptLLM=GPT4, #Call Avg.=4.52024.01 | 93.6 | — | — | — | |
| SC-ReflectLLM=GPT4, #Call Avg.=92024.01 | 93.4 | — | — | — | |
| ExpertPromptLLM=GPT4, #Call Avg.=22024.01 | 93.3 | — | — | — | |
| SC-SelectLLM=GPT4, #Call Avg.=92024.01 | 93.2 | — | — | — | |
| Multi-AgentLLM=GPT4, #Call Avg.=92024.01 | 93.2 | — | — | — | |
| COMT+CCRLBase Model=QWEN3-4B2026.01 | 93.2 | — | — | — | |
| Hint-PromptLLM=GPT4, #Call Avg.=6.72024.01 | 93.1 | — | — | — | |
| FEW-SHOTBase Model=QWEN3-8B2026.01 | 93.1 | — | — | — | |
| CoT PromptLLM=GPT4, #Call Avg.=12024.01 | 93 | — | — | — | |
| CoT + Skill-BasedPrompting=Chain-of-Thought with Skill-Based exemplar selection, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 92.6 | — | — | — | |
| SC-VoteLLM=GPT4, #Call Avg.=82024.01 | 92.5 | — | — | — | |
| Agent-GWOBackbone=GPT-4o-mini2026.04 | 92.3 | — | — | — | |
| PALPrompting=Program-Aided Language Models, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 92.2 | — | — | — | |
| FEW-SHOTBase Model=QWEN2.5-7B2026.01 | 92.2 | — | — | — | |
| Entro-duction2025.03 | 92 | 11.2 | — | — | |
| CoTPrompting=Chain-of-Thought, Model=GPT-4-0613, Shots=4-shot, Decoding Strategy=Greedy decoding2024.05 | 91.9 | — | — | — | |
| COMT+CCRLBase Model=QWEN3-8B2026.01 | 91.9 | — | — | — | |
| AoTBackbone=Gemma-3-12b-it2026.04 | 91.8 | — | — | — | |
| FEW-SHOTBase Model=QWEN3-4B2026.01 | 91.6 | — | — | — | |
| Self-ReflectionLLM=GPT4, #Call Avg.=32024.01 | 91.5 | — | — | — | |
| AoTBackbone=GPT-4o-mini2026.04 | 91.5 | — | — | — | |
| COMTBase Model=QWEN3-8B2026.01 | 91.2 | — | — | — | |
| COT-SFT+RLBase Model=QWEN2.5-7B2026.01 | 90.9 | — | — | — | |
| Agent-GWOBackbone=Gemma-3-12b-it2026.04 | 90.9 | — | — | — | |
| DRR2025.03 | 90.2 | — | — | — | |
| COMTBase Model=QWEN3-4B2026.01 | 90.1 | — | — | — | |
| COMT+CCRLBase Model=QWEN2.5-7B2026.01 | 90.1 | — | — | — | |
| Agent-GWOBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 90.1 | — | — | — | |
| COT-SFT+RLBase Model=QWEN3-4B2026.01 | 90 | — | — | — | |
| AoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 90 | — | — | — | |
| CoT-SC@maj64majority_vote=642025.03 | 89.6 | 24 | — | — | |
| ToTBackbone=GPT-4o-mini2026.04 | 89.6 | — | — | — | |
| hybridModel=Qwen2.5Math-7B2025.02 | 89.5 | — | — | — | |
| GoTBackbone=GPT-4o-mini2026.04 | 89.2 | — | — | — | |
| Self-ContrastLLM=GPT3.5, #Call Avg.=7.82024.01 | 89 | — | — | — | |
| AFlowBackbone=GPT-4o-mini2026.04 | 88.7 | — | — | — | |
| TATAModel=Qwen2.5-14B2025.02 | 88.4 | — | — | — | |
| TATAModel=Qwen2.5Math-7B2025.02 | 88.1 | — | — | — | |
| COT-SFTBase Model=QWEN2.5-7B2026.01 | 87.9 | — | — | — | |
| CoT-SC@maj8majority_vote=82025.03 | 87.5 | 24 | — | — | |
| GoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 87.5 | — | — | — | |
| COMTBase Model=QWEN2.5-7B2026.01 | 87.3 | — | — | — | |
| GoTBackbone=Gemma-3-12b-it2026.04 | 86.8 | — | — | — | |
| COT-SFTBase Model=QWEN3-4B2026.01 | 86.7 | — | — | — | |
| PaLMPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | 86.6 | — | — | — | |
| Complex CoT2025.03 | 86.2 | 8 | — | — | |
| COMT+CCRLBase Model=LLAMA3.1-8B2026.01 | 86.2 | — | — | — | |
| ensembleModel=Qwen2.5-3B2025.02 | 86.2 | — | — | — | |
| TATAModel=Qwen2.5-7B2025.02 | 86.2 | — | — | — | |
| Self-AgreementModel=GPT-3.5-turbo, Scenario=Scenario 32023.11 | 86 | — | — | — | |
| AFlowBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 86 | — | — | — | |
| GPT-3.5-turboPrompting Strategy=Self-Consistency, Scenario=Scenario 32023.11 | 85.9 | — | — | — | |
| CoT-SC/n=5Backbone=GPT-4o-mini2026.04 | 85.8 | — | — | — | |
| TATAModel=Qwen2.5Math-1.5B2025.02 | 85.6 | — | — | — | |
| ToTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 85.6 | — | — | — | |
| QWEN2.5-MATHNote=Specialized model reference2026.01 | 85.5 | — | — | — | |
| TATAModel=Qwen2.5-3B2025.02 | 85.3 | — | — | — | |
| AFlowBackbone=Gemma-3-12b-it2026.04 | 85.2 | — | — | — | |
| COT-SFT+RLBase Model=LLAMA3.1-8B2026.01 | 85.1 | — | — | — | |
| GPT-SelectModel=Qwen2.5Math-7B2025.02 | 85.1 | — | — | — | |
| ZERO-SHOTBase Model=QWEN3-4B2026.01 | 85 | — | — | — | |
| COT-SFT+RLBase Model=QWEN3-8B2026.01 | 85 | — | — | — | |
| Self-RefineBackbone=GPT-4o-mini2026.04 | 85 | — | — | — | |
| CoTBackbone=GPT-4o-mini2026.04 | 84.7 | — | — | — | |
| SC-VoteLLM=GPT3.5, #Call Avg.=82024.01 | 84.6 | — | — | — | |
| hybridModel=Qwen2.5-14B2025.02 | 84.5 | — | — | — | |
| ensembleModel=Qwen2.5Math-7B2025.02 | 84.5 | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-MultipleZero-shot evaluation=true2024.10 | 84.33 | — | — | — | |
| hybridModel=Qwen2.5-7B2025.02 | 84.3 | — | — | — | |
| Multi-AgentLLM=GPT3.5, #Call Avg.=92024.01 | 84.1 | — | — | — | |
| FEW-SHOTBase Model=LLAMA3.1-8B2026.01 | 84 | — | — | — | |
| ensembleModel=Qwen2.5Math-1.5B2025.02 | 83.9 | — | — | — | |
| CoT-SC/n=5Backbone=Qwen2.5-Coder-7B-Instruct2026.04 | 83.9 | — | — | — | |
| ToTBackbone=Gemma-3-12b-it2026.04 | 83.9 | — | — | — | |
| Self-talk2025.03 | 83.7 | — | — | — | |
| GPT-SelectModel=Qwen2.5Math-1.5B2025.02 | 83.7 | — | — | — | |
| hybridModel=Qwen2.5Math-1.5B2025.02 | 83.6 | — | — | — | |
| GPT-3.5-turboPrompting Strategy=USC, Scenario=Scenario 32023.11 | 83.5 | — | — | — | |
| CoT2025.03 | 83.4 | 8 | — | — | |
| GPT-SelectModel=Qwen2.5-7B2025.02 | 83.4 | — | — | — | |
| ToT2025.03 | 83.3 | 121 | — | — | |
| Self-RefineBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 83.3 | — | — | — | |
| COMTBase Model=LLAMA3.1-8B2026.01 | 83.1 | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GPTZero-shot evaluation=true2024.10 | 83 | — | — | — | |
| ZERO-SHOTBase Model=QWEN2.5-7B2026.01 | 83 | — | — | — | |
| ensembleModel=Qwen2.5-14B2025.02 | 82.8 | — | — | — | |
| TATAModel=LLaMA-3-8B2025.02 | 82.7 | — | — | — | |
| CoTBackbone=Qwen2.5-Coder-7B-Instruct2026.04 | 82.7 | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-Multiple (w/o Peer-Review)Zero-shot evaluation=true2024.10 | 82.67 | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-GeminiZero-shot evaluation=true2024.10 | 82.33 | — | — | — | |
| GPT-3.5-TurboZero-shot evaluation=true2024.10 | 82.3 | — | — | — | |
| DEEPSEEK-MATHNote=Specialized model reference2026.01 | 82.2 | — | — | — | |
| GPT-3.5-turboPrompting Strategy=Few-Shot CoT, Scenario=Scenario 32023.11 | 82 | — | — | — | |
| Llama3.1-8B+ReDistillZero-shot evaluation=true2024.10 | 82 | — | — | — | |
| Llama3.1-8B-Instruct (FAIR) + Teacher-MixtralZero-shot evaluation=true2024.10 | 82 | — | — | — | |
| ensembleModel=Qwen2.5-7B2025.02 | 82 | — | — | — | |
| GPT-SelectModel=Qwen2.5-1.5B2025.02 | 81.8 | — | — | — | |
| Llama3.1-8B-InstructZero-shot evaluation=true2024.10 | 81.67 | — | — | — |