Mathematical Reasoning on GSM8K (success rate)
93.6Success RateAFlow
Evaluation Results
| Method | Links | |
|---|---|---|
| AFlowMethod category=Automatically designed multi-agent frameworks, Executor LLM=GPT-4o-mini, Execution protocol=multi-agent2026.01 | 93.6 | |
| OneFlowMethod category=Single-LLM implementation, Executor LLM=GPT-4o-mini, Execution protocol=single-agent execution2026.01 | 93.3 | |
| OneFlowMethod category=Automatically designed multi-agent frameworks, Executor LLM=GPT-4o-mini, Execution protocol=multi-agent2026.01 | 93 | |
| AFlowMethod category=Single-LLM implementation, Executor LLM=GPT-4o-mini, Execution protocol=single-agent execution2026.01 | 92.9 | |
| CoT SCMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard, shots=5-shot2026.01 | 92.6 | |
| Kimi-K2 BaseArchitecture=MoE, Activated Params=32B, Total Params=1043B2026.02 | 92.1 | |
| DeepSeek-V3 BaseArchitecture=MoE, Activated Params=37B, Total Params=671B2026.02 | 87.6 | |
| CoTMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard2026.01 | 87.1 | |
| MultiPersonaMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard2026.01 | 87.1 | |
| IOMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard2026.01 | 87 | |
| OptoPrimeimplementation=Trace, optimizer=GPT-4o-2024-08-06, student model=GPT-3.5-turbo-11062024.06 | 82.5 | |
| TextGradversion=24-10-30, optimizer=GPT-4o-2024-08-06, student model=GPT-3.5-turbo-11062024.06 | 82.4 | |
| TextGradimplementation=Trace, optimizer=GPT-4o-2024-08-06, student model=GPT-3.5-turbo-11062024.06 | 82 | |
| TextGradsource=Reported, optimizer=GPT-4o-2024-08-06, student model=GPT-3.5-turbo-11062024.06 | 81.1 | |
| GLM-4.5 BaseArchitecture=MoE, Activated Params=32B, Total Params=355B2026.02 | 79.4 | |
| GLM-5 BaseArchitecture=MoE, Activated Params=40B, Total Params=744B2026.02 | 68.8 |