Mathematical Reasoning on AGIEval MATH (Accuracy)
95.7AccuracyUPA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| UPAExecutor=GPT-52026.01 | 95.7 | — | — | |
| IOExecutor=GPT-52026.01 | 95.3 | — | — | |
| SPOExecutor=GPT-52026.01 | 94.9 | — | — | |
| CoTExecutor=GPT-52026.01 | 94.8 | — | — | |
| Agent 1 (Qwen3-30B-A3B-Instruct)Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 93.3 | — | — | |
| Agent 1 (Qwen3-30B-A3B-Instruct)Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 93.3 | — | — | |
| UPAExecutor=DeepSeek-V3.22026.01 | 93.1 | — | — | |
| CoTExecutor=DeepSeek-V3.22026.01 | 91.7 | — | — | |
| IOExecutor=DeepSeek-V3.22026.01 | 89.5 | — | — | |
| Iterative Critique-and-Routing ControllerController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 88.1 | — | — | |
| Iterative Critique-and-Routing ControllerController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 87.9 | — | — | |
| UPAExecutor=Claude-4.5-Sonnet2026.01 | 86.6 | — | — | |
| SPOExecutor=DeepSeek-V3.22026.01 | 86.3 | — | — | |
| CoTExecutor=Claude-4.5-Sonnet2026.01 | 86.2 | — | — | |
| IOExecutor=Claude-4.5-Sonnet2026.01 | 85.9 | — | — | |
| SPOExecutor=Claude-4.5-Sonnet2026.01 | 84.7 | — | — | |
| Controller V2Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 80.9 | — | — | |
| Router-R1Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 76.3 | — | — | |
| Qwen3-I(4B)Model Family=Qwen3, Target Model (TL)=ℐ (4B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 75.6 | — | — | |
| Random RouterController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 73.7 | — | — | |
| RouterDCController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 73.6 | — | — | |
| Controller V1Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 73.6 | — | — | |
| Agent 2 (Qwen2.5-7B-Instruct)Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 73.4 | — | — | |
| GAPGModel=Qwen2.5-1.5B2025.09 | 73.4 | — | — | |
| RoBERTa RouterController=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 72.9 | — | — | |
| Controller V2Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 72.4 | — | — | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 4B, Transfer Setting=Second Setting (Red Shading)2026.04 | 71.3 | — | — | |
| Ministral-3-I(14B)Model Family=Ministral-3, Target Model (TL)=ℐ (14B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 70.6 | — | — | |
| Router-R1Model=Qwen2.5-1.5B2025.09 | 70.4 | — | — | |
| Ministral-3-I(3B)Model Family=Ministral-3, Target Model (TL)=ℐ (3B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 68.7 | — | — | |
| Agent 2 (Ministral-3-8B-Instruct)Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 68.1 | — | — | |
| Qwen3-I(14B)Model Family=Qwen3, Target Model (TL)=ℐ (14B), Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 67.8 | — | — | |
| Router-R1Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 66.6 | — | — | |
| RouterDCController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 65.4 | — | — | |
| GAPGModel=LLaMA-3.2-3B2025.09 | 64.5 | — | — | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 4B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 64.1 | — | — | |
| Random RouterController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 64 | — | — | |
| RoBERTa RouterController=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 62.8 | — | — | |
| Qwen3-14BModel Family=Qwen3, Target Model (TL)=14B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 61.1 | — | — | |
| Controller V1Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 59.7 | — | — | |
| Router-R1Model=LLaMA-3.2-3B2025.09 | 59.2 | — | — | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=4B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 14B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 58.9 | — | — | |
| AutoMixModel=Qwen2.5-1.5B2025.09 | 54.9 | — | — | |
| Agent 3 (Qwen2.5-1.5B-Instruct)Controller=Qwen2.5-7B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Qwen2.5-7B-Instruct, Agent 3=Qwen2.5-1.5B-Instruct2026.05 | 54.7 | — | — | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 3B, Transfer Setting=Second Setting (Red Shading)2026.04 | 54 | — | — | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=3B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 8B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 53.4 | — | — | |
| UNLOCKModel Family=Qwen3, Target Model (TL)=4B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 14B, Transfer Setting=Second Setting (Red Shading)2026.04 | 52.4 | — | — | |
| Qwen3-4BModel Family=Qwen3, Target Model (TL)=4B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 52.3 | — | — | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 3B, Transfer Setting=Task-Conditioned Transfer With Limited Data2026.04 | 51.9 | — | — | |
| Ministral-3-8BModel Family=Ministral-3, Target Model (TL)=8B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 50.7 | — | — | |
| AutoMixModel=LLaMA-3.2-3B2025.09 | 50.1 | — | — | |
| UNLOCKModel Family=Ministral-3, Target Model (TL)=3B, Unlock/Instruction-tuning Configuration (SU)=+UNLOCK from 8B, Transfer Setting=Second Setting (Red Shading)2026.04 | 49.9 | — | — | |
| Step-backLLM=GPT, Avg. Cost ($)=-2026.04 | 47.5 | 65 | — | |
| Ministral-3-3BModel Family=Ministral-3, Target Model (TL)=3B, Unlock/Instruction-tuning Configuration (SU)=-2026.04 | 46.9 | — | — | |
| DisambiguSLMLLM=GPT, Avg. Cost ($)=0.022026.04 | 46.7 | 68.6 | — | |
| OPROLLM=GPT, Avg. Cost ($)=4.512026.04 | 46.1 | 66.6 | — | |
| SPOLLM=GPT, Avg. Cost ($)=0.152026.04 | 46.1 | 66.9 | — | |
| Prompt BreederLLM=GPT, Avg. Cost ($)=4.822026.04 | 45.9 | 64.5 | — | |
| T-AllBackbone=LLaMA-3.1-8B-Instruct2026.05 | 45.6 | — | 73.6 | |
| Step-backLLM=DeepSeek, Avg. Cost ($)=-2026.04 | 45.2 | 62.3 | — | |
| T-MLPBackbone=LLaMA-3.1-8B-Instruct2026.05 | 45.1 | — | 78.2 | |
| DisambiguSLMLLM=DeepSeek, Avg. Cost ($)=0.022026.04 | 44.9 | 66.3 | — | |
| CoTLLM=GPT, Avg. Cost ($)=-2026.04 | 44.5 | 63.8 | — | |
| APELLM=GPT, Avg. Cost ($)=9.072026.04 | 44.4 | 64.8 | — | |
| TextGradLLM=GPT, Avg. Cost ($)=13.142026.04 | 44.4 | 63.9 | — | |
| SPOLLM=DeepSeek, Avg. Cost ($)=0.152026.04 | 44.2 | 64.8 | — | |
| DisambiguSLMLLM=LLaMA, Avg. Cost ($)=0.022026.04 | 44.2 | 66 | — | |
| Step-backLLM=LLaMA, Avg. Cost ($)=-2026.04 | 44.1 | 61.4 | — | |
| OPROLLM=DeepSeek, Avg. Cost ($)=4.512026.04 | 44 | 63.9 | — | |
| Prompt BreederLLM=DeepSeek, Avg. Cost ($)=4.822026.04 | 43.8 | 62.2 | — | |
| SPOLLM=LLaMA, Avg. Cost ($)=0.152026.04 | 43.5 | 64 | — | |
| OPROLLM=LLaMA, Avg. Cost ($)=4.512026.04 | 43.3 | 63.1 | — | |
| APELLM=DeepSeek, Avg. Cost ($)=9.072026.04 | 42.9 | 62.4 | — | |
| Prompt BreederLLM=LLaMA, Avg. Cost ($)=4.822026.04 | 42.9 | 61.4 | — | |
| CoTLLM=DeepSeek, Avg. Cost ($)=-2026.04 | 42.6 | 61.2 | — | |
| APELLM=LLaMA, Avg. Cost ($)=9.072026.04 | 42.5 | 61.7 | — | |
| Naïve promptingLLM=GPT, Avg. Cost ($)=-2026.04 | 42.1 | 61.6 | — | |
| RephraseLLM=GPT, Avg. Cost ($)=-2026.04 | 42.1 | 61.1 | — | |
| TextGradLLM=DeepSeek, Avg. Cost ($)=13.142026.04 | 42.1 | 61.6 | — | |
| CoTLLM=LLaMA, Avg. Cost ($)=-2026.04 | 42 | 60.4 | — | |
| TextGradLLM=LLaMA, Avg. Cost ($)=13.142026.04 | 41.8 | 60.9 | — | |
| Prompt AgentLLM=GPT, Avg. Cost ($)=2.712026.04 | 41.4 | 65 | — | |
| Prompt AgentLLM=DeepSeek, Avg. Cost ($)=2.712026.04 | 40.7 | 62.7 | — | |
| RephraseLLM=DeepSeek, Avg. Cost ($)=-2026.04 | 40.3 | 58.8 | — | |
| Naïve promptingLLM=DeepSeek, Avg. Cost ($)=-2026.04 | 40.2 | 59.2 | — | |
| Prompt AgentLLM=LLaMA, Avg. Cost ($)=2.712026.04 | 40.1 | 62 | — | |
| RephraseLLM=LLaMA, Avg. Cost ($)=-2026.04 | 39.8 | 57.93 | — | |
| Naïve promptingLLM=LLaMA, Avg. Cost ($)=-2026.04 | 39.6 | 58.2 | — | |
| BaseBackbone=LLaMA-3.1-8B-Instruct2026.05 | 38.5 | — | 72.3 | |
| T-MLPBackbone=LLaMA-3.2-3B-Instruct2026.05 | 36.2 | — | 67.3 | |
| T-AllBackbone=LLaMA-3.2-3B-Instruct2026.05 | 35.8 | — | 65 | |
| Vanilla SFTBackbone=LLaMA-3.1-8B-Instruct2026.05 | 33.1 | — | 65.5 | |
| BaseBackbone=LLaMA-3.2-3B-Instruct2026.05 | 27.6 | — | 58.6 | |
| Agent 3 (Llama-3-2-1B-Instruct)Controller=Qwen3-4B-Base, Agent 1=Qwen3-30B-A3B-Instruct, Agent 2=Ministral-3-8B-Instruct, Agent 3=Llama-3-2-1B-Instruct2026.05 | 26.5 | — | — | |
| Vanilla SFTBackbone=LLaMA-3.2-3B-Instruct2026.05 | 25.1 | — | 57.3 | |
| T-AllBackbone=Mistral-7B-Instruct2026.05 | 19.3 | — | 50.3 | |
| T-MLPBackbone=Mistral-7B-Instruct2026.05 | 18.5 | — | 51.8 | |
| Vanilla SFTBackbone=Mistral-7B-Instruct2026.05 | 15.1 | — | 38.2 | |
| BaseBackbone=Mistral-7B-Instruct2026.05 | 13 | — | 43.2 |