Mathematical Reasoning on AIME 25 (Avg@8 accuracy)
61.25AIME 25 Avg@8 AccuracyMAS-Orchestra
Evaluation Results
| Method | Links | |
|---|---|---|
| MAS-OrchestraOrchestration Type=Training-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 61.25 | |
| DebateAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 57.5 | |
| AFlowOrchestration Type=Inference-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 53.33 | |
| SCAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 51.67 | |
| ReflexionAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 50.42 | |
| CoTAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 45 | |
| MAS-GPTOrchestration Type=Public Training-time Orchestration, Sub-agent Backbone=GPT-120b (low)2026.01 | 43.33 | |
| MaASOrchestration Type=Inference-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 40.83 | |
| ToolOrchestraOrchestration Type=Public Training-time Orchestration, Sub-agent Backbone=GPT-120b (low)2026.01 | 11.25 |