Mathematical Reasoning on AIME 24 (Avg@8 accuracy)
66.25Avg@8 AccuracyMAS-Orchestra
Evaluation Results
| Method | Links | |
|---|---|---|
| MAS-OrchestraOrchestration Type=Training-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 66.25 | |
| AFlowOrchestration Type=Inference-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 62.5 | |
| DebateAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 62.08 | |
| ReflexionAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 60.83 | |
| MAS-GPTOrchestration Type=Public Training-time Orchestration, Sub-agent Backbone=GPT-120b (low)2026.01 | 58.75 | |
| SCAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 57.5 | |
| CoTAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 50 | |
| MaASOrchestration Type=Inference-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 32.5 | |
| ToolOrchestraOrchestration Type=Public Training-time Orchestration, Sub-agent Backbone=GPT-120b (low)2026.01 | 23.33 |