Out-of-Domain Reasoning on GPQA
65.43Avg@8 AccuracyAFlow
Evaluation Results
| Method | Links | |
|---|---|---|
| AFlowOrchestration Type=Inference-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 65.43 | |
| MAS-OrchestraOrchestration Type=Training-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 65.21 | |
| DebateAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 64.14 | |
| MAS-GPTOrchestration Type=Public Training-time Orchestration, Sub-agent Backbone=GPT-120b (low)2026.01 | 63.51 | |
| SCAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 62.88 | |
| ReflexionAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 62.37 | |
| CoTAgentOrchestration Type=Standalone Agent, Sub-agent Backbone=GPT-120b (low)2026.01 | 60.54 | |
| MaASOrchestration Type=Inference-time Orchestration, Orchestrator LLM=Qwen-7b, Sub-agent Backbone=GPT-120b (low)2026.01 | 40.78 | |
| ToolOrchestraOrchestration Type=Public Training-time Orchestration, Sub-agent Backbone=GPT-120b (low)2026.01 | 29.8 |