Multi-task Reasoning and Efficiency on AQUA-RAT, HumanEval, and GSM8K
60.9Average AccuracyNEXA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| NEXABackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 60.9 | 1.33 | 18,363,825 | |
| SelfOrg*Backbone=Qwen2.5-1.5B-Instruct, Number of agents=10, Communication=Single sequential round2026.05 | 60.36 | 2.67 | 28,405,941 | |
| SingleBackbone=Qwen2.5-1.5B-Instruct, Number of agents=12026.05 | 57.7 | 5.67 | 1,149,321 | |
| CoTBackbone=Qwen2.5-1.5B-Instruct, Prompting strategy=Chain-of-Thought2026.05 | 56.95 | 5.33 | 1,298,471 | |
| GPTSwarmBackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 54.01 | 6.33 | 38,021,213 | |
| AgentPruneBackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 52.71 | 4 | 31,444,709 | |
| GDesignerBackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 51.79 | 5.67 | 31,798,540 | |
| SCBackbone=Qwen2.5-1.5B-Instruct, Prompting strategy=Self-Consistency2026.05 | 46.96 | 5 | 11,215,153 |