Agentic tasks on BrowseComp
9.45AccuracyMAS-ZERO
Evaluation Results
| Method | Links | |
|---|---|---|
| MAS-ZEROBackbone LLM=GPT-4o, Search capability=with search2025.05 | 9.45 | |
| CoT-SCBackbone LLM=GPT-4o, Search capability=with search2025.05 | 8.66 | |
| Self-RefineBackbone LLM=GPT-4o, Search capability=with search2025.05 | 5.51 | |
| CoTBackbone LLM=GPT-4o, Search capability=with search2025.05 | 3.97 | |
| DebateBackbone LLM=GPT-4o, Search capability=with search2025.05 | 3.94 |