Agentic Tasks on Frames
70.45AccuracyDebate
Evaluation Results
| Method | Links | |
|---|---|---|
| DebateBackbone LLM=GPT-4o, Search capability=with search2025.05 | 70.45 | |
| Self-RefineBackbone LLM=GPT-4o, Search capability=with search2025.05 | 67.89 | |
| MAS-ZEROBackbone LLM=GPT-4o, Search capability=with search2025.05 | 65.18 | |
| CoT-SCBackbone LLM=GPT-4o, Search capability=with search2025.05 | 63.58 | |
| CoTBackbone LLM=GPT-4o, Search capability=with search2025.05 | 59.76 |