Pairwise comparison evaluation on Arena Human Preference Data
51.5AccuracyCollabEval
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CollabEvalModel Setting=multi-agents2026.03 | 51.5 | 1.517 | 53.2 | 9.07 | |
| Mistral LargeModel Setting=Single-LLM2026.03 | 50.5 | 1 | 54.95 | 5.25 | |
| Llama3 70bModel Setting=Single-LLM2026.03 | 48.8 | 1 | 55.47 | 0.39 | |
| Round-Table Agents EvalModel Setting=multi-agents2026.03 | 48.7 | 1.258 | 12.7 | 47.37 | |
| SonnetModel Setting=Single-LLM2026.03 | 48.4 | 1 | 48.06 | 13.95 |