Response Generation on AdvBench
0.95Win RateChain-of-Thought
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Chain-of-ThoughtMain Model=Vicuna, Inference Strategy=Chain-of-Thought2025.06 | 0.95 | 0.04 | 0.01 | |
| Best-of-NMain Model=Vicuna, Inference Strategy=Best-of-N2025.06 | 0.95 | 0.03 | 0.02 | |
| Multi-Agent DebateMain Model=Vicuna, Inference Strategy=Multi-Agent Debate2025.06 | 0.92 | 0.06 | 0.02 |