Question Answering on RefuteBench 1.0 (test)
98.5FAClaude-2
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Claude-22024.02 | 98.5 | 97 | 94.5 | 74.49 | 59.66 | 65.86 | |
| GPT-42024.02 | 83 | 95 | 94.5 | 73.45 | 69.68 | 68.89 | |
| LLAMA-2-13B-ChatParameters=13B, Model Type=Chat2024.02 | 75 | 76 | 37 | 54.93 | 24.12 | 31.72 | |
| LLAMA-2-7B-ChatParameters=7B, Model Type=Chat2024.02 | 70 | 65.5 | 11 | 41.4 | 11.73 | 12.86 | |
| ALPACA-7BParameters=7B2024.02 | 64 | 43 | 16 | 34.15 | 24.2 | 26.22 | |
| Mistral-7B-Instruct-v0.2Parameters=7B, Model Type=Instruct2024.02 | 8 | 15 | 16.5 | 17.03 | 14.59 | 12.91 | |
| ChatGPT2024.02 | 6.5 | 17.5 | 13 | 13.93 | 3 | 10.17 |