Generative Multiple-choice Question Answering on TruthfulQA
76.3TA RateLlama 2-Chat
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Llama 2-ChatBase Model=Llama 2-Chat, Selection Strategy=None2024.03 | 76.3 | 13.7 | 45 | |
| Mistral-Instruct-v0.2Base Model=Mistral-Instruct-v0.2, Selection Strategy=None2024.03 | 75.4 | 22.6 | 49 | |
| Mistral-Instruct-v0.2 + TACS-SBase Model=Mistral-Instruct-v0.2, Selection Strategy=TACS-S2024.03 | 46.3 | 91.4 | 68.9 | |
| Mistral-Instruct-v0.2 + TACS-TBase Model=Mistral-Instruct-v0.2, Selection Strategy=TACS-T2024.03 | 44.9 | 89.6 | 67.2 | |
| Llama 2-Chat + TACS-SBase Model=Llama 2-Chat, Selection Strategy=TACS-S2024.03 | 43.7 | 74.9 | 64.3 | |
| Llama 2-Chat + TACS-TBase Model=Llama 2-Chat, Selection Strategy=TACS-T2024.03 | 43.4 | 85.8 | 64.7 |