Generative Multiple-choice Question Answering on ConflictQA
98.8TA RateMistral-Instruct-v0.2
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Mistral-Instruct-v0.2Base Model=Mistral-Instruct-v0.2, Selection Strategy=None2024.03 | 98.8 | 12.8 | 55.3 | — | |
| Mistral-Instruct-v0.2 + TACS-TBase Model=Mistral-Instruct-v0.2, Selection Strategy=TACS-T2024.03 | 98 | 17.3 | 57.7 | — | |
| Llama 2-ChatBase Model=Llama 2-Chat, Selection Strategy=None2024.03 | 97.4 | 12.2 | 54.8 | — | |
| Llama 2-Chat + TACS-TBase Model=Llama 2-Chat, Selection Strategy=TACS-T2024.03 | 95.7 | 24.9 | 60.3 | — | |
| Mistral-Instruct-v0.2 + TACS-SBase Model=Mistral-Instruct-v0.2, Selection Strategy=TACS-S2024.03 | 88.3 | 58.5 | 73.4 | — | |
| Llama 2-Chat + TACS-SBase Model=Llama 2-Chat, Selection Strategy=TACS-S2024.03 | 83.9 | 58.7 | 71.3 | — | |
| Llama 2-ChatBackbone=Llama 2-Chat, Model Size=7B2024.03 | — | — | — | 79.9 | |
| Mistral-Instruct-v0.2Backbone=Mistral-Instruct-v0.2, Model Size=7B2024.03 | — | — | — | 80 | |
| TACS-S (Sentence-level)Backbone=Llama 2-Chat, Model Size=7B, Granularity=Sentence-level2024.03 | — | — | — | 81.2 | |
| TACS-S (Sentence-level)Backbone=Mistral-Instruct-v0.2, Model Size=7B, Granularity=Sentence-level2024.03 | — | — | — | 81 | |
| TACS-T (Token-level)Backbone=Llama 2-Chat, Model Size=7B, Granularity=Token-level2024.03 | — | — | — | 81.3 | |
| TACS-T (Token-level)Backbone=Mistral-Instruct-v0.2, Model Size=7B, Granularity=Token-level2024.03 | — | — | — | 83.2 |