Multi-Classification on LIAR Closed
26.99AccuracyTELLER
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TELLERLLM=FLAN-T5-large (780M)2024.02 | 26.99 | 18.04 | |
| w/ InterventionLLM=FLAN-T5-xxl (11B)2024.02 | 26.91 | 21.3 | |
| TELLERLLM=FLAN-T5-xxl (11B)2024.02 | 26.83 | 19.68 | |
| w/ InterventionLLM=FLAN-T5-large (780M)2024.02 | 26.28 | 18.49 | |
| TELLERLLM=Llama2 (13B)2024.02 | 25.81 | 17.71 | |
| Few-shotLLM=GPT-3.5-turbo2024.02 | 25.65 | 25.56 | |
| w/ InterventionLLM=FLAN-T5-xl (3B)2024.02 | 25.57 | 19.62 | |
| w/ InterventionLLM=Llama2 (13B)2024.02 | 25.1 | 16.78 | |
| TELLERLLM=FLAN-T5-xl (3B)2024.02 | 24.31 | 17.4 | |
| w/ InterventionLLM=Llama2 (7B)2024.02 | 23.92 | 15.14 | |
| TELLERLLM=Llama2 (7B)2024.02 | 23.29 | 15.51 | |
| DirectLLM=FLAN-T5-xxl (11B)2024.02 | 22.42 | 18.31 | |
| Few-shot COTLLM=GPT-3.5-turbo2024.02 | 20.69 | 17.2 | |
| DirectLLM=GPT-3.5-turbo2024.02 | 20.46 | 20.34 | |
| DirectLLM=FLAN-T5-xl (3B)2024.02 | 19.67 | 16.57 | |
| DirectLLM=FLAN-T5-base (250M)2024.02 | 19.43 | 11.79 | |
| DirectLLM=FLAN-T5-large (780M)2024.02 | 19.43 | 17.84 | |
| DirectLLM=FLAN-T5-small (80M)2024.02 | 18.17 | 9.28 | |
| DirectLLM=Llama2 (7B)2024.02 | 18.02 | 9.97 | |
| Few-shot LogicLLM=GPT-3.5-turbo2024.02 | 16.37 | 13.98 | |
| DirectLLM=Llama2 (13B)2024.02 | 7.32 | 2.85 | |
| Zero-shot COTLLM=GPT-3.5-turbo2024.02 | 7.16 | 9.2 |