Question Answering on TruthfulQA
86.6AccuracyLLaMA-3.1-8B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLaMA-3.1-8BApproach=Ours, Learning Paradigm=Fine-tuned2024.11 | 86.6 | — | — | — | |
| GPT-4oApproach=CoT, Learning Paradigm=Vanilla2024.11 | 85.4 | — | — | — | |
| SERCModel=14B2026.04 | 85 | — | — | 86.2 | |
| GPT-4oApproach=Zero-shot, Learning Paradigm=Vanilla2024.11 | 84.8 | — | — | — | |
| Re-ExModel=14B2026.04 | 84.5 | — | — | 85.3 | |
| LLaMA-3.1-8BApproach=Planning-token, Learning Paradigm=Fine-tuned2024.11 | 82.5 | — | — | — | |
| CoVe+RAGModel=14B2026.04 | 82.5 | — | — | 84.7 | |
| Qwen 2.5-7BApproach=Ours, Learning Paradigm=Fine-tuned2024.11 | 81.2 | — | — | — | |
| PSFT2025.08 | 80.19 | — | — | — | |
| SERCModel=8B2026.04 | 80 | — | — | 82.5 | |
| LLaMA-3.1-8BApproach=LoRA, Learning Paradigm=Fine-tuned2024.11 | 79.8 | — | — | — | |
| LLaMA-3.1-70BApproach=Zero-shot, Learning Paradigm=Vanilla2024.11 | 79.3 | — | — | — | |
| LLaMA-2-7BApproach=Ours, Learning Paradigm=Fine-tuned2024.11 | 78.6 | — | — | — | |
| SFT2025.08 | 77.89 | — | — | — | |
| Re-ExModel=8B2026.04 | 77.5 | — | — | 76.5 | |
| LLaMA-2-7BApproach=Planning-token, Learning Paradigm=Fine-tuned2024.11 | 77 | — | — | — | |
| LLaMA-2-7BApproach=LoRA, Learning Paradigm=Fine-tuned2024.11 | 76.7 | — | — | — | |
| LLaMA-3.1-70BApproach=CoT, Learning Paradigm=Vanilla2024.11 | 76.2 | — | — | — | |
| Qwen 2.5-7BApproach=Planning-token, Learning Paradigm=Fine-tuned2024.11 | 76.2 | — | — | — | |
| CoVe+RAGModel=8B2026.04 | 75 | — | — | 75.5 | |
| Qwen 2.5-7BApproach=Zero-shot, Learning Paradigm=Vanilla2024.11 | 72.6 | — | — | — | |
| Qwen 2.5-7BApproach=LoRA, Learning Paradigm=Fine-tuned2024.11 | 72.5 | — | — | — | |
| RARRModel=14B2026.04 | 70.5 | — | — | 76.5 | |
| RARRModel=8B2026.04 | 68.5 | — | — | 72.5 | |
| Initial AnswerModel=14B2026.04 | 67.5 | — | — | 66.7 | |
| MAT-STEERBackbone=Llama-3.1-8B2025.02 | 61.94 | — | — | — | |
| Full Fine-tuningModel Size=13B, Transfer Source=N/A2024.06 | 61.93 | — | — | — | |
| LoRA TuningModel Size=13B, Transfer Source=N/A2024.06 | 61.93 | — | — | — | |
| LLaMA-3.1-8BApproach=Zero-shot, Learning Paradigm=Vanilla2024.11 | 61.6 | — | — | — | |
| OursModel Size=13B, Transfer Source=7B2024.06 | 61.56 | — | — | — | |
| Proxy TuningModel Size=13B, Transfer Source=7B2024.06 | 61.02 | — | — | — | |
| Full Fine-tuningModel Size=13B, Transfer Source=7B2024.06 | 60.02 | — | — | — | |
| CoVeModel=14B2026.04 | 60 | — | — | 63 | |
| LinUCB+KLAlgorithm Family=Contextual Linear2026.02 | 59.5 | — | — | — | |
| LITOBackbone=Llama-3.1-8B2025.02 | 58.63 | — | — | — | |
| ForgettingBase model=LLaMA-3.1-8B, Granularity=token-level2025.08 | 58.39 | — | — | — | |
| Qwen 2.5-7BApproach=CoT, Learning Paradigm=Vanilla2024.11 | 56.7 | — | — | — | |
| NL-ITIBackbone=Llama-3.1-8B2025.02 | 56.67 | — | — | — | |
| DPOBackbone=Llama-3.1-8B, Protocol=Direct Preference Optimization2025.02 | 56.1 | — | — | — | |
| ICLBackbone=Llama-3.1-8B, Protocol=In-context learning2025.02 | 55.32 | — | — | — | |
| ICVBackbone=Llama-3.1-8B2025.02 | 55.21 | — | — | — | |
| RAdaptBackbone=Llama-3.1-8B2025.02 | 55.09 | — | — | — | |
| SFTBackbone=Llama-3.1-8B, Protocol=Supervised Fine-Tuning2025.02 | 54.02 | — | — | — | |
| MergeBackbone=Llama-3.1-8B, Protocol=Model Merging2025.02 | 53.26 | — | — | — | |
| ForgettingBase model=LLaMA-2-13B2025.08 | 52.82 | — | — | — | |
| ITIBackbone=Llama-3.1-8B2025.02 | 52.68 | — | — | — | |
| IgnoringBase model=LLaMA-3.1-8B, Granularity=token-level2025.08 | 52.38 | — | — | — | |
| LLaMA-3.1-8BApproach=CoT, Learning Paradigm=Vanilla2024.11 | 50.6 | — | — | — | |
| ForgettingBase model=LLaMA-3.2-3B, Granularity=token-level2025.08 | 50.32 | — | — | — | |
| Initial AnswerModel=8B2026.04 | 50 | — | — | 54.5 | |
| Llama-3.1-8BBackbone=Llama-3.1-8B, Protocol=Base model2025.02 | 49.91 | — | — | — | |
| No-RewriteAlgorithm Family=Base2026.02 | 49.6 | — | — | — | |
| LLAMA PRO INSTRUCTTraining Stage=SFT comparison2024.01 | 48.8 | — | — | — | |
| LaVINzero-shot=true, setup=Mc1_targets2023.05 | 47.9 | — | — | — | |
| ForgettingBase model=LLaMA-3.1-8B, Granularity=sequence-level2025.08 | 47.83 | — | — | — | |
| HaluSearch (MCTS)Backbone=Llama3.1-8B-Instruct2025.01 | 47.5 | — | — | — | |
| CoVeModel=8B2026.04 | 47.5 | — | — | 54.5 | |
| IgnoringBase model=LLaMA-3.2-3B, Granularity=token-level2025.08 | 47.23 | — | — | — | |
| IgnoringBase model=LLaMA-3.1-8B, Granularity=sequence-level2025.08 | 47.05 | — | — | — | |
| Base ModelModel Size=13B, Transfer Source=N/A2024.06 | 46.13 | — | — | — | |
| LLaMA2-ChatParameters=7B, Training Stage=SFT comparison2024.01 | 45.57 | — | — | — | |
| BaseBase model=LLaMA-3.1-8B2025.08 | 45.08 | — | — | — | |
| HaluSearch (MCTS)Backbone=Qwen2-7B-Instruct2025.01 | 45.07 | — | — | — | |
| ForgettingBase model=LLaMA-3.2-1B, Granularity=token-level2025.08 | 44.83 | — | — | — | |
| LLaMA2-7BNumber of Parameters=7B, Backbone=LLaMA2, Shots=0, Prepended Examples=62024.07 | 44.6 | — | — | — | |
| Full TokensBase model=LLaMA-3.1-8B, Training Protocol=standard SFT2025.08 | 44.51 | — | — | — | |
| WizardMathParameters=7B, Training Stage=SFT comparison2024.01 | 43.65 | — | — | — | |
| BoNBackbone=Llama3.1-8B-Instruct2025.01 | 43.5 | — | — | — | |
| IgnoringBase model=LLaMA-2-13B2025.08 | 43.01 | — | — | — | |
| Full TokensBase model=LLaMA-3.2-3B, Training Protocol=standard SFT2025.08 | 42.95 | — | — | — | |
| Full Tokens (standard SFT)Base model=LLaMA-2-13B2025.08 | 42.65 | — | — | — | |
| IgnoringBase model=LLaMA-3.2-1B, Granularity=token-level2025.08 | 42.4 | — | — | — | |
| StarCoderParameters=15B, Training Stage=Pretrained comparison2024.01 | 41.28 | — | — | — | |
| CodeLLaMA-InstructParameters=7B, Training Stage=SFT comparison2024.01 | 41.25 | — | — | — | |
| ForgettingBase model=LLaMA-3.2-3B, Granularity=sequence-level2025.08 | 40.95 | — | — | — | |
| AdamWRuntime=8.82026.04 | 40.8 | — | — | — | |
| IgnoringBase model=LLaMA-3.2-3B, Granularity=sequence-level2025.08 | 40.58 | — | — | — | |
| 50% KV PredModel=Qwen3 32B2026.03 | 40.02 | — | — | — | |
| TransActNumber of Parameters=1.3B, Backbone=LLaMA, Shots=0, Prepended Examples=62024.07 | 39.6 | — | — | — | |
| IgnoringBase model=LLaMA-3.2-1B, Granularity=sequence-level2025.08 | 39.56 | — | — | — | |
| Base Model2026.04 | 39.5 | — | — | — | |
| BaseBase model=LLaMA-3.2-3B2025.08 | 39.45 | — | — | — | |
| BLURRuntime=9.42026.04 | 39.2 | — | — | — | |
| BaselineModel=Qwen3 32B2026.03 | 39.16 | — | — | — | |
| LLAMA PROParameters=8B, Training Stage=Pretrained comparison2024.01 | 39.04 | — | — | — | |
| SCBackbone=Llama3.1-8B-Instruct2025.01 | 39 | — | — | — | |
| LLaMA-2-13BApproach=CoT, Learning Paradigm=Vanilla2024.11 | 39 | — | — | — | |
| ForgettingBase model=LLaMA-3.2-1B, Granularity=sequence-level2025.08 | 38.93 | — | — | — | |
| LLaMA2Parameters=7B, Training Stage=Pretrained comparison2024.01 | 38.76 | — | — | — | |
| Full TokensBase model=LLaMA-3.2-1B, Training Protocol=standard SFT2025.08 | 38.74 | — | — | — | |
| LLaMA-Basezero-shot=true, setup=Mc1_targets2023.05 | 38.7 | — | — | — | |
| OPT-1.3BNumber of Parameters=1.3B, Backbone=OPT, Shots=0, Prepended Examples=62024.07 | 38.7 | — | — | — | |
| SIFTRuntime=20.22026.04 | 38.6 | — | — | — | |
| POMERuntime=9.92026.04 | 38.4 | — | — | — | |
| CoTBackbone=Qwen2-7B-Instruct2025.01 | 38 | — | — | — | |
| MuonRuntime=11.82026.04 | 37.9 | — | — | — | |
| BaseBase model=LLaMA-3.2-1B2025.08 | 37.83 | — | — | — | |
| CodeLLaMAParameters=7B, Training Stage=Pretrained comparison2024.01 | 37.82 | — | — | — | |
| LLaMA-2-13BApproach=Zero-shot, Learning Paradigm=Vanilla2024.11 | 37.8 | — | — | — | |
| OPT-2.7BNumber of Parameters=2.7B, Backbone=OPT, Shots=0, Prepended Examples=62024.07 | 37.6 | — | — | — |