Truthful Question Answering on TruthfulQA (Accuracy)
56.8Accuracy (TruthfulQA)LLaDA
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaDABudget=1/1, Evaluation Protocol=fully generative2026.05 | 56.8 | |
| ME-DLM Stage 3Budget=1/2, Evaluation Protocol=fully generative2026.05 | 54.6 | |
| IPObackbone=DS-8B2025.09 | 53.9 | |
| ME-DLM Stage 3Budget=1/1, Evaluation Protocol=fully generative2026.05 | 53.5 | |
| RealSafebackbone=DS-8B, alignment=SFT-based2025.09 | 53.4 | |
| ME-DLM Stage 2Budget=1/1, Evaluation Protocol=fully generative2026.05 | 52.5 | |
| LLaDABudget=1/2, Evaluation Protocol=fully generative2026.05 | 49.8 | |
| ME-DLM Stage 3Budget=1/4, Evaluation Protocol=fully generative2026.05 | 49.8 | |
| FPModel=DREAM, Precision=Full Precision2026.06 | 49.76 | |
| STaR-QuantModel=LLADA-1.5, Quantization=W8A82026.06 | 49.72 | |
| STaR-QuantModel=DREAM, Quantization=W8A82026.06 | 49.02 | |
| ME-DLM Stage 2Budget=1/2, Evaluation Protocol=fully generative2026.05 | 48.6 | |
| STaR-QuantModel=LLADA, Quantization=W8A82026.06 | 48.56 | |
| FPModel=LLADA, Precision=Full Precision2026.06 | 47.49 | |
| FPModel=LLADA-1.5, Precision=Full Precision2026.06 | 47.2 | |
| ME-DLM Stage 3Budget=1/8, Evaluation Protocol=fully generative2026.05 | 46.9 | |
| Basebackbone=DS-8B2025.09 | 46.7 | |
| LLaDABudget=1/4, Evaluation Protocol=fully generative2026.05 | 44.4 | |
| Qwen2.5Size=490M2025.04 | 39.74 | |
| CLIMBSize=950M2025.04 | 39.06 | |
| SmolLMSize=360M2025.04 | 37.93 | |
| Llama-3.2Size=1.2B2025.04 | 37.67 | |
| TinyLlamaSize=1.1B2025.04 | 37.6 | |
| CLIMBSize=350M2025.04 | 36.86 | |
| ME-DLM Stage 2Budget=1/4, Evaluation Protocol=fully generative2026.05 | 36.1 | |
| AMD-OLMoSize=1.2B2025.04 | 32.22 | |
| DATEDGPT-INSTRUCT-20240-shot=true, Training data cutoff year=20242026.03 | 28.6 | |
| SmolLM-1.7B-Instruct0-shot=true2026.03 | 27.7 | |
| DATEDGPT-INSTRUCT-20170-shot=true, Training data cutoff year=20172026.03 | 26.8 | |
| DATEDGPT-INSTRUCT-20210-shot=true, Training data cutoff year=20212026.03 | 26.4 | |
| TinyLlama-1.1B0-shot=true2026.03 | 26 | |
| DATEDGPT-INSTRUCT-20130-shot=true, Training data cutoff year=20132026.03 | 25.1 | |
| OPT-1.3B0-shot=true2026.03 | 23.8 | |
| Pythia-1B0-shot=true2026.03 | 23.6 | |
| LLaDABudget=1/8, Evaluation Protocol=fully generative2026.05 | 23.3 | |
| GPT2-XL0-shot=true2026.03 | 22.4 | |
| ME-DLM Stage 2Budget=1/8, Evaluation Protocol=fully generative2026.05 | 19.5 |