Truthfulness on TruthfulQA (Accuracy, Correct out of 30, Ranking)
83.3AccuracyGemma 2 9B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemma 2 9BType=Open-source, Why this score?=Best factual accuracy, strong reasoning per parameter, Format followed?=Yes, Genuine knowledge failure?=No2026.06 | 83.3 | 25 | 1 | |
| Mistral 7BType=Open-source, Why this score?=Strong factual ground relative to its size, Format followed?=Yes, Genuine knowledge failure?=No2026.06 | 63.3 | 19 | 2 | |
| FP16BIT=16, Evaluation Protocol=Few-shot2026.07 | 52.49 | — | — | |
| VQLLMBIT=2, Evaluation Protocol=Few-shot2026.07 | 51.05 | — | — | |
| GSRQBIT=2, Evaluation Protocol=Few-shot2026.07 | 50.12 | — | — | |
| GSRQBIT=1, Evaluation Protocol=Few-shot2026.07 | 48.67 | — | — | |
| GSRQBIT=0.75, Evaluation Protocol=Few-shot2026.07 | 47.26 | — | — | |
| VQLLMBIT=1, Evaluation Protocol=Few-shot2026.07 | 44.25 | — | — | |
| Llama 3 8BType=Open-source, Why this score?=Struggled with deeply embedded cultural myths in training data, Format followed?=Yes, Genuine knowledge failure?=Partial2026.06 | 40 | 12 | 3 | |
| Claude HaikuType=Closed-source, Why this score?=Format failure, wrote explanations instead of single letters, Format followed?=No, Genuine knowledge failure?=No (format issue only)2026.06 | 0 | 0 | 4 |