Correctness Detection on TruthfulQA (in-domain)
0.924AUCLlama-1B
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-1BModel Type=Instruction-tuned2026.02 | 0.924 | |
| Qwen2-7BModel Type=Instruction-tuned2026.02 | 0.901 | |
| Qwen2-1.5BModel Type=Instruction-tuned2026.02 | 0.877 | |
| GPT-2-LargeModel Type=Base models2026.02 | 0.751 | |
| GPT-2-MedModel Type=Base models2026.02 | 0.75 | |
| GPT-2Model Type=Base models2026.02 | 0.732 |