Answer Correctness Prediction on Average (TriviaQA, HotpotQA, MedMCQA) (val)
0.91Head EntropyHEAD ENTROPY
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| HEAD ENTROPYModel=Qwen3 8B2026.02 | 0.91 | — | 0.03 | — | |
| HEAD ENTROPYModel=Qwen3 1.7B2026.02 | 0.9 | — | 0.03 | — | |
| HEAD ENTROPYModel=Qwen3 32B2026.02 | 0.9 | — | 0.02 | — | |
| HEAD ENTROPYModel=Average2026.02 | 0.88 | — | 0.02 | — | |
| HEAD ENTROPYModel=Llama 3.1 8B2026.02 | 0.85 | — | 0 | — | |
| HEAD ENTROPYModel=Llama 3.2 3B2026.02 | 0.84 | — | 0 | — | |
| HEAD ENTROPYModel=Qwen3-32B, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | 0.81 | — | 0.18 | — | |
| HEAD ENTROPYModel=Qwen3-8B, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | 0.76 | — | 0.14 | — | |
| HEAD ENTROPYModel=Average, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | 0.74 | — | 0.12 | — | |
| HEAD ENTROPYModel=Llama-3.1-8B-Instruct, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | 0.73 | — | 0.12 | — | |
| HEAD ENTROPYModel=Qwen3-1.7B, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | 0.72 | — | 0.08 | — | |
| HEAD ENTROPYModel=Llama-3.2-3B-Instruct, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | 0.68 | — | 0.08 | — | |
| Hidden State LRModel=Llama-3.2-3B-Instruct, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | — | — | — | 0.6 | |
| Hidden State LRModel=Llama-3.1-8B-Instruct, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | — | — | — | 0.61 | |
| Hidden State LRModel=Qwen3-1.7B, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | — | — | — | 0.65 | |
| Hidden State LRModel=Qwen3-8B, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | — | — | — | 0.63 | |
| Hidden State LRModel=Qwen3-32B, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | — | — | — | 0.64 | |
| Hidden State LRModel=Average, Input Scope=Question tokens only, Prediction Timing=Pre-generation2026.02 | — | — | — | 0.63 | |
| LookbackLensModel=Llama 3.2 3B2026.02 | — | 0.84 | — | — | |
| LookbackLensModel=Llama 3.1 8B2026.02 | — | 0.85 | — | — | |
| LookbackLensModel=Qwen3 1.7B2026.02 | — | 0.87 | — | — | |
| LookbackLensModel=Qwen3 8B2026.02 | — | 0.88 | — | — | |
| LookbackLensModel=Qwen3 32B2026.02 | — | 0.88 | — | — | |
| LookbackLensModel=Average2026.02 | — | 0.86 | — | — |