Correctness Prediction on TriviaQA
0.999AUROCSelf-consistency (10 samples)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Self-consistency (10 samples)Inference cost=10 forward passes2026.04 | 0.999 | — | |
| ParaGradModel=Llama3.1-Instruct8B2026.05 | 0.8649 | — | |
| G-NLLModel=Llama3.1-Instruct8B2026.05 | 0.8591 | — | |
| ParaGradModel=Mistral-Nemo-Instruct12B2026.05 | 0.8591 | — | |
| HybridGradModel=Llama3.1-Instruct8B2026.05 | 0.8589 | — | |
| SARModel=Llama3.1-Instruct8B2026.05 | 0.8565 | — | |
| SARModel=Mistral-Nemo-Instruct12B2026.05 | 0.8523 | — | |
| ExGradModel=Llama3.1-Instruct8B2026.05 | 0.8522 | — | |
| AssessorModel=Llama 3.1 8B, Assessor type=logistic regression2025.09 | 0.852 | — | |
| SemGradModel=Llama3.1-Instruct8B2026.05 | 0.8472 | — | |
| DegModel=Llama3.1-Instruct8B2026.05 | 0.8467 | — | |
| G-NLLModel=Mistral-Nemo-Instruct12B2026.05 | 0.8461 | — | |
| AssessorModel=Mistral 7B Instruct v0.3, Assessor type=logistic regression2025.09 | 0.846 | — | |
| LN-PEModel=Llama3.1-Instruct8B2026.05 | 0.8453 | — | |
| ExGradModel=Mistral-Nemo-Instruct12B2026.05 | 0.8453 | — | |
| HybridGradModel=Mistral-Nemo-Instruct12B2026.05 | 0.8413 | — | |
| LN-PEModel=Mistral-Nemo-Instruct12B2026.05 | 0.8402 | — | |
| Self-ConModel=Llama3.1-Instruct8B2026.05 | 0.8356 | — | |
| M.I.Model=Llama3.1-Instruct8B2026.05 | 0.8352 | — | |
| S.E.Model=Llama3.1-Instruct8B2026.05 | 0.8312 | — | |
| DegModel=Mistral-Nemo-Instruct12B2026.05 | 0.8311 | — | |
| DirectionModel=Llama 3.3 70B Instruct, Train dataset=TriviaQA2025.09 | 0.826 | — | |
| S.D.Model=Llama3.1-Instruct8B2026.05 | 0.8244 | — | |
| SemGradModel=Mistral-Nemo-Instruct12B2026.05 | 0.8237 | — | |
| ParaGradModel=Qwen3-Instruct4B2026.05 | 0.8202 | — | |
| M.I.Model=Mistral-Nemo-Instruct12B2026.05 | 0.8188 | — | |
| Probe(AVG)Model=Gemma3-12B2025.08 | 0.818 | — | |
| Self-ConModel=Mistral-Nemo-Instruct12B2026.05 | 0.818 | — | |
| HybridGradModel=Qwen3-Instruct4B2026.05 | 0.8169 | — | |
| SARModel=Qwen3-Instruct4B2026.05 | 0.8152 | — | |
| P(True)Model=Mistral-Nemo-Instruct12B2026.05 | 0.8139 | — | |
| Probe(EOS)Model=Qwen2.5-7B2025.08 | 0.812 | — | |
| G-NLLModel=Qwen3-Instruct4B2026.05 | 0.8101 | — | |
| Probe(EOS)Model=Gemma3-12B2025.08 | 0.81 | — | |
| AssessorModel=Qwen 2.5 7B Instruct, Assessor type=logistic regression2025.09 | 0.807 | — | |
| S.E.Model=Mistral-Nemo-Instruct12B2026.05 | 0.8064 | — | |
| LogProbModel=Gemma3-12B2025.08 | 0.806 | — | |
| DirectionModel=Llama 3.1 8B, Train dataset=TriviaQA2025.09 | 0.804 | — | |
| SemGradModel=Qwen3-Instruct4B2026.05 | 0.804 | — | |
| ExGradModel=Qwen3-Instruct4B2026.05 | 0.8037 | — | |
| LN-PEModel=Qwen3-Instruct4B2026.05 | 0.8 | — | |
| DirectionModel=Mistral 7B Instruct v0.3, Train dataset=TriviaQA2025.09 | 0.796 | — | |
| Prob(Exact)Model=Gemma3-12B2025.08 | 0.796 | — | |
| S.D.Model=Mistral-Nemo-Instruct12B2026.05 | 0.7907 | — | |
| AssessorModel=DeepSeek R1 Distill Qwen 32B, Assessor type=logistic regression2025.09 | 0.79 | — | |
| AssessorModel=Ministral 8B Instruct 2410, Assessor type=logistic regression2025.09 | 0.789 | — | |
| Probe(AVG)Model=Qwen2.5-7B2025.08 | 0.786 | — | |
| P(True)Model=Llama3.1-Instruct8B2026.05 | 0.786 | — | |
| Prob(Exact)Model=Fanar1-9b2025.08 | 0.783 | — | |
| DegModel=Qwen3-Instruct4B2026.05 | 0.7821 | — | |
| Prob(Exact)Model=Qwen2.5-7B2025.08 | 0.781 | — | |
| |S| Hybridn (sampling budget)=82026.04 | 0.778 | 0.626 | |
| Hb Pluginn (sampling budget)=82026.04 | 0.776 | 0.635 | |
| NumSetsn (sampling budget)=82026.04 | 0.775 | 0.632 | |
| LogProbModel=Fanar1-9b2025.08 | 0.774 | — | |
| LogProbModel=Qwen2.5-7B2025.08 | 0.774 | — | |
| Distilled verbal (CSFT)Inference cost=1 forward pass2026.04 | 0.774 | — | |
| LogTokUModel=Qwen2.5-7B2025.08 | 0.773 | — | |
| Self-ConModel=Qwen3-Instruct4B2026.05 | 0.7664 | — | |
| Probe(AVG)Model=Fanar1-9b2025.08 | 0.765 | — | |
| SHADEn (sampling budget)=82026.04 | 0.765 | 0.625 | |
| S.D.Model=Qwen3-Instruct4B2026.05 | 0.7641 | — | |
| Hb Hybridn (sampling budget)=82026.04 | 0.764 | 0.624 | |
| P(True)Model=Qwen3-Instruct4B2026.05 | 0.763 | — | |
| M.I.Model=Qwen3-Instruct4B2026.05 | 0.7626 | — | |
| INSIDEModel=Llama3.1-Instruct8B2026.05 | 0.7624 | — | |
| S.E.Model=Qwen3-Instruct4B2026.05 | 0.7616 | — | |
| AssessorModel=Llama 3.3 70B Instruct, Assessor type=logistic regression2025.09 | 0.759 | — | |
| DirectionModel=Qwen 2.5 7B Instruct, Train dataset=TriviaQA2025.09 | 0.758 | — | |
| DSEn (sampling budget)=52026.04 | 0.757 | 0.653 | |
| |S| Hybridn (sampling budget)=102026.04 | 0.757 | 0.62 | |
| Probe(EU)Model=Qwen2.5-7B2025.08 | 0.754 | — | |
| Probe(EU)Model=Fanar1-9b2025.08 | 0.751 | — | |
| Probe(EU)Model=Gemma3-12B2025.08 | 0.751 | — | |
| Probe(EOS)Model=Fanar1-9b2025.08 | 0.739 | — | |
| Hb Pluginn (sampling budget)=102026.04 | 0.738 | 0.625 | |
| NumSetsn (sampling budget)=102026.04 | 0.737 | 0.612 | |
| P(true)Model=Qwen2.5-7B2025.08 | 0.736 | — | |
| DirectionModel=DeepSeek R1 Distill Qwen 32B, Train dataset=TriviaQA2025.09 | 0.735 | — | |
| DirectionModel=Ministral 8B Instruct 2410, Train dataset=TriviaQA2025.09 | 0.734 | — | |
| NumSetsn (sampling budget)=52026.04 | 0.733 | 0.595 | |
| DSEn (sampling budget)=102026.04 | 0.726 | 0.718 | |
| INSIDEModel=Mistral-Nemo-Instruct12B2026.05 | 0.7256 | — | |
| INSIDEModel=Qwen3-Instruct4B2026.05 | 0.7247 | — | |
| Hb Pluginn (sampling budget)=52026.04 | 0.724 | 0.591 | |
| Hb Hybridn (sampling budget)=102026.04 | 0.724 | 0.626 | |
| SHADEn (sampling budget)=102026.04 | 0.721 | 0.627 | |
| |S| Hybridn (sampling budget)=52026.04 | 0.717 | 0.624 | |
| SHADEn (sampling budget)=52026.04 | 0.71 | 0.634 | |
| Hb Hybridn (sampling budget)=52026.04 | 0.709 | 0.633 | |
| Logit entropyInference cost=1 forward pass2026.04 | 0.701 | — | |
| KLEn (sampling budget)=102026.04 | 0.691 | 0.751 | |
| P(True)2026.03 | 0.69 | — | |
| LogTokUModel=Fanar1-9b2025.08 | 0.683 | — | |
| KLEn (sampling budget)=82026.04 | 0.679 | 0.782 | |
| P(true)Model=Fanar1-9b2025.08 | 0.672 | — | |
| KLEn (sampling budget)=52026.04 | 0.664 | 0.741 | |
| DSEn (sampling budget)=82026.04 | 0.659 | 0.658 | |
| Verb. conf.Model=Qwen 2.5 7B Instruct2025.09 | 0.643 | — | |
| P(true)Model=Gemma3-12B2025.08 | 0.631 | — |