Selective Prediction on TriviaQA 200 samples (test)
62.5Rejection Accuracy (80%)Semantic Entropy
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Semantic EntropyEnsemble Composition=Two LLMs (Llama + Qwen)2026.03 | 62.5 | 58.3 | 57.9 | 57.5 | 0.176 | 0.67 | |
| UEEnsemble Composition=Two LLMs (Llama + Qwen)2026.03 | 62.5 | 58.3 | 58.9 | 57.5 | 0.177 | 0.661 | |
| UAEnsemble Composition=Two LLMs (Llama + Qwen)2026.03 | 59.3 | 58.3 | 57.8 | 57.5 | 0.174 | 0.666 | |
| UCoEEnsemble Composition=Two LLMs (Llama + Qwen)2026.03 | 59.3 | 58.3 | 57.8 | 57.5 | 0.174 | 0.683 | |
| PfalseEnsemble Composition=Two LLMs (Llama + Qwen)2026.03 | 57.5 | 57.5 | 57.5 | 57.5 | 0.172 | 0.586 | |
| Regular EntropyEnsemble Composition=Two LLMs (Llama + Qwen)2026.03 | 56.2 | 58.3 | 60.5 | 57.5 | 0.175 | 0.475 | |
| Semantic EntropyEnsemble Composition=Three LLMs (Llama + Qwen + Mistral)2026.03 | 39.4 | 41.6 | 46.8 | 37.5 | 0.123 | 0.687 | |
| UAEnsemble Composition=Three LLMs (Llama + Qwen + Mistral)2026.03 | 39.4 | 41.6 | 43.7 | 37.5 | 0.121 | 0.673 | |
| UCoEEnsemble Composition=Three LLMs (Llama + Qwen + Mistral)2026.03 | 39.4 | 41.6 | 43.7 | 37.5 | 0.121 | 0.772 | |
| PfalseEnsemble Composition=Three LLMs (Llama + Qwen + Mistral)2026.03 | 37.5 | 40.5 | 40.5 | 37.5 | 0.117 | 0.54 | |
| UEEnsemble Composition=Three LLMs (Llama + Qwen + Mistral)2026.03 | 36.8 | 38.8 | 43.7 | 37.5 | 0.116 | 0.716 | |
| Regular EntropyEnsemble Composition=Three LLMs (Llama + Qwen + Mistral)2026.03 | 34.2 | 33.3 | 34.3 | 37.5 | 0.103 | 0.416 |