Factuality Detection on Short-form QA (Average of NQ, PopQA, TriviaQA, SimpleQA) (test)
71.1PR-AUCFRANQ condition-calibrated
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FRANQ condition-calibratedLLM Model=Falcon 3B Base2025.05 | 71.1 | 47.7 | |
| XGBoost (all UQ features)LLM Model=Falcon 3B Base2025.05 | 70.5 | 46.2 | |
| Degree MatrixLLM Model=Falcon 3B Base2025.05 | 70.2 | 46.4 | |
| Sum of EigenvaluesLLM Model=Falcon 3B Base2025.05 | 70 | 46 | |
| FRANQ calibratedLLM Model=Falcon 3B Base2025.05 | 67.2 | 41.1 | |
| XGBoost (FRANQ features)LLM Model=Falcon 3B Base2025.05 | 67 | 36.8 | |
| AlignScoreLLM Model=Falcon 3B Base2025.05 | 66.6 | 37.2 | |
| FRANQ condition-calibratedLLM Model=Llama 8B Instruct2025.05 | 64.7 | 54 | |
| FRANQ calibratedLLM Model=Llama 8B Instruct2025.05 | 64.4 | 53.4 | |
| CCPLLM Model=Falcon 3B Base2025.05 | 64.1 | 30.4 | |
| FRANQ no calibrationLLM Model=Falcon 3B Base2025.05 | 64.1 | 34.5 | |
| Mean Token EntropyLLM Model=Llama 8B Instruct2025.05 | 64 | 49.1 | |
| Lexical SimilarityLLM Model=Llama 8B Instruct2025.05 | 63.9 | 53.2 | |
| Semantic EntropyLLM Model=Llama 8B Instruct2025.05 | 63.7 | 51.9 | |
| XGBoost (all UQ features)LLM Model=Llama 8B Instruct2025.05 | 63.4 | 50.3 | |
| FRANQ condition-calibratedLLM Model=Llama 3B Instruct2025.05 | 63.1 | 54.1 | |
| Degree MatrixLLM Model=Llama 3B Instruct2025.05 | 62.9 | 52 | |
| Max Sequence Prob.LLM Model=Falcon 3B Base2025.05 | 62.8 | 25.6 | |
| Sum of EigenvaluesLLM Model=Llama 3B Instruct2025.05 | 62.8 | 51.8 | |
| Sum of EigenvaluesLLM Model=Llama 8B Instruct2025.05 | 62.8 | 48.9 | |
| FRANQ calibratedLLM Model=Llama 3B Instruct2025.05 | 62.8 | 53.7 | |
| Degree MatrixLLM Model=Llama 8B Instruct2025.05 | 62.7 | 49.2 | |
| Semantic EntropyLLM Model=Falcon 3B Base2025.05 | 62.3 | 27.8 | |
| Lexical SimilarityLLM Model=Falcon 3B Base2025.05 | 61.8 | 27.7 | |
| Mean Token EntropyLLM Model=Falcon 3B Base2025.05 | 61.3 | 24.2 | |
| Semantic EntropyLLM Model=Llama 3B Instruct2025.05 | 61.3 | 52.5 | |
| SentenceSARLLM Model=Falcon 3B Base2025.05 | 60.2 | 26.3 | |
| Mean Token EntropyLLM Model=Llama 3B Instruct2025.05 | 59.4 | 48.1 | |
| XGBoost (all UQ features)LLM Model=Llama 3B Instruct2025.05 | 59.4 | 49.4 | |
| SentenceSARLLM Model=Llama 3B Instruct2025.05 | 57.1 | 48.3 | |
| Max Sequence Prob.LLM Model=Llama 8B Instruct2025.05 | 56.9 | 40.7 | |
| Lexical SimilarityLLM Model=Llama 3B Instruct2025.05 | 56.4 | 47.9 | |
| Max Sequence Prob.LLM Model=Llama 3B Instruct2025.05 | 55.8 | 45.4 | |
| SentenceSARLLM Model=Llama 8B Instruct2025.05 | 55.6 | 41.4 | |
| Parametric KnowledgeLLM Model=Falcon 3B Base2025.05 | 55.6 | 10.4 | |
| CCPLLM Model=Llama 8B Instruct2025.05 | 55.3 | 41.7 | |
| FRANQ no calibrationLLM Model=Llama 3B Instruct2025.05 | 55.3 | 40.3 | |
| CCPLLM Model=Llama 3B Instruct2025.05 | 55.1 | 44.3 | |
| XGBoost (FRANQ features)LLM Model=Llama 3B Instruct2025.05 | 52.6 | 40.9 | |
| XGBoost (FRANQ features)LLM Model=Llama 8B Instruct2025.05 | 52.4 | 38.5 | |
| FRANQ no calibrationLLM Model=Llama 8B Instruct2025.05 | 52.3 | 34 | |
| Parametric KnowledgeLLM Model=Llama 8B Instruct2025.05 | 49.9 | 33 | |
| FRANQ condition-calibratedLLM Model=Gemma 12B Instruct2025.05 | 49.6 | 28.3 | |
| FRANQ calibratedLLM Model=Gemma 12B Instruct2025.05 | 48.1 | 25.8 | |
| XGBoost (all UQ features)LLM Model=Gemma 12B Instruct2025.05 | 47.4 | 30.1 | |
| Sum of EigenvaluesLLM Model=Gemma 12B Instruct2025.05 | 46.7 | 26 | |
| Semantic EntropyLLM Model=Gemma 12B Instruct2025.05 | 46.6 | 26.1 | |
| Degree MatrixLLM Model=Gemma 12B Instruct2025.05 | 46.4 | 26 | |
| FRANQ no calibrationLLM Model=Gemma 12B Instruct2025.05 | 44.7 | 22.5 | |
| AlignScoreLLM Model=Llama 8B Instruct2025.05 | 43.2 | 22.4 | |
| Lexical SimilarityLLM Model=Gemma 12B Instruct2025.05 | 43 | 24 | |
| Parametric KnowledgeLLM Model=Llama 3B Instruct2025.05 | 42.5 | 24.7 | |
| Mean Token EntropyLLM Model=Gemma 12B Instruct2025.05 | 42.3 | 23 | |
| SentenceSARLLM Model=Gemma 12B Instruct2025.05 | 41.6 | 17.4 | |
| AlignScoreLLM Model=Llama 3B Instruct2025.05 | 41.5 | 20.7 | |
| XGBoost (FRANQ features)LLM Model=Gemma 12B Instruct2025.05 | 41.4 | 19.6 | |
| CCPLLM Model=Gemma 12B Instruct2025.05 | 41.2 | 19.8 | |
| Max Sequence Prob.LLM Model=Gemma 12B Instruct2025.05 | 40 | 16.2 | |
| AlignScoreLLM Model=Gemma 12B Instruct2025.05 | 37.6 | 15.8 | |
| Parametric KnowledgeLLM Model=Gemma 12B Instruct2025.05 | 36.4 | 10.5 |