Confidence Estimation (Freeform Tagging) on WildHallu
4.1Brier Score (BS)LOVEC-DPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LOVEC-DPOBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 4.1 | 1.3 | 51.8 | |
| LOVEC-GRPOTagging Format=Freeform, Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 6 | 8.2 | 63.1 | |
| LOVEC-DPOTagging Format=Freeform, Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 6.3 | 5.4 | 62.1 | |
| LOVEC-GRPOBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 7.3 | 5.6 | 52.2 | |
| LOVEC-SFTBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 8 | 12.2 | 36.1 | |
| LOVEC-SFTTagging Format=Freeform, Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 8.9 | 15.1 | 58.8 | |
| luqBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 11.9 | 16.3 | 50 | |
| Self-ConsBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 13.4 | 17.7 | 43.2 | |
| Verb-ConfBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 18.5 | 19.2 | 35.1 | |
| p(true)Base Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 19.3 | 22.8 | 25.4 | |
| VanillaBase Model=Gemma-2-9B-It, Tagging Strategy=Sentence-level iterative tagging2025.05 | 22.5 | 26.3 | 28.9 |