Confidence Estimation (Iterative Tagging) on WildHallu
5.7Brier Score (BS)LOVEC-GRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LOVEC-GRPOTagging Format=Iterative, Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 5.7 | 2.5 | 57 | |
| LOVEC-DPOTagging Format=Iterative, Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 6 | 5 | 60.4 | |
| LOVEC-SFTTagging Format=Iterative, Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 9.1 | 15.2 | 51.1 | |
| VanillaBackbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 10.8 | 6 | 9.1 | |
| LUQBackbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 14.5 | 21.5 | 56.8 | |
| p(true)-ftBackbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 16.4 | 19.5 | 47.5 | |
| Self-ConsBackbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 16.5 | 24.3 | 47.8 | |
| Verb-ConfBackbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 20.3 | 22.1 | 13.4 | |
| p(true)Backbone=Llama3-8B-Instruct, Decoding Strategy=Greedy2025.05 | 23.8 | 23.6 | 15.8 |