Faithful Calibration on Faithful Calibration Dataset Suite (test)
85PQARLMF
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| RLMFBackbone=Llama3.1-8B-Ins, Method=Reinforcement Learning with Metacognitive Feedback (Numerical)2026.06 | 85 | 81 | 83 | 82 | 81 | 84 | 84 | 83 | 86 | 86 | 84 | 41 | 0.26 | |
| RLMFBackbone=Qwen3-8B, Method=Reinforcement Learning with Metacognitive Feedback (Numerical)2026.06 | 85 | 82 | 86 | 82 | 84 | 82 | 83 | 82 | 83 | 84 | 83 | 57 | 0.19 | |
| +RLBackbone=Llama3.1-8B-Ins, Method=RL Ablation (No Metacognitive Scaling)2026.06 | 82 | 78 | 80 | 79 | 75 | 73 | 81 | 80 | 73 | 72 | 77 | 40 | 0.2 | |
| RLMF + Rewr.Backbone=Qwen3-8B, Method=Reinforcement Learning with Metacognitive Feedback (Linguistic Rewriting)2026.06 | 82 | 86 | 80 | 84 | 80 | 80 | 87 | 87 | 82 | 82 | 83 | 57 | 0.19 | |
| RLMF + Rewr.Backbone=Llama3.1-8B-Ins, Method=Reinforcement Learning with Metacognitive Feedback (Linguistic Rewriting)2026.06 | 81 | 86 | 80 | 81 | 80 | 81 | 82 | 81 | 87 | 83 | 82 | 41 | 0.26 | |
| +RLBackbone=Qwen3-8B, Method=RL Ablation (No Metacognitive Scaling)2026.06 | 75 | 66 | 69 | 20 | 55 | 54 | 58 | 44 | 32 | 38 | 51 | 59 | 0.26 | |
| FUTBackbone=Llama3.1-8B-Ins, Method=Faithful Uncertainty Tuning2026.06 | 69 | 67 | 68 | 66 | 63 | 70 | 63 | 63 | 68 | 67 | 66 | 31 | 0.29 | |
| MetaFaithBackbone=Llama3.1-8B-Ins, Method=Metacognitive Prompting2026.06 | 68 | 71 | 65 | 67 | 67 | 64 | 64 | 66 | 68 | 72 | 67 | 28 | 0.36 | |
| Gemini-3.1-ProModel Type=Frontier Model2026.06 | 62 | 71 | 70 | 68 | 72 | 68 | 66 | 71 | 73 | 82 | 70 | 78 | 0.15 | |
| Llama3.1-8B-InsBackbone=Llama3.1-8B-Ins, Configuration=Base Model2026.06 | 60 | 61 | 61 | 50 | 65 | 62 | 48 | 61 | 59 | 71 | 60 | 31 | 0.33 | |
| Gemini-3-FlashModel Type=Frontier Model2026.06 | 59 | 64 | 55 | 66 | 67 | 70 | 65 | 66 | 77 | 71 | 66 | 72 | 0.16 | |
| FUTBackbone=Qwen3-8B, Method=Faithful Uncertainty Tuning2026.06 | 57 | 75 | 48 | 74 | 72 | 71 | 66 | 67 | 71 | 74 | 67 | 38 | 0.41 | |
| Qwen3-8BBackbone=Qwen3-8B, Configuration=Base Model2026.06 | 53 | 63 | 57 | 54 | 63 | 59 | 59 | 59 | 7 | 62 | 54 | 55 | 0.31 | |
| MetaFaithBackbone=Qwen3-8B, Method=Metacognitive Prompting2026.06 | 53 | 66 | 47 | 68 | 67 | 72 | 70 | 49 | 70 | 67 | 63 | 51 | 0.29 | |
| GPT-5Model Type=Frontier Model, Prompting=MetaFaith prompting2026.06 | 50 | 61 | 52 | 66 | 59 | 57 | 60 | 57 | 68 | 77 | 61 | 69 | 0.19 |