Factuality Evaluation on FActScore
84.5Pairwise ScoreGrounded Decoding
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Grounded DecodingStrategy=adaptive2026.05 | 84.5 | 0.89 | |
| Grounded DecodingStrategy=static2026.05 | 81.2 | 0.89 | |
| CoCoA2026.05 | 79.5 | 0.87 | |
| COIECD2026.05 | 78.9 | 0.87 | |
| AdaCAD2026.05 | 78.1 | 0.86 | |
| CAD2026.05 | 74.5 | 0.85 | |
| DoLa2026.05 | 73.2 | 0.86 | |
| kNN-LM2026.05 | 72.5 | 0.88 | |
| Standard RAG2026.05 | 71.2 | 0.88 | |
| Self-Improving PretrainingTraining Data=SlimPajama, Pretraining for=Factuality2026.01 | 69.3 | — | |
| Llama BaseStage=Base Model2026.01 | 50 | — | |
| Llama Pretrain BaselineTraining Data=SlimPajama, Pretraining for=Factuality2026.01 | 48.9 | — | |
| KLCF-zeroModel Backbone=Qwen2.5-14B-Base, Training Stage=Base Model2025.09 | 0.612 | — | |
| GRPO&FActScoreModel Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.576 | — | |
| KLCFModel Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.557 | — | |
| DPO&KLCModel Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.549 | — | |
| GRPO&FActScoreModel Backbone=Qwen2.5-14B-Base, Training Stage=Base Model2025.09 | 0.548 | — | |
| GRPO&FActScore&No.ClaimsModel Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.539 | — | |
| DPO&FActScoreModel Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.53 | — | |
| IntuitorModel Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.527 | — | |
| Self-Eval-P(True)Model Backbone=Qwen2.5-14B-Base, Training Stage=Base Model2025.09 | 0.52 | — | |
| SFT (Distill)Model Backbone=DeepSeek-R1-Distill-Qwen-14B, Training Stage=SFT Model2025.09 | 0.487 | — | |
| Base (10-shot)Model Backbone=Qwen2.5-14B-Base, Training Stage=Base Model2025.09 | 0.468 | — | |
| CoVeModel Backbone=Qwen2.5-14B-Base, Training Stage=Base Model2025.09 | 0.46 | — |