Claim Verification on Human evaluation dataset 400 claims
0.937AccuracyGPT-4o-mini
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4o-mini2025.12 | 0.937 | 0.01 | 0.875 | |
| NewsScopebase_model=LLaMA 3.1 8B Instruct, fine_tuning=LoRA (r=16, alpha=16), numeric_grounding_filter=enabled2025.12 | 0.916 | 0.02 | 0.83 | |
| NewsScopebase_model=LLaMA 3.1 8B Instruct, fine_tuning=LoRA (r=16, alpha=16), numeric_grounding_filter=disabled2025.12 | 0.894 | 0.025 | 0.85 |