Claim Verification on CoverBench, LLM-AggreFact Out-of-Domain
68.6CoverBench ScoreSelf-Ask
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Self-AskModel Scale=GPT-4.1-mini2026.05 | 68.6 | 78.9 | 73.8 | |
| Decomposed PromptingModel Scale=32B2026.05 | 64.2 | 79.4 | 71.8 | |
| DECOMPOSERLModel Scale=7B2026.05 | 62.5 | 77 | 69.8 | |
| Decomposed PromptingModel Scale=14B2026.05 | 61.3 | 79.3 | 70.3 | |
| SimpleModel Scale=7B2026.05 | 52.5 | 74.9 | 63.7 | |
| SimpleModel Scale=3B2026.05 | 51.3 | 74 | 62.7 |