Scientific Claim Verification on SCIFACT abstract-level (test)
89.9Support F1Qwen-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen-7BEvaluation Protocol=Fine-Tuning2026.06 | 89.9 | 82.5 | 90.61 | 87.37 | |
| Qwen-MaxEvaluation Protocol=Zero-shot2026.06 | 89.55 | 85.62 | 79.65 | 84.94 | |
| DeepSeek-V3Evaluation Protocol=Zero-shot2026.06 | 89.17 | 84.71 | 79.25 | 84.38 | |
| Qwen-MaxEvaluation Protocol=3-shot prompt learning2026.06 | 87.57 | 82.42 | 80.19 | 83.39 | |
| DeepSeek-R1Evaluation Protocol=Zero-shot2026.06 | 87.14 | 82.22 | 72.81 | 80.72 | |
| Qwen-plusEvaluation Protocol=3-shot prompt learning2026.06 | 85.92 | 79.04 | 72.51 | 79.16 | |
| DeepSeek-V3Evaluation Protocol=3-shot prompt learning2026.06 | 85.33 | 80.05 | 78.78 | 81.39 | |
| Qwen-plusEvaluation Protocol=Zero-shot2026.06 | 82.84 | 79.66 | 63.97 | 75.49 | |
| DeepSeek-R1-Distill-Llama-8BEvaluation Protocol=Fine-Tuning2026.06 | 78.81 | 67.32 | 84.23 | 76.79 | |
| DeepSeek-R1-Distill-Llama-8BEvaluation Protocol=Zero-shot2026.06 | 77.56 | 73.83 | 56.82 | 69.4 | |
| DeepSeek-R1Evaluation Protocol=3-shot prompt learning2026.06 | 65.88 | 39.37 | 32.63 | 45.96 | |
| Qwen-7BEvaluation Protocol=Zero-shot2026.06 | 63.24 | 9.08 | 34.55 | 35.62 |