Pairwise Citation Prediction on SciJudgeBench (test)
87.9CS Pairwise AccuracySciJudge-Qwen2.5-14B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SciJudge-Qwen2.5-14BTraining method=GRPO, Backbone=Qwen2.5-14B2026.03 | 87.9 | 78.7 | 74.4 | 84.9 | 80.6 | |
| SciJudge-Qwen2.5-32BTraining method=GRPO, Backbone=Qwen2.5-32B2026.03 | 85.4 | 77.9 | 82.2 | 89.9 | 83.7 | |
| SciJudge-Qwen3-30BTraining method=Group Relative Policy Optimization (GRPO), Evaluation protocol=Position-swap consistency, Backbone=Qwen3-30B2026.03 | 83.5 | 78.7 | 78.7 | 82.3 | 80.6 | |
| SciJudge-Qwen2.5-7BTraining method=GRPO, Backbone=Qwen2.5-7B2026.03 | 83 | 68.8 | 71.5 | 87.4 | 76.9 | |
| Gemini-3.0-Pro-PreviewModel family=SOTA Models2026.03 | 81.1 | 73 | 72.6 | 76.5 | 75.7 | |
| GPT-5.2-ThinkingModel family=SOTA Models, Reasoning trace=Enabled2026.03 | 79.1 | 68.8 | 69.4 | 73.1 | 72.7 | |
| GLM-5Model family=SOTA Models2026.03 | 79.1 | 75.4 | 69.4 | 72.3 | 73.6 | |
| SciJudge-Qwen3-4BTraining method=Group Relative Policy Optimization (GRPO), Evaluation protocol=Position-swap consistency, Backbone=Qwen3-4B2026.03 | 78.6 | 74.6 | 71.2 | 79.8 | 75.3 | |
| DeepSeek-V3.2-ThinkingModel family=SOTA Models, Reasoning trace=Enabled2026.03 | 78.2 | 67.2 | 64.1 | 72.3 | 69.9 | |
| SciJudge-Qwen2.5-3BTraining method=GRPO, Backbone=Qwen2.5-3B2026.03 | 76.2 | 76.2 | 66.2 | 81.5 | 73.2 | |
| MiniMax-M2.5Model family=SOTA Models2026.03 | 75.7 | 64.8 | 64.1 | 71.4 | 68.7 | |
| Qwen3-30B-A3B-InstructModel family=Open-source Models, Evaluation protocol=Position-swap consistency2026.03 | 73.8 | 70.5 | 59.4 | 65.5 | 66.3 | |
| SciJudge-Qwen2.5-1.5BTraining method=GRPO, Backbone=Qwen2.5-1.5B2026.03 | 72.3 | 73 | 69.4 | 77.3 | 72.1 | |
| Qwen2.5-32B-InstructModel family=Open-source Models2026.03 | 71.4 | 61.5 | 55.9 | 62.2 | 62.2 | |
| DeepSeek-V3.2Model family=SOTA Models2026.03 | 67 | 68 | 57.3 | 62.2 | 62.6 | |
| Qwen3-4B-InstructModel family=Open-source Models, Evaluation protocol=Position-swap consistency2026.03 | 66.5 | 65.6 | 54.8 | 57.1 | 60.3 | |
| Qwen2.5-14B-InstructModel family=Open-source Models2026.03 | 64.1 | 63.9 | 54.5 | 56.3 | 59.1 | |
| Qwen2.5-7B-InstructModel family=Open-source Models2026.03 | 57.3 | 37.7 | 37 | 51.3 | 45.2 | |
| SciJudge-Llama3.1-8BTraining method=GRPO, Backbone=Llama3.1-8B2026.03 | 56.8 | 59.8 | 55.2 | 60.5 | 57.3 | |
| Llama3.1-8B-InstructModel family=Open-source Models2026.03 | 34.5 | 44.3 | 35.9 | 35.3 | 36.8 | |
| Qwen2.5-3B-InstructModel family=Open-source Models2026.03 | 16.5 | 36.9 | 23.8 | 21 | 23.5 | |
| Qwen2.5-1.5B-InstructModel family=Open-source Models2026.03 | 6.3 | 10.7 | 6 | 6.7 | 7 |