Source Prediction on Eval-Actions 1.0 (test)
99.6AccuracyAutoEval-S
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AutoEval-SEvaluation Protocol=Rank-Guided (RG)2026.01 | 99.6 | 99.5 | 99.5 | |
| QwenVL3-4BEvaluation Protocol=Rank-Guided (RG)2026.01 | 99.1 | 98.9 | 96.4 | |
| AutoEval-SEvaluation Protocol=Expert Grading (EG)2026.01 | 99.1 | 98.7 | 99 | |
| QwenVL2.5-3BEvaluation Protocol=Expert Grading (EG)2026.01 | 98.7 | 98.4 | 98.7 | |
| QwenVL2.5-3BEvaluation Protocol=Rank-Guided (RG)2026.01 | 98.7 | 98.4 | 98.7 | |
| InternVL3.5-4BEvaluation Protocol=Rank-Guided (RG)2026.01 | 98.7 | 98.3 | 98.4 | |
| QwenVL3-4BEvaluation Protocol=Expert Grading (EG)2026.01 | 96.8 | 95.9 | 96.4 | |
| InternVL3.5-4BEvaluation Protocol=Expert Grading (EG)2026.01 | 94.9 | 93.3 | 94.9 | |
| AutoEval-PEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 86.9 | 88.7 | 86.2 | |
| QwenVL3-4BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 85.8 | 83.5 | 85.4 | |
| InternVL3.5-4BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 85 | 83.6 | 85 | |
| SmolVLM2.2BEvaluation Protocol=Expert Grading (EG)2026.01 | 83.9 | 81.2 | 83.4 | |
| QwenVL2.5-3BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 80.6 | 78.1 | 80.4 | |
| SmolVLM2.2BEvaluation Protocol=Rank-Guided (RG)2026.01 | 76.5 | 73.1 | 76.1 | |
| SmolVLM2.2BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 68.6 | 66.5 | 68.8 | |
| InternVL3.5-4BEvaluation Protocol=Expert Grading (EG), Supervised Fine-Tuning=No (Zero-Shot)2026.01 | 46.6 | 51.3 | 50.5 | |
| InternVL3.5-4BEvaluation Protocol=Rank-Guided (RG), Supervised Fine-Tuning=No (Zero-Shot)2026.01 | 38.8 | 56 | 50 |