Success Prediction on Eval-Actions 1.0 (test)
91AccuracyQwenVL3-4B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| QwenVL3-4BEvaluation Protocol=Rank-Guided (RG)2026.01 | 91 | 92.9 | 88.8 | |
| AutoEval-SEvaluation Protocol=Rank-Guided (RG)2026.01 | 91 | 93 | 90.1 | |
| InternVL3.5-4BEvaluation Protocol=Rank-Guided (RG)2026.01 | 90.6 | 92.5 | 89.6 | |
| AutoEval-SEvaluation Protocol=Expert Grading (EG)2026.01 | 90.6 | 92.8 | 88.5 | |
| QwenVL3-4BEvaluation Protocol=Expert Grading (EG)2026.01 | 90.2 | 92.4 | 88.2 | |
| InternVL3.5-4BEvaluation Protocol=Expert Grading (EG)2026.01 | 90 | 92.1 | 88.1 | |
| AutoEval-PEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 83 | 86.4 | 81.2 | |
| InternVL3.5-4BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 81.7 | 84.2 | 81 | |
| QwenVL3-4BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 81 | 84 | 80.3 | |
| QwenVL2.5-3BEvaluation Protocol=Rank-Guided (RG)2026.01 | 78.4 | 83.2 | 76 | |
| QwenVL2.5-3BEvaluation Protocol=Expert Grading (EG)2026.01 | 76.1 | 81.2 | 73.8 | |
| SmolVLM2.2BEvaluation Protocol=Expert Grading (EG)2026.01 | 69.3 | 73.9 | 68.2 | |
| QwenVL2.5-3BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 69.1 | 74.3 | 67.6 | |
| SmolVLM2.2BEvaluation Protocol=Rank-Guided (RG)2026.01 | 66.4 | 71.3 | 65.5 | |
| InternVL3.5-4BEvaluation Protocol=Rank-Guided (RG), Supervised Fine-Tuning=No (Zero-Shot)2026.01 | 62.3 | 76.8 | 50 | |
| SmolVLM2.2BEvaluation Protocol=Chain-of-Thought (CoT)2026.01 | 62.1 | 67.5 | 61 | |
| InternVL3.5-4BEvaluation Protocol=Expert Grading (EG), Supervised Fine-Tuning=No (Zero-Shot)2026.01 | 56.2 | 66.9 | 51.2 |