Problem Solving and Unsolvability Detection on Overall
97.4Solvable AccuracyGemini-3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-3Model Scale=32025.12 | 97.4 | 84.1 | 90.8 | |
| Deepseek-V3.2-RModel Scale=V3.2-R2025.12 | 88.4 | 84.4 | 86.1 | |
| Qwen3-4B + UnsolvableRLModel Scale=4B, Training Protocol=UnsolvableRL2025.12 | 69.4 | 87.5 | 78.6 | |
| GPT-5.1-LowModel Scale=5.1-Low2025.12 | 45.9 | 66.6 | 56.2 | |
| Qwen3-4B InstructModel Scale=4B, Training Protocol=Instruct2025.12 | 43.4 | 38.8 | 41.1 | |
| Qwen3-1.7B + UnsolvableRLModel Scale=1.7B, Training Protocol=UnsolvableRL2025.12 | 25.5 | 76.4 | 50.9 | |
| Qwen3-1.7B InstructModel Scale=1.7B, Training Protocol=Instruct2025.12 | 23 | 41.7 | 32.4 |