Problem Solving and Unsolvability Detection on Game24
98Solvable AccuracyDeepseek-V3.2-R
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Deepseek-V3.2-RModel Scale=V3.2-R2025.12 | 98 | 77.5 | 87.8 | |
| Qwen3-4B + UnsolvableRLModel Scale=4B, Training Protocol=UnsolvableRL2025.12 | 94.5 | 99 | 96.8 | |
| Gemini-3Model Scale=32025.12 | 94 | 94 | 94 | |
| Qwen3-1.7B + UnsolvableRLModel Scale=1.7B, Training Protocol=UnsolvableRL2025.12 | 84 | 100 | 92 | |
| Qwen3-4B InstructModel Scale=4B, Training Protocol=Instruct2025.12 | 81 | 17 | 49 | |
| Qwen3-1.7B InstructModel Scale=1.7B, Training Protocol=Instruct2025.12 | 78 | 23 | 50.5 | |
| GPT-5.1-LowModel Scale=5.1-Low2025.12 | 14 | 50 | 32 |