Problem Solving and Unsolvability Detection on HamCycle
100Solvable AccuracyGemini-3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-3Model Scale=32025.12 | 100 | 98 | 99 | |
| Deepseek-V3.2-RModel Scale=V3.2-R2025.12 | 83.5 | 94 | 88.8 | |
| Qwen3-4B + UnsolvableRLModel Scale=4B, Training Protocol=UnsolvableRL2025.12 | 41.1 | 94.5 | 67.8 | |
| Qwen3-4B InstructModel Scale=4B, Training Protocol=Instruct2025.12 | 37.5 | 57 | 47.3 | |
| Qwen3-1.7B InstructModel Scale=1.7B, Training Protocol=Instruct2025.12 | 22.9 | 28 | 25.5 | |
| Qwen3-1.7B + UnsolvableRLModel Scale=1.7B, Training Protocol=UnsolvableRL2025.12 | 20.8 | 88 | 54.4 | |
| GPT-5.1-LowModel Scale=5.1-Low2025.12 | 8.3 | 82 | 45.1 |