Logic Reasoning on Zebralogic (Score)
96.1ScoreQwen 3 VL 32B Think
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 3 VL 32B ThinkModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=true2025.12 | 96.1 | |
| Qwen 3 VL 8B Think2025.12 | 91.2 | |
| Qwen 3 32BModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=false2025.12 | 88.3 | |
| Qwen 3 8B2025.12 | 85.2 | |
| Olmo 3.1 Think 32BTraining Stage=Final Think 3.1, Model Family=Olmo 3.1, Parameter Count=32B, Thinking Capability=true2025.12 | 80.1 | |
| K2-V2 70B InstructModel Family=K2, Parameter Count=70B, Thinking Capability=false2025.12 | 79.2 | |
| Olmo 3 Think (Final 3.0)Training Stage=Final Think 3.0, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 76 | |
| Olmo 3 Think (DPO)Training Stage=DPO, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 74.5 | |
| Olmo 3 Think (SFT)Training Stage=SFT, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 70.5 | |
| DS-R1 32BModel Family=DeepSeek-R1, Parameter Count=32B, Thinking Capability=true2025.12 | 69.4 | |
| Olmo 3 7B ThinkStage=Final Think2025.12 | 66.5 | |
| Qwen 3 VL 8B Inststage=Instruct2025.12 | 64.3 | |
| Nemotron Nano 9B v22025.12 | 60.8 | |
| Olmo 3 7B ThinkStage=DPO2025.12 | 60.6 | |
| Olmo 3 7B ThinkStage=SFT2025.12 | 57.9 | |
| OpenThinker3 7B2025.12 | 34.9 | |
| Olmo 3 7B Instructstage=Final Instruct2025.12 | 32.9 | |
| Olmo 3 7B Instructstage=DPO2025.12 | 28.4 | |
| DS-R1 Qwen 7B2025.12 | 26.1 | |
| Qwen 3 8Bstage=Instruct2025.12 | 25.4 | |
| OR Nemotron 7B2025.12 | 22.4 | |
| Olmo 3 7B Instructstage=SFT2025.12 | 18 | |
| Granite 3.3 8B Inststage=Instruct2025.12 | 17.6 | |
| DARLBase=Inst, Verifier=✗2026.01 | 14.2 | |
| Qwen 2.5 7Bstage=Instruct2025.12 | 10.7 | |
| RLVRBase=Inst, Verifier=Rule2026.01 | 10.1 | |
| Qwen2.5-7B-InstBase=Base, Verifier=-2026.01 | 10.1 | |
| DARLBase=Base, Verifier=✗2026.01 | 9.2 | |
| Llama3.1-8B-InstBase=Inst, Verifier=-2026.01 | 8.9 | |
| General ReasonerBase=Base, Verifier=Model2026.01 | 8.9 | |
| RLPRBase=Inst, Verifier=✗2026.01 | 8.3 | |
| RLVRBase=Base, Verifier=Rule2026.01 | 7.7 | |
| VeriFreeBase=Base, Verifier=✗2026.01 | 7.6 | |
| RLPRBase=Base, Verifier=✗2026.01 | 6.4 | |
| SimpleRL-ZooBase=Base, Verifier=Rule2026.01 | 6.3 | |
| OLMo 2 7B Inststage=Instruct2025.12 | 5.3 | |
| Apertus 8B Inststage=Instruct2025.12 | 5.3 | |
| TTRLBase=Base, Verifier=Rule2026.01 | 2.1 | |
| Qwen2.5-7BBase=-, Verifier=-2026.01 | 0.8 | |
| SimpleRL-ZooBase=Math, Verifier=Rule2026.01 | 0.7 | |
| Oat-ZeroBase=Math, Verifier=Rule2026.01 | 0.2 | |
| PRIMEBase=Math, Verifier=Rule2026.01 | 0 |