Logic Puzzle Solving on ZebraLogic
60.8AccuracyGreedy (BL)
Evaluation Results
| Method | Links | |
|---|---|---|
| Greedy (BL)Model=Phi-4-reasoning-plus (13B)2026.04 | 60.8 | |
| BranchingModel=Phi-4-reasoning-plus (13B)2026.04 | 60.4 | |
| SteeringModel=Phi-4-reasoning-plus (13B)2026.04 | 60.2 | |
| TTPOModel=Phi-4-reasoning-plus (13B)2026.04 | 58.4 | |
| Beam Search (BL)Model=Phi-4-reasoning-plus (13B)2026.04 | 57.8 | |
| TTPOModel=Gemma-3-12b-it2026.04 | 55.6 | |
| Greedy (BL)Model=Gemma-3-12b-it2026.04 | 53.2 | |
| SteeringModel=Gemma-3-12b-it2026.04 | 53.2 | |
| BranchingModel=Gemma-3-12b-it2026.04 | 52.4 | |
| Beam Search (BL)Model=Gemma-3-12b-it2026.04 | 51.8 | |
| BranchingModel=Gemma-3-4b-it2026.04 | 42 | |
| TTPOModel=Gemma-3-4b-it2026.04 | 40.4 | |
| SteeringModel=Phi-4-mini-instruct (4B)2026.04 | 39.6 | |
| BranchingModel=Phi-4-mini-instruct (4B)2026.04 | 39.6 | |
| SteeringModel=Gemma-3-4b-it2026.04 | 39 | |
| Greedy (BL)Model=Gemma-3-4b-it2026.04 | 38.8 | |
| Greedy (BL)Model=Phi-4-mini-instruct (4B)2026.04 | 38.8 | |
| TTPOModel=Phi-4-mini-instruct (4B)2026.04 | 38.6 | |
| Beam Search (BL)Model=Gemma-3-4b-it2026.04 | 33.4 | |
| Beam Search (BL)Model=Phi-4-mini-instruct (4B)2026.04 | 31.8 | |
| OPDLM-8BScale=8B, Tokens=0.066B, FLOPs=4.2, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 12.9 | |
| OPDLM-4BScale=4B, Tokens=0.076B, FLOPs=2.4, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 10.5 | |
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 7.8 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 6.3 | |
| Fast-dLLM-v2-7BScale=8B, Tokens=1B, FLOPs=42, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 3.5 |