Logic puzzle reasoning on BABABENCH Tier 1
95.56Success RateHuman (with tutorial video)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human (with tutorial video)2026.01 | 95.56 | — | — | |
| LCVMethod Variant=Contrastive LCD2026.01 | 93.33 | — | 798.5 | |
| LCVMethod Variant=Vanilla LSFT2026.01 | 88.89 | — | 763.9 | |
| Gemini 3 Pro PreviewPrompting Strategy=Code-as-Policy2026.01 | 71.11 | 1.27 | 953.4 | |
| Claude Sonnet 4.5Prompting Strategy=Code-as-Policy2026.01 | 68.89 | 2.52 | 834.2 | |
| Qwen2.5-72B-InstructPrompting Strategy=Code-as-Policy2026.01 | 62.22 | — | — | |
| Gemini 3 Pro PreviewPrompting Strategy=Direct Policy2026.01 | 62.22 | 1.43 | 13.5 | |
| TheoryCoderBackbone=GPT-4o2026.01 | 62.22 | — | 2.46 | |
| Claude Sonnet 4.5Prompting Strategy=Direct Policy2026.01 | 57.78 | 1.24 | 14.7 | |
| Llama-3-70B-InstructPrompting Strategy=Code-as-Policy2026.01 | 53.33 | — | — | |
| Deepseek-v3.2Prompting Strategy=Code-as-Policy2026.01 | 40 | 4.24 | 1.16 | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 28.89 | — | — | |
| TheoryCoderBackbone=GPT-4o, Number of API calls=12026.01 | 24.44 | — | 1.13 | |
| Llama-3-70B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 17.78 | — | — | |
| Deepseek-v3.2Prompting Strategy=Direct Policy2026.01 | 17.78 | 7.38 | 15.2 | |
| Llama-3-70B-InstructPrompting Strategy=Direct Policy2026.01 | 15.56 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 13.33 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Code-as-Policy2026.01 | 13.33 | — | — | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct Policy2026.01 | 13.33 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct Policy2026.01 | 8.89 | — | — | |
| GPT-OSS-120BPrompting Strategy=Direct Policy2026.01 | 8.89 | — | — | |
| GPT-OSS-120BPrompting Strategy=Direct Policy (CoT)2026.01 | 8.89 | — | — | |
| GPT-OSS-120BPrompting Strategy=Code-as-Policy2026.01 | 6.67 | — | — |