Logic puzzle reasoning on BABABENCH Tier 2
82.22Success RateHuman (with tutorial video)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human (with tutorial video)2026.01 | 82.22 | — | — | |
| LCVMethod Variant=Contrastive LCD2026.01 | 75.56 | — | 1.16 | |
| LCVMethod Variant=Vanilla LSFT2026.01 | 60 | — | 1.25 | |
| Qwen2.5-72B-InstructPrompting Strategy=Code-as-Policy2026.01 | 53.33 | — | — | |
| TheoryCoderBackbone=GPT-4o2026.01 | 53.33 | — | 3.22 | |
| Deepseek-v3.2Prompting Strategy=Code-as-Policy2026.01 | 31.11 | 5.67 | 892.4 | |
| Claude Sonnet 4.5Prompting Strategy=Code-as-Policy2026.01 | 31.11 | 2.87 | 943.6 | |
| Gemini 3 Pro PreviewPrompting Strategy=Code-as-Policy2026.01 | 31.11 | 1.65 | 1.05 | |
| Llama-3-70B-InstructPrompting Strategy=Code-as-Policy2026.01 | 26.67 | — | — | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 20 | — | — | |
| Deepseek-v3.2Prompting Strategy=Direct Policy2026.01 | 13.33 | 16 | 12.8 | |
| Claude Sonnet 4.5Prompting Strategy=Direct Policy2026.01 | 13.33 | 3.31 | 11.3 | |
| Gemini 3 Pro PreviewPrompting Strategy=Direct Policy2026.01 | 13.33 | 2.07 | 19.1 | |
| TheoryCoderBackbone=GPT-4o, Number of API calls=12026.01 | 13.33 | — | 1.31 | |
| Qwen2.5-7B-InstructPrompting Strategy=Code-as-Policy2026.01 | 11.11 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 6.67 | — | — | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct Policy2026.01 | 6.67 | — | — | |
| GPT-OSS-120BPrompting Strategy=Code-as-Policy2026.01 | 6.67 | — | — | |
| Llama-3-70B-InstructPrompting Strategy=Direct Policy2026.01 | 6.67 | — | — | |
| Llama-3-70B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 6.67 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct Policy2026.01 | 2.22 | — | — | |
| GPT-OSS-120BPrompting Strategy=Direct Policy2026.01 | 2.22 | — | — | |
| GPT-OSS-120BPrompting Strategy=Direct Policy (CoT)2026.01 | 2.22 | — | — |