Logic puzzle reasoning on BABABENCH Tier 3
7,800Success Rate (SR)Human (with tutorial video)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human (with tutorial video)2026.01 | 7,800 | — | — | |
| LCVMethod Variant=Contrastive LCD2026.01 | 6,200 | — | 1.31 | |
| TheoryCoderBackbone=GPT-4o2026.01 | 5,200 | — | 3.08 | |
| LCVMethod Variant=Vanilla LSFT2026.01 | 4,800 | — | 1.27 | |
| Qwen2.5-72B-InstructPrompting Strategy=Code-as-Policy2026.01 | 4,200 | — | — | |
| Deepseek-v3.2Prompting Strategy=Code-as-Policy2026.01 | 3,600 | 3.93 | 962.9 | |
| Claude Sonnet 4.5Prompting Strategy=Code-as-Policy2026.01 | 3,000 | 2.22 | 677.4 | |
| Gemini 3 Pro PreviewPrompting Strategy=Code-as-Policy2026.01 | 2,800 | 1.18 | 1.14 | |
| Llama-3-70B-InstructPrompting Strategy=Code-as-Policy2026.01 | 1,800 | — | — | |
| Gemini 3 Pro PreviewPrompting Strategy=Direct Policy2026.01 | 1,000 | 2 | 10.4 | |
| GPT-OSS-120BPrompting Strategy=Code-as-Policy2026.01 | 800 | — | — | |
| Llama-3-70B-InstructPrompting Strategy=Direct Policy2026.01 | 800 | — | — | |
| Llama-3-70B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 800 | — | — | |
| Deepseek-v3.2Prompting Strategy=Direct Policy2026.01 | 800 | 16 | 18.5 | |
| Claude Sonnet 4.5Prompting Strategy=Direct Policy2026.01 | 800 | 3.15 | 16.9 | |
| TheoryCoderBackbone=GPT-4o, Number of API calls=12026.01 | 800 | — | 1.31 | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 667 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct Policy (CoT)2026.01 | 400 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Code-as-Policy2026.01 | 400 | — | — | |
| GPT-OSS-120BPrompting Strategy=Direct Policy (CoT)2026.01 | 400 | — | — | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct Policy2026.01 | 222 | — | — | |
| GPT-OSS-120BPrompting Strategy=Direct Policy2026.01 | 200 | — | — | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct Policy2026.01 | 0 | — | — |