Logical Reasoning on LogiHard-2k Full-set Baseline
84Original AccuracyHuman
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Humansource=30 graduate evaluators2026.05 | 84 | — | — | |
| GLM-4.72026.05 | 83 | 81.5 | 82.1 | |
| GLM-52026.05 | 82 | 80.15 | 81.3 | |
| DeepSeek-V3.22026.05 | 82 | 81.3 | 80.5 | |
| DeepSeek-V4-Pro2026.05 | 81.5 | 80.2 | 82.4 | |
| DeepSeek-R12026.05 | 81.41 | 82.1 | 80.5 | |
| Gemini-3.1-pro2026.05 | 81 | 79.8 | 80.1 | |
| o32026.05 | 81 | 80.5 | 81.8 | |
| GPT-5.42026.05 | 80.5 | 79.62 | 78.8 | |
| Kimi-k2.52026.05 | 79.5 | 80.2 | 78.5 | |
| Claude-Opus-4-62026.05 | 79.06 | 78.5 | 79.8 | |
| Qwen3.5-plus2026.05 | 79 | 78.5 | 79.6 | |
| Qwen3.6-plustemperature=1.0, max_tokens=65,5362026.05 | 78.5 | 77.2 | 76.8 |