Logical Reasoning on LogiHard-2k Hard-Mode IRT-CAT
79.5H-Comb AccuracyHuman
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Humansource=30 graduate evaluators2026.05 | 79.5 | — | — | — | 0.28 | |
| GLM-52026.05 | 38.33 | 69.7 | 0.494 | -0.554 | 1.048 | |
| GPT-5.42026.05 | 31.67 | 57.45 | -0.346 | -1.458 | 1.112 | |
| Kimi-k2.52026.05 | 31.15 | 63.89 | 0.119 | -1.581 | 1.7 | |
| DeepSeek-V4-Pro2026.05 | 30 | 61.76 | 0.006 | -1.48 | 1.486 | |
| GLM-4.72026.05 | 30 | 74.2 | 1.738 | -1.402 | 3.141 | |
| DeepSeek-R12026.05 | 26.67 | 54.39 | -0.446 | -1.831 | 1.386 | |
| Claude-Opus-4-62026.05 | 26.67 | 56.82 | -0.217 | -1.741 | 1.524 | |
| o32026.05 | 25 | 62.79 | 0.045 | -1.805 | 1.85 | |
| Gemini-3.1-pro2026.05 | 23.33 | 59.52 | -0.077 | -1.744 | 1.667 | |
| Qwen3.5-plus2026.05 | 18.33 | 65.38 | 0.287 | -2.323 | 2.609 | |
| DeepSeek-V3.22026.05 | 16.67 | 72.73 | 0.549 | -2.367 | 2.916 | |
| Qwen3.6-plustemperature=1.0, max_tokens=65,5362026.05 | 5 | 51.67 | -0.534 | -2.871 | 2.336 |