Causal Reasoning on CLadder 1.0 (test)
94.8Overall AccHuman
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| HumanModel Configuration=Human Expert2025.11 | 94.8 | — | — | — | — | — | — | |
| Qwen3 (Final)Model Configuration=Final (CRAwDAD Debate)2025.11 | 89.41 | 96.24 | 93.51 | 80.35 | 89.52 | 87.91 | 91.2 | |
| DeepSeek-R1 (Final)Model Configuration=Final (CRAwDAD Debate)2025.11 | 87.45 | 94.62 | 89.06 | 80.04 | 87.83 | 86.94 | 87.67 | |
| Qwen3Model Configuration=Initial2025.11 | 84.16 | 93.77 | 89.8 | 71.53 | 84.96 | 83.1 | 84.69 | |
| DeepSeek-R1Model Configuration=Initial2025.11 | 78.03 | 90.67 | 77.29 | 67.94 | 77.63 | 77.24 | 79.42 | |
| GPT-4 + CausalCoTPrompting Strategy=CausalCoT2025.11 | 70.4 | 83.35 | 67.47 | 62.05 | 69.25 | 71.58 | 70.12 | |
| GPT-4Model Configuration=Initial2025.11 | 62.03 | 63.01 | 62.82 | 60.55 | 62.27 | 63.09 | 60.47 |