Sensitivity to Logical Boundaries on QuestBench
0.4391Logic-QALIVE-Self
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ALIVE-SelfBackbone=Qwen3-30B-A3B-Instr2026.02 | 0.4391 | 0.3135 | |
| Qwen3-30B-A3B-Instr2026.02 | 0.4018 | 0.085 | |
| GPT-4o2026.02 | 0.3278 | 0.1451 | |
| DeepSeek-V3.22026.02 | 0.2713 | 0.2365 | |
| Kimi-K22026.02 | 0.1513 | 0.2103 |