Knowledge-based Agent Reasoning on KARLBench 1.0 (test)
75.9BrowseComp-Plus ScoreClaude 4.6 Opus
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Claude 4.6 Opusreasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 75.9 | 79.9 | 61.4 | 83 | 58.6 | 46.1 | 77.9 | 62.3 | 67.5 | |
| KARL-BCP (VGS N = 17)Training Setup=Single-Task RL, Value-Guided Search (VGS)=true, N=172026.03 | 70.4 | — | — | — | — | — | — | — | — | |
| KARL (par. N = 20)Training Setup=Multi-Task RL, Parallel thinking=true, N=202026.03 | 69.5 | 86.7 | 58.1 | 84.2 | 60.8 | 49 | 78.1 | 63 | 68.1 | |
| GPT 5reasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 68.3 | 68.2 | 55.6 | 86.7 | 44.4 | 37.5 | 68.3 | 56.1 | 60.1 | |
| KARL (par. N = 10)Training Setup=Multi-Task RL, Parallel thinking=true, N=102026.03 | 67.5 | 86.7 | 58.6 | 84.5 | 59.7 | 47.8 | 77.1 | 62.7 | 67.5 | |
| Claude 4.5 Opusreasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 62.5 | 74.7 | 57.4 | 80.7 | 54.9 | 39.1 | 68.6 | 58 | 61.6 | |
| KARL (par. N = 3)Training Setup=Multi-Task RL, Parallel thinking=true, N=32026.03 | 62.2 | 83.7 | 57.7 | 80.8 | 55.1 | 44.8 | 73 | 59.6 | 64.1 | |
| KARL-BCPTraining Setup=Single-Task RL2026.03 | 59.6 | 68 | 51.6 | 77 | 44.1 | 32.4 | 62.3 | 51.3 | 55.5 | |
| KARLTraining Setup=Multi-Task RL2026.03 | 58.5 | 80.2 | 55.2 | 76 | 47.8 | 35.7 | 69.4 | 53.7 | 58.9 | |
| Claude 4.6 Sonnetreasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 57.9 | 77.7 | 62.6 | 81.3 | 50.2 | 43.8 | 67.8 | 59.5 | 62.3 | |
| Minimax m2.5context management=best with/without compression2026.03 | 56.5 | 69.3 | 53.3 | 78 | 39.3 | 34.5 | 62.9 | 51.3 | 55.2 | |
| Qwen 3.5 397B A17Bcontext management=best with/without compression2026.03 | 55.8 | 68.2 | 51.9 | 79.3 | 42.8 | 34.7 | 62 | 52.2 | 55.5 | |
| Claude 4.5 Sonnetreasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 54.6 | 75.2 | 55 | 79.3 | 54.8 | 32.6 | 64.9 | 55.4 | 58.6 | |
| GPT 5.2reasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 47.8 | 62 | 47.9 | 80.3 | 41.1 | 37.9 | 54.9 | 51.8 | 52.8 | |
| Claude 4.5 Haikureasoning effort=best of low, medium, and high, context management=best with/without compression2026.03 | 45.8 | 72.4 | 48.7 | 73.7 | 48 | 35 | 59.1 | 51.4 | 53.9 | |
| GLM 4.5 Aircontext management=best with/without compression2026.03 | 44.7 | 66 | 52.9 | 72.7 | 45.9 | 33.4 | 55.4 | 51.2 | 52.6 | |
| KARL-TRECTraining Setup=Single-Task RL2026.03 | 42.2 | 85 | 56.7 | 68.3 | 50.8 | 37.5 | 63.6 | 53.3 | 56.8 |