Long-context reasoning on Oolong-Synthetic 199 samples (evaluation set)
89.77Oolong ScoreRAH
Evaluation Results
| Method | Links | |
|---|---|---|
| RAHBackbone=Claude Sonnet 4.52026.06 | 89.77 | |
| RAHBackbone=GPT-52026.06 | 81.36 | |
| Codex, No RetrieverBackbone=GPT-52026.06 | 71.75 | |
| RLMBackbone=GPT-52026.06 | 64.38 | |
| Full-context baselineBackbone=GPT-52026.06 | 59.22 |