Long-context Question Answering on Oolong Synthetic
71.75ScoreCodex (No Retriever)
Evaluation Results
| Method | Links | |
|---|---|---|
| Codex (No Retriever)Retriever=None, Base Model=GPT-52026.03 | 71.75 | |
| Codex + BM25Retriever=BM25, Base Model=GPT-52026.03 | 71.07 | |
| Codex + Gemini Emb.Retriever=Gemini Embeddings, Base Model=GPT-52026.03 | 68.03 | |
| RLM2026.03 | 64.38 | |
| Best Published2026.03 | 64.38 | |
| GPT-5 Full ContextContext Type=Full Context2026.03 | 59.22 | |
| RAGRetriever=Gemini Embeddings2026.03 | 45.53 | |
| ReAct AgentAgent Loop=ReAct2026.03 | 31.39 |