Long-context Question Answering on BrowseComp-Plus
88.5AccuracyCodex (No Retriever)
Evaluation Results
| Method | Links | |
|---|---|---|
| Codex (No Retriever)Retriever=None, Base Model=GPT-52026.03 | 88.5 | |
| Codex + Gemini Emb.Retriever=Gemini Embeddings, Base Model=GPT-52026.03 | 84 | |
| Best Published2026.03 | 80 | |
| Codex + BM25Retriever=BM25, Base Model=GPT-52026.03 | 78.5 | |
| ReAct AgentAgent Loop=ReAct2026.03 | 72.5 | |
| RAGRetriever=Gemini Embeddings2026.03 | 65 | |
| GPT-5 Full ContextContext Type=Full Context2026.03 | 20 |