Question Answering on FinanceBench N=150
98.7AccuracyMafin 2.5 (Vectify AI)
Evaluation Results
| Method | Links | |
|---|---|---|
| Mafin 2.5 (Vectify AI)Context Width=–, Model=GPT-4o2026.05 | 98.7 | |
| Golden Evidence + GPT-5-miniLLM Backbone=GPT-5-mini, Retrieval Strategy=oracle retrieval (full page)2026.05 | 94 | |
| Our Oracle (evidence pages)Context Width=Evidence pages, Model=GPT-4.12026.05 | 93.3 | |
| AgenticRAGLLM Backbone=GPT-5-mini, Tool Suite=search, find, open, summ.2026.05 | 92 | |
| AgenticRAGLLM Backbone=Claude Sonnet 4.5, Tool Suite=search, find, open, summ.2026.05 | 91.78 | |
| Oracle (evidence pages)Context Width=Evidence pages, Model=GPT-4-Turbo2026.05 | 85 | |
| OODAContext Width=–, Model=GPT-3.5-turbo2026.05 | 82 | |
| Full filing in contextContext Width=95K tokens, Model=GPT-4-Turbo2026.05 | 79 | |
| DatabricksContext Width=64K tokens, Model=o1-preview-20242026.05 | 75 | |
| Single vector store per filingContext Width=–, Model=GPT-4-Turbo2026.05 | 50 | |
| Agentic w. keyword search toolsTool Suite=pdfgrep, rga, linux cmd2026.05 | 32.71 | |
| Traditional RAGRetrieval Strategy=Traditional RAG2026.05 | 24.24 | |
| Shared vector storeContext Width=–, Model=GPT-4-Turbo2026.05 | 19 | |
| Closed bookContext Width=–, Model=GPT-4-Turbo2026.05 | 9 |