Long-term Memory Question Answering on LongMemEval-S (500 questions)
98.7KU AccuracyByteRover
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ByteRover2026.04 | 98.7 | 98.6 | 98.2 | 96.7 | 91.7 | 84.2 | 92.8 | — | — | — | |
| Memorasource=results from respective papers, configuration=differs from current paper harness2026.04 | 97.4 | 98.6 | 78.6 | 83.3 | 89.5 | 78.2 | 87.4 | — | — | — | |
| UaCGenerator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge2026.06 | 97 | 94 | 96 | 83 | 65 | 81 | — | 83 | 79.5 | — | |
| Chronossource=results from respective papers, configuration=differs from current paper harness2026.04 | 96.2 | 94.3 | 100 | 80 | 90.2 | 91.7 | 92.6 | — | — | — | |
| MemMachineGenerator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge2026.06 | 96 | 96 | 96 | 63 | 69 | 88 | — | 84.8 | 81.4 | 0.33 | |
| Hindsightevaluation harness=Gemini 3 Flash judge2026.04 | 94.9 | 97.1 | 96.4 | 80 | 91 | 87.2 | 91.4 | — | — | — | |
| HonChoevaluation harness=Gemini 3 Flash judge2026.04 | 94.9 | 94.3 | 96.4 | 90 | 88.7 | 85 | 90.4 | — | — | — | |
| SmartSearchsource=results from respective papers, configuration=differs from current paper harness2026.04 | 93.6 | 100 | 85.7 | 96.7 | 82.7 | 84.2 | 88.4 | — | — | — | |
| Full ContextGenerator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge2026.06 | 91 | 96 | 96 | 70 | 74 | 87 | — | 85.4 | 82 | 0.19 | |
| TiMemsource=results from respective papers, configuration=differs from current paper harness2026.04 | 87.7 | 96.3 | 85.7 | 55.3 | 73.4 | 72.8 | 79 | — | — | — | |
| EverMemOSGenerator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge, variant=lite2026.06 | 87 | 87 | 73 | 87 | 68 | 72 | — | 76.4 | 72.5 | 0.002 | |
| Zepsource=results from respective papers, configuration=differs from current paper harness2026.04 | 83.3 | 92.9 | 80.4 | 56.7 | 62.4 | 57.9 | 71.2 | — | — | — | |
| Full-contextsource=results from respective papers, configuration=differs from current paper harness2026.04 | 78.2 | 81.4 | 94.6 | 20 | 45.1 | 44.3 | 60.2 | — | — | — | |
| HindsightGenerator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge, variant=lite2026.06 | 78 | 93 | 66 | 77 | 62 | 72 | — | 73 | 68.9 | 10 | |
| A-MEMGenerator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge2026.06 | 54 | 49 | 93 | 40 | 37 | 44 | — | 49.6 | 45.2 | 10 | |
| Mem0Generator Model=Gemini 3 Flash, Evaluator=LLM-as-Judge2026.06 | 32 | 36 | 7 | 33 | 18 | 23 | — | 23.8 | 20.3 | 10 |