Long-conversation Question Answering on LoCoMo categories 1-4
91.6Overall ScoreMem0 platform
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Mem0 platformEvaluation Protocol=Different-Protocol, Answer Model=Managed Commercial Pipeline, Ctx tokens/query=∼7,0002026.05 | 91.6 | — | — | — | — | 0.131 | |
| PRISMEvaluation Protocol=Different-Protocol, Answer Model=gpt-5.5, Ctx tokens/query=2,0232026.05 | 89.1 | 89 | 87.9 | 92.7 | 89.2 | 0.44 | |
| PRISMEvaluation Protocol=Same-Protocol, Answer Model=gpt-4o-mini, Ctx tokens/query=2,0232026.05 | 83.1 | 78.7 | 78.8 | 81.3 | 86.3 | 0.411 | |
| M-FlowEvaluation Protocol=Different-Protocol, Answer Model=gpt-5-mini, Ctx tokens/query=2,5882026.05 | 81.8 | 75.2 | 79.4 | 58.3 | 87.6 | 0.316 | |
| MAGMAEvaluation Protocol=Same-Protocol, Answer Model=gpt-4o-mini, Ctx tokens/query=3,3702026.05 | 68.8 | 52.8 | 65 | 51.7 | 77.6 | 0.204 | |
| Mem0gEvaluation Protocol=Same-Protocol, Answer Model=gpt-4o-mini, Ctx tokens/query=3,6162026.05 | 68.4 | 47.2 | 58.1 | 75.7 | 65.7 | 0.189 | |
| Mem0Evaluation Protocol=Same-Protocol, Answer Model=gpt-4o-mini, Ctx tokens/query=1,7642026.05 | 66.9 | 51.2 | 55.5 | 72.9 | 67.1 | 0.379 | |
| Full ContextEvaluation Protocol=Same-Protocol, Answer Model=gpt-4o-mini, Ctx tokens/query=26,0312026.05 | 48.1 | 46.8 | 56.2 | 48.6 | 63 | 0.018 |