Conversational Memory Retrieval on LoCoMo V3.3 (Mode A)
80Single-Hop AccuracyPaper 2
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Paper 2LLM judge=Azure GPT-5.4-mini, Ingestion chunk size=5-turn2026.04 | 80 | 60 | 60 | 85 | 70 | 74.8 | |
| SuperLocalMemory V3.3 (R3)LLM judge=Azure GPT-5.4-mini, Ingestion chunk size=5-turn, Round=Round 3, Optimization=Best of 5 rounds2026.04 | 65.1 | 49.2 | 53.8 | 82.5 | 76.1 | 70.4 | |
| SuperLocalMemory V3.3 (Baseline)LLM judge=Azure GPT-5.4-mini, Ingestion chunk size=5-turn, Status=Baseline2026.04 | 60.5 | 25.4 | 38.5 | 86.8 | 63.4 | 62.8 |