Personalized retrieval and QA over heterogeneous user corpora on PersonaBench Noise Level 0.5
25.99F1 ScoreMemCoE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MemCoEBackbone=Qwen2.5-7B-Instruct, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 25.99 | 52.02 | |
| A-MemBackbone=Qwen2.5-7B-Instruct, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 25.19 | 42.64 | |
| RAGBackbone=Qwen2.5-7B-Instruct, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 24.31 | 36.68 | |
| LightMemBackbone=Qwen2.5-7B-Instruct, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 19.65 | 41.21 | |
| Mem0Backbone=Qwen2.5-7B-Instruct, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 19.22 | 38.23 | |
| Long ContextBackbone=Qwen2.5-7B-Instruct2026.05 | 17.83 | 26.9 | |
| MemAgentBackbone=Qwen2.5-7B-Instruct, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 16.51 | 45 | |
| Mem-αBackbone=Qwen3-4B, Retrieval model=all-MiniLM-L6-v2, Top-K retrieval=102026.05 | 16.43 | 44.19 |