Long-term conversation recall on LOCOMO Cross-LLM portability
80.8LLM-as-Judge AccuracyUaC (GPT-5.4)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| UaC (GPT-5.4)N (Sample Size)=120, Judge Model=Gemini 3 Flash, Answer Generation Model=GPT-5.42026.06 | 80.8 | 72.9 | — |
| Method | Links | |||
|---|---|---|---|---|
| UaC (GPT-5.4)N (Sample Size)=120, Judge Model=Gemini 3 Flash, Answer Generation Model=GPT-5.42026.06 | 80.8 | 72.9 | — |