Multi-turn conversation on Long-MT-Bench+
7.36AccuracyRhea
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Rhea2025.12 | 7.36 | 29.08 | |
| BM25(RAG)2025.12 | 6.65 | 10.81 | |
| Vanilla2025.12 | 6.32 | 27.29 | |
| MemGAS2025.12 | 6.07 | — | |
| Recent-k2025.12 | 5.03 | 13.89 | |
| Reply-Soft-Compress2025.12 | 4.55 | 31.79 | |
| LongAlpacaBase Model=Vicuna-7B2025.12 | 2.43 | 23.73 | |
| MemochaBase Model=Vicuna-7B2025.12 | 1.88 | 11.87 | |
| LlmLingua22025.12 | 1.5 | 29.73 | |
| Summary2025.12 | 1.49 | 33.55 |