Response Selection on PersonaMem
64.36AccuracyTALLRec
Evaluation Results
| Method | Links | |
|---|---|---|
| TALLRecInference Setting=Direct Full-history Sequence Models w/o Preference Inference2026.01 | 64.36 | |
| GPT-OSS-20BInference Setting=Full-history Preference Inference2026.01 | 61.74 | |
| DeepSeek-R1-671BInference Setting=Full-history Preference Inference2026.01 | 61.44 | |
| Qwen3-8Bnon-thinkingInference Setting=Direct Full-history Sequence Models w/o Preference Inference2026.01 | 61.4 | |
| DeepSeek-R1-671BInference Setting=Streaming Preference Inference2026.01 | 58.96 | |
| ALIGNXPLORE+Inference Setting=Full-history Preference Inference2026.01 | 58.08 | |
| Qwen3-32BthinkingInference Setting=Full-history Preference Inference2026.01 | 57.36 | |
| GPT-OSS-20BInference Setting=Streaming Preference Inference2026.01 | 54.82 | |
| ALIGNXPLORE+Inference Setting=Streaming Preference Inference2026.01 | 54.58 | |
| Qwen3-8BthinkingInference Setting=Full-history Preference Inference2026.01 | 54.36 | |
| ALIGNXPLOREInference Setting=Full-history Preference Inference2026.01 | 53.98 | |
| Qwen3-32BthinkingInference Setting=Streaming Preference Inference2026.01 | 53.26 | |
| Qwen3-8BthinkingInference Setting=Streaming Preference Inference2026.01 | 51.68 | |
| DS-R1-Distill-Qwen-7BInference Setting=Full-history Preference Inference2026.01 | 49.28 | |
| ALIGNXPLOREInference Setting=Streaming Preference Inference2026.01 | 48.42 | |
| DS-R1-Distill-Qwen-7BInference Setting=Streaming Preference Inference2026.01 | 46.64 |