Response Selection on AlignX
75.03AccuracyALIGNXPLORE+
Evaluation Results
| Method | Links | |
|---|---|---|
| ALIGNXPLORE+Inference Setting=Full-history Preference Inference2026.01 | 75.03 | |
| ALIGNXPLORE+Inference Setting=Streaming Preference Inference2026.01 | 73.67 | |
| ALIGNXPLOREInference Setting=Streaming Preference Inference2026.01 | 69.9 | |
| ALIGNXPLOREInference Setting=Full-history Preference Inference2026.01 | 66.6 | |
| TALLRecInference Setting=Direct Full-history Sequence Models w/o Preference Inference2026.01 | 66.3 | |
| DeepSeek-R1-671BInference Setting=Full-history Preference Inference2026.01 | 65.9 | |
| Qwen3-32BthinkingInference Setting=Full-history Preference Inference2026.01 | 64.93 | |
| Qwen3-32BthinkingInference Setting=Streaming Preference Inference2026.01 | 64.6 | |
| DeepSeek-R1-671BInference Setting=Streaming Preference Inference2026.01 | 64.06 | |
| Qwen3-8BthinkingInference Setting=Streaming Preference Inference2026.01 | 62.9 | |
| Qwen3-8BthinkingInference Setting=Full-history Preference Inference2026.01 | 62.73 | |
| Qwen3-8Bnon-thinkingInference Setting=Direct Full-history Sequence Models w/o Preference Inference2026.01 | 59.63 | |
| GPT-OSS-20BInference Setting=Streaming Preference Inference2026.01 | 56.86 | |
| DS-R1-Distill-Qwen-7BInference Setting=Streaming Preference Inference2026.01 | 56.4 | |
| GPT-OSS-20BInference Setting=Full-history Preference Inference2026.01 | 55.63 | |
| DS-R1-Distill-Qwen-7BInference Setting=Full-history Preference Inference2026.01 | 54.03 |