Response Selection on P-Soups Style
0.88AccuracyQwen3-32Bthinking
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-32BthinkingInference Setting=Full-history Preference Inference2026.01 | 0.88 | |
| Qwen3-8BthinkingInference Setting=Full-history Preference Inference2026.01 | 0.875 | |
| ALIGNXPLORE+Inference Setting=Full-history Preference Inference2026.01 | 0.8633 | |
| GPT-OSS-20BInference Setting=Full-history Preference Inference2026.01 | 0.86 | |
| DeepSeek-R1-671BInference Setting=Full-history Preference Inference2026.01 | 0.8566 | |
| DeepSeek-R1-671BInference Setting=Streaming Preference Inference2026.01 | 0.85 | |
| Qwen3-8BthinkingInference Setting=Streaming Preference Inference2026.01 | 0.85 | |
| Qwen3-32BthinkingInference Setting=Streaming Preference Inference2026.01 | 0.8366 | |
| GPT-OSS-20BInference Setting=Streaming Preference Inference2026.01 | 0.83 | |
| ALIGNXPLORE+Inference Setting=Streaming Preference Inference2026.01 | 0.8033 | |
| ALIGNXPLOREInference Setting=Full-history Preference Inference2026.01 | 0.78 | |
| ALIGNXPLOREInference Setting=Streaming Preference Inference2026.01 | 0.7483 | |
| TALLRecInference Setting=Direct Full-history Sequence Models w/o Preference Inference2026.01 | 0.7016 | |
| DS-R1-Distill-Qwen-7BInference Setting=Full-history Preference Inference2026.01 | 0.6583 | |
| DS-R1-Distill-Qwen-7BInference Setting=Streaming Preference Inference2026.01 | 0.6083 | |
| Qwen3-8Bnon-thinkingInference Setting=Direct Full-history Sequence Models w/o Preference Inference2026.01 | 0.4233 |