Personalized Agentic Social Support on ComPASS-Bench History-based setting
2.83Preference Alignment Score (w/o History)GPT-5.1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5.1Mode=Instant2026.04 | 2.83 | 2.94 | 0.11 | |
| GPT-5.1Mode=Reasoning2026.04 | 2.77 | 2.91 | 0.14 | |
| Qwen3-32BMode=Reasoning2026.04 | 2.6 | 2.34 | -0.26 | |
| Gemini-3-ProMode=Instant2026.04 | 2.46 | 2.62 | 0.16 | |
| Qwen3-MaxMode=Instant2026.04 | 2.46 | 2.65 | 0.19 | |
| Qwen3-32BMode=Instant2026.04 | 2.39 | 2.22 | -0.17 | |
| Qwen3-8BMode=Reasoning2026.04 | 2.37 | 2.28 | -0.09 | |
| Claude-Sonnet-4.5Mode=Instant2026.04 | 2.36 | 2.51 | 0.15 | |
| DeepSeek-V3.2Mode=Instant2026.04 | 2.36 | 2.46 | 0.1 | |
| ComPASS-QwenMode=Instant2026.04 | 2.36 | 2.5 | 0.14 | |
| Qwen3-8BMode=Instant2026.04 | 2.32 | 2.15 | -0.17 | |
| Llama-3.1-8BMode=Instant2026.04 | 1.93 | 1.75 | -0.18 |