Compound LLM Collaboration System on Human-preferred responses dataset
19.8Win Rate (Chosen)SysDPO-Sampling
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SysDPO-SamplingBackbone=Qwen1.5-1.8B-Chat2025.02 | 19.8 | 66.4 | |
| SysDPO-ψ2Stage=ψ2 training only2025.02 | 18.1 | 63.9 | |
| Separate-DPOBackbone=Qwen1.5-1.8B-Chat2025.02 | 16.6 | 57.3 | |
| SysDPO-ψ1Stage=ψ1 training only2025.02 | 16 | 60.4 | |
| Prompted SystemBackbone=Qwen1.5-1.8B-Chat2025.02 | 12.8 | — |