Scientific Reasoning on GPQA Diamond (Avg@8)
55.6Avg@8R1-Distill-Qwen-32B
Evaluation Results
| Method | Links | |
|---|---|---|
| R1-Distill-Qwen-32B# Trainable Params=32B2025.10 | 55.6 | |
| s1.1-32B# Trainable Params=32B2025.10 | 51.9 | |
| Target + THINKLOGIT-DPO# Trainable Params=78M2025.10 | 42.4 | |
| Target + THINKLOGIT# Trainable Params=02025.10 | 41.8 | |
| Qwen2.5-32B# Trainable Params=-2025.10 | 36.9 | |
| R1-Distill-Qwen-1.5B# Trainable Params=-2025.10 | 28.9 |