Chain-of-thought reasoning on EUREQA (held-out half of hard_5)
21.6Best@3Random-w
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Random-wModel=Qwen3-8B, Prompting strategy=single-answer prompt2026.05 | 21.6 | 22.2 | 22.9 | 23.7 | 10.5 | |
| VPOModel=Qwen3-8B, Prompting strategy=multi-answer prompt2026.05 | 21.3 | 23.6 | 25.7 | 27.9 | 51.2 | |
| GRPOModel=Qwen3-8B, Prompting strategy=single-answer prompt2026.05 | 21.2 | 21.9 | 22.6 | 23.6 | 11.9 | |
| Multi-RLVRModel=Qwen3-8B, Prompting strategy=multi-answer prompt2026.05 | 21 | 23 | 24.9 | 26.7 | 52.6 | |
| MaxRLModel=Qwen3-8B, Prompting strategy=single-answer prompt2026.05 | 20.9 | 21.6 | 22.4 | 23.7 | 11.7 | |
| Max-at-KModel=Qwen3-8B, Prompting strategy=single-answer prompt2026.05 | 20.6 | 21.4 | 22.4 | 23.7 | 14 | |
| Qwen3-8BPrompting strategy=multi-answer prompt2026.05 | 10.2 | 11.3 | 12.6 | 14 | 17.1 | |
| Qwen3-8BPrompting strategy=single-answer prompt2026.05 | 8.1 | 9.6 | 11.7 | 15.3 | 17.1 |