General Reasoning on GPQA diamond
75.18Avg@8 AccuracyROSA (+LM + M)
Evaluation Results
| Method | Links | |
|---|---|---|
| ROSA (+LM + M)Model=Qwen3-8B2025.09 | 75.18 | |
| ROSA (+HS + R)Model=Qwen3-8B2025.09 | 70.27 | |
| ROSA (+LM + R)Model=Qwen3-8B2025.09 | 69.11 | |
| ROSA (+LM + M)Model=Qwen2.5-7B-Instruct2025.09 | 45.83 | |
| ROSA (+HS + R)Model=Qwen2.5-7B-Instruct2025.09 | 43.16 | |
| ROSA (+LM + R)Model=Qwen2.5-7B-Instruct2025.09 | 42.24 | |
| BaselineModel=Qwen3-8B2025.09 | 41.16 | |
| GPSBackbone=DeepSeek-R1-Distill-1.5B, Strategy=GPS, Runtime=16h2026.02 | 30.4 | |
| PCLBackbone=DeepSeek-R1-Distill-1.5B, Strategy=PCL, Runtime=17h2026.02 | 28.5 | |
| Uniform SamplingBackbone=DeepSeek-R1-Distill-1.5B, Strategy=Uniform, Runtime=16h2026.02 | 27.5 | |
| MoPPSBackbone=DeepSeek-R1-Distill-1.5B, Strategy=MoPPS, Runtime=17h2026.02 | 27.5 | |
| Dynamic Sampling (Oracle)Backbone=DeepSeek-R1-Distill-1.5B, Strategy=DS, Runtime=30h2026.02 | 26.8 | |
| GRESOBackbone=DeepSeek-R1-Distill-1.5B, Strategy=GRESO, Runtime=27h2026.02 | 26.4 | |
| BaselineModel=Qwen2.5-7B-Instruct2025.09 | 26.14 | |
| ROSA (+LM + M)Model=DeepSeek-R1-Distill-Llama-8B2025.09 | 25.36 | |
| GRESOBackbone=DeepSeek-R1-Distill-7B, Strategy=GRESO, Runtime=53h2026.02 | 25 | |
| MoPPSBackbone=DeepSeek-R1-Distill-7B, Strategy=MoPPS, Runtime=42h2026.02 | 23.4 | |
| DeepSeek-R1-Distill-1.5BBackbone=DeepSeek-R1-Distill-1.5B, Strategy=Base2026.02 | 22.8 | |
| PCLBackbone=DeepSeek-R1-Distill-7B, Strategy=PCL, Runtime=50h2026.02 | 22.4 | |
| ROSA (+HS + R)Model=DeepSeek-R1-Distill-Llama-8B2025.09 | 22.23 | |
| GPSBackbone=DeepSeek-R1-Distill-7B, Strategy=GPS, Runtime=49h2026.02 | 22.2 | |
| ROSA (+LM + R)Model=DeepSeek-R1-Distill-Llama-8B2025.09 | 21.14 | |
| Dynamic Sampling (Oracle)Backbone=DeepSeek-R1-Distill-7B, Strategy=DS, Runtime=77h2026.02 | 19.6 | |
| DeepSeek-R1-Distill-7BBackbone=DeepSeek-R1-Distill-7B, Strategy=Base2026.02 | 19.2 | |
| BaselineModel=DeepSeek-R1-Distill-Llama-8B2025.09 | 19.03 | |
| Uniform SamplingBackbone=DeepSeek-R1-Distill-7B, Strategy=Uniform, Runtime=40h2026.02 | 17.3 | |
| ROSA (+LM + M)Model=Qwen3-0.6B2025.09 | 13.16 | |
| BaselineModel=Qwen3-0.6B2025.09 | 12.2 | |
| ROSA (+HS + R)Model=Qwen3-0.6B2025.09 | 10.54 | |
| ROSA (+LM + M)Model=Qwen2.5-0.5B-Instruct2025.09 | 10.27 | |
| ROSA (+LM + R)Model=Qwen3-0.6B2025.09 | 9.09 | |
| ROSA (+HS + R)Model=Qwen2.5-0.5B-Instruct2025.09 | 8.53 | |
| ROSA (+LM + R)Model=Qwen2.5-0.5B-Instruct2025.09 | 7.07 | |
| BaselineModel=Qwen2.5-0.5B-Instruct2025.09 | 3.54 |