Reasoning on In-Domain Reasoning Benchmarks (AIME, AMC, MATH-500, Minerva, Olympiad)
40AIME 24 ScoreGeoMin
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GeoMinBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 40 | 23.3 | 73.5 | 83.6 | 32.7 | 55.3 | 51.4 | |
| TTRLBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 33.3 | 20 | 73.5 | 80.4 | 30.9 | 47.4 | 47.6 | |
| TraPOBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 33.3 | 23.3 | 67.5 | 81 | 32 | 52.7 | 48.3 | |
| Co-rewardingBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 30 | 20 | 69.9 | 79.6 | 28.7 | 48.1 | 46.1 | |
| Self-certaintyBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 16.7 | 16.7 | 51.8 | 73 | 26.1 | 43.8 | 38 | |
| Tok-entropyBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 6.7 | 6.7 | 36.1 | 59.4 | 19.1 | 29.3 | 26.2 | |
| Seq-entropyBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 3.3 | 10 | 42.2 | 56.2 | 22.4 | 29.4 | 27.3 |