Reasoning on Out-of-Domain Reasoning Suite (ARC-c, GPQA*, MMLU-Pro) (test)
93.9ARC-c Accuracy (avg@4)Fully Supervised
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Fully SupervisedBackbone=Qwen3-8B-Base, Training Setting=Fully Supervised, Labeled Data Ratio=100%2026.06 | 93.9 | 50.5 | 65.5 | 70 | |
| GeoMinBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 93.7 | 48.2 | 66.5 | 69.5 | |
| TTRLBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 92.9 | 46.1 | 64.3 | 67.8 | |
| TTRLBackbone=Qwen3-8B-Base, Training Setting=Unsupervised RLVR, Labeled Data Ratio=0%2026.06 | 92.3 | 44.6 | 64 | 67 | |
| TraPOBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 92.3 | 45.8 | 63.5 | 67.2 | |
| Co-rewardingBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 91.7 | 31.2 | 63.4 | 62.1 | |
| Co-rewardingBackbone=Qwen3-8B-Base, Training Setting=Unsupervised RLVR, Labeled Data Ratio=0%2026.06 | 91 | 27.8 | 62.7 | 60.5 | |
| Seq-entropyBackbone=Qwen3-8B-Base, Training Setting=Unsupervised RLVR, Labeled Data Ratio=0%2026.06 | 81.3 | 20.8 | 59.5 | 53.9 | |
| Seq-entropyBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 78.1 | 17.3 | 58.8 | 51.4 | |
| Tok-entropyBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 73.9 | 17.8 | 57.4 | 49.7 | |
| Self-certaintyBackbone=Qwen3-8B-Base, Training Setting=Semi-Supervised RLVR, Labeled Data Ratio=10%2026.06 | 70.8 | 22.9 | 59.7 | 51.1 | |
| GeoMinBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 65.1 | 36 | 51.6 | 50.9 | |
| Self-certaintyBackbone=Qwen3-8B-Base, Training Setting=Unsupervised RLVR, Labeled Data Ratio=0%2026.06 | 64.4 | 18.9 | 59.2 | 47.5 | |
| Tok-entropyBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 48.9 | 11.7 | 38 | 32.9 | |
| Tok-entropyBackbone=Qwen3-8B-Base, Training Setting=Unsupervised RLVR, Labeled Data Ratio=0%2026.06 | 43.1 | 10.2 | 56.7 | 36.7 | |
| Seq-entropyBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 40.3 | 15.2 | 38.2 | 31.2 | |
| TTRLBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 34.8 | 37.6 | 51.4 | 41.3 | |
| TraPOBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 34.6 | 34.5 | 51.1 | 40.1 | |
| Co-rewardingBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 31.3 | 29.9 | 51.2 | 37.5 | |
| Before RLBackbone=Qwen3-8B-Base, Training Setting=Pre-training / Baseline2026.06 | 29.7 | 11.5 | 46.6 | 29.3 | |
| Self-certaintyBackbone=Deepseek-R1-Distill-Llama-8B, Learning Setting=semi-supervised, Labeled Data Ratio=10%2026.06 | 25.3 | 16.2 | 44.1 | 28.5 |