Reasoning on Out-of-Domain Reasoning Suite
94.5ARC-c ScoreLIE
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LIEModel Size=8B, Training Recipe=GSPO + LIE2026.02 | 94.5 | 55.1 | 70.7 | 73.4 | |
| GSPOModel Size=8B, Training Recipe=GSPO2026.02 | 93.2 | 50 | 68.3 | 70.5 | |
| LIEModel Size=4B, Training Recipe=GSPO + LIE2026.02 | 91.4 | 47.5 | 63.8 | 67.6 | |
| GSPOModel Size=4B, Training Recipe=GSPO2026.02 | 88.4 | 48.5 | 61.5 | 66.1 | |
| TRAPOTraining Paradigm=Semi-supervised, Labeled Samples=4K, Unlabeled Samples=12K2025.12 | 84.6 | 43.9 | 50.7 | 59.7 | |
| TRAPOTraining Paradigm=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 83.6 | 38.9 | 48.1 | 56.9 | |
| On-Policy RLTraining Paradigm=Fully Supervised, Labeled Samples=45K2025.12 | 82.3 | 40.4 | 49.3 | 57.3 | |
| Fully SupervisedTraining Paradigm=Fully Supervised, Labeled Samples=2K2025.12 | 82 | 38.9 | 52.4 | 57.8 | |
| GSPOModel Size=1.7B, Training Recipe=GSPO2026.02 | 79.4 | 28.2 | 41 | 49.5 | |
| LIEModel Size=1.7B, Training Recipe=GSPO + LIE2026.02 | 79.1 | 33.8 | 45.6 | 52.8 | |
| Self-certaintyTraining Paradigm=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 77.1 | 32.8 | 45.7 | 51.9 | |
| TTRLTraining Paradigm=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 76.7 | 33.8 | 36.2 | 48.9 | |
| Self-certaintyTraining Paradigm=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 76.7 | 37.9 | 45.6 | 53.4 | |
| Token-level EntropyTraining Paradigm=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 76.5 | 30.8 | 44.7 | 50.7 | |
| Sentence-level EntropyTraining Paradigm=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 75.1 | 31.3 | 44.3 | 50.2 | |
| Token-level EntropyTraining Paradigm=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 74.5 | 36.4 | 35.8 | 48.9 | |
| Sentence-level EntropyTraining Paradigm=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 74.5 | 34.8 | 43.3 | 50.9 | |
| PRIME-ZeroTraining Paradigm=Fully Supervised, Labeled Samples=45K2025.12 | 73.3 | 18.2 | 32.7 | 41.4 | |
| Qwen-InstructTraining Paradigm=Instruction-tuned2025.12 | 70.3 | 24.7 | 34.1 | 43 | |
| Qwen-InstructTraining Paradigm=Original Model2025.12 | 70.3 | 24.7 | 34.1 | 43 | |
| Oat-ZeroTraining Paradigm=Fully Supervised, Labeled Samples=45K2025.12 | 70.1 | 23.7 | 41.7 | 45.2 | |
| Qwen3-BaseModel Size=4B, Training Recipe=Base2026.02 | 66.9 | 26.3 | 30.9 | 41.4 | |
| OpenReasoner-ZeroTraining Paradigm=Fully Supervised, Labeled Samples=45K2025.12 | 66.2 | 29.8 | 58.7 | 51.6 | |
| TTRLTraining Paradigm=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 62 | 31.8 | 43.5 | 45.8 | |
| Qwen3-BaseModel Size=8B, Training Recipe=Base2026.02 | 58.5 | 32.3 | 51.2 | 47.3 | |
| Qwen3-BaseModel Size=1.7B, Training Recipe=Base2026.02 | 54.1 | 20.2 | 27.5 | 33.9 | |
| SimpleRL-ZeroTraining Paradigm=Fully Supervised, Labeled Samples=45K2025.12 | 30.2 | 23.2 | 34.5 | 29.3 | |
| Qwen-BaseTraining Paradigm=Pre-trained2025.12 | 18.2 | 11.1 | 16.9 | 15.4 | |
| Qwen-BaseTraining Paradigm=Original Model2025.12 | 18.2 | 11.1 | 16.9 | 15.4 |