Scientific and General Reasoning on Nine Competition-level Benchmarks Out-of-Distribution
57.3ARC-c ScoreSupervised
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SupervisedLearning Paradigm=Supervised, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 57.3 | 0 | 38.9 | 32.1 | |
| TRAPOLearning Paradigm=Semi-supervised, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 34.4 | 0 | 33.5 | 22.6 | |
| TTRLLearning Paradigm=Unsupervised, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 25.7 | 0 | 31.9 | 19.2 | |
| Original ModelBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Base2025.12 | 24.2 | 0.5 | 38.6 | 21.1 | |
| Self-certaintyBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 13.3 | 0 | 39.5 | 17.6 | |
| Self-certaintyBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 12.7 | 0 | 40.3 | 17.7 | |
| Sentence-level EntropyBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 12.3 | 0 | 41.9 | 18.1 | |
| TRAPOBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 12.1 | 0 | 43.4 | 18.5 | |
| Sentence-level EntropyBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 11.7 | 0 | 41.5 | 17.7 | |
| TTRLBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 11.5 | 0 | 40.9 | 17.5 | |
| Token-level EntropyBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Semi-supervised, Labeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 11.3 | 0 | 41.6 | 17.6 | |
| TTRLBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 11 | 0 | 41.8 | 17.6 | |
| Token-level EntropyBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Unsupervised, Unlabeled ID Samples=1K, Unlabeled OOD Samples=1K2025.12 | 10.5 | 0 | 38.7 | 16.4 | |
| Fully SupervisedBackbone=LLaMA-3.1-8B-Instruct, Training Protocol=Fully Supervised, Labeled ID Samples=2K2025.12 | 10.4 | 0 | 47.5 | 19.3 | |
| Original ModelLearning Paradigm=Baseline, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 3.7 | 0 | 11 | 4.9 |