Mathematical Reasoning on In-Distribution Avg
45.6Average ScoreTRAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| TRAPOTraining Paradigm=Semi-supervised, Labeled Samples Count=4K, Unlabeled Samples Count=12K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 45.6 | |
| Fully SupervisedTraining Paradigm=Supervised, Labeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 45.5 | |
| HiLLBackbone=Qwen2.5-7B-Instruct2026.04 | 44.2 | |
| Fully SupervisedTraining Paradigm=Supervised, Labeled Samples Count=4K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 43.1 | |
| HiLLw/o TWBackbone=Qwen2.5-7B-Instruct2026.04 | 42.7 | |
| TRAPOTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 42.6 | |
| SAGEBackbone=Qwen2.5-7B-Instruct2026.04 | 42.3 | |
| LUFFYBackbone=Qwen2.5-7B-Instruct2026.04 | 41.7 | |
| GRPOBackbone=Qwen2.5-7B-Instruct2026.04 | 41.1 | |
| Scaf-GRPOBackbone=Qwen2.5-7B-Instruct2026.04 | 41 | |
| Self-certaintyTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 40 | |
| Token-level EntropyTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 40 | |
| Sentence-level EntropyTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 39.7 | |
| Fully SupervisedTraining Paradigm=Supervised, Labeled Samples Count=1K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 39.4 | |
| TTRLTraining Paradigm=Semi-supervised, Labeled Samples Count=1K, Unlabeled Samples Count=3K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 39.2 | |
| Self-certaintyTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 38.3 | |
| TTRLTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 38.2 | |
| BaseBackbone=Qwen2.5-7B-Instruct2026.04 | 37.8 | |
| Qwen-InstructTraining Paradigm=Original, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 37.6 | |
| Token-level EntropyTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 37.6 | |
| Sentence-level EntropyTraining Paradigm=Unsupervised, Unlabeled Samples Count=45K, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 32.6 | |
| HiLLBackbone=Llama-3.2-3B-Instruct2026.04 | 24.6 | |
| SAGEBackbone=Llama-3.2-3B-Instruct2026.04 | 23.9 | |
| HiLLw/o TWBackbone=Llama-3.2-3B-Instruct2026.04 | 23.7 | |
| GRPOBackbone=Llama-3.2-3B-Instruct2026.04 | 21.9 | |
| Scaf-GRPOBackbone=Llama-3.2-3B-Instruct2026.04 | 21.5 | |
| Qwen-BaseTraining Paradigm=Original, Backbone Model=Qwen2.5-Math-7B, Sampling Temperature (T=0.6)=0.62025.12 | 19 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.04 | 17.8 | |
| LUFFYBackbone=Llama-3.2-3B-Instruct2026.04 | 14.7 |