Multi-task Language Understanding on MMLU-Pro (pass@1)
92.85Pass@1GT-Reward
Evaluation Results
| Method | Links | |
|---|---|---|
| GT-RewardBase Model=Qwen3-4B-Base2025.11 | 92.85 | |
| GT-RewardBase Model=Qwen3-8B-Base2025.11 | 92.85 | |
| GT-RewardBase Model=Llama-3.1-8B-Instruct2025.11 | 91.42 | |
| GT-RewardBase Model=Qwen3-1.7B-Base2025.11 | 85.71 | |
| Self-HarmonyBase Model=Qwen3-8B-Base2025.11 | 77.68 | |
| Majority-VotingBase Model=Qwen3-8B-Base2025.11 | 77.58 | |
| GT-RewardBase Model=Llama-3.2-3B-Instruct2025.11 | 72.85 | |
| RentBase Model=Qwen3-8B-Base2025.11 | 69.91 | |
| Self-HarmonyBase Model=Qwen3-4B-Base2025.11 | 67.68 | |
| RentBase Model=Qwen3-4B-Base2025.11 | 66.16 | |
| IntuitorBase Model=Qwen3-4B-Base2025.11 | 64.55 | |
| Qwen2.5-7B-InstructTraining Pipeline=PSFT → GRPO2025.08 | 63.65 | |
| Qwen2.5-7B-InstructTraining Pipeline=SFT → GRPO2025.08 | 62.61 | |
| Llama3.1-8B-InstructTraining Pipeline=PSFT → GRPO2025.08 | 60.92 | |
| IntuitorBase Model=Qwen3-8B-Base2025.11 | 60.54 | |
| Qwen2.5-7B-InstructTraining Pipeline=PSFT2025.08 | 59.18 | |
| Qwen2.5-7B-InstructTraining Pipeline=SFT2025.08 | 58.98 | |
| SFT + GRPOBackbone=Qwen2.5-7B-Instruct2026.01 | 57.8 | |
| Co-RewardBase Model=Qwen3-8B-Base2025.11 | 57.59 | |
| GRPOBackbone=Qwen2.5-7B-Instruct2026.01 | 57 | |
| Llama3.1-8B-InstructTraining Pipeline=PSFT2025.08 | 56.58 | |
| SCR (Ours)Backbone=Qwen2.5-7B-Instruct2026.01 | 56.3 | |
| Self-RefineBackbone=Qwen2.5-7B-Instruct2026.01 | 55.9 | |
| BaseBackbone=Qwen2.5-7B-Instruct2026.01 | 55.8 | |
| Llama3.1-8B-InstructTraining Pipeline=SFT → GRPO2025.08 | 54.39 | |
| Self-HarmonyBase Model=Qwen3-1.7B-Base2025.11 | 53.66 | |
| Majority-VotingBase Model=Qwen3-4B-Base2025.11 | 52.95 | |
| Co-RewardBase Model=Qwen3-4B-Base2025.11 | 51.79 | |
| SCR-Stage IBackbone=Qwen2.5-7B-Instruct2026.01 | 51.5 | |
| Llama3.1-8B-InstructTraining Pipeline=SFT2025.08 | 50.29 | |
| Before RLBase Model=Qwen3-8B-Base2025.11 | 50.09 | |
| Self-HarmonyBase Model=Llama-3.1-8B-Instruct2025.11 | 50 | |
| SCR-SFTBackbone=Qwen2.5-7B-Instruct2026.01 | 49.1 | |
| Co-RewardBase Model=Qwen3-1.7B-Base2025.11 | 47.14 | |
| GRPOBackbone=Llama3.1-8B-Instruct2026.01 | 45.5 | |
| Majority-VotingBase Model=Llama-3.1-8B-Instruct2025.11 | 45.36 | |
| SFT + GRPOBackbone=Qwen2.5-3B-Instruct2026.01 | 45.3 | |
| Majority-VotingBase Model=Qwen3-1.7B-Base2025.11 | 44.82 | |
| GRPOBackbone=Qwen2.5-3B-Instruct2026.01 | 44.6 | |
| SCR (Ours)Backbone=Qwen2.5-3B-Instruct2026.01 | 44.5 | |
| Self-HarmonyBase Model=Llama-3.2-3B-Instruct2025.11 | 44.29 | |
| Before RLBase Model=Llama-3.1-8B-Instruct2025.11 | 43.75 | |
| Self-RefineBackbone=Qwen2.5-3B-Instruct2026.01 | 43.5 | |
| SCR (Ours)Backbone=Llama3.1-8B-Instruct2026.01 | 43.2 | |
| SFT + GRPOBackbone=Llama3.1-8B-Instruct2026.01 | 43 | |
| Co-RewardBase Model=Llama-3.1-8B-Instruct2025.11 | 42.77 | |
| BaseBackbone=Llama3.1-8B-Instruct2026.01 | 42.4 | |
| SCR-Stage IBackbone=Llama3.1-8B-Instruct2026.01 | 42.2 | |
| SCR-Stage IBackbone=Qwen2.5-3B-Instruct2026.01 | 41.3 | |
| RentBase Model=Llama-3.1-8B-Instruct2025.11 | 40.8 | |
| IntuitorBase Model=Llama-3.1-8B-Instruct2025.11 | 40 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.01 | 39.6 | |
| SCR-SFTBackbone=Llama3.1-8B-Instruct2026.01 | 39.1 | |
| SCR-SFTBackbone=Qwen2.5-3B-Instruct2026.01 | 38.4 | |
| IntuitorBase Model=Llama-3.2-3B-Instruct2025.11 | 34.64 | |
| RentBase Model=Llama-3.2-3B-Instruct2025.11 | 34.38 | |
| Before RLBase Model=Llama-3.2-3B-Instruct2025.11 | 34.11 | |
| Co-RewardBase Model=Llama-3.2-3B-Instruct2025.11 | 34.02 | |
| Majority-VotingBase Model=Llama-3.2-3B-Instruct2025.11 | 31.43 | |
| IntuitorBase Model=Qwen3-1.7B-Base2025.11 | 31.25 | |
| Before RLBase Model=Qwen3-4B-Base2025.11 | 27.59 | |
| RentBase Model=Qwen3-1.7B-Base2025.11 | 18.04 | |
| Before RLBase Model=Qwen3-1.7B-Base2025.11 | 16.61 | |
| Self-RefineBackbone=Llama3.1-8B-Instruct2026.01 | 9.9 |