General Domain Reasoning on ARC-c and MMLU-Pro
83.7ARC-cTRAPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TRAPOBackbone=Qwen2.5-Math-7B2025.12 | 83.7 | 52.8 | 68.3 | |
| GRPOBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 80.5 | 47.2 | 63.9 | |
| LUFFYBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 80.5 | 53 | 66.7 | |
| ReLIFTBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 76.2 | 52.5 | 64.4 | |
| PRIME-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 73.3 | 32.7 | 53 | |
| Qwen2.5-Math-7B-InstructType=Instruct Model2025.12 | 72.4 | 38.3 | 55.4 | |
| Oat-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 70.1 | 41.7 | 55.9 | |
| OpenReasoner-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 66.2 | 58.7 | 62.5 | |
| SFT-then-RLBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 52.9 | 36.1 | 44.5 | |
| SFTBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 51.5 | 33.1 | 42.3 | |
| Qwen2.5-Math-7BType=Base Model2025.12 | 34.5 | 23.3 | 28.9 | |
| SimpleRL-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 30.2 | 34.5 | 32.4 |