Math Tutoring on BigMath In-Domain
57.4RsolOPRO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OPROenables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Ped.2026.05 | 57.4 | 38.2 | 84.6 | 68.4 | |
| EvoPromptenables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Gen.2026.05 | 56.6 | 28.6 | 84.7 | 71.1 | |
| MIPROv2enables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Gen.2026.05 | 56.3 | 29.8 | 84.8 | 70.7 | |
| ParetoGradenables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Gen.2026.05 | 56.3 | 25.2 | 84.5 | 71.9 | |
| MetaBlendenables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Gen.2026.05 | 56.2 | 39.2 | 84.3 | 67.6 | |
| TF-GRPOenables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Gen.2026.05 | 56.1 | 31.2 | 84.5 | 70 | |
| Frameenables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Ped.2026.05 | 55.8 | 35.4 | 84.5 | 68.3 | |
| CondBridgeenables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Gen.2026.05 | 55.2 | 33.6 | 84.3 | 69.1 | |
| ACEenables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Ped.2026.05 | 55.1 | 46 | 83.3 | 64.2 | |
| GEPAenables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Gen.2026.05 | 54.2 | 32.4 | 84.1 | 68.6 | |
| TextGradenables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Ped.2026.05 | 52.7 | 37.4 | 84.5 | 66.6 | |
| LeakShieldenables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Gen.2026.05 | 49.6 | 26.4 | 84.7 | 69.3 | |
| DeepSeek-V3.2 (Ped.)enables thinking (Th.)=true, prompt seed type=Ped.2026.05 | 39 | 11 | 82 | 70 | |
| Claude-4-Opus (Ped.)enables thinking (Th.)=true, prompt seed type=Ped.2026.05 | 35 | 9 | 76 | 67.3 | |
| GPT-5.2 (Ped.)enables thinking (Th.)=true, prompt seed type=Ped.2026.05 | 34 | 0 | 44 | 59.3 | |
| Ped. Think R (RL)enables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Ped.2026.05 | 29.4 | 17.2 | 77.6 | 63.3 | |
| Think R (RL)enables thinking (Th.)=true, applies thinking reward (Th.R)=true, prompt seed type=Gen.2026.05 | 28.4 | 18.2 | 76.4 | 62.1 | |
| Think NR (RL)enables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Gen.2026.05 | 28.1 | 18 | 73 | 60.4 | |
| Ped. Think NR (RL)enables thinking (Th.)=true, applies thinking reward (Th.R)=false, prompt seed type=Ped.2026.05 | 27.5 | 21.4 | 76.6 | 60.7 | |
| NoThink (RL)enables thinking (Th.)=false, applies thinking reward (Th.R)=false, prompt seed type=Gen.2026.05 | 12 | 30 | 18 | 33.3 | |
| No optimizationenables thinking (Th.)=false, applies thinking reward (Th.R)=false, prompt seed type=Gen.2026.05 | 12 | 30 | 18 | 33.3 |