Mathematical Reasoning on HMMT 25 (Accuracy)
45.42Accuracy (HMMT 25)p1
Evaluation Results
| Method | Links | |
|---|---|---|
| p1Backbone=Qwen3-4B-Instruct-2507, S=[1, 23], M=16, Training Reward Imp.=0.212026.04 | 45.42 | |
| General Teacher2026.05 | 45 | |
| p1Backbone=Qwen3-4B-Instruct-2507, S=[17, 27], M=16, Training Reward Imp.=0.092026.04 | 42.34 | |
| p1Backbone=Qwen3-4B-Instruct-2507, S=[4, 5, 17, 20], M=8, Training Reward Imp.=0.042026.04 | 42.08 | |
| p1Backbone=Qwen3-4B-Instruct-2507, S=[23], M=32, Training Reward Imp.=0.072026.04 | 41.25 | |
| baseBackbone=Qwen3-4B-Instruct-2507, S=/, M=/, Training Reward Imp.=/2026.04 | 40.68 | |
| GEPABackbone=Qwen3-4B-Instruct-2507, S=15/15, M=/, Training Reward Imp.=/2026.04 | 40.57 | |
| RLBackbone=Qwen3-4B-Instruct-2507, S=all 30, M=1, Training Reward Imp.=02026.04 | 40.26 | |
| GEPABackbone=Qwen3-4B-Instruct-2507, S=28/[1,23], M=/, Training Reward Imp.=/2026.04 | 40.21 | |
| GEPABackbone=Qwen3-4B-Instruct-2507, S=[1,23]/28, M=/, Training Reward Imp.=/2026.04 | 40.1 | |
| Uni-OPDDistillation Scenario=Single-Teacher Distillation, Model Scale=4B2026.05 | 39.8 | |
| Uni-OPDDistillation Scenario=Multi-Teacher Distillation, Model Scale=4B2026.05 | 39.6 | |
| ExOPDDistillation Scenario=Single-Teacher Distillation, Model Scale=4B2026.05 | 39.3 | |
| ExOPDDistillation Scenario=Multi-Teacher Distillation, Model Scale=4B2026.05 | 39.2 | |
| p1Backbone=Qwen3-4B-Instruct-2507, S=[25], M=32, Training Reward Imp.=0.072026.04 | 38.59 | |
| TeacherTraining Strategy=RL2026.05 | 38.5 | |
| OPDDistillation Scenario=Multi-Teacher Distillation, Model Scale=4B2026.05 | 38.3 | |
| OPDDistillation Scenario=Single-Teacher Distillation, Model Scale=4B2026.05 | 37.8 | |
| ExPODistillation Scenario=Single-Teacher Distillation, Model Scale=4B2026.05 | 37 | |
| ExPODistillation Scenario=Multi-Teacher Distillation, Model Scale=4B2026.05 | 36.3 | |
| SFTDistillation Scenario=Multi-Teacher Distillation, Model Scale=4B2026.05 | 34.8 | |
| CaMOPD2026.05 | 34.58 | |
| Vanilla MOPD2026.05 | 33.96 | |
| SelecTKD2026.05 | 33.54 | |
| Relaxed OPD2026.05 | 32.92 | |
| Medical Teacher2026.05 | 28.33 | |
| p1Backbone=Qwen3-1.7B, S=[5, 26], M=16, Training Reward Imp.=0.132026.04 | 27.08 | |
| p1Backbone=Qwen3-1.7B, S=[10, 11], M=16, Training Reward Imp.=0.142026.04 | 26.93 | |
| GEPABackbone=Qwen3-1.7B, S=28/[5,26], M=/, Training Reward Imp.=/2026.04 | 25.52 | |
| baseBackbone=Qwen3-1.7B, S=/, M=/, Training Reward Imp.=/2026.04 | 25.36 | |
| GEPABackbone=Qwen3-1.7B, S=15/15, M=/, Training Reward Imp.=/2026.04 | 25.16 | |
| p1Backbone=Qwen3-1.7B, S=[5], M=32, Training Reward Imp.=0.252026.04 | 25.05 | |
| GEPABackbone=Qwen3-1.7B, S=[5,26]/28, M=/, Training Reward Imp.=/2026.04 | 25 | |
| RLBackbone=Qwen3-1.7B, S=all 30, M=1, Training Reward Imp.=02026.04 | 24.9 | |
| p1Backbone=Qwen3-1.7B, S=[5, 14, 20, 26], M=8, Training Reward Imp.=0.052026.04 | 24.79 | |
| p1Backbone=Qwen3-1.7B, S=[26], M=32, Training Reward Imp.=0.292026.04 | 24.38 | |
| DyJR (M = 16)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=162026.03 | 13.6 | |
| DyJRBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.05, M=82026.03 | 12.7 | |
| DyJR (alpha_JS = 0.01)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.012026.03 | 12.5 | |
| DyJR (alpha_JS = 0.2)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.22026.03 | 10.9 | |
| RLEPBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.8 | |
| DPH-RLBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.5 | |
| DyJR (Forward KL)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, Divergence Type=Forward KL2026.03 | 10.5 | |
| Ex-GRPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.4 | |
| DyJR (M = 32)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=322026.03 | 10.3 | |
| GRPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 9.8 | |
| DyJR (M = 64)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=642026.03 | 9.7 | |
| DAPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 9.2 | |
| StudentModel Scale=4B2026.05 | 9.2 | |
| Base ModelBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 1.9 |