Generative Language Modeling and Problem Solving on IFEval, AIME25, GSM8K, GPQA, HumanEval, LCB Suite
90.4IFEval ScoreOriginal
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| OriginalExperts (N)=128, Calibration Mixture (C4:Math:Code)=–, Base Model=Qwen3-30B-A3B-Instruct-25072026.04 | 90.4 | 56.7 | 89.3 | 47 | 93.3 | 48.6 | 70.9 | |
| REAMExperts (N)=96, Calibration Mixture (C4:Math:Code)=0:0.5:0.5, Base Model=Qwen3-30B-A3B-Instruct-25072026.04 | 89.9 | 60 | 86.3 | 38.4 | 93.3 | 51 | 69.8 | |
| REAPExperts (N)=96, Calibration Mixture (C4:Math:Code)=0.2:0.25:0.55, Base Model=Qwen3-30B-A3B-Instruct-25072026.04 | 89.6 | 50 | 87.9 | 39.4 | 94.5 | 50.3 | 68.6 | |
| HC-SMoEExperts (N)=96, Calibration Mixture (C4:Math:Code)=0.5:0:0.5, Base Model=Qwen3-30B-A3B-Instruct-25072026.04 | 88.2 | 60 | 84.7 | 34.3 | 91.5 | 45.9 | 67.4 | |
| FreqExperts (N)=96, Calibration Mixture (C4:Math:Code)=0:0.3:0.7, Base Model=Qwen3-30B-A3B-Instruct-25072026.04 | 87.8 | 60 | 82.9 | 36.9 | 93.9 | 44 | 67.6 |