Mathematical Reasoning on OlympiadBench (Last/Best Scores)
51.63Last ScoreDAPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DAPOContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 51.63 | 57.27 | |
| RPG-REINFORCE-UFKLContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 50.74 | 55.64 | |
| REINFORCE++-BaselineContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 49.26 | 58.75 | |
| GRPOContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 49.26 | 55.94 | |
| RPG-URKLContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 48.37 | 58.01 | |
| REINFORCE++Context Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 47.78 | 62.02 | |
| RPG-UFKLContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 46.88 | 55.64 | |
| RPG-REINFORCE-URKLContext Length=4k, Base Model=Qwen-2.5-7B-Instruct2025.05 | 46.74 | 58.16 |