Accuracy on AIME 2023 (Mathematical Reasoning)
13.33Accuracy (AIME 2023)LoRA-RLPO
Evaluation Results
| Method | Links | |
|---|---|---|
| LoRA-RLPOBackbone=Qwen2.5-7B-Instruct, Fine-tuning algorithm=GRPO, Training dataset=DAPO-Math-17k, Rank=r=32, Evaluation protocol=pass@16, B0=0, A0=V ⊤ r2026.06 | 13.33 | |
| LoRABackbone=Qwen2.5-7B-Instruct, Fine-tuning algorithm=GRPO, Training dataset=DAPO-Math-17k, Rank=r=32, Evaluation protocol=pass@16, B0=0, A0=N (0, 1/n)2026.06 | 11.11 | |
| MiLoRABackbone=Qwen2.5-7B-Instruct, Fine-tuning algorithm=GRPO, Training dataset=DAPO-Math-17k, Rank=r=32, Evaluation protocol=pass@16, B0=U−rΣ 1/2 −r, A0=Σ1/2 −r V ⊤ −r2026.06 | 11.11 | |
| LoRA-RLMOBackbone=Qwen2.5-7B-Instruct, Fine-tuning algorithm=GRPO, Training dataset=DAPO-Math-17k, Rank=r=32, Evaluation protocol=pass@16, B0=0, A0=V ⊤ −r2026.06 | 10 | |
| PiSSABackbone=Qwen2.5-7B-Instruct, Fine-tuning algorithm=GRPO, Training dataset=DAPO-Math-17k, Rank=r=32, Evaluation protocol=pass@16, B0=UrΣ 1/2 r, A0=Σ1/2 r V ⊤ r2026.06 | 6.67 |