Multilingual Mathematical Reasoning on MT Math100
93.6AccuracyROSA2
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ROSA2Model=Qwen3-8B2026.03 | 93.6 | — | — | |
| ROSAModel=Qwen3-8B2026.03 | 88.4 | — | — | |
| ROSAModel=Qwen3-8B, Update Location=LM Head, Reward Model=Model-based2025.09 | 88.37 | — | — | |
| ROSAModel=Qwen3-8B, Update Location=Hidden States, Reward Model=Rule-based2025.09 | 86.93 | — | — | |
| ROSAModel=Qwen3-8B, Update Location=LM Head, Reward Model=Rule-based2025.09 | 85.16 | — | — | |
| TextGradModel=Qwen3-8B2026.03 | 81.2 | — | — | |
| ROSA2Model=Qwen2.5-7B-Base2026.03 | 78.2 | — | — | |
| TextGradModel=Qwen2.5-7B-Base2026.03 | 75.4 | — | — | |
| BaselineModel=Qwen3-8B2026.03 | 75.2 | — | — | |
| ROSAModel=Qwen2.5-7B-Instruct, Update Location=LM Head, Reward Model=Model-based2025.09 | 75.13 | — | — | |
| BaselineModel=Qwen3-8B2025.09 | 74.74 | — | — | |
| ROSA2Model=Qwen3-0.6B-Instruct2026.03 | 73.4 | — | — | |
| ROSAModel=Qwen2.5-7B-Instruct, Update Location=LM Head, Reward Model=Rule-based2025.09 | 73.16 | — | — | |
| ROSAModel=Qwen2.5-7B-Instruct, Update Location=Hidden States, Reward Model=Rule-based2025.09 | 72.27 | — | — | |
| ROSAModel=Qwen2.5-7B-Base2026.03 | 70.4 | — | — | |
| TextGradModel=Qwen3-0.6B-Instruct2026.03 | 62.2 | — | — | |
| ROSAModel=Qwen3-0.6B-Instruct2026.03 | 62 | — | — | |
| BaselineModel=Qwen2.5-7B-Base2026.03 | 60.4 | — | — | |
| BaselineModel=Qwen2.5-7B-Instruct2025.09 | 60.34 | — | — | |
| SP3F-7BLanguage=nl2026.01 | 60.1 | 87 | — | |
| Qwen2.5-7B-Instruct + Translate TestTraining Stage=Instruct, Evaluation Protocol=Translate Test2026.01 | 60.08 | — | 59.34 | |
| SP3F-7BLanguage=vi2026.01 | 59.8 | 94.4 | — | |
| ROSAModel=Qwen3-0.6B, Update Location=LM Head, Reward Model=Model-based2025.09 | 59.4 | — | — | |
| SP3F-7BLanguage=af2026.01 | 59.1 | 83.7 | — | |
| Qwen2.5-7B-InstructLanguage=nl2026.01 | 58.5 | 58.8 | — | |
| Qwen2.5-7B-InstructLanguage=vi2026.01 | 56.9 | 75.4 | — | |
| SP3F-7BTraining Stage=Full Pipeline2026.01 | 56.84 | — | 82.93 | |
| ROSAModel=Qwen3-0.6B, Update Location=LM Head, Reward Model=Rule-based2025.09 | 56.6 | — | — | |
| SP3F-7BLanguage=he2026.01 | 56.4 | 75.5 | — | |
| Qwen2.5-7B-InstructLanguage=af2026.01 | 55.9 | 62.9 | — | |
| SP3F-7BLanguage=tr2026.01 | 53.3 | 92 | — | |
| Qwen2.5-7B-InstructLanguage=he2026.01 | 52.7 | 36.4 | — | |
| Qwen2.5-7B-InstructTraining Stage=Instruct2026.01 | 52.12 | — | 65.66 | |
| ROSAModel=Qwen3-0.6B, Update Location=Hidden States, Reward Model=Rule-based2025.09 | 51.9 | — | — | |
| SP3F-7BLanguage=tl2026.01 | 51.8 | 87.1 | — | |
| SP3F-7BLanguage=Average2026.01 | 51.7 | 83.6 | — | |
| Qwen2.5-7B-InstructLanguage=tr2026.01 | 50.6 | 70.7 | — | |
| ROSA2Model=DeepSeek-R1-Distill-Llama-8B2026.03 | 50.6 | — | — | |
| Qwen2.5-7B-InstructLanguage=Average2026.01 | 48.3 | 62.7 | — | |
| Qwen2.5-7B + RLVRTraining Stage=SFT + RLVR2026.01 | 44.5 | — | 86.1 | |
| Qwen2.5-7B-InstructLanguage=tl2026.01 | 44.4 | 56.7 | — | |
| ROSAModel=DeepSeek-R1-Distill-Llama-8B2026.03 | 38.6 | — | — | |
| SP3F-7BLanguage=gu2026.01 | 38 | 68.9 | — | |
| Qwen2.5-7B-InstructLanguage=gu2026.01 | 35.5 | 65 | — | |
| SP3F-7BLanguage=pa2026.01 | 35.1 | 80.3 | — | |
| Qwen2.5-7B-InstructLanguage=pa2026.01 | 32.1 | 75.5 | — | |
| BaselineModel=Qwen3-0.6B2025.09 | 31.3 | — | — | |
| TextGradModel=DeepSeek-R1-Distill-Llama-8B2026.03 | 30.4 | — | — | |
| Qwen2.5-7B + SFTTraining Stage=SFT2026.01 | 26.72 | — | 58.26 | |
| BaselineModel=Qwen3-0.6B-Instruct2026.03 | 26.2 | — | — | |
| ROSAModel=Qwen2.5-0.5B-Instruct, Update Location=LM Head, Reward Model=Model-based2025.09 | 25.2 | — | — | |
| ROSA2Model=Qwen2.5-0.5B-Instruct2026.03 | 25.2 | — | — | |
| ROSAModel=DeepSeek-R1-Distill-Llama-8B, Update Location=LM Head, Reward Model=Model-based2025.09 | 24.67 | — | — | |
| ROSAModel=DeepSeek-R1-Distill-Llama-8B, Update Location=Hidden States, Reward Model=Rule-based2025.09 | 23.85 | — | — | |
| ROSAModel=Qwen2.5-0.5B-Instruct, Update Location=LM Head, Reward Model=Rule-based2025.09 | 22.8 | — | — | |
| ROSAModel=DeepSeek-R1-Distill-Llama-8B, Update Location=LM Head, Reward Model=Rule-based2025.09 | 21.17 | — | — | |
| Qwen2.5-7BTraining Stage=Base2026.01 | 21.16 | — | 58.22 | |
| ROSAModel=Qwen2.5-0.5B-Instruct, Update Location=Hidden States, Reward Model=Rule-based2025.09 | 20.9 | — | — | |
| ROSAModel=Qwen2.5-0.5B-Instruct2026.03 | 19.6 | — | — | |
| TextGradModel=Qwen2.5-0.5B-Instruct2026.03 | 18.4 | — | — | |
| BaselineModel=DeepSeek-R1-Distill-Llama-8B2026.03 | 17.8 | — | — | |
| BaselineModel=DeepSeek-R1-Distill-Llama-8B2025.09 | 17.35 | — | — | |
| BaselineModel=Qwen2.5-0.5B-Instruct2025.09 | 15.4 | — | — | |
| BaselineModel=Qwen2.5-0.5B-Instruct2026.03 | 15.4 | — | — |