Machine Translation on Low-resource Translation Macro-Average (chrF)
33.35chrFOur RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=full2026.06 | 33.35 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=full2026.06 | 30.2 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=full2026.06 | 23 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=full2026.06 | 22.55 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=none2026.06 | 22.42 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=full2026.06 | 21.25 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=full2026.06 | 18.43 | |
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=none2026.06 | 18.15 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=none2026.06 | 17.96 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=none2026.06 | 17.69 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=none2026.06 | 16.06 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=none2026.06 | 11.39 |