Machine Translation on English-to-Kalamang
34.64chrF ScoreOur RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=full2026.06 | 34.64 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=full2026.06 | 30.05 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=full2026.06 | 28.6 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=full2026.06 | 25.58 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=full2026.06 | 24.44 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=full2026.06 | 20.14 | |
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=none2026.06 | 16.87 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=none2026.06 | 16.69 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=none2026.06 | 14.85 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=none2026.06 | 14.7 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=none2026.06 | 13.58 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=none2026.06 | 12.25 |