Machine Translation on English-to-Dinka
22.91chrFOur RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=full2026.06 | 22.91 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=full2026.06 | 19.49 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=full2026.06 | 16.06 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=none2026.06 | 14.52 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=none2026.06 | 10.75 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=full2026.06 | 10.44 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=none2026.06 | 9.68 | |
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=none2026.06 | 6.43 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=none2026.06 | 6.4 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=full2026.06 | 5.06 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=none2026.06 | 4.1 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=full2026.06 | 3.29 |