Machine Translation on English-to-Guarani
26.79chrFOur RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=full2026.06 | 26.79 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=full2026.06 | 24.86 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=none2026.06 | 20.83 | |
| Our RLBase Model=Qwen3-4B-Base, Training Method=RL, Context=none2026.06 | 18.56 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=full2026.06 | 17.28 | |
| Our RLBase Model=Llama-3.2-3B-Instruct, Training Method=RL, Context=none2026.06 | 15.77 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=full2026.06 | 13.07 | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base, Training Method=Base (untuned), Context=none2026.06 | 12.56 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=none2026.06 | 12.31 | |
| Llama-3.2-3B-InstBase Model=Llama-3.2-3B-Instruct, Training Method=Instruct-tuned, Context=none2026.06 | 7.84 | |
| Our SFTBase Model=Llama-3.2-3B-Instruct, Training Method=SFT, Context=full2026.06 | 7.13 | |
| Our SFTBase Model=Qwen3-4B-Base, Training Method=SFT, Context=full2026.06 | 6.39 |