Machine Translation on WMT24++ (test)
77.8en-ar ScoreStructured Reasoning
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Structured ReasoningSetup=Structured Reasoning, Reasoning Source=Internal (Proposed)2026.02 | 77.8 | 87.1 | 83.6 | 86.4 | 66.3 | 83.8 | 84.7 | 86.2 | 82.1 | 82 | — | — | |
| Structured ReasoningSetup=Ablation, Training Data Size=10k Samples2026.02 | 77.6 | 86.7 | 83.4 | 85.9 | 65.9 | 83.6 | 84.7 | 86.1 | 81.8 | 81.7 | — | — | |
| Command-A-ReasoningSetup=Injected Reasoning, Reasoning Trajectory Source=Command-A-Reasoning2026.02 | 76.9 | 86.7 | 82.9 | 85.7 | 66.1 | 83.4 | 84.5 | 85.8 | 82.3 | 81.6 | — | — | |
| Direct TranslationSetup=Baseline, Training=Standard MT SFT2026.02 | 76.7 | 86.3 | 82.7 | 85.7 | 66.2 | 83.7 | 84.5 | 85.5 | 81.9 | 81.5 | — | — | |
| DeepSeek-R1Setup=Injected Reasoning, Reasoning Trajectory Source=DeepSeek-R12026.02 | 76.5 | 86.5 | 83.1 | 85.7 | 66.1 | 83.7 | 84.5 | 85.6 | 82.3 | 81.6 | — | — | |
| Claude-4-Opusreasoning_traces=without, metric=XCOMET-XL2026.02 | 75.1 | 85.2 | 81 | 83.4 | 64.6 | 84 | 83.8 | 85.1 | 81.6 | 80.4 | — | — | |
| Claude-4-Opusreasoning_traces=with, metric=XCOMET-XL2026.02 | 74.9 | 84.9 | 81.5 | 83.2 | 64.4 | 83.9 | 83.5 | 84.2 | 80.7 | 80.1 | — | — | |
| Structured ReasoningSetup=Ablation, Training Data Size=1k Samples2026.02 | 74.3 | 80.2 | 81 | 83.5 | 60.9 | 80 | 80.6 | 83.8 | 79.4 | 78.2 | — | — | |
| Gemini-2.5-Flashreasoning_traces=without, metric=XCOMET-XL2026.02 | 74 | 85.5 | 81.2 | 83.8 | 63.8 | 82.7 | 83.6 | 84.5 | 80.1 | 79.9 | — | — | |
| DeepSeek-R1reasoning_traces=without, metric=XCOMET-XL2026.02 | 73 | 83.7 | 79.2 | 82.7 | 63.1 | 81.3 | 81.8 | 83.8 | 79.3 | 78.7 | — | — | |
| Command-A-Reasoningreasoning_traces=without, metric=XCOMET-XL2026.02 | 72.6 | 83.5 | 81.2 | 83 | 60.4 | 80.7 | 81.7 | 82.6 | 78.6 | 78.3 | — | — | |
| Gemini-2.5-Flashreasoning_traces=with, metric=XCOMET-XL2026.02 | 71.2 | 83.9 | 79.9 | 82.5 | 63.6 | 81.9 | 81.3 | 83.4 | 79.4 | 78.6 | — | — | |
| Base modelSetup=Baseline, Backbone=Command-A-Reasoning, Training=MT without reasoning2026.02 | 70.4 | 83 | 78.6 | 82.5 | 59.8 | 80.6 | 81.4 | 82 | 77.2 | 77.3 | — | — | |
| Command-A-Reasoningreasoning_traces=with, metric=XCOMET-XL2026.02 | 70.4 | 83 | 78.6 | 82.5 | 59.8 | 80.6 | 81.4 | 82 | 77.2 | 77.3 | — | — | |
| DeepSeek-R1reasoning_traces=with, metric=XCOMET-XL2026.02 | 69.5 | 82.5 | 76.8 | 82 | 61 | 74.8 | 74.2 | 83 | 70.4 | 74.9 | — | — | |
| Gemini 2.5 Flash†Model Type=LLM2026.03 | — | — | — | — | — | — | — | — | — | — | 24.2 | — | |
| Gemini 2.5 Flash†Model Type=LLM2026.03 | — | — | — | — | — | — | — | — | — | — | 51.9 | 92.2 | |
| Gemini 3 Flash (preview)Model Type=LLM2026.03 | — | — | — | — | — | — | — | — | — | — | 27.5 | — | |
| Gemini 3 Flash (preview)Model Type=LLM2026.03 | — | — | — | — | — | — | — | — | — | — | 53.1 | 93.5 | |
| Gemini 3 Pro (preview)Model Type=LLM2026.03 | — | — | — | — | — | — | — | — | — | — | 32.9 | — | |
| Gemini 3 Pro (preview)Model Type=LLM2026.03 | — | — | — | — | — | — | — | — | — | — | 53.7 | 93.4 | |
| NLLBData Augmentation Strategy=No data augmentation2026.03 | — | — | — | — | — | — | — | — | — | — | 29.5 | — | |
| NLLBData Augmentation Strategy=HR→LR augmentation2026.03 | — | — | — | — | — | — | — | — | — | — | 26.3 | — | |
| NLLBData Augmentation Strategy=LR→HR augmentation2026.03 | — | — | — | — | — | — | — | — | — | — | 44.1 | — | |
| NLLBData Augmentation Strategy=LR→HR augmentation + dictionary prompting2026.03 | — | — | — | — | — | — | — | — | — | — | 44.3 | — | |
| NLLBData Augmentation Strategy=No data augmentation2026.03 | — | — | — | — | — | — | — | — | — | — | 35.2 | 80.1 | |
| NLLBData Augmentation Strategy=HR→LR augmentation2026.03 | — | — | — | — | — | — | — | — | — | — | 44.7 | 88.9 | |
| NLLBData Augmentation Strategy=LR→HR augmentation2026.03 | — | — | — | — | — | — | — | — | — | — | 48.5 | 91.6 | |
| NLLBData Augmentation Strategy=LR→HR augmentation + dictionary prompting2026.03 | — | — | — | — | — | — | — | — | — | — | 48.8 | 91.8 |