Machine Translation on lexical choice benchmark
54BLEUgpt-4o
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| gpt-4oReasoning=without2025.10 | 54 | 70 | 92 | 89 | — | |
| gpt-4oReasoning=with2025.10 | 54 | 69 | 92 | 88 | 0.31 | |
| gpt-4Reasoning=with2025.10 | 54 | 69 | 92 | 87 | 2.58 | |
| gpt-4Reasoning=without2025.10 | 51 | 67 | 91 | 86 | — | |
| gpt-4-turboReasoning=without2025.10 | 50 | 66 | 91 | 87 | — | |
| gpt-4-turboReasoning=with2025.10 | 49 | 65 | 89 | 86 | -0.31 | |
| gpt-3.5-turboReasoning=with2025.10 | 49 | 66 | 91 | 86 | 2.11 | |
| gpt-3.5-turboReasoning=without2025.10 | 47 | 65 | 91 | 86 | — | |
| Llama3.3Reasoning=with2025.10 | 47 | 63 | 90 | 84 | 0.15 | |
| Llama3.3Reasoning=without2025.10 | 46 | 62 | 90 | 85 | — | |
| Phi-4Reasoning=with2025.10 | 44 | 61 | 90 | 84 | 1.77 | |
| Phi-4Reasoning=without2025.10 | 43 | 60 | 89 | 83 | — | |
| DeepSeek-R1 32BReasoning=with2025.10 | 41 | 58 | 88 | 81 | 2.07 | |
| DeepSeek-R1 32BReasoning=without2025.10 | 39 | 56 | 88 | 82 | — | |
| DeepSeek-R1 14BReasoning=without2025.10 | 38 | 55 | 87 | 80 | — | |
| Llama3.1Reasoning=without2025.10 | 35 | 52 | 87 | 79 | — | |
| DeepSeek-R1 14BReasoning=with2025.10 | 33 | 51 | 85 | 77 | -5.1 | |
| nllb-200Reasoning=without2025.10 | 31 | 51 | 86 | 77 | — | |
| DeepSeek-R1 8BReasoning=without2025.10 | 31 | 48 | 86 | 76 | — | |
| DeepSeek-R1 8BReasoning=with2025.10 | 30 | 48 | 83 | 75 | -0.63 | |
| Llama3.1Reasoning=with2025.10 | 29 | 49 | 86 | 77 | -5.71 | |
| MistralReasoning=with2025.10 | 28 | 45 | 85 | 75 | 1.17 | |
| MistralReasoning=without2025.10 | 27 | 45 | 84 | 74 | — | |
| Llama3.2Reasoning=without2025.10 | 25 | 42 | 83 | 73 | — | |
| Llama3.2Reasoning=with2025.10 | 22 | 43 | 82 | 70 | -2.81 |