Machine Translation Evaluation on WMT MQM 2022 (test)
91.6Accuracy (System, 3 LPs)Remedy-R
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Remedy-RType=LLM Judges, Parameters=32B2025.12 | 91.6 | 55.2 | 57.8 | 55.7 | 52.2 | 73.4 | |
| ReMedyType=Scalar Metrics, Parameters=9B2025.12 | 91.2 | 58.9 | 61 | 60.4 | 55.4 | 75.1 | |
| EAPrompt (GPT3.5-Turbo)Type=LLM Judges, Parameters=>100B2025.12 | 91.2 | 53.3 | 56.7 | 53.3 | 50 | 72.3 | |
| PaLMType=LLM Judges, Parameters=540B2025.12 | 90.1 | 50.8 | 55.4 | 48.6 | 48.5 | 70.5 | |
| GEMBA-DA (GPT4)Type=LLM Judges, Parameters=>100B2025.12 | 89.8 | 55.6 | 58.2 | 55 | 53.4 | 72.7 | |
| Remedy-RType=LLM Judges, Parameters=7B2025.12 | 89.1 | 54.8 | 58 | 56 | 50.4 | 71.9 | |
| Remedy-RType=LLM Judges, Parameters=14B2025.12 | 88.7 | 56 | 58 | 55.8 | 54.2 | 72.4 | |
| MQM-APE (Mistral)Type=LLM Judges, Parameters=8x22B2025.12 | 88.3 | 54.2 | 56.9 | 55.1 | 50.6 | 71.3 | |
| PaLM-2 BISON FTType=Scalar Metrics, Parameters=>100B2025.12 | 88 | 57.3 | 61 | 51.5 | 59.5 | 72.7 | |
| MQM-APE (Qwen)Type=LLM Judges, Parameters=72B2025.12 | 85.8 | 54.5 | 56.4 | 55.7 | 51.4 | 70.2 | |
| EAPrompt (Llama2)Type=LLM Judges, Parameters=70B2025.12 | 85.4 | 52.3 | 55.2 | 51.4 | 50.2 | 68.9 | |
| MetricX-XXLType=Scalar Metrics, Parameters=13B2025.12 | 85 | 58.8 | 61.1 | 54.6 | 60.6 | 71.9 | |
| GEMBA-MQM (Qwen)Type=LLM Judges, Parameters=72B2025.12 | 84.7 | 53.8 | 56 | 54.7 | 50.6 | 69.3 | |
| EAPrompt (Mistral)Type=LLM Judges, Parameters=8x7B2025.12 | 84 | 50.9 | 53.8 | 50.6 | 48.2 | 67.5 | |
| COMET-22 (ensemble)Type=Scalar Metrics, Parameters=5x0.5B2025.12 | 83.9 | 57.3 | 60.2 | 54.1 | 57.7 | 70.6 | |
| COMET-22-DAType=Scalar Metrics, Parameters=0.5B2025.12 | 82.8 | 54.5 | 58.2 | 49.5 | 55.7 | 68.7 |