Document Translation on 200-page document translation benchmark
4.59LFBabelDOC
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| BabelDOCEvaluation Protocol=Human Evaluation2026.05 | 4.59 | — | 4.28 | 4.46 | 4.47 | 2.85 | |
| BabelDOCEvaluation Protocol=LLM-as-a-Judge, LLM Judge Model=Gemini-2.5-Flash2026.05 | 4.46 | — | 4.19 | 4.49 | 4.43 | 1.7 | |
| DeepL (Doc)Evaluation Protocol=LLM-as-a-Judge, LLM Judge Model=Gemini-2.5-Flash2026.05 | 4.2 | — | 4.19 | 4.24 | 4.38 | 2.03 | |
| DeepL (Doc)Evaluation Protocol=Human Evaluation2026.05 | 3.44 | — | 3.62 | 3.63 | 4.21 | 2.33 | |
| PDFMathTrans.Evaluation Protocol=Human Evaluation2026.05 | 3.29 | — | 3.4 | 3.28 | 3.34 | 6.25 | |
| PDFMathTrans.Evaluation Protocol=LLM-as-a-Judge, LLM Judge Model=Gemini-2.5-Flash2026.05 | 2.55 | — | 2.78 | 2.54 | 3.02 | 5.55 | |
| BabelDOCEvaluation Protocol=Automatic Layout Metric2026.05 | — | 50 | — | — | — | — | |
| DeepL (Doc)Evaluation Protocol=Automatic Layout Metric2026.05 | — | 19.8 | — | — | — | — | |
| PDFMathTrans.Evaluation Protocol=Automatic Layout Metric2026.05 | — | 48.7 | — | — | — | — |