Human Evaluation Translation on Sand-Glass Zh-to-En
4.97AccuracyDeepSeek-V3
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DeepSeek-V3Setting=w/o syll.2026.01 | 4.97 | — | 5 | 4.99 | |
| GPT-5Setting=w/o syll.2026.01 | 4.96 | — | 5 | 4.98 | |
| Gemini-2.5-ProSetting=w/o syll.2026.01 | 4.95 | — | 5 | 4.97 | |
| Claude-4.1-OpusSetting=w/o syll.2026.01 | 4.91 | — | 4.98 | 4.94 | |
| HOMURA_RubricSetting=Ours2026.01 | 4.91 | — | 4.83 | 4.87 | |
| HOMURA_ReasonSetting=Ours2026.01 | 4.85 | — | 4.93 | 4.89 | |
| GPT-5Setting=Best-of-N2026.01 | 4.71 | — | 4.8 | 4.75 | |
| Gemini-2.5-ProSetting=Best-of-N2026.01 | 4.57 | — | 4.57 | 4.57 | |
| DeepSeek-V3Setting=Best-of-N2026.01 | 4.56 | — | 4.58 | 4.57 | |
| Claude-4.1-OpusSetting=Best-of-N2026.01 | 4.51 | — | 4.49 | 4.5 | |
| Gemini-2.5-ProSetting=w/ syll.2026.01 | 4.44 | — | 4.9 | 4.67 | |
| Claude-4.1-OpusSetting=w/ syll.2026.01 | 4.42 | — | 4.93 | 4.67 | |
| GPT-5Setting=w/ syll.2026.01 | 4.42 | — | 4.87 | 4.65 | |
| DeepSeek-V3Setting=w/ syll.2026.01 | 4.42 | — | 4.87 | 4.65 |