Machine Translation on NLLB 53 languages subset (test)
45.41Score (High->High)54.5B MOE model
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 54.5B MOE modelEnc experts=768, Dec experts=768, Decoding hardware=1-GPU2022.12 | 45.41 | 38.98 | 31.89 | 39.72 | 35.4 | 28.83 | 37.29 | 33.23 | 26.95 | 36.81 | |
| Fixed per layer (lang-pair)Enc experts=216, Dec experts=72, Pruning level=80%, Pruning metric=importance, Pruning granularity=language-pair-specific, Decoding hardware=1-GPU2022.12 | 45.37 | 39.06 | 31.79 | 39.2 | 35.03 | 28.47 | 37.05 | 33.16 | 26.63 | 36.59 | |
| Fixed per layer (lang)Enc experts=216, Dec experts=72, Pruning level=80%, Pruning metric=importance, Pruning granularity=language-specific, Decoding hardware=1-GPU2022.12 | 45.35 | 39.1 | 31.82 | 39.18 | 35.1 | 28.51 | 37.02 | 33.19 | 26.62 | 36.61 | |
| 3.3B dense modelEnc experts=6, Dec experts=6, Decoding hardware=1-GPU2022.12 | 44.18 | 38.3 | 31.45 | 38.24 | 34.6 | 27.93 | 35.93 | 32.02 | 26.47 | 35.81 | |
| Fixed per layer (global)Enc experts=216, Dec experts=72, Pruning level=80%, Pruning metric=importance, Pruning granularity=global, Decoding hardware=1-GPU2022.12 | 43.2 | 37.6 | 31.68 | 37.37 | 33.94 | 28.4 | 35.38 | 31.97 | 26.84 | 35.34 |