Language Modeling on Pile (Loss)
1.876LossDeepSeekMoE 145B
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeekMoE 145B# Shot=N/A, # Total Params=144.6B, # Activated Params=22.2B, Relative Expert Size=0.125, # Experts=4 + 128, # Activated Experts=4 + 12, FLOPs per 4K Tokens=585.6T, # Training Tokens=245B2024.01 | 1.876 | |
| DeepSeekMoE 142B (Half Activated)# Shot=N/A, # Total Params=142.3B, # Activated Params=12.2B, Relative Expert Size=0.125, # Experts=2 + 128, # Activated Experts=2 + 6, FLOPs per 4K Tokens=374.6T, # Training Tokens=245B2024.01 | 1.888 | |
| DeepSeek 67B (Dense)# Shot=N/A, # Total Params=67.4B, # Activated Params=67.4B, Relative Expert Size=N/A, # Experts=N/A, # Activated Experts=N/A, FLOPs per 4K Tokens=2057.5T, # Training Tokens=245B2024.01 | 1.905 | |
| Engram-40BShots=-, Total Params=39.5B, Activated Params=3.8B, Trained Tokens=262B, Experts=2 + 55 (top-6), Engram Params=18.5B2026.01 | 1.942 | |
| Engram-27BShots=-, Total Params=26.7B, Activated Params=3.8B, Trained Tokens=262B, Experts=2 + 55 (top-6), Engram Params=5.7B2026.01 | 1.95 | |
| MoE-27BShots=-, Total Params=26.7B, Activated Params=3.8B, Trained Tokens=262B, Experts=2 + 72 (top-6)2026.01 | 1.96 | |
| GShard 137B# Shot=N/A, # Total Params=136.5B, # Activated Params=21.6B, Relative Expert Size=1, # Experts=0 + 16, # Activated Experts=0 + 2, FLOPs per 4K Tokens=572.7T, # Training Tokens=245B2024.01 | 1.961 | |
| Dense-4BShots=-, Total Params=4.1B, Activated Params=3.8B, Trained Tokens=262B2026.01 | 2.091 |