Cloze Test on CHID
90.3AccuracyDeepSeekMoE 145B
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeekMoE 145B# Shot=0-shot2024.01 | 90.3 | |
| DeepSeekMoE 16B# Shot=0-shot, # Total Params=16.4B, # Activated Params=2.8B, FLOPs per 4K Tokens=74.4T, # Training Tokens=2T2024.01 | 89.4 | |
| DeepSeekMoE 16B# Shot=0-shot, # Total Params=16.4B, # Activated Params=2.8B, FLOPs per 4K Tokens=74.4T, # Training Tokens=2T2024.01 | 89.4 | |
| DeepSeek 7B (Dense)# Shot=0-shot, # Total Params=6.9B, # Activated Params=6.9B, FLOPs per 4K Tokens=183.5T, # Training Tokens=2T2024.01 | 89.3 | |
| DeepSeek 67B (Dense)# Shot=0-shot2024.01 | 88.5 | |
| DeepSeekMoE 142B (Half Activated)# Shot=0-shot2024.01 | 88.3 | |
| GShard 137B# Shot=0-shot2024.01 | 86.9 | |
| DenseCompression Ratio=0%, Post-training compensation=No, Evaluation Protocol=Zero-shot2024.12 | 41.6 | |
| LLaMA2 7B# Shot=0-shot, # Total Params=6.7B, # Activated Params=6.7B, FLOPs per 4K Tokens=187.9T, # Training Tokens=2T2024.01 | 37.9 | |
| LaCoCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 36.1 | |
| LLMPrunerCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 28.4 | |
| GRASPCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 26.2 | |
| LLM-Streamline-LayerCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 24.1 | |
| LLM-Streamline-FFNCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 22.8 | |
| ShortGPTCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 21.5 | |
| SliceGPTCompression Ratio=25%, Post-training compensation=Yes, Evaluation Protocol=Zero-shot2024.12 | 18.5 |