Reading Comprehension on RACE-h (val)
84.3AccuracyUnified PPT
Evaluation Results
| Method | Links | |
|---|---|---|
| Unified PPTtuning_strategy=unified pre-trained prompt tuning2021.09 | 84.3 | |
| PPTtuning_strategy=pre-trained prompt tuning2021.09 | 83.9 | |
| FTtuning_strategy=fine-tuning2021.09 | 83.7 | |
| PTtuning_strategy=prompt-tuning2021.09 | 82.5 | |
| DeepSeekMoE# Shot=5-shot, # Total Params=2.0B, # Activated Params=0.3B, FLOPs per 2K Tokens=4.3T, # Training Tokens=100B2024.01 | 31.7 | |
| Switch# Shot=5-shot, # Total Params=2.0B, # Activated Params=0.2B, FLOPs per 2K Tokens=2.9T, # Training Tokens=100B2024.01 | 30.9 | |
| GShard# Shot=5-shot, # Total Params=2.0B, # Activated Params=0.3B, FLOPs per 2K Tokens=4.3T, # Training Tokens=100B2024.01 | 30.4 | |
| Hash Layer# Shot=5-shot, # Total Params=2.0B, # Activated Params=0.2B, FLOPs per 2K Tokens=2.9T, # Training Tokens=100B2024.01 | 30 | |
| Dense# Shot=5-shot, # Total Params=0.2B, # Activated Params=0.2B, FLOPs per 2K Tokens=2.9T, # Training Tokens=100B2024.01 | 29 |