Language understanding and knowledge on MMLU-Pro
57.2AccuracyEvo 8B
Evaluation Results
| Method | Links | |
|---|---|---|
| Evo 8BPost-training=SFT, Number of shots=0, E2E Latency (s)=8.6, Inference Speed (tokens/s)=522026.02 | 57.2 | |
| Qwen2.5 7BPost-training=SFT+RL, Number of shots=5, E2E Latency (s)=8.1, Inference Speed (tokens/s)=462026.02 | 56.3 | |
| MDLM 9BPost-training=SFT+RL, Number of shots=5, E2E Latency (s)=18.9, Inference Speed (tokens/s)=222026.02 | 52.1 | |
| BD3-LM 7BPost-training=SFT+RL, Number of shots=5, E2E Latency (s)=14.2, Inference Speed (tokens/s)=282026.02 | 48.1 | |
| LLaMA3 8BPost-training=SFT+RL, Number of shots=0, E2E Latency (s)=7.4, Inference Speed (tokens/s)=582026.02 | 41.9 | |
| LLaDA 8BPost-training=SFT, Number of shots=0, E2E Latency (s)=21.8, Inference Speed (tokens/s)=162026.02 | 37 | |
| Full-Attn# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Full-Attn2026.02 | 33.8 | |
| HySparse# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=HySparse2026.02 | 32.6 | |
| HySparse# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=HySparse2026.02 | 29 | |
| Hybrid SWA# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Hybrid SWA2026.02 | 27.2 | |
| Full-Attn# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Full-Attn2026.02 | 26.8 | |
| Hybrid SWA# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Hybrid SWA2026.02 | 26.5 | |
| Qwen2-1.5B# Non-Emb Params=1.2B2024.07 | 21.8 | |
| Gemma-2B# Non-Emb Params=2.0B2024.07 | 15.9 | |
| Qwen2-0.5B# Non-Emb Params=0.3B2024.07 | 14.7 | |
| ARD 7BPost-training=SFT+RL, Number of shots=0, E2E Latency (s)=32.5, Inference Speed (tokens/s)=122026.02 | 12.8 |