General Knowledge Evaluation on CEVAL
85.52AccuracyQwen3-14B-Base
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-14B-BaseBackbone=Qwen3-14B-Base, CPT Strategy=None2025.12 | 85.52 | |
| DOS-CPTBackbone=Qwen3-14B-Base, CPT Strategy=Distance-to-Optimum Selection2025.12 | 85.44 | |
| HPS-CPTBackbone=Qwen3-14B-Base, CPT Strategy=High-PPL Sampling2025.12 | 85.17 | |
| RS-CPTBackbone=Qwen3-14B-Base, CPT Strategy=Random Data Sampling2025.12 | 85.14 | |
| LPS-CPTBackbone=Qwen3-14B-Base, CPT Strategy=Low-PPL Sampling2025.12 | 85.14 | |
| Qwen3-30B-A3BSparsity Level=Base, K=82026.05 | 83.56 | |
| BEAMSparsity Level=Mid Sparsity, beta=0.012026.05 | 81.46 | |
| Top-K ReducedSparsity Level=Mid Sparsity, K=42026.05 | 80.4 | |
| MoE-DynamicSparsity Level=Mid Sparsity, phi=0.32026.05 | 78.17 | |
| Top-K PruningSparsity Level=Mid Sparsity, K=42026.05 | 75.85 | |
| BEAMSparsity Level=High Sparsity, beta=0.12026.05 | 74.04 | |
| OPDLM-8BScale=8B, Tokens=0.066B, FLOPs=4.2, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 73.3 | |
| Top-K ReducedSparsity Level=High Sparsity, K=22026.05 | 72.63 | |
| Fast-dLLM-v2-7BScale=8B, Tokens=1B, FLOPs=42, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 70.3 | |
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 70.2 | |
| OPDLM-4BScale=4B, Tokens=0.076B, FLOPs=2.4, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 66.9 | |
| BEAMSparsity Level=Extreme Sparsity, beta=1.02026.05 | 66.28 | |
| MoE-DynamicSparsity Level=High Sparsity, phi=0.12026.05 | 64.38 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 62.9 | |
| AdaMoESparsity Level=Mid Sparsity, Null=1282026.05 | 62.71 | |
| Top-K ReducedSparsity Level=Extreme Sparsity, K=12026.05 | 52.97 | |
| AdaMoESparsity Level=High Sparsity, Null=2562026.05 | 38.75 | |
| Top-K PruningSparsity Level=High Sparsity, K=22026.05 | 17.51 |