General Knowledge and Language Understanding on MMLU (Accuracy)
92.39AccuracySafeSieve
Evaluation Results
| Method | Links | |
|---|---|---|
| SafeSieveParadigm=post-prune, Base model=deepseek-V3-671B2025.08 | 92.39 | |
| G-DesignerParadigm=pre-design, Base model=deepseek-V3-671B2025.08 | 91.13 | |
| AgentPruneParadigm=post-prune, Base model=deepseek-V3-671B2025.08 | 90.99 | |
| GPTSwarmParadigm=pre-design, Base model=deepseek-V3-671B2025.08 | 90.52 | |
| AgentDropoutParadigm=post-prune, Base model=deepseek-V3-671B2025.08 | 90.17 | |
| CoTParadigm=single, Base model=deepseek-V3-671B2025.08 | 89.31 | |
| VanillaParadigm=single, Base model=deepseek-V3-671B2025.08 | 87.97 | |
| G-DesignerParadigm=pre-design, Base model=gpt-4o-mini (~8B)2025.08 | 87.2 | |
| AgentPruneParadigm=post-prune, Base model=gpt-4o-mini (~8B)2025.08 | 83.3 | |
| GPTSwarmParadigm=pre-design, Base model=gpt-4o-mini (~8B)2025.08 | 82.8 | |
| SafeSieveParadigm=post-prune, Base model=gpt-4o-mini (~8B)2025.08 | 82.32 | |
| AgentDropoutParadigm=post-prune, Base model=gpt-4o-mini (~8B)2025.08 | 80.01 | |
| CoTParadigm=single, Base model=gpt-4o-mini (~8B)2025.08 | 78.43 | |
| VanillaParadigm=single, Base model=gpt-4o-mini (~8B)2025.08 | 77.81 | |
| RTBackbone=Llama3-8B-instruct, Relative Compute=1x2026.03 | 63.9 | |
| Llama3-8B-instructBackbone=Llama3-8B-instruct, Relative Compute=0x2026.03 | 63.8 | |
| RT-EAT-LATBackbone=Llama3-8B-instruct, Relative Compute=9x2026.03 | 61.3 | |
| Llama2-7B-chatBackbone=Llama2-7B-chat, Relative Compute=0x2026.03 | 46.4 | |
| RTBackbone=Llama2-7B-chat, Relative Compute=1x2026.03 | 45.6 | |
| RT-EAT-LATBackbone=Llama2-7B-chat, Relative Compute=9x2026.03 | 45.4 | |
| RT-EATBackbone=Llama2-7B-chat, Relative Compute=9x2026.03 | 44.8 | |
| R2D2Backbone=Llama2-7B-chat, Relative Compute=6558x2026.03 | 44.1 |