Knowledge-based reasoning on MMLU College Medicine 1.0 (test)
86.13AccuracyQwQ-32B (Full COT)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| QwQ-32B (Full COT)Model=QwQ-32B, Pruning Level=None (Full COT)2025.08 | 86.13 | 2,912.3 | |
| QwQ-32B (80%)Model=QwQ-32B, Pruning Level=80% step-entropy pruning2025.08 | 84.97 | 2,475.9 | |
| QwQ-32B (90%)Model=QwQ-32B, Pruning Level=90% step-entropy pruning2025.08 | 84.97 | 2,326.4 | |
| QwQ-32B (No Thinking)Model=QwQ-32B, Pruning Level=No Thinking (Complete elimination)2025.08 | 84.3 | — | |
| DeepSeek-R1-7B (80%)Pruning level=80%2025.08 | 62.34 | 2,127.9 | |
| DeepSeek-R1-7B (90%)Pruning level=90%2025.08 | 62.34 | 2,069.4 | |
| DeepSeek-R1-7BPruning level=Full COT2025.08 | 61.73 | 2,612.7 | |
| DeepSeek-R1-7BPruning level=No Thinking2025.08 | 52.46 | — |