Graduate-level STEM Reasoning on GPQA Diamond (Pass@1, #Tokens)
63.54Pass@1 AccuracyQwen3-14B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-14BMode=Original, Base Model=Qwen3-14B2025.10 | 63.54 | 7,161 | |
| Qwen3-14B MixReasoningMode=MixReasoning, Base Model=Qwen3-14B2025.10 | 62.87 | 6,225 | |
| Qwen3-14B Concise ModeMode=Concise Mode, alpha=32, Base Model=Qwen3-14B2025.10 | 61.43 | 6,502 | |
| Qwen3-8BMode=Original, Base Model=Qwen3-8B2025.10 | 57.32 | 8,624 | |
| Qwen3-8B MixReasoningMode=MixReasoning, Base Model=Qwen3-8B2025.10 | 57.17 | 6,734 | |
| QwQ-32B MixReasoningMode=MixReasoning, Base Model=QwQ-32B2025.10 | 56.75 | 5,783 | |
| QwQ-32BMode=Original, Base Model=QwQ-32B2025.10 | 56.47 | 7,923 | |
| QwQ-32B Concise ModeMode=Concise Mode, alpha=32, Base Model=QwQ-32B2025.10 | 54.65 | 6,078 | |
| Qwen3-14B Concise ModeMode=Concise Mode, alpha=64, Base Model=Qwen3-14B2025.10 | 51.01 | 2,156 | |
| Qwen3-8B Concise ModeMode=Concise Mode, alpha=32, Base Model=Qwen3-8B2025.10 | 50.38 | 5,918 | |
| QwQ-32B Concise ModeMode=Concise Mode, alpha=64, Base Model=QwQ-32B2025.10 | 48.89 | 3,248 | |
| Qwen3-8B Concise ModeMode=Concise Mode, alpha=64, Base Model=Qwen3-8B2025.10 | 45.56 | 5,173 | |
| SuperThoughts (Transformer Compressor)Backbone=Qwen-2.5-Instruct-14B, Adaptive Threshold (τ)=0.99992026.06 | 43 | 575.3 | |
| SuperThoughts (Projection Compressor)Backbone=Qwen-2.5-Instruct-14B, Adaptive Threshold (τ)=0.99992026.06 | 42.9 | 554 | |
| Standard CoTBackbone=Qwen-2.5-Instruct-14B2026.06 | 41.5 | 675.3 | |
| SuperThoughts (Projection Compressor)Backbone=Qwen-2.5-Instruct-14B, Adaptive Threshold (τ)=0.9992026.06 | 41.5 | 510.2 | |
| SuperThoughts (Transformer Compressor)Backbone=Qwen-2.5-Instruct-14B, Adaptive Threshold (τ)=0.999992026.06 | 41.2 | 607.7 | |
| SuperThoughts (Projection Compressor)Backbone=Qwen-2.5-Instruct-14B, Adaptive Threshold (τ)=0.999992026.06 | 40 | 615.2 | |
| SuperThoughts (Transformer Compressor)Backbone=Qwen-2.5-Instruct-14B, Adaptive Threshold (τ)=0.9992026.06 | 39.1 | 502.1 | |
| DeepSeek-R1-7B MixReasoningMode=MixReasoning, Base Model=DeepSeek-R1-7B2025.10 | 35.66 | 4,991 | |
| DeepSeek-R1-7BMode=Original, Base Model=DeepSeek-R1-7B2025.10 | 35.05 | 6,030 | |
| DeepSeek-R1-7B Concise ModeMode=Concise Mode, alpha=64, Base Model=DeepSeek-R1-7B2025.10 | 30.32 | 4,897 | |
| DeepSeek-R1-7B Concise ModeMode=Concise Mode, alpha=32, Base Model=DeepSeek-R1-7B2025.10 | 25.76 | 4,536 |