Code Generation on HumanEval (Acc, AMSR)
98.2Accuracy (HumanEval)Dense
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DenseModel=DeepSeek R1 Distill Qwen 32B2025.07 | 98.2 | 0 | |
| ReasonCacheModel=QwQ 32B2025.07 | 97.2 | 22.1 | |
| ReasonCacheModel=DeepSeek R1 Distill Qwen 32B2025.07 | 96.9 | 21.5 | |
| DenseModel=QwQ 32B2025.07 | 96.8 | 0 | |
| ReasonCacheModel=Phi 4 reasoning plus2025.07 | 94.5 | 16.8 | |
| DenseModel=Phi 4 reasoning plus2025.07 | 93.3 | 0 | |
| QuestModel=QwQ 32B2025.07 | 91.5 | 20.8 | |
| StreamingLLMModel=DeepSeek R1 Distill Qwen 32B2025.07 | 88.4 | 20.1 | |
| SnapKVModel=DeepSeek R1 Distill Qwen 32B2025.07 | 87.2 | 9.6 | |
| QuestModel=DeepSeek R1 Distill Qwen 32B2025.07 | 86.6 | 20.4 | |
| QuestModel=Phi 4 reasoning plus2025.07 | 85.4 | 15.3 | |
| StreamingLLMModel=Phi 4 reasoning plus2025.07 | 81.2 | 14.8 | |
| StreamingLLMModel=QwQ 32B2025.07 | 79.9 | 21.7 | |
| SnapKVModel=Phi 4 reasoning plus2025.07 | 79.3 | 11.5 | |
| SnapKVModel=QwQ 32B2025.07 | 75.6 | 6.2 |