Mathematical Reasoning on AIME (Accuracy, Peak Memory)
35.6AccuracyFullKV
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FullKVModel=Qwen-2.5-7B, rho=1.002026.05 | 35.6 | 22.5 | |
| ArborKVModel=Qwen-2.5-7B, rho=0.502026.05 | 33.8 | 11.6 | |
| ThinKVModel=Qwen-2.5-7B, rho=0.502026.05 | 33.2 | 11.6 | |
| StreamingLLMModel=Qwen-2.5-7B, rho=0.502026.05 | 31.9 | 11.4 | |
| H2OModel=Qwen-2.5-7B, rho=0.502026.05 | 31.5 | 10.3 | |
| FullKVModel=Llama-3.1-8B, rho=1.002026.05 | 28.3 | 23 | |
| ArborKVModel=Llama-3.1-8B, rho=0.502026.05 | 27.2 | 11.5 | |
| ThinKVModel=Llama-3.1-8B, rho=0.502026.05 | 26.5 | 11.2 | |
| StreamingLLMModel=Llama-3.1-8B, rho=0.502026.05 | 26.1 | 13.5 | |
| ArborKVModel=Qwen-2.5-7B, rho=0.252026.05 | 26.1 | 5.5 | |
| H2OModel=Llama-3.1-8B, rho=0.502026.05 | 25.6 | 10.2 | |
| ThinKVModel=Qwen-2.5-7B, rho=0.252026.05 | 25 | 5.7 | |
| ArborKVModel=Llama-3.1-8B, rho=0.252026.05 | 24.3 | 6.1 | |
| ThinKVModel=Llama-3.1-8B, rho=0.252026.05 | 23.1 | 6.2 | |
| StreamingLLMModel=Qwen-2.5-7B, rho=0.252026.05 | 22.4 | 5.3 | |
| StreamingLLMModel=Llama-3.1-8B, rho=0.252026.05 | 21.6 | 6.1 | |
| H2OModel=Llama-3.1-8B, rho=0.252026.05 | 20.5 | 4.9 | |
| H2OModel=Qwen-2.5-7B, rho=0.252026.05 | 19.6 | 5.3 |