Mathematical Reasoning on MATH-OAI
90.7AccuracyLitTrans
Evaluation Results
| Method | Links | |
|---|---|---|
| LitTransBase model=QwQ-32B, Decoding=greedy, Maximum tokens=32K, Training method=LightTransfer-Train, Layer Replacement=50% layers replaced with streaming attention, Optimization=Flex Attention2024.10 | 90.7 | |
| LightTransfer-TrainBackbone=QwQ-32B, Mode=Train2024.10 | 90.7 | |
| QwQ-STILLBase model=QwQ-32B, Decoding=greedy, Maximum tokens=32K2024.10 | 90.2 | |
| QwQ-STILLBackbone=QwQ-32B2024.10 | 90.2 | |
| LightTransfer-TestBackbone=QwQ-32B, Mode=Test2024.10 | 85 | |
| SqueezeAttention-TestBackbone=QwQ-32B, Mode=Test2024.10 | 82.6 | |
| DuoAttn.-TrainBackbone=QwQ-32B, Mode=Train2024.10 | 79.2 | |
| LongGenBase model=QwQ-32B, Decoding=greedy, Maximum tokens=32K2024.10 | 78.2 | |
| MiniCache-TestBackbone=QwQ-32B, Mode=Test2024.10 | 20 |