Mathematical Reasoning on Math Benchmarks (test)
28.9GSM8K AccuracyCodeLlama-Baseline
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| CodeLlama-BaselineModel=CodeLlama-7B, Training=Static domain mixture2024.11 | 28.9 | 11.1 | 9.8 | 24 | 61.1 | 56.4 | 36 | 46.9 | 34.3 | |
| CodeLlama-VelocituneModel=CodeLlama-7B, Training=Dynamic domain weights adjustment2024.11 | 28.4 | 11.7 | 11.4 | 25.1 | 60.9 | 56.1 | 37.3 | 56.2 | 35.9 | |
| CodeLlamaModel=CodeLlama-7B2024.11 | 12.4 | 6 | 5.2 | 14.1 | 50.5 | 44.5 | 20.9 | 18.8 | 21.6 |