Code Generation on HumanEval (Pass@1, HE+)
55.4Pass@1GRIP 16B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRIP 16BModel Size=16B, Total Budget=300B tokens2026.02 | 55.4 | 48.2 | |
| GRIP 8BModel Size=8B, Total Budget=300B tokens2026.02 | 52.4 | 46 | |
| 6. GRIP (Full)Model=GRIP-8B, Total Budget=100B tokens2026.02 | 47.2 | 43.8 | |
| 5. + DiversityModel=GRIP-8B, Total Budget=100B tokens2026.02 | 45.4 | 42 | |
| 4. + Loss ReplayModel=GRIP-8B, Total Budget=100B tokens2026.02 | 44.8 | 41.5 | |
| 3. + Static ReplayModel=GRIP-8B, Total Budget=100B tokens2026.02 | 41.2 | 38.5 | |
| Random 16BModel Size=16B, Total Budget=300B tokens2026.02 | 40.8 | 38.5 | |
| 2. + Static BudgetModel=GRIP-8B, Total Budget=100B tokens2026.02 | 38.6 | 35.4 | |
| Random 8BModel Size=8B, Total Budget=300B tokens2026.02 | 38.5 | 36.2 | |
| 1. RandomModel=GRIP-8B, Total Budget=100B tokens2026.02 | 35.2 | 32.1 |