Large Language Model Evaluation on Math Specialized Target (test)
49.7Weighted Average ScoreCAMEL
Evaluation Results
| Method | Links | |
|---|---|---|
| CAMELSampling Strategy=Hourglass2026.03 | 49.7 | |
| SODMSampling Strategy=Rectangle2026.03 | 49.4 | |
| DMLSampling Strategy=Rectangle2026.03 | 47.9 | |
| Model-size agnosticSampling Strategy=Rectangle2026.03 | 44.8 |