Large Language Model Evaluation on Code Specialized Target (test)
52.8Weighted Average ScoreCAMEL
Evaluation Results
| Method | Links | |
|---|---|---|
| CAMELSampling Strategy=Hourglass2026.03 | 52.8 | |
| SODMSampling Strategy=Rectangle2026.03 | 52 | |
| DMLSampling Strategy=Rectangle2026.03 | 50 | |
| Model-size agnosticSampling Strategy=Rectangle2026.03 | 49.7 |