Aggregated LLM Evaluation on Balanced Objective Aggregate Suite
53.2Weighted Average ScoreCAMEL
Evaluation Results
| Method | Links | |
|---|---|---|
| CAMELTraining Objective=Balanced, Sampling Strategy=Hourglass2026.03 | 53.2 | |
| SODMTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 52.6 | |
| DMLTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 51.9 | |
| Model-size agnosticTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 51.4 | |
| Human DesignedTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 49.6 |