Large Language Model Evaluation on Knowledge Specialized Target (test)
56.5Weighted Average ScoreCAMEL
Evaluation Results
| Method | Links | |
|---|---|---|
| CAMELSampling Strategy=Hourglass2026.03 | 56.5 | |
| SODMSampling Strategy=Rectangle2026.03 | 55.9 | |
| Model-size agnosticSampling Strategy=Rectangle2026.03 | 55.1 | |
| DMLSampling Strategy=Rectangle2026.03 | 55 |