In-Context Learning Aggregate Evaluation on ICL composite Standard Benchmarks
24.3Macro AccuracyBalanced
Evaluation Results
| Method | Links | |
|---|---|---|
| BalancedModel Scale=1B2025.09 | 24.3 | |
| InductionModel Scale=0.5B2025.09 | 23.9 | |
| Anti-inductionModel Scale=1B2025.09 | 23.6 | |
| BaselineModel Scale=0.13B2025.09 | 22.7 |