General Language Understanding on BIG-Bench
83.6PerformanceBest Model
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Best ModelBackbone=LLaMA 3.1 8B, Shots=5-shot, #D=52025.10 | 83.6 | 1.2 | |
| BSBABackbone=LLaMA 3.1 8B, Shots=5-shot, #D=152025.10 | 81.2 | 1.83 | |
| BaselineBackbone=LLaMA 3.1 8B, Shots=5-shot2025.10 | 70 | — |