General Language Modeling on BIG-Bench (test)
83.6AccuracyBest Model
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Best ModelBackbone LLM=LLaMA 3.1 8B, Shot count=0-shot2025.10 | 83.6 | 5 | -14.4 | |
| Best ModelBackbone LLM=Mistral 7B, Shot count=0-shot2025.10 | 76.4 | 9 | -28 |
| Method | Links | |||
|---|---|---|---|---|
| Best ModelBackbone LLM=LLaMA 3.1 8B, Shot count=0-shot2025.10 | 83.6 | 5 | -14.4 | |
| Best ModelBackbone LLM=Mistral 7B, Shot count=0-shot2025.10 | 76.4 | 9 | -28 |