Natural Language Understanding on BIG-Bench Hard (BBH)
42.1AccuracyArcana
Evaluation Results
| Method | Links | |
|---|---|---|
| Arcanazero-shot=true2024.10 | 42.1 | |
| Vicuna-v1.5zero-shot=true2024.10 | 41.2 | |
| LLaMA-2zero-shot=true2024.10 | 38.2 | |
| LLaMA-2-Chatzero-shot=true2024.10 | 35.6 | |
| WizardLMzero-shot=true2024.10 | 34.7 |