Tool Use on API-Bench (test)
60AccuracyClaude 3.5 Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude 3.5 Sonnetevaluation_mode=zero-shot2024.07 | 60 | |
| GPT-4oevaluation_mode=zero-shot2024.07 | 41.4 | |
| GPT-3.5 Turboevaluation_mode=zero-shot2024.07 | 36.3 | |
| Llama 3 405Bevaluation_mode=zero-shot2024.07 | 35.3 | |
| Llama 3 70Bevaluation_mode=zero-shot2024.07 | 29.7 | |
| Mixtral 8x22Bevaluation_mode=zero-shot2024.07 | 26 | |
| GPT-4evaluation_mode=zero-shot2024.07 | 22.5 | |
| Gemma 2 9Bevaluation_mode=zero-shot2024.07 | 11.6 | |
| Llama 3 8Bevaluation_mode=zero-shot2024.07 | 8.2 | |
| Mistral 7Bevaluation_mode=zero-shot2024.07 | 4.7 |