Next-state prediction on StableToolBench (STB)
49.25EM AccuracyLlama3.1-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama3.1-8BEvaluation Protocol=SFT2025.12 | 49.25 | 78.97 | |
| Qwen2.5-7BEvaluation Protocol=SFT2025.12 | 48.9 | 79.15 | |
| GPT-4o-miniEvaluation Protocol=Zero-shot2025.12 | 0 | 13.94 | |
| GPT-4oEvaluation Protocol=Zero-shot2025.12 | 0 | 11.88 | |
| GPT-4-turboEvaluation Protocol=Zero-shot2025.12 | 0 | 12.64 | |
| GPT-4.1Evaluation Protocol=Zero-shot2025.12 | 0 | 12.83 | |
| GPT-5Evaluation Protocol=Zero-shot2025.12 | 0 | 8.02 | |
| Gemini-2.5-flashEvaluation Protocol=Zero-shot2025.12 | 0 | 8.74 | |
| Claude-sonnet-4.5Evaluation Protocol=Zero-shot2025.12 | 0 | 11.36 | |
| GPT-4o-miniEvaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 13.44 | |
| GPT-4oEvaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 11.08 | |
| GPT-4-turboEvaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 10.72 | |
| GPT-4.1Evaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 10.33 | |
| GPT-5Evaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 6.28 | |
| Gemini-2.5-flashEvaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 8.47 | |
| Claude-sonnet-4.5Evaluation Protocol=Few-shot (3 shot)2025.12 | 0 | 13.11 |