Next-state prediction on WebShop
79.05EM AccuracyQwen2.5-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7BEvaluation Protocol=SFT2025.12 | 79.05 | |
| Llama3.1-8BEvaluation Protocol=SFT2025.12 | 77.24 | |
| Gemini-2.5-flashEvaluation Protocol=Few-shot (3 shot)2025.12 | 66.09 | |
| GPT-5Evaluation Protocol=Few-shot (3 shot)2025.12 | 65.9 | |
| GPT-4oEvaluation Protocol=Few-shot (3 shot)2025.12 | 64.62 | |
| GPT-4.1Evaluation Protocol=Few-shot (3 shot)2025.12 | 64.23 | |
| GPT-4-turboEvaluation Protocol=Few-shot (3 shot)2025.12 | 62.76 | |
| GPT-4o-miniEvaluation Protocol=Few-shot (3 shot)2025.12 | 61.93 | |
| Claude-sonnet-4.5Evaluation Protocol=Zero-shot2025.12 | 58.8 | |
| GPT-4oEvaluation Protocol=Zero-shot2025.12 | 58.2 | |
| GPT-4.1Evaluation Protocol=Zero-shot2025.12 | 58.07 | |
| Gemini-2.5-flashEvaluation Protocol=Zero-shot2025.12 | 57.64 | |
| Claude-sonnet-4.5Evaluation Protocol=Few-shot (3 shot)2025.12 | 56.65 | |
| GPT-4o-miniEvaluation Protocol=Zero-shot2025.12 | 56.59 | |
| GPT-4-turboEvaluation Protocol=Zero-shot2025.12 | 52.45 | |
| GPT-5Evaluation Protocol=Zero-shot2025.12 | 46.12 |