Next-state prediction on TextWorld (TW)
70.6EM AccuracyQwen2.5-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-7BEvaluation Protocol=SFT2025.12 | 70.6 | |
| Llama3.1-8BEvaluation Protocol=SFT2025.12 | 70.45 | |
| Claude-sonnet-4.5Evaluation Protocol=Few-shot (3 shot)2025.12 | 49.12 | |
| GPT-5Evaluation Protocol=Few-shot (3 shot)2025.12 | 44.27 | |
| Gemini-2.5-flashEvaluation Protocol=Few-shot (3 shot)2025.12 | 40.35 | |
| Claude-sonnet-4.5Evaluation Protocol=Zero-shot2025.12 | 17.7 | |
| GPT-4oEvaluation Protocol=Few-shot (3 shot)2025.12 | 14.11 | |
| GPT-4.1Evaluation Protocol=Few-shot (3 shot)2025.12 | 13.39 | |
| GPT-4-turboEvaluation Protocol=Few-shot (3 shot)2025.12 | 11.66 | |
| GPT-4o-miniEvaluation Protocol=Few-shot (3 shot)2025.12 | 11.43 | |
| GPT-5Evaluation Protocol=Zero-shot2025.12 | 9.2 | |
| GPT-4oEvaluation Protocol=Zero-shot2025.12 | 7.86 | |
| Gemini-2.5-flashEvaluation Protocol=Zero-shot2025.12 | 3.51 | |
| GPT-4o-miniEvaluation Protocol=Zero-shot2025.12 | 0.36 | |
| GPT-4-turboEvaluation Protocol=Zero-shot2025.12 | 0 | |
| GPT-4.1Evaluation Protocol=Zero-shot2025.12 | 0 |