LLM Agent Reasoning on WebArena
46.5AccuracyARES
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ARESStrategy=Adaptive Selection, Backbone LLM=gpt-oss-20b2026.03 | 46.5 | — | 8.9 | 1,512,000 | — | 11,723 | — | 1,324 | — | |
| Rule-Based (High)Strategy=Rule-Based, Effort Level=High, Backbone LLM=gpt-oss-20b2026.03 | 45 | 1.5 | 10 | 2,763,000 | 1,251,000 | 21,424 | 9,701 | 2,154 | 830 | |
| Rule-Based (Medium)Strategy=Rule-Based, Effort Level=Medium, Backbone LLM=gpt-oss-20b2026.03 | 42.6 | 3.9 | 9.9 | 538,000 | 974,000 | 4,170 | 7,553 | 420 | 904 | |
| Prompting-Based (Gemini 3 Pro)Strategy=Prompting-Based, Model=Gemini 3 Pro, Backbone LLM=gpt-oss-20b2026.03 | 41.9 | 4.6 | 9.1 | 1,164,000 | 348,000 | 9,023 | 2,670 | 995 | 329 | |
| Prompting-Based (GPT 5)Strategy=Prompting-Based, Model=GPT 5, Backbone LLM=gpt-oss-20b2026.03 | 41.1 | 5.4 | 9 | 1,159,000 | 353,000 | 8,990 | 2,733 | 1,001 | 323 | |
| Rule-Based (Random)Strategy=Rule-Based, Effort Level=Random, Backbone LLM=gpt-oss-20b2026.03 | 40.3 | 6.2 | 8 | 857,000 | 655,000 | 6,643 | 5,080 | 830 | 494 | |
| Rule-Based (Low)Strategy=Rule-Based, Effort Level=Low, Backbone LLM=gpt-oss-20b2026.03 | 37.4 | 9.1 | 8.9 | 67,000 | 1,445,000 | 520 | 11,203 | 58 | 1,266 |