LLM Agent Reasoning on BrowseComp-Plus
42.7AccuracyRule-Based (High)
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Rule-Based (High)Strategy=Rule-Based, Effort Level=High, Backbone LLM=gpt-oss-20b2026.03 | 42.7 | 1.4 | 55.4 | 1,841 | 770 | 12,276 | 5,495 | 220 | 35 | |
| ARESStrategy=Adaptive Selection, Backbone LLM=gpt-oss-20b2026.03 | 41.3 | — | 36.5 | 1,071 | — | 6,781 | — | 185 | — | |
| Prompting-Based (GPT 5)Strategy=Prompting-Based, Model=GPT 5, Backbone LLM=gpt-oss-20b2026.03 | 38.7 | 2.6 | 45.1 | 1,398 | 327 | 9,321 | 2,540 | 206 | 21 | |
| Prompting-Based (Gemini 3 Pro)Strategy=Prompting-Based, Model=Gemini 3 Pro, Backbone LLM=gpt-oss-20b2026.03 | 37.3 | 4 | 41.6 | 1,144 | 73 | 7,628 | 847 | 183 | 2 | |
| Rule-Based (Medium)Strategy=Rule-Based, Effort Level=Medium, Backbone LLM=gpt-oss-20b2026.03 | 34 | 7.3 | 26.2 | 538 | 533 | 3,590 | 3,191 | 137 | 48 | |
| Rule-Based (Random)Strategy=Rule-Based, Effort Level=Random, Backbone LLM=gpt-oss-20b2026.03 | 30.7 | 10.6 | 18 | 392 | 679 | 2,616 | 4,165 | 145 | 40 | |
| Rule-Based (Low)Strategy=Rule-Based, Effort Level=Low, Backbone LLM=gpt-oss-20b2026.03 | 8 | 33.3 | 4.6 | 5 | 1,066 | 332 | 6,449 | 72 | 113 |