Web Search on BrowseComp-Plus
53.02Pass@3AdaCoM
Evaluation Results
| Method | Links | |
|---|---|---|
| AdaCoMBackbone=Kimi-K2-Instruct, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 53.02 | |
| AdaCoMBackbone=Qwen3-Max, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 52 | |
| MemActBackbone=Qwen3-Max, CM Model=agent itself, CM Trained=×2026.05 | 50.67 | |
| AdaCoMBackbone=GLM-4.5-Air, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 50.67 | |
| AdaCoMBackbone=Average, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 48.92 | |
| MemActBackbone=GLM-4.5-Air, CM Model=agent itself, CM Trained=×2026.05 | 46 | |
| SumCoMBackbone=Kimi-K2-Instruct, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 45.33 | |
| SumAgentBackbone=Kimi-K2-Instruct, CM Model=agent itself, CM Trained=×2026.05 | 44.67 | |
| ReActBackbone=GLM-4.5-Air, CM Model=-, CM Trained=-2026.05 | 44.67 | |
| AdaCoM w/o train.Backbone=Kimi-K2-Instruct, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 42.75 | |
| SumCoMBackbone=Qwen3-Max, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 42.67 | |
| SumCoMBackbone=Average, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 40.33 | |
| AdaCoMBackbone=DeepSeek-V3, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 39.97 | |
| SumAgentBackbone=Qwen3-Max, CM Model=agent itself, CM Trained=×2026.05 | 38.67 | |
| ReActBackbone=Qwen3-Max, CM Model=-, CM Trained=-2026.05 | 38 | |
| SumCoMBackbone=DeepSeek-V3, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 37.33 | |
| MemActBackbone=Average, CM Model=agent itself, CM Trained=×2026.05 | 37.33 | |
| MemActAgent=Qwen3-Max, CM Model=agent itself, CM Trained=×2026.05 | 37.33 | |
| AdaCoMAgent=Qwen3-Max, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 36.67 | |
| AdaCoMAgent=Kimi-K2-Instruct, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 36.2 | |
| SumCoMBackbone=GLM-4.5-Air, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 36 | |
| AdaCoMAgent=GLM-4.5-Air, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 35.33 | |
| ReActBackbone=Average, CM Model=-, CM Trained=-2026.05 | 34.67 | |
| SumCoMAgent=Kimi-K2-Instruct, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 33.78 | |
| AdaCoMAgent=Avg., CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 33.6 | |
| AdaCoM w/o train.Backbone=Average, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 33.45 | |
| SumAgentBackbone=Average, CM Model=agent itself, CM Trained=×2026.05 | 33 | |
| AdaCoM w/o train.Backbone=GLM-4.5-Air, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 32.67 | |
| ReActAgent=GLM-4.5-Air, CM Model=–, CM Trained=–2026.05 | 32.56 | |
| SumCoMAgent=Qwen3-Max, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 32.22 | |
| MemActAgent=GLM-4.5-Air, CM Model=agent itself, CM Trained=×2026.05 | 32 | |
| AdaCoM w/o train.Backbone=Qwen3-Max, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 30.67 | |
| SumAgentAgent=Kimi-K2-Instruct, CM Model=agent itself, CM Trained=×2026.05 | 30.44 | |
| MemActBackbone=Kimi-K2-Instruct, CM Model=agent itself, CM Trained=×2026.05 | 30 | |
| SumCoMAgent=Avg., CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 29.39 | |
| ReActBackbone=Kimi-K2-Instruct, CM Model=-, CM Trained=-2026.05 | 29.33 | |
| SumAgentBackbone=DeepSeek-V3, CM Model=agent itself, CM Trained=×2026.05 | 28.67 | |
| AdaCoM w/o train.Agent=Kimi-K2-Instruct, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 28.02 | |
| ReActAgent=Qwen3-Max, CM Model=–, CM Trained=–2026.05 | 27.78 | |
| AdaCoM w/o train.Backbone=DeepSeek-V3, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 27.7 | |
| ReActBackbone=DeepSeek-V3, CM Model=-, CM Trained=-2026.05 | 26.67 | |
| SumAgentAgent=Qwen3-Max, CM Model=agent itself, CM Trained=×2026.05 | 26.67 | |
| SumCoMAgent=GLM-4.5-Air, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 26.44 | |
| AdaCoMAgent=DeepSeek-V3, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 26.19 | |
| MemActAgent=Avg., CM Model=agent itself, CM Trained=×2026.05 | 25.72 | |
| SumCoMAgent=DeepSeek-V3, CM Model=Qwen3-4B-Inst., CM Trained=✓2026.05 | 25.11 | |
| ReActAgent=Avg., CM Model=–, CM Trained=–2026.05 | 24.17 | |
| MemActBackbone=DeepSeek-V3, CM Model=agent itself, CM Trained=×2026.05 | 22.67 | |
| SumAgentAgent=Avg., CM Model=agent itself, CM Trained=×2026.05 | 22.06 | |
| AdaCoM w/o train.Agent=Avg., CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 22.01 | |
| AdaCoM w/o train.Agent=GLM-4.5-Air, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 21.33 | |
| AdaCoM w/o train.Agent=Qwen3-Max, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 21.11 | |
| SumAgentBackbone=GLM-4.5-Air, CM Model=agent itself, CM Trained=×2026.05 | 20 | |
| SumAgentAgent=DeepSeek-V3, CM Model=agent itself, CM Trained=×2026.05 | 19.56 | |
| ReActAgent=Kimi-K2-Instruct, CM Model=–, CM Trained=–2026.05 | 18.56 | |
| ReActAgent=DeepSeek-V3, CM Model=–, CM Trained=–2026.05 | 17.78 | |
| AdaCoM w/o train.Agent=DeepSeek-V3, CM Model=Qwen3-4B-Inst., CM Trained=×2026.05 | 17.57 | |
| MemActAgent=Kimi-K2-Instruct, CM Model=agent itself, CM Trained=×2026.05 | 16.89 | |
| MemActAgent=DeepSeek-V3, CM Model=agent itself, CM Trained=×2026.05 | 16.67 | |
| SumAgentAgent=GLM-4.5-Air, CM Model=agent itself, CM Trained=×2026.05 | 11.56 |