Code-Centric Agent Interaction on ColBench
55.3Pass RateInfoPO (Ours)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InfoPO (Ours)Type=RL Training, Backbone=Qwen3-4B2026.02 | 55.3 | 43.9 | |
| InfoPO (Ours)Type=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 53.4 | 42.6 | |
| GPT-4.1Type=Prompting, Backbone=Closed-Source Model2026.02 | 52.9 | 40.3 | |
| Gemini-3-FlashType=Prompting, Backbone=Closed-Source Model2026.02 | 51.5 | 38.2 | |
| InfoPO w/o stdType=RL Training, Backbone=Qwen3-4B2026.02 | 49.8 | 39.5 | |
| RAGENType=RL Training, Backbone=Qwen3-4B2026.02 | 47.9 | 36.1 | |
| InfoPO w/o GateType=RL Training, Backbone=Qwen3-4B2026.02 | 47.5 | 37.2 | |
| UserRLType=RL Training, Backbone=Qwen3-4B2026.02 | 46.8 | 34.2 | |
| Search-R1Type=RL Training, Backbone=Qwen3-4B2026.02 | 46.7 | 35.5 | |
| InfoPO w/o GateType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 46.6 | 36.8 | |
| GPT-4o-miniType=Prompting, Backbone=Closed-Source Model2026.02 | 46.3 | 34.2 | |
| Search-R1Type=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 45.7 | 35.2 | |
| RAGENType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 44.9 | 34.8 | |
| UserRLType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 43.6 | 32.7 | |
| InfoPO w/o stdType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 39.5 | 29.8 | |
| InfoPO w/o RextType=RL Training, Backbone=Qwen3-4B2026.02 | 34.5 | 28.5 | |
| ReflexionType=Prompting, Backbone=Qwen3-4B2026.02 | 30.3 | 18.4 | |
| InfoPO w/o RextType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 28.5 | 34.2 | |
| ReflexionType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 27.6 | 16.8 | |
| Qwen3Type=Prompting, Backbone=Qwen3-4B2026.02 | 27.2 | 15.3 | |
| ReActType=Prompting, Backbone=Qwen3-4B2026.02 | 26.9 | 14.5 | |
| Qwen2.5Type=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 24.2 | 14 | |
| ReActType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 23.8 | 13.5 |