User-Centric Agent Interaction on UserGym
59.6Travel ScoreGPT-4o-mini
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| GPT-4o-miniType=Prompting, Backbone=Closed-Source Model2026.02 | 59.6 | 8.9 | 68.3 | 9.1 | 11.7 | 56.8 | 1.277 | 51.2 | |
| InfoPO (Ours)Type=RL Training, Backbone=Qwen3-4B2026.02 | 58.9 | 11.5 | 55.6 | 9.7 | 15.4 | 84.9 | 1.862 | 54.2 | |
| InfoPO (Ours)Type=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 58.8 | 16.7 | 53.5 | 9.1 | 17.8 | 48 | 1.892 | 48.8 | |
| Gemini-3-FlashType=Prompting, Backbone=Closed-Source Model2026.02 | 57.4 | 42.3 | 69.5 | 16.7 | 15.3 | 96.8 | 1.718 | 82.9 | |
| Search-R1Type=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 56.5 | 11.3 | 41.2 | 4.3 | 15.4 | 43.5 | 1.805 | 41.6 | |
| InfoPO w/o stdType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 56.5 | 14.2 | 49.8 | 3.5 | 6.5 | 45.5 | 1.845 | 45.2 | |
| GPT-4.1Type=Prompting, Backbone=Closed-Source Model2026.02 | 55.4 | 5.1 | 59.9 | 10.9 | 26.7 | 48 | 1.867 | 73.2 | |
| InfoPO w/o GateType=RL Training, Backbone=Qwen3-4B2026.02 | 54.8 | 8.5 | 48.2 | 6.2 | 11.5 | 75.8 | 1.765 | 46 | |
| UserRLType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 54.6 | 11.5 | 44.4 | 4.8 | 15.2 | 42.9 | 1.826 | 42.4 | |
| InfoPO w/o GateType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 54.2 | 12.5 | 46.5 | 8.2 | 4.2 | 43.2 | 1.81 | 46.5 | |
| RAGENType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 53.8 | 12.4 | 54.8 | 0 | 14.8 | 44.6 | 1.815 | 41.6 | |
| UserRLType=RL Training, Backbone=Qwen3-4B2026.02 | 53.8 | 9.5 | 50.7 | 5.3 | 12.1 | 69.7 | 1.732 | 52 | |
| InfoPO w/o stdType=RL Training, Backbone=Qwen3-4B2026.02 | 51.2 | 10.2 | 51.8 | 7.5 | 11.8 | 79.5 | 1.812 | 48.5 | |
| InfoPO w/o RextType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 48.5 | 5.5 | 35.2 | 3.5 | 1.5 | 38.5 | 1.45 | 32.5 | |
| Search-R1Type=RL Training, Backbone=Qwen3-4B2026.02 | 48.2 | 10.3 | 42.5 | 5.9 | 11.3 | 70.2 | 1.711 | 41.2 | |
| RAGENType=RL Training, Backbone=Qwen3-4B2026.02 | 47.7 | 10.4 | 51.1 | 0.6 | 11.7 | 71.4 | 1.572 | 48.8 | |
| ReActType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 45.2 | 6.4 | 32.5 | 3.7 | 7.8 | 35.4 | 1.378 | 31.6 | |
| ReflexionType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 44.5 | 5.2 | 31.2 | 2.2 | 7.4 | 36.4 | 1.32 | 30.5 | |
| Qwen2.5Type=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 44.1 | 2.6 | 28.9 | 0 | 6.2 | 37.6 | 1.254 | 29.2 | |
| InfoPO w/o RextType=RL Training, Backbone=Qwen3-4B2026.02 | 35.2 | 4.5 | 38.5 | 3.2 | 8.8 | 42.5 | 1.512 | 38.2 | |
| ReflexionType=Prompting, Backbone=Qwen3-4B2026.02 | 29.1 | 4.5 | 47.4 | 3.5 | 12.4 | 43.8 | 1.717 | 50.1 | |
| ReActType=Prompting, Backbone=Qwen3-4B2026.02 | 27.9 | 5.3 | 47.6 | 0.5 | 14.7 | 43.5 | 1.782 | 45.4 | |
| Qwen3Type=Prompting, Backbone=Qwen3-4B2026.02 | 27.7 | 2.6 | 45.2 | 0.6 | 7.1 | 44.4 | 1.66 | 48.8 |