Multimodal tool-use reasoning on TRACE-BENCH (test)
78.23AccuracyTRACER-RL
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| TRACER-RL2026.05 | 78.23 | 3,486 | 573 | 95.72 | 93.61 | 90.52 | |
| TRACER-SFT2026.05 | 73.56 | 4,034 | 654 | 93.29 | 77.53 | 70.7 | |
| Qwen3-VL-8B-Instructtool_use=true, tuning=SFT2026.05 | 70.94 | 4,949 | 590 | 84.39 | — | — | |
| Gemini-2.5-Protool_use=true2026.05 | 54.43 | 1,569 | 438 | 70.27 | — | — | |
| ChatGPT-4o-latesttool_use=false2026.05 | 38.29 | — | — | — | — | — | |
| MiMo-VL-7B-RLtool_use=true2026.05 | 37.09 | 8,491 | 1,513 | 57.31 | — | — | |
| ChatGPT-4o-latesttool_use=true2026.05 | 34.96 | 2,003 | 1,804 | 56.1 | — | — | |
| Claude-3-5-sonnettool_use=true2026.05 | 30.52 | 585 | 524 | 75.52 | — | — | |
| Claude-3-5-sonnettool_use=false2026.05 | 30.33 | — | — | — | — | — | |
| Gemini-2.5-Protool_use=false2026.05 | 27.35 | — | — | — | — | — | |
| Qwen3-VL-8B-Instructtool_use=true2026.05 | 26.62 | 5,580 | 3,730 | 72.07 | — | — | |
| Qwen3-VL-8B-Instructtool_use=false2026.05 | 23.52 | — | — | — | — | — | |
| MiMo-VL-7B-RLtool_use=false2026.05 | 22.66 | — | — | — | — | — | |
| Tuned LLaVA-7Btool_use=true2026.05 | 18.8 | 4,311 | 617 | 30.91 | — | — | |
| InternVL3.5-7Btool_use=true2026.05 | 13.68 | 249 | 10 | 48.39 | — | — | |
| InternVL3.5-7Btool_use=false2026.05 | 12.27 | — | — | — | — | — | |
| LLaVA-v1.5-7B VLMtool_use=false2026.05 | 8.57 | — | — | — | — | — | |
| Qwen2-VL-2B-Instructtool_use=false2026.05 | 7.71 | — | — | — | — | — | |
| Tuned LLaVA-7Btool_use=false2026.05 | 7.21 | — | — | — | — | — | |
| Qwen2-VL-7B-Instructtool_use=true2026.05 | 5.86 | 1,373 | 1,051 | 38.18 | — | — | |
| Qwen2-VL-7B-Instructtool_use=false2026.05 | 3.64 | — | — | — | — | — | |
| Qwen2-VL-2B-Instructtool_use=true2026.05 | 2.1 | 4,067 | 3,915 | 21.82 | — | — | |
| LLaVA-v1.5-7B VLMtool_use=true2026.05 | 1.17 | 9,684 | 9,684 | 0.01 | — | — |