Reasoning Segmentation on ReasonSeg (test)
73.8gIoUTikArt
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TikArtParams=8B2026.02 | 73.8 | — | — | — | — | — | — | 73.2 | — | — | — | — | |
| VisHarness2026.05 | 73.2 | — | — | — | — | — | — | 67 | — | — | — | — | |
| EVOL-SAM3Version=Qwen2.5-VL 7B, Training-RES=false, Training-ReasonSeg=false2025.12 | 72.5 | — | — | — | — | — | — | 67.4 | — | — | — | — | |
| StARBackbone=Qwen 3, Model Components=+ MV, Zero-shot=true2026.05 | 71.8 | — | — | — | — | — | — | 66.5 | — | — | — | — | |
| VisHarness-SFTMode=SFT2026.05 | 71 | — | — | — | — | — | — | 66.6 | — | — | — | — | |
| SAM 3 AgentVersion=Qwen2.5-VL 72B, Training-RES=false, Training-ReasonSeg=false2025.12 | 70.8 | — | — | — | — | — | — | 64 | — | — | — | — | |
| StARBackbone=Qwen 3, Zero-shot=true2026.05 | 70.6 | — | — | — | — | — | — | 63.8 | — | — | — | — | |
| TikArt w/o ObservationParams=8B2026.02 | 70 | — | — | — | — | — | — | 68.7 | — | — | — | — | |
| TikArt w/o ApertureParams=8B2026.02 | 69.6 | — | — | — | — | — | — | 68.1 | — | — | — | — | |
| B-GRTOZero-shot=true2026.05 | 68.9 | — | — | — | — | — | — | 60.2 | — | — | — | — | |
| StARBackbone=Qwen 2.5, Model Components=+ MV, Zero-shot=true2026.05 | 68.5 | — | — | — | — | — | — | 65 | — | — | — | — | |
| GRPOZero-shot=true2026.05 | 68.5 | — | — | — | — | — | — | 62.7 | — | — | — | — | |
| GenSeg-R1-4BParameters=4B2026.02 | 68.4 | — | — | — | — | — | — | 62.73 | — | — | — | — | |
| StARBackbone=Qwen 2.5, Zero-shot=true2026.05 | 67.5 | — | — | — | — | — | — | 61.3 | — | — | — | — | |
| Seg-ReSearch-8BNumber of parameters=8B2026.02 | 67.4 | — | — | — | — | — | — | 59 | — | — | — | — | |
| GRPOTraining Stage=stage-1, Zero-shot=true2026.05 | 67.4 | — | — | — | — | — | — | 59.9 | — | — | — | — | |
| DGSegMLLM=Qwen2.5-VL-7B, setting=zero-shot2026.07 | 67.3 | — | — | — | — | — | — | 63.6 | — | — | — | — | |
| TikArt w/o Zoom ActionParams=8B2026.02 | 67.2 | — | — | — | — | — | — | 65.8 | — | — | — | — | |
| SAM 3 AgentVersion=Llama4 Maverick, Training-RES=false, Training-ReasonSeg=false2025.12 | 67.1 | — | — | — | — | — | — | 60.9 | — | — | — | — | |
| GRTOZero-shot=true2026.05 | 66.8 | — | — | — | — | — | — | 55.4 | — | — | — | — | |
| Ours (multi-object)consistent-checkpoint=true2026.05 | 66.73 | — | — | — | — | — | — | — | — | — | — | — | |
| CoT-SegZero-shot=true, Self-correction=true2026.01 | 66.7 | — | — | — | — | — | — | 60.4 | — | — | — | — | |
| IC-Seg-8BParameters=8B2026.05 | 66.7 | — | — | — | — | — | — | 60.2 | — | — | — | — | |
| StARBackbone=Qwen 2.5, Training Stage=stage-1, Zero-shot=true2026.05 | 66.7 | — | — | — | — | — | — | 60.8 | — | — | — | — | |
| Rea2SegModel=Qwen2.5-VL-3B, Selection=Top 32026.06 | 66.6 | — | — | — | — | — | — | 65.5 | — | — | — | — | |
| DR2SegLanguage Model=Qwen2.5VL-7B, Segmentation Model=SAM32026.01 | 66.5 | — | — | — | — | — | — | 61.7 | 30.7 | — | — | — | |
| RSAgentModel Version=Qwen2.5-VL, Params=7B, RES Training=true, ReasonSeg Training=false2025.12 | 66.5 | — | — | — | — | — | — | 57.9 | — | — | — | — | |
| DR2SegLanguage Model=Qwen2.5VL-7B, Trained on ReasonSeg train split=true2026.01 | 66.1 | — | — | — | — | — | — | 63.6 | 27.2 | — | — | — | |
| Qwen3-VL-8B*Number of parameters=8B2026.02 | 66 | — | — | — | — | — | — | 53.7 | — | — | — | — | |
| CoT-SegZero-shot=true2026.01 | 66 | — | — | — | — | — | — | 58.8 | — | — | — | — | |
| GenSeg-R1-GVariant=G2026.02 | 65.93 | — | — | — | — | — | — | 57.83 | — | — | — | — | |
| EVOL-SAM3Version=Qwen2.5-VL 3B, Training-RES=false, Training-ReasonSeg=false2025.12 | 65.9 | — | — | — | — | — | — | 58.9 | — | — | — | — | |
| SELF1E-8B2026.03 | 65.7 | — | — | — | — | — | — | 67 | — | — | — | — | |
| Dr. SegModel Category=RL-based 7B VLLMs2026.02 | 65.6 | — | — | — | — | — | — | — | — | 71 | 74 | — | |
| Dr. SegMethod Category=Explicit-Prompt-Based Reasoning Segmentation (ERS), Experimental Condition=Independent experiments under the same experimental conditions2026.06 | 65.6 | — | — | — | — | — | — | 58 | — | — | — | — | |
| VisionReasonerLanguage Model=Qwen2.5VL-7B, Segmentation Model=SAM32026.01 | 65.5 | — | — | — | — | — | — | 59.2 | 64.7 | — | — | — | |
| Rea2SegModel=LLaVA-v1.5-7B, Selection=Top 32026.06 | 65.5 | — | — | — | — | — | — | 64.8 | — | — | — | — | |
| Qwen3-VL-8B*Parameters=8B2026.05 | 65.4 | — | — | — | — | — | — | 54.7 | — | — | — | — | |
| ViSurf (Qwen2.5VL-7B + SAM2)Backbone=Qwen2.5VL-7B, Segmentation Model=SAM22025.10 | 65 | — | — | — | — | — | — | — | — | — | — | — | |
| DR2SegLanguage Model=Qwen2.5VL-7B, Trained on ReasonSeg train split=false2026.01 | 64.8 | — | — | — | — | — | — | 62.8 | 55.4 | — | — | — | |
| CR-SegMethod Category=Internal-Representation-Based Reasoning Segmentation (IRS)2026.06 | 64.8 | — | — | — | — | — | — | 62.6 | — | — | — | — | |
| Dr. SegMethod Category=Explicit-Prompt-Based Reasoning Segmentation (ERS), Experimental Condition=Reproduced using LoRA fine-tuning on Qwen3-VL-4B and SAM32026.06 | 64.5 | — | — | — | — | — | — | 58 | — | — | — | — | |
| SegCompass-13BBackbone=LLaVA-1.5, Zero-shot=true2026.05 | 64.2 | — | — | — | — | — | — | 66.5 | — | — | — | — | |
| SegCompass-7BBackbone=Qwen2.5-VL, Zero-shot=true2026.05 | 64 | — | — | — | — | — | — | 64.8 | — | — | — | — | |
| Vision-Reasoner-7B2026.01 | 63.6 | — | — | — | — | — | — | — | — | — | — | — | |
| VisionReasonerLanguage Model=Qwen2.5VL-7B, Trained on ReasonSeg train split=false2026.01 | 63.6 | — | — | — | — | — | — | 58.2 | 84.8 | — | — | — | |
| VisionReasonerParams=7B2026.02 | 63.6 | — | — | — | — | — | — | — | — | — | — | — | |
| VisionReasonerModel Category=RL-based 7B VLLMs2026.02 | 63.6 | — | — | — | — | — | — | — | — | 69.3 | 72.2 | — | |
| VisionReasoner-7BBackbone=7B2025.10 | 63.6 | — | — | — | — | — | — | — | — | — | — | — | |
| VisionReasoner-7Bconsistent-checkpoint=true2026.05 | 63.6 | — | — | — | — | — | — | — | — | — | — | — | |
| VisionReasoner-7BYear=2025, Backbone=Qwen2.5-VL, Zero-shot=true2026.05 | 63.6 | — | — | — | — | — | — | — | — | — | — | — | |
| VisionReasonerZero-shot=true2026.05 | 63.6 | — | — | — | — | — | — | 55.7 | — | — | — | — | |
| GELLA2024.02 | 63.3 | — | — | — | — | — | — | 64.1 | — | — | — | — | |
| SAM 3 AgentVersion=Qwen2.5-VL 7B, Training-RES=false, Training-ReasonSeg=false2025.12 | 63 | — | — | — | — | — | — | 53.5 | — | — | — | — | |
| SAM3-Agent-7BPublication=ICLR’262026.05 | 63 | — | — | — | — | — | — | 53.5 | — | — | — | — | |
| ConceptSeg-R1-7Bfine-tuned=false2026.05 | 63 | — | — | — | — | — | — | 59.3 | — | — | — | — | |
| SAM3-Agent2026.05 | 63 | — | — | — | — | — | — | — | — | — | — | — | |
| SAM-VeteranVersion=Qwen2.5-VL 32B, Training-RES=true, Training-ReasonSeg=true2025.12 | 62.9 | — | — | — | — | — | — | 58.2 | — | — | — | — | |
| VisionReasonerMethod Category=Explicit-Prompt-Based Reasoning Segmentation (ERS), Experimental Condition=Independent experiments under the same experimental conditions2026.06 | 62.8 | — | — | — | — | — | — | 55.1 | — | — | — | — | |
| SAM3-Agent-7BNumber of parameters=7B2026.02 | 62.6 | — | — | — | — | — | — | 56.2 | — | — | — | — | |
| SAM-VeteranVersion=Qwen2.5-VL 7B, Training-RES=true, Training-ReasonSeg=true2025.12 | 62.6 | — | — | — | — | — | — | 56.1 | — | — | — | — | |
| SAM3-AgentModel Version=Qwen2.5-VL, Params=7B, RES Training=false, ReasonSeg Training=false2025.12 | 62.6 | — | — | — | — | — | — | 56.2 | — | — | — | — | |
| SAM3-Agent-7BParameters=7B2026.05 | 62.6 | — | — | — | — | — | — | 56.2 | — | — | — | — | |
| SAM 3 AgentBackbone=7B, Zero-shot=true2026.05 | 62.6 | — | — | — | — | — | — | 56.2 | — | — | — | — | |
| SAM-VeteranMLLM=Qwen2.5-VL-7B, setting=zero-shot2026.07 | 62.6 | — | — | — | — | — | — | 56.1 | — | — | — | — | |
| SAM3 AgentMLLM=Qwen2.5-VL-7B, setting=zero-shot2026.07 | 62.6 | — | — | — | — | — | — | 56.2 | — | — | — | — | |
| VisionReasonerLanguage Model=Qwen2.5VL-7B, Trained on ReasonSeg train split=true2026.01 | 62.3 | — | — | — | — | — | — | 54.6 | 81.4 | — | — | — | |
| VGent2025.12 | 62.2 | — | — | — | — | — | — | 64 | — | — | — | — | |
| Rea2SegModel=Qwen2.5-VL-3B, Selection=Top 12026.06 | 62.1 | — | — | — | — | — | — | 62.3 | — | — | — | — | |
| VisionReasonerModel Category=RL-based 7B VLLMs, Independent Experiment=true2026.02 | 61.5 | — | — | — | — | — | — | — | — | 68.4 | 72 | — | |
| GLAMM2024.02 | 61.5 | — | — | — | — | — | — | 62.4 | — | — | — | — | |
| Seg-Zero-7BParameters=7B2026.02 | 61.41 | — | — | — | — | — | — | 55.63 | — | — | — | — | |
| LISA-13B-LLaVA1.5Fine-tuning=true2026.01 | 61.3 | — | — | — | — | — | — | 62.2 | — | — | — | — | |
| LISAVersion=LLaVA1.5 13B, Training-RES=true, Training-ReasonSeg=true2025.12 | 61.3 | — | — | — | — | — | — | 62.2 | — | — | — | — | |
| LISAParameters=13B, Training=Fine-tuned2025.12 | 61.3 | — | — | — | — | — | — | 62.2 | — | — | — | — | |
| LISAParams=13B2026.02 | 61.3 | — | — | — | — | — | — | 62.2 | — | — | — | — | |
| LISA2024.02 | 61.3 | — | — | — | — | — | — | 62.2 | — | — | — | — | |
| ConceptSeg-R1-3Bfine-tuned=false2026.05 | 61.2 | — | — | — | — | — | — | 49.3 | — | — | — | — | |
| Ours (single-object)zero-shot=true2026.05 | 61.11 | — | — | — | — | — | — | — | — | — | — | — | |
| SegCompass-7BBackbone=LLaVA-1.5, Zero-shot=true2026.05 | 61 | — | — | — | — | — | — | 63 | — | — | — | — | |
| DPAD-7B2026.03 | 60.8 | — | — | — | — | — | — | 57.5 | — | — | — | — | |
| HiMTok-8B2026.03 | 60.8 | — | — | — | — | — | — | 66.2 | — | — | — | — | |
| HiMTok-8BYear=2025, Backbone=InternVL-2.5, Zero-shot=true2026.05 | 60.8 | — | — | — | — | — | — | 66.2 | — | — | — | — | |
| SenseNova-Vision2026.07 | 60.7 | — | — | — | — | — | — | — | — | — | — | — | |
| FlowSeg2026.05 | 60.5 | — | — | — | — | — | — | 54.7 | — | — | — | — | |
| RSVPVersion=GPT-4o, Training-RES=false, Training-ReasonSeg=false2025.12 | 60.3 | — | — | — | — | — | — | 60 | — | — | — | — | |
| RSVPModel Version=GPT-4o, RES Training=false, ReasonSeg Training=false2025.12 | 60.3 | — | — | — | — | — | — | 60 | — | — | — | — | |
| RSVPMethod Category=Explicit-Prompt-Based Reasoning Segmentation (ERS)2026.06 | 60.3 | — | — | — | — | — | — | 60 | — | — | — | — | |
| Rea2SegModel=LLaVA-v1.5-7B, Selection=Top 12026.06 | 60.3 | — | — | — | — | — | — | 59.2 | — | — | — | — | |
| POPEN2025.04 | 60.2 | — | — | — | — | — | — | 64.5 | — | — | — | — | |
| SAM-R1-7BNumber of parameters=7B2026.02 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| SAM-R1Language Model=Qwen2.5VL-7B, Trained on ReasonSeg train split=false2026.01 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| DR2SegLanguage Model=Qwen2.5VL-3B, Segmentation Model=SAM22026.01 | 60.2 | — | — | — | — | — | — | 55 | 33.3 | — | — | — | |
| SAM-R1Model Version=Qwen2.5-VL, Params=7B, RES Training=true, ReasonSeg Training=false2025.12 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| SAM-R1Params=7B2026.02 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| Pixel-ThinkModel Category=RL-based 7B VLLMs, Source Status=not open-sourced2026.02 | 60.2 | — | — | — | — | — | — | — | — | — | — | — | |
| SAM-R1-7BParameters=7B2026.05 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| SAM-R1-7BPublication=NeurIPS’252026.05 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| SAM-R1-7BYear=2025, Backbone=Qwen2.5-VL, Zero-shot=true2026.05 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — | |
| SAM-R1Model=Qwen2.5-VL-7B2026.06 | 60.2 | — | — | — | — | — | — | 54.3 | — | — | — | — |