Cross-view Correspondence on RE10K ReasonMatch-Bench high-divergence samples
96.1PrecisionHuman
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HumanAnnotator=Human2026.06 | 96.1 | 93.9 | 94.7 | |
| DCRLModel Category=Open-Source Models2026.06 | 65.7 | 76.6 | 70.6 | |
| GPT-5-miniModel Category=Closed-Source Models2026.06 | 47.4 | 52.7 | 49.7 | |
| Gemini-2.5-ProModel Category=Closed-Source Models2026.06 | 42.8 | 45.6 | 44.1 | |
| Qwen3-VL-235BModel Category=Open-Source Models2026.06 | 42.3 | 50.2 | 45.7 | |
| Claude-4.5-SonnetModel Category=Closed-Source Models2026.06 | 28.1 | 33.6 | 30.5 |