Perception on CVBench (test)
87.6AccuracyVaLR-M
Evaluation Results
| Method | Links | |
|---|---|---|
| VaLR-MModel Category=Latent Reasoning models, Encoder=Multiple (DINOv3, SigLIPv2, pi^3)2026.02 | 87.6 | |
| VaLR-SModel Category=Latent Reasoning models, Encoder=Single (DINOv3)2026.02 | 83.1 | |
| CoVTModel Category=Latent Reasoning models2026.02 | 80 | |
| GPT-4oModel Category=API models2026.02 | 79.2 | |
| Ocean-R1-7BModel Category=Reasoning models, Parameters=7B2026.02 | 78.1 | |
| Qwen2.5-VL-7B + vanilla SFTModel Category=Base model, Training=vanilla SFT2026.02 | 77 | |
| LVTModel Category=Latent Reasoning models2026.02 | 76.9 | |
| Claude-3-SonnetModel Category=API models2026.02 | 76.3 | |
| Qwen2.5-VL-7BModel Category=Base model, Parameters=7B2026.02 | 74.5 | |
| MonetModel Category=Latent Reasoning models2026.02 | 71.1 | |
| R1-OneVision-7BModel Category=Reasoning models, Parameters=7B2026.02 | 67.2 |