Active Hazard Perception on CAMMA-MVOR Expert-Authored Challenge Subset
98.75Perception Rate (Common)Gemini-3.1-flash
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemini-3.1-flashVariant=flash, Decoding Strategy=Greedy2025.06 | 98.75 | 86.41 | 87.29 | 90.13 | |
| GPT-5.4Decoding Strategy=Greedy2025.06 | 98 | 85.4 | 91.1 | 90.68 | |
| Ensemble (Protocol-to-Pixel)Training Strategy=Fine-Tuned (FT), Mode=Ensemble, Alignment=Synthetic Data-Guided2025.06 | 97.22 | 96.29 | 98.97 | 97.23 | |
| GLM-4.6V-FlashVariant=Flash2025.06 | 95.53 | 85.75 | 59.35 | 80.33 | |
| Gemini-3-flashVariant=flash, Decoding Strategy=Greedy2025.06 | 90.92 | 69.07 | 79.05 | 78.26 | |
| Qwen-Screen (Protocol-to-Pixel)Training Strategy=Fine-Tuned (FT), Backbone=Qwen-Screen, Alignment=Synthetic Data-Guided2025.06 | 81.99 | 90.55 | 97.14 | 90.19 | |
| BaichuanMed-OCR-7BParameter Scale=7B, Domain=Medical2025.06 | 68.08 | 83.23 | 9.65 | 64.31 | |
| GPT-5.4-nanoVariant=nano, Decoding Strategy=Greedy2025.06 | 37.42 | 34.33 | 22.69 | 31.6 | |
| Qwen2.5-VL-7B-InstructParameter Scale=7B, Variant=Instruct2025.06 | 4.94 | 87.37 | 6.04 | 30.92 | |
| Llama-3.2-11B-Vision-InstructParameter Scale=11B, Variant=Vision-Instruct2025.06 | 3.8 | 87.62 | 23.75 | 34.55 |