Robot Failure Detection on RLBench Fail
83Execution AccuracyGuardian-8B-Thinking
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Guardian-8B-ThinkingModel Category=Specialized / small models, Reasoning Mode=Thinking (explicit reasoning)2025.12 | 83 | 87 | |
| CLIP+MLPModel Category=Specialized / small models2025.12 | 65 | 53 | |
| GPT4.1Model Category=Large-scale generalist models2025.12 | 63 | 87 | |
| Qwen3-VL-235B-A22BModel Category=Large-scale generalist models2025.12 | 59 | 83 | |
| InternVL3-8BModel Category=Specialized / small models2025.12 | 59 | 70 | |
| SentinelModel Category=Specialized / small models2025.12 | 57 | — | |
| Cosmos-Reason1-7BModel Category=Specialized / small models2025.12 | 54 | 60 |