Long-term Video Object Segmentation on LVOS (val)
63.5J&F ScoreCutie-base
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 63.5 | 59.1 | 67.9 | — | 30.1 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 60.7 | 55.6 | 65.8 | — | 34 | |
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS and DAVIS2024.11 | 60.1 | 55.9 | 64.2 | — | 30.1 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS and DAVIS2024.11 | 58.8 | 54.6 | 62.9 | — | 34 | |
| DEVAKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS and DAVIS2024.11 | 58.3 | 52.8 | 63.8 | — | 48.3 | |
| DEVAKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 55.9 | 51.1 | 60.7 | — | 48.3 | |
| LiVOSKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS and DAVIS, extra_iterations=30k2024.11 | 51.2 | 46.8 | 55.6 | — | 47.3 | |
| LiVOSKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 51.2 | 47.3 | 55.1 | — | 47.3 | |
| LiVOSKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS and DAVIS2024.11 | 50.6 | 46.5 | 54.7 | — | 47.3 | |
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS and DAVIS2024.11 | 49.1 | 44.7 | 53.4 | — | 45.5 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS and DAVIS2024.11 | 48.2 | 43.9 | 52.5 | — | 56.1 | |
| RDEKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS and DAVIS2024.11 | 47.2 | 41.7 | 52.7 | — | 40.6 | |
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 46.8 | 43.1 | 50.6 | — | 45.5 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 46.1 | 42.1 | 50.1 | — | 56.1 |