Long-term Video Object Segmentation on LVOS (test)
63.6J&FCutie-base
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 63.6 | 59.1 | 68 | — | 29.7 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS and DAVIS2024.11 | 57.2 | 53.7 | 60.7 | — | 32.8 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 56.9 | 53.5 | 60.2 | — | 32.8 | |
| DEVAKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 56.5 | 52.2 | 60.8 | — | 46.6 | |
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS and DAVIS2024.11 | 56.2 | 51.8 | 60.5 | — | 29.7 | |
| DEVAKey Encoder=RN-50, Value Encoder=RN-18, with STM=true, training_set=YouTube VOS and DAVIS2024.11 | 54 | 49 | 59 | — | 46.6 | |
| LiVOSKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS and DAVIS, extra_iterations=30k2024.11 | 50.9 | 47 | 54.7 | — | 45.2 | |
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS and DAVIS2024.11 | 48.7 | 44.6 | 52.7 | — | 43 | |
| LiVOSKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 47 | 44 | 50 | — | 45.2 | |
| Cutie-baseKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 45.6 | 41.6 | 49.7 | — | 43 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS and DAVIS2024.11 | 45.2 | 41.3 | 49 | — | 55.2 | |
| RDEKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS and DAVIS2024.11 | 44.7 | 39.2 | 50.2 | — | 39.8 | |
| LiVOSKey Encoder=RN-50, Value Encoder=RN-18, with STM=false, training_set=YouTube VOS and DAVIS2024.11 | 44.6 | 41.2 | 47.9 | — | 45.2 | |
| Cutie-smallKey Encoder=RN-18, Value Encoder=RN-18, with STM=false, memory_mode=downgraded (one memory frame), training_set=YouTube VOS, DAVIS, and MOSE2024.11 | 44 | 40.1 | 48 | — | 55.2 |