Multi-Object Tracking on DanceTrack (test)
65.7HOTAHybrid-SORT
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Hybrid-SORTtuning on test set=true, object detection model=public2024.06 | 65.7 | 67.4 | — | 91.8 | — | — | — | — | — | — | |
| UCMCTracktuning on test set=true, object detection model=public2024.06 | 63.6 | 65 | 51.3 | 88.9 | — | — | — | — | — | — | |
| DeepMoveSORT-TransFiltertuning on test set=false, object detection model=public2024.06 | 63 | 65 | 48.6 | 92.6 | 82 | — | — | — | — | — | |
| Deep OC_SORTtuning on test set=false, object detection model=public2024.06 | 61.3 | 61.5 | 45.8 | 91.8 | 81.6 | — | — | — | — | — | |
| DeepMoveSORT-RNNFiltertuning on test set=false, object detection model=public2024.06 | 59.4 | 61 | 43.1 | 92.5 | 82.1 | — | — | — | — | — | |
| MotionTracktuning on test set=false, object detection model=public2024.06 | 58.2 | 58.6 | 41.7 | 91.3 | 81.4 | — | — | — | — | — | |
| MoveSORTtuning on test set=false, object detection model=public2024.06 | 56.1 | 56 | 38.7 | 91.8 | 81.6 | — | — | — | — | — | |
| SparseTracktuning on test set=false, object detection model=public2024.06 | 55.5 | 58.3 | 39.1 | 91.3 | 78.9 | — | — | — | — | — | |
| OC_SORTtuning on test set=false, object detection model=public2024.06 | 55.1 | 54.9 | 38 | 89.4 | 80.4 | — | — | — | — | — | |
| BOT-SORT-ReIDtuning on test set=false, object detection model=public2024.06 | 54.2 | 57 | 36.8 | 92.1 | 80 | — | — | — | — | — | |
| BOT-SORTtuning on test set=false, object detection model=public2024.06 | 53.6 | 55.3 | 36.1 | 92.3 | 79.9 | — | — | — | — | — | |
| MotionTrack (DanceTrack)tuning on test set=false, object detection model=public2024.06 | 52.9 | 53.8 | 34.7 | 91.3 | 80.9 | — | — | — | — | — | |
| SORTtuning on test set=false, object detection model=public2024.06 | 51.2 | 49.7 | 32.6 | 91.5 | 80.5 | — | — | — | — | — | |
| ByteTracktuning on test set=false, object detection model=public2024.06 | 47.3 | 52.5 | 31.4 | 89.5 | 71.6 | — | — | — | — | — | |
| DeepSORTtuning on test set=false, object detection model=public2024.06 | 45.6 | 47.9 | 29.7 | 87.8 | 71 | — | — | — | — | — | |
| DualMOTMethod Category=query-based, Motion Estimation=Learnable2026.04 | 0.762 | 0.799 | 0.683 | 0.85 | 0.85 | — | — | — | — | — | |
| SAM2MOTBackbone=Grounding-DINO-L, Fine-tuning status=no fine-tuning2025.04 | 0.758 | 0.839 | 0.722 | 0.885 | 0.797 | — | — | — | — | — | |
| SAM2MOTMethod Category=query-based, Motion Estimation=Learnable2026.04 | 0.758 | 0.839 | 0.722 | 0.885 | 0.797 | — | — | — | — | — | |
| SAM2MOTBackbone=Co-DINO-L, Fine-tuning status=no fine-tuning2025.04 | 0.755 | 0.834 | 0.713 | 0.892 | 0.803 | — | — | — | — | — | |
| MATREnd-to-End=true, Parameters (M)=43, Backbone=Swin-Tiny, Additional training on validation set=true2025.09 | 0.739 | 0.765 | 0.648 | 0.929 | 0.845 | — | — | — | — | — | |
| MOTIPPublication=CVPR20252025.04 | 0.737 | 0.794 | 0.659 | 0.927 | 0.826 | — | — | — | — | — | |
| MOTRv2*extra association=true, training data=validation set included, ensemble=test ensemble2022.11 | 0.734 | 0.76 | 0.644 | 0.921 | 0.837 | — | — | — | — | — | |
| MOTRv2Method Category=query-based, Motion Estimation=Learnable2026.04 | 0.734 | 0.76 | 0.644 | 0.921 | 0.837 | — | — | — | — | — | |
| ColTrackMethod Category=Tracking-by-Query, Publication=ICCV20232024.09 | 0.726 | 0.74 | 0.623 | 0.921 | — | — | — | — | — | — | |
| ColTrackPublication=ICCV20232025.04 | 0.726 | 0.74 | 0.623 | 0.921 | — | — | — | — | — | — | |
| MOTIPMethod Category=query-based, Motion Estimation=Learnable2026.04 | 0.72 | 0.768 | 0.635 | 0.919 | 0.818 | — | — | — | — | — | |
| MATREnd-to-End=true, Parameters (M)=43, Backbone=Swin-Tiny2025.09 | 0.713 | 0.753 | 0.616 | 0.919 | 0.826 | — | — | — | — | — | |
| TCEIArchitecture Family=Transformer based, Framework=Deformable DETR, Backbone=ResNet-50, Baseline=MOTIP2026.03 | 0.706 | 0.756 | 0.623 | — | 0.802 | — | — | — | — | — | |
| MOTRv3End-to-End=true, Parameters (M)=1032025.09 | 0.704 | 0.723 | 0.593 | 0.929 | 0.838 | — | — | — | — | — | |
| MOTRv3Method Category=query-based, Motion Estimation=Learnable2026.04 | 0.704 | 0.723 | 0.593 | 0.929 | 0.838 | — | — | — | — | — | |
| HieDGMethod category=Ours, continuous geometric cues=false2026.07 | 0.703 | 0.754 | 0.618 | 0.909 | 0.807 | — | — | — | — | — | |
| TDLPParadigm=tbd, Features=extra features, Detector=YOLOX2025.12 | 0.701 | 0.758 | 0.596 | 0.918 | 0.826 | — | — | — | — | — | |
| MOTIPTracking approach=Tracking-by-attention, Backbone architecture=DAB-Deformable DETR2024.09 | 0.7 | 0.751 | 0.608 | 0.91 | — | — | — | — | — | — | |
| MOTIPYear=2025, Tracking Paradigm=TBQ, YOLOX Detector=false2024.11 | 0.7 | 0.751 | 0.608 | 0.91 | 0.808 | — | — | — | — | — | |
| MOTRv2Method Category=Tracking-by-Query, Publication=CVPR20222024.09 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv22022.11 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Architecture=End-to-end2023.05 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Matcher Type=Learnable Matcher2023.08 | 0.699 | 0.717 | — | 0.919 | — | — | — | — | — | — | |
| MOTRv2extra_data=true2023.07 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Architecture=Hybird based, Venue=CVPR’232024.06 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Publication=CVPR20222025.04 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2End-to-End=false, Parameters (M)=942025.09 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Category=Non End-to-End2025.11 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Year=2023, Tracking Paradigm=TBQ, YOLOX Detector=true2024.11 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTRv2Processing Mode=Online, Extra Training Data=true2025.09 | 0.699 | 0.717 | 0.59 | 0.919 | — | — | — | — | — | — | |
| MOTRv2Detection Setting=Different Detection Settings, Venue=CVPR232026.05 | 0.699 | 0.717 | 0.59 | 0.919 | 0.83 | — | — | — | — | — | |
| MOTIPDetection Setting=Different Detection Settings, Venue=CVPR252026.05 | 0.696 | 0.747 | 0.604 | 0.906 | 0.804 | — | — | — | — | — | |
| MOTIPVenue=CVPR, Year=20252026.06 | 0.696 | 0.747 | 0.604 | 0.906 | 0.804 | — | — | — | — | — | |
| MOTIPMethod category=Query-based2026.07 | 0.696 | 0.747 | 0.604 | 0.906 | 0.804 | — | — | — | — | — | |
| MOTIPArchitecture Family=Transformer based, Framework=Deformable DETR, Backbone=ResNet-502026.03 | 0.695 | 0.746 | 0.602 | — | 0.804 | — | — | — | — | — | |
| CO-MOTArchitecture=End-to-end2023.05 | 0.694 | 0.719 | 0.589 | 0.912 | 0.821 | — | — | — | — | — | |
| CO-MOTEnd-to-End=true, Parameters (M)=402025.09 | 0.694 | 0.719 | 0.589 | 0.912 | 0.821 | — | — | — | — | — | |
| MATR-R50End-to-End=true, Parameters (M)=42, Backbone=ResNet-502025.09 | 0.694 | 0.726 | 0.591 | 0.91 | 0.815 | — | — | — | — | — | |
| CO-MOTCategory=End-to-End2025.11 | 0.694 | 0.719 | 0.589 | 0.912 | 0.821 | — | — | — | — | — | |
| CO-MOTYear=2025, Tracking Paradigm=TBQ, YOLOX Detector=false2024.11 | 0.694 | 0.719 | 0.589 | 0.912 | 0.821 | — | — | — | — | — | |
| TBDQ-NetYear=ours, Tracking Paradigm=TBQ, YOLOX Detector=true2024.11 | 0.694 | 0.734 | 0.598 | 0.895 | 0.808 | — | — | — | — | — | |
| CAMELTrackParadigm=tbd, Features=extra features, Detector=YOLOX2025.12 | 0.693 | 0.749 | 0.589 | 0.914 | 0.818 | — | — | — | — | — | |
| SelfMOTRCategory=End-to-End2025.11 | 0.692 | 0.725 | 0.593 | 0.899 | 0.809 | — | — | — | — | — | |
| MeMOTRextra_data=false, backbone_variant=DAB-Deformable-DETR2023.07 | 0.685 | 0.712 | 0.584 | 0.899 | 0.805 | — | — | — | — | — | |
| MeMOTRTracking approach=Tracking-by-attention, Backbone architecture=DAB-Deformable DETR2024.09 | 0.685 | 0.712 | 0.584 | 0.899 | — | — | — | — | — | — | |
| MeMOTRParadigm=e2e, Detector=own2025.12 | 0.685 | 0.712 | 0.584 | 0.899 | 0.716 | — | — | — | — | — | |
| MeMOTREnd-to-End=true, Parameters (M)=482025.09 | 0.685 | 0.712 | 0.584 | 0.899 | 0.805 | — | — | — | — | — | |
| MeMOTRCategory=End-to-End2025.11 | 0.685 | 0.712 | 0.584 | 0.899 | 0.805 | — | — | — | — | — | |
| MeMOTRYear=2023, Tracking Paradigm=TBQ, YOLOX Detector=false2024.11 | 0.685 | 0.712 | 0.584 | 0.899 | 0.805 | — | — | — | — | — | |
| MeMOTRMethod Category=query-based, Motion Estimation=Learnable2026.04 | 0.685 | 0.712 | 0.584 | 0.899 | 0.805 | — | — | — | — | — | |
| MeMOTRProcessing Mode=Online, Extra Training Data=true2025.09 | 0.685 | 0.712 | 0.584 | 0.899 | — | — | — | — | — | — | |
| HieDG*Method category=Ours, continuous geometric cues=true2026.07 | 0.685 | 0.731 | 0.6 | 0.914 | 0.808 | — | — | — | — | — | |
| NOOUGATCategory=Graph-based2025.11 | 0.684 | 0.727 | 0.587 | 0.889 | — | — | — | — | — | — | |
| NOOUGATProcessing Mode=Offline, Shared Detections=true2025.09 | 0.684 | 0.727 | 0.587 | 0.889 | — | — | — | — | — | — | |
| MOTRv3Category=End-to-End2025.11 | 0.683 | 0.701 | — | 0.917 | — | — | — | — | — | — | |
| Hybrid-SORT+SAMOFTDetection Setting=Same Detection, Venue=Ours2026.05 | 0.682 | 0.715 | 0.571 | 0.915 | 0.817 | — | — | — | — | — | |
| TDLP-bboxParadigm=tbd, Features=bbox features, Detector=YOLOX2025.12 | 0.678 | 0.727 | 0.561 | 0.919 | 0.822 | — | — | — | — | — | |
| MOTIPTracking approach=Tracking-by-attention, Backbone architecture=Deformable DETR2024.09 | 0.675 | 0.722 | 0.576 | 0.903 | — | — | — | — | — | — | |
| MOTIPParadigm=e2e, Detector=own2025.12 | 0.675 | 0.722 | 0.576 | 0.903 | 0.716 | — | — | — | — | — | |
| MOTIP2026.03 | 0.675 | 0.722 | 0.576 | 0.903 | — | — | — | — | — | — | |
| SambaMOTRArchitecture Family=SSM based2026.03 | 0.672 | 0.705 | 0.575 | — | 0.788 | — | — | — | — | — | |
| SambaMOTRCategory=End-to-End2025.11 | 0.672 | 0.7 | 0.575 | 0.881 | 0.788 | — | — | — | — | — | |
| SAMBAMethod Category=query-based, Motion Estimation=Learnable2026.04 | 0.672 | 0.705 | 0.575 | 0.881 | 0.788 | — | — | — | — | — | |
| MT_IoTextra_data=true2023.07 | 0.667 | 0.706 | 0.53 | 0.94 | 0.841 | — | — | — | — | — | |
| MeMOTRTracking approach=Tracking-by-attention, Backbone architecture=Deformable DETR2024.09 | 0.667 | 0.706 | 0.53 | 0.94 | — | — | — | — | — | — | |
| AEDMethod Category=Tracking-by-Detection, Publication=Ours, Detector=YOLOX2024.09 | 0.666 | 0.697 | 0.543 | 0.922 | 0.82 | — | — | — | — | — | |
| AEDPublication=TIP20252025.04 | 0.666 | 0.697 | 0.543 | 0.922 | 0.82 | — | — | — | — | — | |
| TrackTrackCategory=Non End-to-End2025.11 | 0.665 | 0.678 | 0.529 | 0.936 | — | — | — | — | — | — | |
| TrackTrackDetection Setting=Same Detection, Venue=CVPR252026.05 | 0.665 | 0.678 | 0.529 | 0.936 | — | — | — | — | — | — | |
| TrackTrackVenue=CVPR, Year=20252026.06 | 0.665 | 0.678 | 0.529 | 0.936 | — | — | — | — | — | — | |
| IMM-JHSETracking approach=Tracking-by-detection, Detection source=(Zhang et al., 2022)2024.09 | 0.6624 | 0.7172 | 0.5541 | 0.8995 | — | — | — | — | — | — | |
| QTrack2026.03 | 0.66 | — | — | 0.63 | — | — | — | — | 0.83 | 0.35 | |
| NOOUGATProcessing Mode=Online, Shared Detections=true2025.09 | 0.659 | 0.706 | 0.549 | 0.889 | — | — | — | — | — | — | |
| Hybrid-SORT-ReIDMatcher Type=Heuristic Matcher, shared_detections=true2023.08 | 0.657 | 0.674 | — | 0.918 | — | — | — | — | — | — | |
| Hybrid-SORTPublication=AAAI20242025.04 | 0.657 | 0.674 | — | 0.918 | — | — | — | — | — | — | |
| Hybrid-SORTParadigm=tbd, Features=extra features, Detector=YOLOX2025.12 | 0.657 | 0.674 | — | 0.918 | — | — | — | — | — | — | |
| Hybrid-SORTCategory=Non End-to-End2025.11 | 0.657 | 0.63 | 0.526 | 0.918 | — | — | — | — | — | — | |
| Hybrid-SORTYear=2024, Tracking Paradigm=TBD, YOLOX Detector=true2024.11 | 0.657 | 0.674 | — | 0.918 | — | — | — | — | — | — | |
| HybridSORTMethod Category=motion-based, Motion Estimation=Kalman Filter2026.04 | 0.657 | 0.674 | — | 0.918 | — | — | — | — | — | — | |
| HyperSSMMethod Category=motion-based, Motion Estimation=Learnable2026.04 | 0.657 | 0.68 | 0.526 | 0.927 | 0.824 | — | — | — | — | — | |
| Hybrid-SORTProcessing Mode=Online, Shared Detections=true2025.09 | 0.657 | 0.674 | 0.526 | 0.918 | — | — | — | — | — | — | |
| Hybrid-SORT-ReIDDetection Setting=Same Detection, Venue=AAAI242026.05 | 0.657 | 0.674 | — | 0.918 | — | — | — | — | — | — | |
| DfTrack-HybridDetection Setting=Same Detection, Venue=TCSVT252026.05 | 0.655 | 0.682 | 0.524 | 0.927 | 0.821 | — | — | — | — | — | |
| FusionTrackTracking approach=Tracking-by-attention2024.09 | 0.653 | 0.733 | 0.575 | 0.901 | — | — | — | — | — | — | |
| CO-MOTMethod category=Query-based2026.07 | 0.653 | 0.665 | 0.535 | 0.893 | 0.801 | — | — | — | — | — |