AI-generated video detection on GenVidBench matched 27k
100Audited LOGO OOD (AUC)Single-feature (ntotal, clip length)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Single-feature (ntotal, clip length)dim=1, K=12 filter (C2)=collapse2026.06 | 100 | 2 | — | — | |
| WaveRep DINOv2 raw-feature probedim=768, K=12 filter (C2)=pass2026.06 | 99.9 | 0.777 | 0.222 | — | |
| WaveRep (same backbone, raw-feature probe)Subset=27k, dim=768, ms/clip=4202026.06 | 99.9 | 0.777 | 0.222 | 100 | |
| WaveRep native LLRdim=3, K=12 filter (C2)=pass2026.06 | 99.6 | 0.534 | 0.462 | — | |
| WaveRep (forensic LLR, DINOv2 ViT-B/14)Subset=27k, dim=3, ms/clip=4202026.06 | 99.6 | 0.534 | 0.462 | 99.6 | |
| XSFF: MV ⊕ ReStraV (feature-level 34-d)Subset=27k, dim=34, ms/clip=14+10*2026.06 | 94.6 | 0.604 | 0.342 | 99.6 | |
| MV ⊕ ReStraV score blendAlpha=0.252026.06 | 94.4 | — | — | — | |
| fixed-α blend (best linear, baseline)Subset=27k, dim=n/a, ms/clip=14+10*2026.06 | 93.5 | 0.604 | 0.331 | 99.5 | |
| ReStraV DINOv2 ViT-S/14 (matched harness)dim=21, K=12 filter (C2)=pass2026.06 | 93.1 | 0.586 | 0.345 | — | |
| ReStraV (perceptual-straightening)Subset=27k, dim=21, ms/clip=10*2026.06 | 93.1 | 0.586 | 0.345 | 99.5 | |
| ReStraVFeature Dimensions=21-d, Classifier=matched L2-LR2026.06 | 93 | — | — | — | |
| D3 (frame-pixel, native head)Subset=27k, dim=1, ms/clip=50*2026.06 | 88.7 | 0.421 | 0.466 | 88.7 | |
| FVMD (motion-distance, full 1024-d)Subset=27k, dim=1024, ms/clip=59*2026.06 | 88 | 0.574 | 0.306 | 92.4 | |
| TemporalSpec+aug (20-d, LightGBM)dim=20, K=12 filter (C2)=pass2026.06 | 87.1 | 0.634 | 0.237 | — | |
| TemporalSpec+augFeature Dimensions=20-d, Classifier=LightGBM2026.06 | 87.1 | — | — | — | |
| TemporalSpec+aug (20-d, LightGBM)Subset=27k, dim=20, ms/clip=15 (CPU)2026.06 | 87.1 | 0.634 | 0.237 | 96 | |
| FVMD (dimensionality-matched, PCA-13)Subset=27k, dim=13, ms/clip=59*2026.06 | 86.3 | 0.555 | 0.308 | 88 | |
| RAFT (optical-flow -> Tier 0/1)Subset=27k, dim=13, ms/clip=4002026.06 | 85.5 | 0.627 | 0.228 | 90.1 | |
| CLIP-ViT-B/32 mean-pool (appearance)dim=13, K=12 filter (C2)=pass2026.06 | 85.2 | 0.766 | 0.086 | — | |
| CLIP-ViT-B/32 (appearance)Subset=27k, dim=13, ms/clip=67*2026.06 | 85.2 | 0.766 | 0.086 | 94.5 | |
| TemporalSpec (13-d codec-motion)dim=13, K=12 filter (C2)=pass2026.06 | 83.2 | 0.643 | 0.189 | — | |
| TemporalSpecFeature Dimensions=13-d, Classifier=L2-LR2026.06 | 83.2 | — | — | — | |
| TemporalSpec (13-d, L2-LR)Subset=27k, dim=13, ms/clip=14 (CPU)2026.06 | 83.2 | 0.643 | 0.189 | 90.9 | |
| D3 (frame-pixel, features -> L2-LR)Subset=27k, dim=768, ms/clip=50*2026.06 | 55.7 | 0.546 | 0.011 | 66.8 | |
| Three-feature (nI, nP, nB, frame counts)dim=3, K=12 filter (C2)=collapse2026.06 | 52.9 | 2 | — | — |