Cross-modal Generation on VGGSound
87.23Average ScoreGRAM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GRAMDecoder=Kandinsky2026.02 | 87.23 | 62.11 | 106.97 | — | |
| ImageBindDecoder=Stable UnCLIP2026.02 | 50.19 | 50.17 | 53.59 | — | |
| GRAMDecoder=Stable UnCLIP2026.02 | 49.36 | 45.53 | 55.4 | — | |
| UniAlign (Geodesic)Decoder=Kandinsky2026.02 | 48.09 | 45.35 | 50.75 | — | |
| UniAlign (Euclidean)Decoder=Kandinsky2026.02 | 46.6 | 42.72 | 50.51 | — | |
| UniAlign (Geodesic)Decoder=Stable UnCLIP2026.02 | 40.23 | 39.88 | 40.16 | — | |
| UniAlign (Euclidean)Decoder=Stable UnCLIP2026.02 | 40.2 | 39.63 | 39.95 | — | |
| Stable UnCLIP (Self-reconstruction)Decoder=Stable UnCLIP, Reference=Upper-bound reference2026.02 | 34.61 | — | — | — | |
| Kandinsky (Self-reconstruction)Decoder=Kandinsky, Reference=Upper-bound reference2026.02 | 32.99 | — | — | — |