Image Reconstruction on ImageNet 50k 1k (val)
0.63rFIDTokenFlow
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TokenFlowResolution=384x384, Downsample Ratio=14.2, Residual Levels=152024.12 | 0.63 | 22.77 | 73.1 | — | |
| Open-MAGVIT2-I-PTResolution=Original Resolution, Quantizer Type=LFQ, Training Data=100M, Ratio=16, Codebook Size=2621442024.09 | 0.78 | 22.24 | 59 | — | |
| LlamaGenResolution=384x384, Downsample Ratio=14.2, Residual Levels=12024.12 | 0.94 | 21.94 | 72.6 | — | |
| VARResolution=256x256, Downsample Ratio=16, Residual Levels=102024.12 | 1 | 22.63 | 75.5 | — | |
| Mo-VQGANLatent Size=16x16x4, Num Z=10242022.09 | 1.12 | 22.42 | 67.31 | 0.1132 | |
| VILA-UResolution=384x384, Downsample Ratio=14.2, Residual Levels=162024.12 | 1.25 | — | — | — | |
| ViT-VQGANLatent Size=32x32, Num Z=81922022.09 | 1.28 | — | — | — | |
| RQ-VAEResolution=256x256, Downsample Ratio=16, Residual Levels=42024.12 | 1.3 | — | — | — | |
| TokenFlowResolution=256x256, Downsample Ratio=16, Residual Levels=92024.12 | 1.37 | 21.41 | 68.7 | — | |
| Open-MAGVIT2-I-PTResolution=Original Resolution, Quantizer Type=LFQ, Training Data=100M, Ratio=16, Codebook Size=163842024.09 | 1.39 | 21.74 | 56 | — | |
| Open-MAGVIT2-I-PTResolution=Resize 256 x 256, Quantizer Type=LFQ, Training Data=100M, Ratio=16, Codebook Size=2621442024.09 | 1.67 | 22.7 | 64 | — | |
| VILA-UResolution=256x256, Downsample Ratio=16, Residual Levels=42024.12 | 1.8 | — | — | — | |
| RQ-VAELatent Size=8x8x16, Num Z=163842022.09 | 1.83 | — | — | — | |
| CosmosResolution=Original Resolution, Quantizer Type=FSQ, Ratio=16, Codebook Size=640002024.09 | 1.93 | 20.56 | 51 | — | |
| VARResolution=384x384, Downsample Ratio=16, Residual Levels=132024.12 | 2.09 | 22.73 | 77.4 | — | |
| LlamaGenResolution=256x256, Downsample Ratio=16, Residual Levels=12024.12 | 2.19 | 20.79 | 67.5 | — | |
| LlamaGenResolution=Resize 256 x 256, Quantizer Type=VQ, Training Data=70M, Ratio=16, Codebook Size=16384, Initial Training=Pre-trained on ImageNet2024.09 | 2.47 | 20.65 | 54 | — | |
| CosmosResolution=Original Resolution, Quantizer Type=FSQ, Ratio=16, Codebook Size=64000, Source=Reported by Cosmos2024.09 | 2.52 | 20.49 | 52 | — | |
| Open-MAGVIT2-I-PTResolution=Resize 256 x 256, Quantizer Type=LFQ, Training Data=100M, Ratio=16, Codebook Size=163842024.09 | 2.55 | 22.21 | 62 | — | |
| RQ-VAEResolution=256x256, Downsample Ratio=32, Residual Levels=42024.12 | 3.2 | — | — | — | |
| Show-oResolution=Resize 256 x 256, Quantizer Type=LFQ, Training Data=35M, Ratio=16, Codebook Size=81922024.09 | 3.5 | 21.34 | 59 | — | |
| VQGANLatent Size=16x16, Num Z=163842022.09 | 3.64 | 19.93 | 54.24 | 0.1766 | |
| CosmosResolution=Resize 256 x 256, Quantizer Type=FSQ, Ratio=16, Codebook Size=640002024.09 | 4.57 | 19.93 | 49 | — | |
| VQ-GANResolution=256x256, Downsample Ratio=16, Residual Levels=12024.12 | 4.98 | 20 | 62.9 | — | |
| VQGANLatent Size=16x16, Num Z=10242022.09 | 6.25 | 19.47 | 52.14 | 0.195 |