ResearchTasksinstance-aware caption generationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedMyVLM benchmarkIIR-VLM27.06CLIP Image Similarity4Feb 26, 2026