ResearchDatasetsAURORA-BENCHFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsVisual World ModellingAURORA-BENCH Average7.36GPT-4o Score18Instruction-guided image editing preference predictionAURORA-Bench63.62Accuracy12Action-centric EditingAURORA-BENCH All (test)-0.23Human Eval Score4