Source-linked AI summary
JourneyDB: A Benchmark for Generative Image Understanding
Keqiang Sun, Junting Pan, Yuying Ge, Hao Li, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang, Jifeng Dai, Yu Qiao, Limin Wang, Hongsheng Li
TL;DR
Vision-language models are mainly trained on real data, leaving their ability to understand diverse generated images uncertain. JourneyDB addresses this gap with a large paired dataset and four content- and style-oriented benchmarks. Evaluations find weaker performance than on real datasets, while fine-tuning on JourneyDB significantly improves performance.
Problem
Models primarily pretrained on real data may not fully understand the distinctive content and style of generated images.
Method
The paper introduces JourneyDB, a 4-million-image prompt-paired dataset with benchmarks for prompt inversion, style retrieval, captioning, and visual question answering.
Results
State-of-the-art multimodal models perform less effectively on generated images than on real datasets, while fine-tuning on JourneyDB significantly enhances performance.
Takeaways & Limitations
JourneyDB provides a training and evaluation resource for advancing visual understanding of generated content.
Abstract
from arXiv · showhide
While recent advancements in vision-language models have had a transformative impact on multi-modal comprehension, the extent to which these models possess the ability to comprehend generated images remains uncertain. Synthetic images, in comparison to real data, encompass a higher level of diversity in terms of both content and style, thereby presenting significant challenges for the models to fully grasp. In light of this challenge, we introduce a comprehensive dataset, referred to as JourneyDB, that caters to the domain of generative images within the context of multi-modal visual understanding. Our meticulously curated dataset comprises 4 million distinct and high-quality generated images, each paired with the corresponding text prompts that were employed in their creation. Furthermore, we additionally introduce an external subset with results of another 22 text-to-image generative models, which makes JourneyDB a comprehensive benchmark for evaluating the comprehension of generated images. On our dataset, we have devised four benchmarks to assess the performance of generated image comprehension in relation to both content and style interpretation. These benchmarks encompass prompt inversion, style retrieval, image captioning, and visual question answering. Lastly, we evaluate the performance of state-of-the-art multi-modal models when applied to the JourneyDB dataset, providing a comprehensive analysis of their strengths and limitations in comprehending generated content. We anticipate that the proposed dataset and benchmarks will facilitate further research in the field of generative content understanding. The dataset is publicly available at https://journeydb.github.io.
1 Introduction
JourneyDB addresses the uncertain ability of vision-language models to understand generated images, whose content and styles are diverse and often fictional. It provides a large benchmark with paired prompts and four tasks, then evaluates existing models and fine-tuning.
- Dataset and benchmark: The authors collect Midjourney images and use GPT-3.5 to annotate prompt components, captions, and content- and style-relevant questions.The annotations include four answer options and the correct answer for each question.
- Motivation: Generated images combine intricate style descriptions with fictional scenes and compositions, challenging models mainly pretrained on real data.Prompts can describe lighting, camera angle, artistic style, and medium, while generated content may lack real-world counterparts.
- Dataset and benchmark: JourneyDB contains 4 million generated images paired with text prompts for comprehensive generative-content understanding evaluation.The dataset is designed as a benchmark and training resource.
- Dataset and benchmark: The benchmark covers prompt inversion, style retrieval, image captioning, and visual question answering.These tasks assess content and style interpretation from complementary perspectives.
- Evaluation: State-of-the-art models perform less effectively on generated images than on real datasets, while fine-tuning on JourneyDB significantly improves performance.The paper evaluates current multimodal models and analyzes their strengths and limitations.
2 Related Works
JourneyDB builds on image-text and multimodal research while targeting generated-image understanding with model-generated annotations. Its scale and task coverage distinguish it from commonly used datasets.
- Image-text datasets: Existing image-text datasets support captioning and visual question answering, but commonly rely on costly human annotation and focus on real images.Examples include Flickr Caption, COCO Caption, VQA v2.0, and A-OKVQA.
- JourneyDB: JourneyDB collects Midjourney image-prompt pairs and uses GPT-3.5 annotations to support four downstream visual understanding tasks.The dataset is presented as a versatile resource for generated-image understanding.
- Text-to-image generation: Text-to-image models generate images from natural-language conditions, enabling image creation through textual specifications.This work situates JourneyDB within the rapid development of text-to-image generation.
- Multimodal models: Multimodal foundation models connect image and language modalities through large-scale pretraining and contrastive or alignment-based methods.The related work discusses CLIP, ALIGN, Flamingo, and BLIP-2.
- JourneyDB: JourneyDB is intended to enhance image-related tasks and advance understanding of generated content.The paper presents this as a supported consequence of the demonstrated results.
3 Dataset
JourneyDB is assembled from publicly accessible Midjourney prompt-image pairs, expanded with images from 22 other models, and annotated for four visual understanding tasks. Style prompts are hierarchically clustered, while test pairs undergo consistency filtering.
- Data collection: The collection uses publicly accessible Midjourney Discord chat history, where users submit prompts and receive selected upscaled images.These prompt-image pairs form the primary source of JourneyDB.
- Data collection: Twenty-two additional text-to-image models are included to increase dataset diversity and create a cross-model test set.Examples include VQ-Diffusion, DALL·E 2, and StableDiffusion-XL.
- Data annotation: GPT-3.5 annotations segment prompts into Style, Content, Atmosphere, and Others, then generate captions and content- and style-relevant multiple-choice questions.Each question is accompanied by four answer choices.
- Style organization: A hierarchical clustering approach organizes intricate style prompts into a style tree to simplify style retrieval.GPT-3.5 clusters prompt patches, after which categories are manually merged.
- Style organization: The style space distribution and samples are visualized in Figure 2.The figure presents the organized style-prompt space.
- Quality control: Human annotators remove prompt words that do not appear in corresponding test images to improve image-prompt consistency.The consistency filtering addresses errors from imperfect text-to-image generation.
- Dataset statistics: The dataset contains 4,692,751 collected image-prompt pairs, including 4,189,737 training images, 234,156 validation images, and a manually filtered test set.The test set contains 5,402 images and 5,171 prompts after filtering.
4 Benchmarks
JourneyDB defines four benchmarks for understanding generated images across prompt content, style, captions, and visual questions. Evaluations show that existing models struggle with generated content, while fine-tuning on JourneyDB improves prompt inversion.
- 4.1 Prompt Inversion: Prompt inversion predicts the text prompts used to generate an image, testing comprehension of both content and style.The benchmark extends BLEU, METEOR, ROUGE, CIDEr, and sentence-transformer cosine similarity metrics.
- 4.1 Prompt Inversion: Existing models struggle to capture intricate details and style information in generated images, producing lower prompt-inversion performance than on conventional datasets.The evaluated zero-shot models include BLIP-2, Flamingo9B, MiniGPT-4, and Uni-Perceiver v2.
- 4.1 Prompt Inversion: Fine-tuning Uni-Perceiver v2 for 20 epochs significantly improves prompt inversion without hyperparameter tuning or data augmentation.The result supports JourneyDB as a complement to existing image-text datasets, although robust prompt inversion remains difficult.
- 4.2 Image Caption: JourneyDB captioning combines detailed descriptions with high-level summaries, testing fine-grained recognition and holistic understanding.Its generated-image captions differ from natural-image captions in length and in concepts such as emotions and human or object attributes.
- 4.2 Image Caption: Existing captioning models miss key concepts in generated images and may hallucinate objects or text that are absent.Examples include missing children in astronaut suits or a sad potato, while Open-Flamingo invents content in some cases.
- 4.3 Style Retrieval: Style retrieval clusters style prompts into 344 categories before using CLIP for zero-shot retrieval evaluation.The categories include camera parameters, lighting, artist style, and colour schemes, narrowing retrieval within the broad style space.
- 4.4 Visual Question Answering (VQA): Multiple-choice visual question answering evaluates content-relevant and style-relevant understanding of generated images.The benchmark uses GPT-3.5-generated questions to assess both visual content and stylistic attributes.
- 4.4 Visual Question Answering (VQA): Existing multimodal models perform unsatisfactorily on both VQA tasks, with BLIP-2 below 70% accuracy despite outperforming Flamingo9B and MiniGPT-4.The authors attribute difficulty to generated scenes and object compositions that are absent from reality, such as a tree growing out of a piano.
5 Conclusion
JourneyDB is introduced as a four-task benchmark intended to advance understanding of generative content.
- JourneyDB provides four downstream tasks for advancing comprehension of generated images.
A Data samples
The dataset samples show generated images alongside prompts, captions, and content- and style-focused visual question-answering annotations.
- Randomly sampled Midjourney instances are displayed with their corresponding prompts and captions.
- The samples also include style-relative and content-relative VQA annotations, color-coded orange and green respectively.
B Details of Data Annotation
JourneyDB annotates image-prompt consistency, visual understanding tasks, and style clusters using human and GPT-3.5-assisted procedures.
- Image-Prompt Consistency Filtering: Forty professional annotators identify prompt words or phrases that are absent from, or inconsistent with, the corresponding image.
- Visual Understanding Annotation: GPT-3.5 separates prompt terms into Style, Content, Atmosphere, and Other categories for downstream annotations.
- Visual Understanding Annotation: GPT-3.5 generates content-based captions and definite multiple-choice questions with answers and distractors from the categorized prompts.
- Style Clustering: Style prompts are organized into hierarchical categories and subcategories, allowing one prompt to belong to multiple categories.
C Additional Experiments
The paper introduces Question Answering Score to evaluate prompt inversion through interpretable question-answering accuracy rather than direct prompt similarity.
- Question Answering Score (QAS): Question Answering Score evaluates whether an inverted prompt supports correct answers to style- and content-related questions about the generated image.
- Question Answering Score (QAS): QAS separately averages accuracy over N style questions and M content questions, then averages across K images.
- Question Answering Score (QAS): The method converts prompt similarity into a more interpretable question-answering accuracy measure reported in Table 8.
- Figure 6 compares Stable Diffusion v1.4 images with JourneyDB images using HPS, pairing images in each row generated from the same prompt.
C.2 Analysis of Image Quality
JourneyDB combines high-quality generated images with broad benchmark resources and compares image quality across generation models. Its evaluation includes prompt inversion and image captioning on an extension test set.
- Images generated by Midjourney exhibit better visual quality than Stable Diffusion images according to Human Preference Score.
- JourneyDB provides 4 million generated image-prompt pairs, 1 million captions, and over 8 million VQA annotations.
- The extension test set evaluates prompt inversion and image captioning.
D Cross-model Test Set
JourneyDB extends its evaluation across 22 additional text-to-image models, producing a manually cleaned cross-model test set. This expansion increases model diversity and supports evaluation of generated-image visual understanding.
- 45,803 images form the final cross-model test set after annotators removed inconsistent image-prompt pairs.Each model initially contributed 3,200 generated images, and 60 annotators helped clean the pairs.
- The extension includes results from 22 additional text-to-image generative models, including VQ-Diffusion, DALL·E 2, and StableDiffusion-XL.
- The manually cleaned extension contains text prompts, image captions, and VQA annotations for evaluating visual understanding models.