Source-linked AI summary

A very preliminary analysis of DALL-E 2

Gary Marcus, Ernest Davis, Scott Aaronson

arXiv:2204.13807v2cs.CVcs.AI

TL;DR

The paper probes suspected weaknesses in DALL-E 2 using deliberately challenging captions and default image-generation settings. Across examples, the system often captured broad scene elements but missed specified relations, commonsense constraints, or parts of complex descriptions.

  • Problem

    The paper examines suspected weaknesses in DALL-E 2's ability to handle challenging captions involving commonsense and complex descriptions.

  • Method

    The authors designed challenging examples, screened proposed captions with Google image search, and generated ten images per caption at default settings.

  • Results

    Across the examples, DALL-E 2 sometimes captured broad scene elements but frequently missed specified spatial relations, commonsense constraints, or descriptive details.

  • Takeaways & Limitations

    The examples indicate that plausible image generation can coexist with failures to satisfy precise relational and commonsense requirements.

  • Takeaways & Limitations

    The authors lacked specific knowledge of DALL-E 2's internals, training set, or abilities beyond published information and circulated examples.

Abstract

from arXiv · show

The DALL-E 2 system generates original synthetic images corresponding to an input text as caption. We report here on the outcome of fourteen tests of this system designed to assess its common sense, reasoning and ability to understand complex texts. All of our prompts were intentionally much more challenging than the typical ones that have been showcased in recent weeks. Nevertheless, for 5 out of the 14 prompts, at least one of the ten images fully satisfied our requests. On the other hand, on no prompt did all of the ten images satisfy our requests.

Methods

The study tested DALL-E 2 with fourteen deliberately challenging captions, generating ten images per caption at default settings. The examples probed complex relations, commonsense reasoning, spatial specifications, and stylistic or compositional constraints.

  • Methods: Ten images were generated for each caption at DALL-E 2's default settings.Examples 6 and 12 were rerun after caption changes caused by technical issues.
  • Methods: The authors designed examples to probe suspected weaknesses and screened proposed captions with Google Images before modifying or rejecting some.They also lacked specific knowledge of DALL-E 2's internals, training set, and abilities beyond published material and circulated examples.
  • Complex specifications: Complex spatial specifications were often missed: none of the red-ball, blue-pyramid, car, and toaster images was correct.This example tested several nested object relations in one caption.
  • Complex specifications: The system often included named objects but failed their requested relations, as with the donkey, octopus, rope, and cat scenario.All images contained the caption's main components, but none got more than one stated relation right.
  • Spatial specifications: Viewpoint specifications were handled relatively well, but object placement remained unreliable in the tomato-and-pumpkin scene and the elephant-behind-tree scene.The requested viewpoint was consistently achieved for the tomato example, while only one elephant image showed the specified position.
  • Prompt variations: Five of ten oil-painting images of the couple in a downpour were exactly correct, while the other five incorrectly included an umbrella.A photographic version produced six conforming images, but raised a policy-related issue involving potentially recognizable people.
Loading 2204.13807v2…