Source-linked AI summary
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, Lijuan Wang
TL;DR
The paper asks how GPT-4V expands multimodal capabilities beyond existing LMM evaluations. It probes the model with curated qualitative samples across inputs, tasks, and prompting methods, finding broad capabilities and flexible interleaved multimodal processing while noting reliability and evaluation limits.
Problem
Existing benchmarks may not adequately assess LMMs such as GPT-4V because their task formulations and training-data assumptions can be unsuitable.
Method
The report organizes carefully designed qualitative samples spanning diverse domains, tasks, inputs, working modes, and prompting techniques.
Results
GPT-4V shows broad capabilities across many domains and supports arbitrary mixtures of multimodal inputs, while visual referring prompting offers a new image-based instruction method.
Takeaways & Limitations
The explorations provide a reference for emerging applications, next-generation multimodal task formulations, and future GPT-4V-based systems.
Takeaways & Limitations
The showcased capabilities may require carefully designed prompts and may not work consistently across different samples.
Abstract
from arXiv · showhide
Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the understanding of LMMs. The analysis focuses on the intriguing tasks that GPT-4V can perform, containing test samples to probe the quality and genericity of GPT-4V's capabilities, its supported inputs and working modes, and the effective ways to prompt the model. In our approach to exploring GPT-4V, we curate and organize a collection of carefully designed qualitative samples spanning a variety of domains and tasks. Observations from these samples demonstrate that GPT-4V's unprecedented ability in processing arbitrarily interleaved multimodal inputs and the genericity of its capabilities together make GPT-4V a powerful multimodal generalist system. Furthermore, GPT-4V's unique capability of understanding visual markers drawn on input images can give rise to new human-computer interaction methods such as visual referring prompting. We conclude the report with in-depth discussions on the emerging application scenarios and the future research directions for GPT-4V-based systems. We hope that this preliminary exploration will inspire future research on the next-generation multimodal task formulation, new ways to exploit and enhance LMMs to solve real-world problems, and gaining better understanding of multimodal foundation models. Finally, we acknowledge that the model under our study is solely the product of OpenAI's innovative work, and they should be fully credited for its development. Please see the GPT-4V contributions paper for the authorship and credit attribution: https://cdn.openai.com/contributions/gpt-4v.pdf
1 Introduction
The report explores GPT-4V through qualitative samples designed to assess its multimodal inputs, capabilities across tasks, prompting methods, and future applications. It emphasizes broad capability coverage while noting limitations in benchmark suitability and sample reliability.
- Supported inputs and working modes: GPT-4V processes arbitrary mixtures of images, sub-images, text, scene text, and visual pointers while supporting instruction following, chain-of-thought, and few-shot learning.
- Capabilities across domains and tasks: GPT-4V demonstrates impressive human-level capabilities across a wide range of domains, including visual understanding, document reasoning, coding, temporal reasoning, and abstract reasoning.
- Prompting methods: Visual referring prompting uses visual pointers and scene text drawn on input images to provide instructions and example demonstrations.
- Future directions: The report organizes qualitative explorations to inspire new multimodal applications, task formulations, benchmarks, and GPT-4V-based systems.
- Evaluation limitations: Existing benchmarks may inadequately evaluate LMMs because their captions can be less detailed than model outputs and GPT-4V’s pretraining data are not publicly known.
- Evaluation limitations: The showcased capabilities may depend on carefully designed prompts and may not work consistently across different samples.
2 GPT-4V’s Input Modes
GPT-4V accepts text-only inputs, single image-text pairs, and flexibly interleaved image-text inputs. Interleaved inputs support multimodal tasks that combine information across multiple images and texts.
- Text-only inputs: GPT-4V operates as a unimodal language model with text-only inputs and performs language and coding tasks.
- Single image-text pair: GPT-4V accepts a single image-text pair or a single image for vision and vision-language tasks such as recognition, localization, captioning, and question answering.
- Interleaved image-text inputs: Interleaved inputs can be visually centric, text-centric, or balanced mixtures of images and text.
- Interleaved image-text inputs: Interleaved inputs let GPT-4V associate information across images and text, including aggregating receipt taxes or calculating menu-based beer costs.
- Interleaved image-text inputs: Interleaved image-text processing supports in-context few-shot learning and other advanced test-time prompting techniques.
3 GPT-4V’s Working Modes and Prompting Techniques
GPT-4V supports diverse working modes and prompting techniques, including text instruction following, visual referring prompting, integrated multimodal inputs, and in-context few-shot learning. These methods provide flexible ways to format tasks, demonstrate desired behavior, and adapt to novel multimodal problems.
- 3.1 Following Text Instructions: GPT-4V follows text instructions to define desired outputs and perform arbitrary vision-language tasks.Instructions can shape both the requested output and the task performed by the model.
- 3.1 Following Text Instructions: Constrained prompting preserves a specified response format, such as JSON, even when information extraction contains mistakes.The paper demonstrates this approach on driver’s-license information extraction.
- 3.1 Following Text Instructions: Explicitly conditioning on good performance can improve GPT-4V’s counting responses compared with a basic counting instruction.The authors compare different text instructions for counting apples.
- 3.2 Visual Pointing and Visual Referring Prompting: GPT-4V understands visual pointers drawn on images, including arrows, boxes, circles, hand drawings, coordinates, and image crops.Visual pointing provides a channel for referring to arbitrary spatial regions of interest.
- 3.2 Visual Pointing and Visual Referring Prompting: Visual referring prompting edits image pixels with pointers or scene text to focus descriptions, associate objects with indices, or pose localized questions.These image edits can replace conventional text prompts for specific tasks.
- 3.3 Visual + Text Prompting: GPT-4V integrates arbitrary mixes of images, sub-images, texts, scene texts, and visual pointers as instructions, examples, or queries.This flexibility supports multimodal example-grounded instruction and adaptation to unseen tasks.
- 3.3 Visual + Text Prompting: Multimodal example-grounded instruction provides task demonstrations that are more directly related to visual queries than purely textual instructions.The paper contrasts this approach with instruction following and conventional in-context few-shot learning.
- 3.4 In-context Few-shot Learning: In-context few-shot learning can improve performance without parameter updates, but this report limits its use to avoid information leakage and leaves quantitative gains for future work.A two-shot line-plot example reaches the correct answer after an additional in-context example.
4 Vision-Language Capability
GPT-4V demonstrates broad vision-language capabilities across recognition, reasoning, localization, captioning, and document understanding. The qualitative examples show strong performance on diverse tasks, alongside limitations in localization accuracy and occasional document-detail omissions.
- Recognition: GPT-4V accurately recognizes celebrities, dishes, scene text, and multiple car-brand logos across varied visual inputs.It also identifies ingredients, garnishes, cooking techniques, and both handwritten and printed scene text.
- Specialized understanding: GPT-4V shows basic understanding of medical images, identifying conditions such as a Jones fracture and flagging potential concerns in lung CT scans.The report presents these observations as preliminary medical-image capabilities.
- Reasoning: GPT-4V handles multimodal reasoning tasks, including counterfactual questions, spatial relationships, visual mathematics, charts, and commonsense interpretation.Examples include identifying a right triangle and dimensions, explaining chart processes, and reasoning about people, objects, and camera perspective.
- Localization and captioning: GPT-4V generates bounding-box coordinates for object localization, but the coordinates are inaccurate and performance degrades in complex, crowded scenes.More promising results appear when the scene or background is relatively simple; further prompting is needed.
- Localization and captioning: GPT-4V produces encouraging dense-captioning results for detailed descriptions of image regions.The explored task ordinarily requires a complex system integrating components such as object detection, celebrity recognition, and image captioning.
- Document understanding: GPT-4V provides reasonable understanding of floor plans, posters, exams, and multi-page technical reports.It can identify locations, connect scene text with geographic knowledge, and describe a report’s main idea and proposed method across pages, though it may miss implementation details.
5 Interaction with Humans: Visual Referring Prompting
GPT-4V understands visual markers and scene text overlaid on images, enabling visual referring prompting as a nuanced human-computer interaction method. It can also generate coordinate-based pointers for a closed interaction loop, although spatial precision remains limited.
- 5 Interaction with Humans: Visual Referring Prompting: Visual pointers provide an intuitive interaction channel that both humans and machines can generate and understand.The report frames visual referring prompting as a method for seamless interaction with computer and mobile GUIs, documents, and slides.
- 5.1 Understand Pointing Inputs: GPT-4V understands circles, boxes, and hand-drawn markers overlaid on images, preserving global context while producing grounded descriptions of pointed regions.Visual pointers avoid the loss of global context associated with cropped boxes or mask regions.
- 5.1 Understand Pointing Inputs: GPT-4V can interpret text-format region coordinates without box-token finetuning, but the current prompting approach is less spatially precise than visual pointers.Example coordinates identify objects such as a beer bottle, string lights, and a table set, with lower reliability than visual pointers.
- 5.2 Visual Referring Prompting: Visual referring prompting edits image pixels by adding visual pointers or scene text as complementary instructions to conventional text prompts.The method supports object indexing, image-pointed questions, document and chart highlighting, and visual pattern specification.
- 5.2 Visual Referring Prompting: Visual referring prompting can associate pointed objects with indexes, answer questions attached to image regions, guide document reasoning, and represent patterns concisely.These examples illustrate interaction through arrows, scene text, and arbitrary pointed regions.
- 5.3 Generate Pointing Outputs: The section asks whether model-generated pointers can support a closed-loop interaction process between GPT-4V and humans.This extends human-provided visual referring prompts toward model-generated visual outputs.
- 5.3 Generate Pointing Outputs: GPT-4V can generate region-coordinate outputs to ground objects referred to by text or reference images, but the resulting locations are coarse and inaccurate.The model is prompted to ground objects such as a blue Subaru SUV or a black Audi sedan.
- 5.3 Generate Pointing Outputs: Generated pointers remain useful for human interpretation and multi-step reasoning even when they do not perfectly cover the queried region.GPT-4V can interpret pointers that it generates and provide grounded descriptions from them.
6 Temporal and Video Understanding
GPT-4V analyzes sequences of video frames by interpreting scenes, poses, temporal order, anticipated events, and interactions. Visual pointing further grounds temporal understanding in a person of interest and the social nature of events.
- 6.1 Multi-image Sequencing: GPT-4V comprehends video-frame sequences by recognizing scenes, interpreting human poses, and relating movements to ongoing activities.The analysis goes beyond object and scene identification to derive meaning from pose variations and actions.
- 6.2 Video Understanding: GPT-4V can reorder shuffled long-term and short-term image sequences according to temporal progressions and cause-and-effect relationships.Examples include a sushi-making event and specified actions such as opening or closing a door.
- 6.2 Video Understanding: GPT-4V anticipates both short-term and long-term future events from initial frames, including soccer actions and subsequent sushi-preparation steps.The examples involve activities with different temporal structures and complexities.
- 6.2 Video Understanding: GPT-4V localizes when a player strikes a ball and reasons about whether a goalkeeper blocks it from their interaction.The latter requires interpreting spatial positions, interaction dynamics, and the resulting outcome.
- 6.3 Visual Referring Prompting for Grounded Temporal Understanding: The report extends visual referring prompting to temporal understanding to provide enhanced control over video comprehension tasks.This connects spatially grounded prompting with analysis of image-frame sequences.
- 6.3 Visual Referring Prompting for Grounded Temporal Understanding: Pointing inputs let GPT-4V focus temporal descriptions on a circled person while preserving event order and interpreting interaction tone.The examples distinguish friendly interactions from violent incidents.
7 Abstract Visual Reasoning and Intelligence Quotient Test
GPT-4V is examined on abstract visual interpretation, part-object association, and human intelligence tests. The reported examples show interpretation of ambiguous shapes and performance on visual and verbal reasoning tasks.
- 7 Abstract Visual Reasoning and Intelligence Quotient Test: The section evaluates whether GPT-4V can abstract semantics from visual signals and perform different types of human intelligence tests.The evaluation treats abstract visual reasoning as a fundamental ability associated with human intelligence.
- 7.1 Abstract Visual Understanding: GPT-4V interprets abstract visual stimuli such as tangrams and provides semantic descriptions and reasoning for candidate shapes.It identifies a sub-figure as best illustrating a flying goose and describes others as a person or robot, or a boat or hat.
- 7.2 Association of Parts and Objects: GPT-4V associates object parts with semantically meaningful objects and processes both natural images and parts segmented by SAM.The probes ask it to localize parts by meaning and combine segmented parts.
- 7.3 Wechsler Adult Intelligence Scale: The report tests GPT-4V on abstract reasoning tasks from the Wechsler Adult Intelligence Scale, described as a gold-standard IQ test with multiple sub-tests.The examples assess cognitive abilities through representative questions.
- 7.4 Raven's Progressive Matrices: Raven’s Progressive Matrices tests abstract reasoning and problem-solving using image matrices with one missing figure and multiple candidate answers.The test is designed to reduce the influence of language, culture, and formal education.
- 7.4 Raven's Progressive Matrices: The report presents IQ-test questions either as a complete page or as multiple sub-figures with optional instructions and examples.The latter processing strategy is described as a way to further boost answer accuracy.
8 Emotional Quotient Test
GPT-4V interprets emotions from facial expressions and visual content, judges image aesthetics, and generates text conditioned on desired emotional effects. These capabilities are discussed in relation to emotion-aware interaction.
- 8 Emotional Quotient Test: The report frames emotion understanding as involving facial-expression reading, visual sentiment analysis, and emotion-conditioned text generation.These three capabilities define the EQ-oriented examination in the section.
- 8.1 Reading Human Emotions: GPT-4V reliably identifies emotions from facial expressions and provides rationales based on observed visual cues.The reported examples indicate understanding of facial emotions rather than only emotion labels.
- 8.2 Understand How Visual Content Arouses Emotions: GPT-4V interprets visual sentiments such as contentment, anger, awe, and fear using both semantic content and image style.The report connects this capability to anticipating emotional responses to visual contents.
- 8.2 Understand How Visual Content Arouses Emotions: GPT-4V judges image aesthetics in alignment with societal standards and norms.The section presents aesthetic judgment as another form of alignment with human subjective judgments.
- 8.3 Emotion Conditioned Output: GPT-4V generates text conditioned on perceived or desired emotions, including descriptions made more horrifying or comforting.The examples use humorous, uneasy, anxious, scary, and comforting communication goals.
9 Emerging Application Highlights
GPT-4V is explored across emerging applications including defect detection, safety inspection, and grocery checkout. Results show strong performance in some scenarios, but reliability depends on product familiarity, prompting, and task decomposition.
- 9.1 Spot the Difference: GPT-4V compares visually similar images to identify differences, but may misdescribe the precise nature of those differences.It identifies several differing regions while confusing the number of hairband cuts with hair shade.
- 9.2 Industry: GPT-4V confidently detects defects in familiar products but may hesitate, refuse, or overlook major defects in uncommon or variable products.In one tire example, it mentions dirt while missing damage to the rim requiring repair.
- 9.2 Industry: Adding a defect-free reference image and refining the prompt enables GPT-4V to identify defects in all three single-image failure cases.
- 9.2 Industry: GPT-4V fails to count three people without helmets in the full image, while cropped person regions divide detection and PPE assessment into separate steps.The cropped-input approach relies on an external person detector and GPT-4V for visual reasoning about safety issues.
- 9.2 Industry: GPT-4V misidentifies several grocery items in a basket, but catalog reference images enable it to identify all five items and support price retrieval.
9.3 Medical
The report examines GPT-4V for radiology reporting and auto-insurance applications using domain-aware evaluation and structured prompts. Examples show useful identification and reporting abilities alongside clinically important errors and image-availability constraints.
- 9.3 Medical: Medical professionals evaluate GPT-4V-generated radiology reports because assessing their accuracy requires domain knowledge.
- 9.3 Medical: GPT-4V produces accurate reports for some abdominal X-ray and knee MRI cases, but misses a distal radial fracture in a hand X-ray.The reports retain a high-quality format that can serve as a drafting template for medical professionals.
- 9.3 Medical: GPT-4V can mislocalize findings and hallucinate measurements in radiology reports, while interleaved image-text inputs support use of prior scans and diagnosis histories.
- 9.4 Auto Insurance: For auto-damage evaluation, GPT-4V identifies and localizes damage across four images, describes each instance, and sometimes estimates repair cost.
- 9.4 Auto Insurance: For insurance reporting, GPT-4V attempts to return vehicle details in JSON, but license plates can be difficult to discern under occlusion.Real-world reporting typically uses multiple vehicle views, which are often unavailable in public examples.
9.5 Customized Captioner
GPT-4V supports personalized photo organization and dense captioning by using visual references, names, segmentation cut-outs, and global image context. These inputs produce captions that identify people and richly describe objects with contextual references.
- Customized Captioner: Visual prompts paired with family-member names enable GPT-4V to generate captions that explicitly identify people in family photos.The approach is presented for more precise and tailored photo organization.
- Customized Captioner: Providing SAM-generated object cut-outs alongside the original image enables highly intricate captions for individual objects.
- Customized Captioner: Dense captions can reference context-image details absent from an object cut-out, such as a snail on a frog or a turtle floating in the scene.
9.6 Image Generation
The report uses GPT-4V to evaluate generated images and improve image-editing prompts. The examples show prompt-alignment scoring, OCR-based text comparison, and iterative prompt refinement for editing.
- Image Generation: GPT-4V rates generated-image similarity to a text prompt from 1 to 10, assigning 1 to an irrelevant image and 9 to the most relevant example.The examples include progressively improved generations from RL-Diffusion.
- Image Generation: GPT-4V uses OCR to recognize rendered text in generated images and compare it with the requested text prompt.The example compares variants such as “Azuze Research,” “ARAUIE,” and “Azure Azure” with “Azure Research.”
- Image Generation: GPT-4V can generate an editing prompt tailored to an original image and textual editing requirement.
- Image Generation: GPT-4V can rewrite an editing prompt using the original image, initial prompt, and edited image, enabling repeated refinement cycles.
9.7 Embodied Agent
GPT-4V is explored as an embodied agent operating household appliances and navigating a virtual house. The examples show that restructuring visual instructions can correct failures and that iterative action prediction reaches a task goal.
- Operating machine: GPT-4V identifies the coffee machine button for “8 OZ coffee” from a menu image, but initially confuses the “6 OZ coffee” option with the power button.The failure is attributed to visual confusion caused by the option’s positioning on the menu and machine.
- Operating machine: Presenting isolated operating-menu images enables GPT-4V to recognize the precise “6 OZ coffee” button after the full-menu approach fails.
- Navigation: GPT-4V predicts successive actions in a virtual house tour to reach the kitchen and retrieve an item from the fridge.The task uses screenshots from a Redfin virtual house tour and manually executes each predicted action before the next turn.
- Navigation: Within the third turn, GPT-4V reaches the fridge and predicts moving into alignment, opening the door, and retrieving the requested item.
9.8 GUI Navigation
GPT-4V is tested as an agent navigating computer and smartphone GUIs for web browsing, recipe retrieval, news reading, and online shopping. The examples show reasonable task-directed actions alongside some inaccurate predictions and descriptions.
- Computer GUI navigation: GPT-4V predicts GUI actions from a current screenshot, an end goal, and a list of possible mouse, click, and keyboard actions.The setup covers task-oriented computer navigation such as finding recipes and reading today’s news.
- Web browsing: GPT-4V completes a Mapo Tofu recipe search and print task, then recognizes details in the printed recipe, including cooking time, ingredients, author, and link.
- Online shopping: GPT-4V predicts smartphone actions to shop for an ergonomic keyboard within a $50–$100 budget, including opening the Amazon app.
- Notification understanding: GPT-4V interprets computer notifications, including proposing Maps for a Seattle meeting and handling call and message notifications.
- Watching videos: GPT-4V describes videos from screenshot sequences with or without subtitles, suggesting potential for automatic transcript generation for user-generated videos.
- Observed errors: The GUI examples also contain inaccurate predictions, including an incorrect tab-closing action, an inaccurate Amazon-icon location, and an inaccurate description of a printed recipe image.
10 LMM Powered Agents
The report discusses multimodal extensions of agent techniques, including plugins, chains, self-reflection, self-consistency, and retrieval augmentation. Examples suggest these mechanisms can provide multimodal information, correct outputs, and aggregate repeated answers.
- Multimodal Chains: Existing plugin chains typically exchange text, while generated images are generally not fed back into language models for further analysis.
- Multimodal Chains: Multimodal chains extend ReAct-style systems so plugins provide multimodal information that GPT-4V collectively processes for reasoning tasks such as PPE counting.The illustrated process uses two rounds of thought, action, and observation, each activating a specific plugin.
- Self-Reflection: Self-reflection improves generated figures by correcting the number of data points and restoring a percentage, although the result remains not exactly identical to the reference.
- Self-Reflection: Self-reflection can also revise a text-to-image prompt after GPT-4V identifies that the dog’s breed was omitted.
- Self-Consistency: Self-consistency aggregates multiple sampled counting outputs, using majority vote to produce the final answer “4 boats.”Samples are obtained through repeated runs or rephrased input instructions on the same image.
- Retrieval-Augmented LMMs: Retrieval-augmented LMMs are presented as a way to integrate specialized, recent, or user-customized information into prompts.
11 Conclusions
The report presents GPT-4V explorations as a reference for discovering additional uses and informing future research. It points toward richer multimodal generation, broader data sources, and continued attention to model limitations.
- The report probes GPT-4V across application scenarios and aims to support future research exploring additional uses and deeper understanding of its capabilities.
- Scope: The report focuses on future research directions because GPT models’ weaknesses and limitations have been discussed extensively elsewhere.
- Towards Future LMMs: Future LMMs could generate interleaved image-text content and incorporate video, audio, and other sensor data.
- Towards Future LMMs: More versatile models could learn from online web content and real-world physical environments to support continuous self-evolution.