Source-linked AI summary

Visual General Intelligence: A White Paper

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, Robert Geirhos, Aditi Raghunathan, Yuki M. Asano, Deva Ramanan, David Fouhey, Andrew J. Davison, Yilun Du, Jiajun Wu, Zhuang Liu

arXiv:2608.25924v1cs.CV

TL;DR

The paper asks what capabilities can emerge from visual experience and learning as a pathway toward AGI. It synthesizes diverse perspectives on VGI, finding partial evidence across visual capabilities while emphasizing that their integration remains incomplete and evaluation is difficult in data-limited settings.

  • Problem

    The paper asks what forms of intelligence can emerge from visual modalities such as images, videos, and geometry, given that vision lacks agreed symbolic units and visual capabilities remain difficult to isolate from language.

  • Method

    The paper brings together diverse perspectives on visual modalities, benchmarks, learning paradigms, and the relationship between vision and language to examine pathways toward VGI.

  • Results

    Current systems provide partial evidence across emergent video capabilities, geometric reconstruction, persistent spatial representations, creative alternatives, and embodied interaction, but these abilities remain fragmented.

  • Takeaways & Limitations

    VGI should remain a broad capacity involving visual knowledge, world structure, alternatives, adaptation, prediction, discovery, and action rather than a single model or benchmark.

  • Takeaways & Limitations

    Scientific visual problems are constrained by limited data and the absence of convenient ground truth for validation.

Abstract

from arXiv · show

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.

1. Introduction

The paper asks what forms of generality can emerge from visual learning and presents VGI as an open, vision-centered research agenda rather than a settled concept.

  • Intelligence is framed as acquiring representations from environmental signals, capturing regularities, generalizing to unseen situations, and linking representations to prediction or action.
  • Vision may constitute a foundational intelligent function for capturing world structure, rather than merely serving as a sensor or input modality.
  • Unexpected abilities emerging from simple objectives, large-scale experience, and scaling provide a clue for investigating visual intelligence, although language-model success may not transfer directly.
  • A central open question is under what learning strategies, data structures, architectures, and evaluation settings visual representations can generalize.
  • Vision-language and multimodal systems connect visual features with language, but intertwined data, tuning, and reasoning make visual-origin capabilities difficult to isolate.
  • The paper treats VGI as an exploratory agenda that brings together diverse hypotheses and examines visual understanding, prediction, and generalization before or alongside language.

2. Perspectives on Visual General Intelligence

The paper presents multiple perspectives on VGI and emphasizes that the path from visual intelligence to general intelligence remains an open set of questions.

  • The paper brings together diverse research perspectives to clarify open questions that computer vision may explore toward general intelligence.
  • The views are not intended to provide a single definition of VGI, but to summarize multiple possibilities for vision research toward general intelligence.

2.1. Video models are visual foundation models (Robert Geirhos)

This perspective proposes large-scale generative video models as visual foundation models, arguing that they can unify fragmented visual tasks through broad capabilities learned without task-specific training.

  • 2.1. Video models are visual foundation models: Visual intelligence is characterized as solving visual tasks from perception to reasoning without explicit training on each new task.
  • 2.1. Video models are visual foundation models: Large-scale training and generative objectives are proposed as principles for transferring the language-model foundation-model paradigm to video models.
  • 2.1. Video models are visual foundation models: Video generation is presented as a general framework that includes still images, while recent capability growth makes video models candidates for visual foundation models.
  • 2.1. Video models are visual foundation models: Without task-specific training, Veo 3 performed image-to-video tasks including edge detection, segmentation, keypoint localization, super-resolution, editing, style transfer, maze solving, and graph traversal.
  • 2.1. Video models are visual foundation models: The proposed direction could unify much of computer vision’s fragmented landscape of task-specific models.

2.2. Creativity as a Test of Visual General Intelligence (Aditi Raghunathan)

The paper treats creativity as a test of visual general intelligence, requiring coherent, structurally diverse, and original alternatives produced through globally coordinated generation.

  • 2.2. Creativity as a Test of Visual General Intelligence: Open-ended visual intelligence includes creative tasks such as designing constrained objects, proposing robot plans, constructing diagrams, and imagining scene evolutions.
  • 2.2. Creativity as a Test of Visual General Intelligence: Creative visual outputs may require hidden plans coordinating scene layout, 3D structure, object relationships, mechanisms, camera trajectories, or actions.
  • 2.2. Creativity as a Test of Visual General Intelligence: Algorithmic creativity requires solutions that are coherent, distinct, and not simply reproduced from training examples.
  • 2.2. Creativity as a Test of Visual General Intelligence: Evaluation should measure coherence, structural diversity, and originality alongside task-relevant physical, geometric, semantic, or temporal constraints.
  • 2.2. Creativity as a Test of Visual General Intelligence: On controlled tasks, teacherless multitoken and diffusion approaches generated more diverse and original solutions than conventional next-token learning.
  • 2.2. Creativity as a Test of Visual General Intelligence: Visual models should construct multiple plausible alternatives, preserve each alternative’s internal logic, and search for useful unfamiliar possibilities.

2.3. A learning system that learns from one datum (Yuki M. Asano)

Visual General Intelligence is framed as a lifelong visual learning system that builds coherent world understanding from experience, rather than a fixed model trained once and reset or retrained for new needs.

  • Current visual systems recognize, reconstruct, answer questions, and follow instructions, but do not coherently reorganize their understanding as the world changes.
  • A VGI system should learn continuously from visual experience without labels, shuffled replay, or assuming future tasks in advance.
  • The vision-first starting point requires discovering objects, persistence, observer motion, independent scene changes, and long-term regularities before language describes them.
  • VGI concerns learning during a system’s own visual lifetime: recognizing novelty, deciding what to update, and retaining useful knowledge.
  • Rather than a larger Transformer alone, the proposed artificial visual brain combines perceptual, spatial-dynamic world-model, memory, and task-action systems with distinct roles and timescales.
  • The long-term objective is a visual system that becomes more competent and structured with experience without being reset, retrained from scratch, or told all future requirements.

2.4. Intelligence will be multimodal, generative, and efficient (Deva Ramanan et al.)

The section argues that visual intelligence should be multimodal, generative, and efficient, with video and generative modeling serving as central directions while computation remains a practical constraint.

  • Visual intelligence should integrate vision with language, audio, tactile sensing, and proprioception, especially for embodied robotics.
  • Generative modeling is presented as a route to visual representation learning because producing images or videos imposes harder demands than exploiting superficial recognition shortcuts.
  • Efficiency is necessary because video generation and inference remain costly, limiting serious efforts largely to frontier laboratories and potentially excluding academia.
  • Recurrence, 3D structure, and multiscale processing are proposed as classic representations that could improve generative-model efficiency, while long-form video requires consistent memory.

2.5. Visual Intelligence and Discovery (David Fouhey)

Visual intelligence for scientific discovery must operate under limited data, weak ground truth, imperfect simulations, and instrument systematics, making validation and domain integration central challenges.

  • Visual intelligence is defined as understanding and possibly generating visual data at human or better levels without unreasonable task-specific supervision.
  • Supervision scale is identified as separating current systems from VGI because strong task performance still requires substantial hand-holding.
  • Scientific applications often cannot collect more data, and models designed for larger datasets may fit small domain-specific datasets poorly.
  • Scientific validation is difficult because many fields lack convenient ground truth, shifting effort toward reconciling predictions with existing knowledge rather than leaderboard comparison.
  • Models can reproduce both physical signals and instrument artifacts, so scientific vision systems must disentangle them using source-data controls, other instruments, and domain knowledge.
  • Simulations are limited because frontier scientific reality is poorly understood, models are approximate or competing, and rendering often has clean solutions only for restricted problems.

2.6. What is the Computational Structure of Spatial AI? (Andrew J. Davison)

Spatial AI is presented as efficient, persistent, task-independent representation of surroundings that supports interaction, prediction, and planning, with computational structure shaped by continual learning and hardware constraints.

  • Spatial AI builds rich but efficient representations of surroundings so artificial devices can interact with them in generally intelligent ways.
  • A persistent yet adaptable representation supports prediction and planning, with visual SLAM exemplifying incremental world modeling for consistent long-horizon behavior.
  • Spatial AI inherits SLAM’s concerns with continual construction from multisensor streams, drift avoidance, real-time operation, and fixed computation and memory budgets.
  • Current localization and sparse mapping still rely mainly on hand-designed geometry and probabilistic state estimation, while dense and semantic SLAM balances these with machine learning.
  • Future Spatial AI may tightly integrate learned and designed components, making the distinction between neural networks and hand-designed algorithms less important than system-level integration.
  • Efficient future processors are expected to use fine-grained parallelism, distributed local memory, and sparse locality-respecting communication, favoring algorithms that map their graphs onto hardware.
  • Spatial AI implementations must provide efficient graph-structured interfaces between scene representations, real-time sensor processing, and actuator outputs.

2.7. Visual Intelligence is Embodied Intelligence (Yilun Du)

Visual intelligence is framed as embodied intelligence: robots should perceive, reason, act, explore, and continually revise visual models through interaction with the physical world.

  • Embodiment is a central test of visual general intelligence because robots must perceive, reason, and act reliably in the physical world.
  • Vision-guided action forms a closed loop in which consequences provide evidence for improving visual understanding.
  • Robust visual systems must generalize across unfamiliar viewpoints, environments, object configurations, and interactions.
  • Generative world models can explain unfamiliar scenes by inferring latent states and recombining geometric and compositional structure.
  • Persistent scene representations should integrate partial observations over time while supporting prediction, action, and targeted information gathering.
  • Active interaction resolves perceptual ambiguities, while continual adaptation incorporates unfamiliar experience after deployment.

2.8. Seeing the Physical World via Code (Shangzhe Wu & Jiajun Wu)

This section presents visual intelligence as recovering physical-world structure from visual observations, with code providing an executable and verifiable representation of scenes.

  • Visual intelligence can be viewed as recovering the physical world’s structure from images that record only viewpoint-dependent light.
  • Recovered structure supports counterfactual questions about unseen views, object manipulation, and changes to the scene.
  • Visual observations preserve physical information that language may omit, supporting vision as a source of world structure in its own right.
  • The proposed decomposition includes entities, object intrinsics, relations, dynamics, and other properties relevant to physical prediction.
  • Whether physical quantities emerge from pixel prediction alone remains open; the authors instead treat structure as an explicit learning target.
  • An effective representation combines discrete relational structure with continuous, difficult-to-name properties such as exact geometry and appearance.
  • Programs can encode hierarchical and constrained physical structure, then be executed, rendered, simulated, and compared with observations.

2.9. From Visual Models to Vision-Native Intelligence (Zhuang Liu)

Vision-native intelligence treats vision as a continuous, active source of context rather than merely an attachment to language, emphasizing adaptive visual abstraction and interaction.

  • The section offers a perspective rather than a complete theory, advocating vision-native systems that acquire context and provide timely feedback.
  • Vision is important because humans use visual and spatial processes to imagine scenes, reason about relations, and plan actions.
  • Unlike language, vision lacks pre-existing symbolic units and compression layers, making visual learning information-intensive.
  • Useful visual granularity is task-dependent: systems should represent irrelevant regions coarsely and task-relevant regions finely.
  • The authors favor scaling with suitable data and objectives rather than hand-designing visual units, attention policies, or active-perception modules.
  • Visual scale must be accompanied by diversity, because repetitive, narrow, biased data may provide limited visual experience.
  • A vision-native system would treat visual context as part of natural interaction and learn when to look, where to look, and how much detail to use.

2.10. Visual intelligence towards general intelligence (Hirokatsu Kataoka et al.)

The section defines visual intelligence as learning usable world structure from visual experience and proposes sequential prediction, generation, and reconstruction as complementary routes toward general intelligence.

  • Visual intelligence is the capacity to acquire usable world knowledge from visual experience and apply it to prediction, imagination, reconstruction, and new tasks.
  • Richer-than-human visual information could support machine intelligence beyond human-centered perception.
  • Visual in-context learning may elicit latent capabilities from models shown a few visual examples of a new task.
  • The central hypothesis combines sequential prediction, open-ended generation, and reconstruction as complementary learning objectives.
  • Autoregressive visual learning pressures models to infer objects, geometry, motion, physical constraints, causality, and uncertainty to predict future states.
  • Generative learning requires models to represent what could exist, including appearance, geometry, materials, lighting, motion, and interaction.
  • Reconstruction recovers hidden causes from incomplete observations, with stronger reconstruction potentially revealing regularities of the physical world.

3. Summary and Discussion

The paper presents diverse, complementary perspectives on VGI rather than a single definition, model, objective, or benchmark. These perspectives broaden visual intelligence toward generation, physical structure, memory, embodiment, multimodal experience, continual learning, and cross-task generalization.

  • The perspectives span scaling and generative modeling, creativity, prediction, reconstruction, continual learning, spatial representations, embodiment, multimodal learning, scientific discovery, and physical structure.
  • Generative learning is emphasized because generation may require models to account for more of the visual world than task-specific recognition.
  • Photorealistic prediction alone does not establish physical understanding; useful visual intelligence should recover entities, properties, relations, and dynamics for intervention, simulation, verification, and action.
  • Visual knowledge should be acquired and maintained over time through continual learning, persistent spatial representations, state estimation, memory revision, active perception, and action.
  • VGI extends beyond ordinary recognition to multimodal experience, scientific validation without conventional ground truth, creativity, and possibilities absent from explicit training experience.
  • VGI is not simply a stronger classifier or language-model attachment; it concerns discovering, verifying, maintaining, and using world structure across capabilities.
  • Benchmarks should assess transfer, continual adaptation, active observation, persistent knowledge, creativity, physical consistency, reliable action, and efficiency rather than isolated static tasks alone.

4. Conclusion

The conclusion frames VGI as an open research agenda about how intelligence may emerge from visual experience and potentially contribute to general intelligence. Existing systems provide partial evidence, but their capabilities remain fragmented and do not yet form a generally adaptable visual system.

  • The paper concludes that VGI has no single definition, architecture, learning objective, or route, with plausible mechanisms including generation, prediction, reconstruction, memory, continual learning, embodiment, active perception, and multimodal experience.
  • VGI should describe broad capacities to acquire knowledge from visual experience, maintain world structure, imagine alternatives, adapt to observations, and use knowledge for prediction, discovery, and action.
  • Existing systems demonstrate components such as cross-task generation, geometric reconstruction, persistent spatial representations, creative alternatives, and interaction-based improvement, but not their integration into a general continually adaptable system.
  • Evidence for intelligence emerging from vision should include transfer, revisable world knowledge, counterfactual reasoning, active information seeking, continual adaptation, spatial and physical consistency, and reliable action in unfamiliar situations.
  • Computer vision may be entering a phase focused on the emergence of visual intelligence, while leaving open which path or convergence will prevail and what evidence will establish it.
Loading 2608.25924v1…