Source-linked AI summary
Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
Yixin Zhu, Tao Gao, Lifeng Fan, Siyuan Huang, Mark Edmonds, Hangxin Liu, Feng Gao, Chi Zhang, Siyuan Qi, Ying Nian Wu, Joshua B. Tenenbaum, Song-Chun Zhu
TL;DR
Current computer vision relies heavily on large datasets for narrow tasks and lacks humanlike common sense for physical and social worlds. The paper proposes “small data for big tasks” by integrating functionality, physics, intent, causality, and utility, and reviews evidence that these dimensions support broader reasoning and generalization. It concludes that cognitive AI should treat these unseen factors as foundational while recognizing limitations in current causal learning and predefined-action systems.
Problem
Current vision systems require large labeled datasets for specialized tasks and lack a general understanding of common physical and social facts.
Method
The paper reviews five common-sense dimensions—functionality, physics, intent, causality, and utility—and advocates joint reasoning over visible and unseen scene concepts.
Results
Across reviewed studies, FPICU is presented as supporting visual tasks beyond classification, including tool use, planning, utility inference, social learning, and reasoning.
Takeaways & Limitations
The paper argues that cognitive AI should incorporate humanlike common sense and unseen causal, physical, functional, intentional, and utility-based factors to address novel tasks.
Takeaways & Limitations
Causal structure learning remains unresolved because observational approaches often identify only Markov equivalence classes, while the space of structures and parameters is exponential.
Abstract
from arXiv · showhide
Recent progress in deep learning is essentially based on a "big data for small tasks" paradigm, under which massive amounts of data are used to train a classifier for a single narrow task. In this paper, we call for a shift that flips this paradigm upside down. Specifically, we propose a "small data for big tasks" paradigm, wherein a single artificial intelligence (AI) system is challenged to develop "common sense", enabling it to solve a wide range of tasks with little training data. We illustrate the potential power of this new paradigm by reviewing models of common sense that synthesize recent breakthroughs in both machine and human vision. We identify functionality, physics, intent, causality, and utility (FPICU) as the five core domains of cognitive AI with humanlike common sense. When taken as a unified concept, FPICU is concerned with the questions of "why" and "how", beyond the dominant "what" and "where" framework for understanding vision. They are invisible in terms of pixels but nevertheless drive the creation, maintenance, and development of visual scenes. We therefore coin them the "dark matter" of vision. Just as our universe cannot be understood by merely studying observable matter, we argue that vision cannot be understood without studying FPICU. We demonstrate the power of this perspective to develop cognitive AI systems with humanlike common sense by showing how to observe and apply FPICU with little training data to solve a wide range of challenging tasks, including tool use, planning, utility inference, and social learning. In summary, we argue that the next generation of AI must embrace "dark" humanlike common sense for solving novel tasks.
1. A Call for a Paradigm Shift in Vision and AI
The paper calls for shifting vision and AI from “big data for small tasks” toward “small data for big tasks” grounded in humanlike common sense. It frames functionality, physics, intent, causality, and utility as invisible factors that support understanding beyond pixels.
- “Big data for small tasks” systems rely on large labeled datasets for specialized tasks yet lack general understanding of common physical and social facts.
- Humanlike scene understanding combines visible analysis with latent functionality, physical laws, causal relationships, and agent intentions that are not represented directly by pixels.
- The proposed “small data for big tasks” paradigm uses reasoning about unobservable factors to generalize across classification, localization, reconstruction, causal reasoning, intuitive physics, affordances, intent, and utility.
- The paper identifies functionality, physics, intent, causality, and utility as five axes of visual common sense for reasoning about why and how scenes and behaviors unfold.
- The authors argue that these unseen ingredients should become central to future AI rather than relying only on increasing data-driven performance and model complexity.
2. Vision: From Data-driven to Task-driven
The section contrasts single-task computer vision with human task-centered vision, in which one visual system supports many activities. It highlights evidence that visual processing reflects object functionality and action demands, not only appearance.
- Humans use one vision system across thousands of tasks, whereas contemporary computer vision commonly designs one model for one task.
- Task-centered vision links visual representation to generalization, adaptation, and transfer across activities rather than to isolated task-specific prediction.
- Human visual responses distinguish manipulable tools from faces, indicating that recognizing how objects support actions is embedded in vision.
- Different grasping strategies require different functional capabilities, so an object’s representation can depend on the intended task.
2.1. “What”: Task-centered Visual Recognition
Task demands shape what people recognize and where they look, challenging purely feed-forward, data-driven accounts of visual categorization. The same scene can elicit different processing depending on the requested category level.
- Scene categorization involves bidirectional interplay between visual input and the viewer’s needs or task, rather than a purely feed-forward process.
- Although data-driven methods achieve strong scene-recognition accuracy on public datasets, they do not fully account for primate image-level behavior patterns.
- The same image produces different gaze patterns when viewers categorize it at a basic level such as restaurant versus a subordinate level such as cafeteria.
- Task-centered representations allow the same object, such as a mug, to receive different visual treatment according to the planned grasp.
2.2. “Where”: Constructing 3D Scenes as a Series of Tasks
The section presents 3D vision as task-dependent rather than requiring a precise, globally accurate scene model. Human spatial representations can be distorted and task-specific, while computational reconstruction can iteratively combine multiple visual cues.
- Single-image 3D reconstruction is ill-posed because infinitely many 3D configurations can match the same 2D projection.
- Human observers represent spatial layout in ways that differ from precise global metric reconstruction and can make systematic localization errors.
- Human spatial representations may be flat and task-dependent rather than based on a stable or distorted 3D scene model.
- Analysis-by-synthesis initializes a 3D representation from individual vision tasks, compares rendered normal, depth, and segmentation maps with image estimates, and iteratively adjusts structure.
2.3. Beyond “What” and “Where”: Towards Scene Understanding with Humanlike Common Sense
Humanlike scene understanding extends beyond visible objects and geometry to latent causality, physics, functionality, intentions, and utility. The paper proposes jointly representing and reasoning over these “dark” concepts with visible scene information to support broader inference and learning.
- Joint reasoning should combine visible recognition with dark concepts including fluents, causality, physics, functionality, intentions, goals, and utility.The paper organizes these concepts into five axes: fluent and perceived causality, intuitive physics, functionality, intentions and goals, and utility and preference.
- 2.3.1. Fluent and Perceived Causality: Fluents are transient object states linked to perceived causality, allowing systems to relate actions to changes that may be invisible in pixels.Examples include a cup being empty or filled, a door being locked, and a car signaling a turn.
- 2.3.2. Intuitive Physics: Intuitive physics constrains scene interpretations by modeling stability and disturbance potentials from gravity, human actions, and environmental forces.The integrated human action field combines primary human motions with secondary motions to estimate which objects are more exposed to accidental collisions.
- 2.3.3. Functionality and 2.3.4. Intentions and Goals: Functionality and human goals shape scene layouts and behavior, while trajectories can reveal the attraction fields of functional objects and the intentions driving agents.People’s actions can be interpreted as goal-directed sequences, such as moving toward objects that satisfy thirst or hunger.
- 2.3.5. Utility and Preference: Observing rational behavior and choices can estimate human utility, which the paper argues may be more invariant for learning than traditional supervised computer-vision training.The proposed utility perspective is connected to planning while distinguishing human preference from action-dependent MDP value.
- 2.3.6. Summary: FPICU is proposed to improve generalization, small-sample learning, and bidirectional inference by providing invariant prior knowledge alongside pixel-based evidence.The paper argues that top-down common-sense inference and bottom-up visual inference can mutually reinforce one another.
3. Causal Perception and Reasoning: The Basis for Understanding
Causality provides a foundation for humanlike visual understanding by linking observed events to inferred causes, counterfactual alternatives, and transferable causal structure. The reviewed evidence shows that humans perceive causal interactions and transfer causal knowledge more effectively than typical reinforcement-learning agents.
- 3.1. Human Causal Perception and Reasoning: Human vision automatically perceives causal interactions from motion patterns, including launching and entraining, even when the stimuli are only pixel patches.Small temporal gaps disrupt launching perception, while different motion relationships produce distinct causal effects.
- 3.1. Human Causal Perception and Reasoning: Retinotopically specific adaptation does not transfer between launching and entraining, supporting distinct categories of causal perception.The result indicates that causal event types are represented as fundamentally different perceptual categories.
- 3.1. Human Causal Perception and Reasoning: Counterfactual simulation strengthens causal judgments by estimating what would have happened if a candidate cause had been removed.Participants judged whether a billiard ball caused or prevented another ball’s motion, using the counterfactual outcome in their viewing and judgments.
- 3.2. Causal Transfer: Challenges for Machine Intelligence: Human causal learning approaches optimal transfer across structurally equivalent but visually different environments, whereas deep RL methods show negative transfer.The comparison concerns rooms with additional levers under both congruent and incongruent conditions.
- 3.3. Causality in Statistical Learning: Causal structure remains difficult to learn statistically because observational data often identify only a Markov equivalence class, while candidate structures and parameters grow exponentially.The paper argues that humanlike AI should constrain possible relationships using heuristically reasonable world knowledge rather than strict exhaustive formalism.
- 3.4. Causality in Computer Vision: Information projection recovers human judgments of causal and noncausal relationships in everyday-action videos, while observational vision cannot guarantee the complete true causal structure.The tested scenes included opening doors, refilling water, turning on lights, and working at a computer.
4. Intuitive Physics: Cues of the Physical World
Intuitive physics is a core component of human commonsense understanding, enabling rapid inference about stability, support, dynamics, and physical interaction. Computer-vision systems increasingly incorporate probabilistic simulation and physics engines to produce physically plausible scene interpretations and reason about future dynamics.
- 4.1. Intuitive Physics in Human Cognition: Humans rapidly infer whether objects will topple, support weight, or undergo other physical changes, making intuitive physics central to object and scene understanding.These judgments support appropriate interaction with dynamic physical environments.
- 4.1. Intuitive Physics in Human Cognition: Mental simulations support physical prediction and manipulation, and the brain regions engaged by physical inference overlap with mechanisms for action planning and tool use.This links intuitive physical understanding to preparation for interacting with the environment.
- 4.1. Intuitive Physics in Human Cognition: Human physical inference can be modeled as a stochastic, coarse forward simulation implemented by an approximate Bayesian physics engine.The proposed account distinguishes rough probabilistic simulation from precise explicit application of Newtonian mechanics.
- 4.2. Physics-based Reasoning in Computer Vision: Physics-based scene parsing improves physical plausibility by identifying support and stability relationships that appearance-based methods can miss.Without physics, objects may appear to float; incorporating it yields stable 3D scene interpretations and improves performance over contemporary data-driven classification methods.
- 4.2. Physics-based Reasoning in Computer Vision: Modern physics-based vision models represent objects with physical properties and use 3D engines to predict future dynamic evolution from images and videos.Galileo operates on mass, position, 3D shape, and friction, while explicitly encoding basic physical laws and learning object properties from videos.
5. Functionality and Affordance: The Opportunity for Task and Action
Functionality and affordance shift visual understanding from fixed object categories toward the tasks and actions objects support. This perspective enables tool selection, affordance recognition, and causal-equivalent imitation across varied appearances and situations.
- 5. Functionality and Affordance: The Opportunity for Task and Action: Objects and scenes are defined by possible functions and actions, while affordances depend on the actor and functionality remains object-invariant.Functional understanding requires reasoning about state changes produced by interacting with an object.
- 5.1. Revelation from Tool Use in Animal Cognition: A single nut-cracking demonstration supports selecting a visually different object because tool use requires reasoning about functionality, physics, and causal relationships.The same object may serve different functions depending on task context, making conventional category labels insufficient.
- 5.2. Perceiving Functionality and Affordance: Function-based representations generalize object recognition across novel functions and situations instead of memorizing category-specific visual examples.This approach combines physical and geometric aspects to reason about underlying task mechanisms.
- 5.3. Mirroring: Causal-equivalent Functionality & Affordance: Evaluation remains difficult because humans and robots have different morphologies, so the same object or environment need not afford identical actions.Measuring natural human manipulation forces is also constrained by instrument accuracy, occlusion, and sensor rigidity.
- 5.3. Mirroring: Causal-equivalent Functionality & Affordance: Causal-equivalent manipulation replaces trajectory copying with force- and goal-based actions that achieve the same object-state change.The mirroring approach uses physics-based simulation and human-object-interaction units to address correspondence across robots and situations.
6. Perceiving Intent: The Sense of Agency
Human vision interprets animate behavior through goals, beliefs, and intentions rather than surface motion alone. Research on infants, simple geometric displays, and inverse planning motivates computational models that infer latent intent from observed actions.
- 6. Perceiving Intent: The Sense of Agency: Intentional behavior is modeled as goal-directed action with equifinal variability, rationality, and plans that adapt to environmental constraints.Agency therefore includes representing future goals and achieving them through different actions across contexts.
- 6.1. The Sense of Agency: Infants recognize goal-directed activity by six months, segment behavior into goal-directed acts by ten months, and later infer alternative plans and intended goals.By eighteen months, children can imitate an intended goal even when the observed action repeatedly fails.
- 6.2. From Animacy to Rationality: Human observers extract goals, emotions, personalities, and social relationships from the motion of simple geometric shapes with minimal visual semantics.The Heider-Simmel display elicited story-like interpretations such as a hero saving a victim, while static frames sharply reduced anthropomorphic descriptions.
- 6.2. From Animacy to Rationality: Motion dynamics strengthen perceived intentionality, whereas degrading structural display features leaves anthropomorphic interpretation comparatively intact.Experiments with moving letters similarly found that biologically meaningful motion supports perceptions of chasing and intentional action.
- 6.2. From Animacy to Rationality: Inverse planning uses Bayesian inference to recover latent mental states by combining observed-action likelihoods with prior beliefs about agents.This formalizes the rationality principle as the inversion of a planning process in which intent causes action.
- 6.2. From Animacy to Rationality: The paper organizes agency inference around a psychological space separating physical-law violations from inferred agent intent.This framework links perception of inanimate physical events with perception of social events involving agents.
7. Learning Utility: The Preference of Choices
Utility models preferences as context-dependent guides to choice, allowing vision systems to infer human values from behavior and apply them to planning and social learning.
- Observing rational choices can reveal an agent’s utility even when the agent does not explicitly deliberate using that function.
- Utility ranks context-dependent preferences, distinguishing subjective desire from more objective measurable value.
- Video-based chair-selection observations let a vision system learn preferred comfort intervals for forces on different body parts.
- Human utility can guide robotics planning: demonstrations supported learning external utility and planning a cloth-folding task under a goal-state preference assumption.
- Utility-based reasoning also connects to social learning and goal-directed language use by modeling values, permissible actions, and rational communication.
8. Summary and Discussions
The discussion presents FPICU as a foundation for cognitive AI, supported by realistic simulation, abstract reasoning benchmarks, and models of physical and social interaction.
- 8. Summary and Discussions: FPICU—functionality, physics, intent, causality, and utility—is proposed as a foundation for cognitive architectures that address robots’ limited practical usefulness.
- 8.1. Physically-Realistic VR/MR Platform: From Big-Data to Big-Tasks: Synthetic virtual environments support holistic evaluation of agents’ ability to adapt across tasks and environments rather than measuring only single-task performance.
- 8.1. Physically-Realistic VR/MR Platform: From Big-Data to Big-Tasks: Physics-based simulation is presented as central for task-driven evaluation, with fidelity and supported material properties determining the scope and complexity of AI tasks.
- 8.1. Physically-Realistic VR/MR Platform: From Big-Data to Big-Tasks: Simulation methods span grid-based, mesh-based, and mesh-free approaches, each with distinct difficulties involving interfaces, deformation, computation, or numerical stability.
- 8. Summary and Discussions: Multi-agent communication research models communication as action and studies learned continuous or discrete protocols for decentralized collaboration.
- 8. Summary and Discussions: Moral judgment remains variable across individuals, groups, cultures, and the principles being weighed.
- 8.3. Measuring the Limits of Intelligence System: IQ tests: RAVEN links rendered visual problems to attributed stochastic image grammars, and structural reasoning modules improved performance on Raven’s Progressive Matrices.