Source-linked AI summary

Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

Jason Y. Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, Angjoo Kanazawa

arXiv:2007.15649v2cs.CV

TL;DR

The paper addresses 3D scene understanding when humans and objects are estimated in isolation or controlled conditions, making their spatial arrangement inconsistent in-the-wild images. It jointly recovers human-object arrangements and shapes from a single image without scene- or object-level 3D supervision, using contextual constraints and losses for scale, silhouette, and interaction. Experiments evaluate holistic processing against independent composition and find the occlusion-aware silhouette and interaction losses most consequential, while detailed object-shape recovery remains challenging.

  • Problem

    Existing 3D estimation often treats bodies and objects in isolation or controlled conditions, motivating holistic in-the-wild scene understanding using contextual cues between them.

  • Method

    The method jointly recovers 3D spatial arrangements and shapes of humans and objects from a single image, learning object-size distributions without 3D supervision and using scale, occlusion-aware silhouette, and interaction constraints.

  • Results

    Holistic processing is evaluated against independent composition, with occlusion-aware silhouette and interaction losses having the most significant ablation effects.

  • Takeaways & Limitations

    Human contextual constraints can help reason about objects in 3D, while 3D shape recovery remains less developed and exemplar-based shapes are effective but impoverished.

  • Takeaways & Limitations

    Detailed 3D object-shape recovery remains a challenge, and expected scale distributions can fail for real-world objects outside them, such as a small bicycle.

Abstract

from arXiv · show

We present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment. Notably, our method runs on datasets without any scene- or object-level 3D supervision. Our key insight is that considering humans and objects jointly gives rise to "3D common sense" constraints that can be used to resolve ambiguity. In particular, we introduce a scale loss that learns the distribution of object size from data; an occlusion-aware silhouette re-projection loss to optimize object pose; and a human-object interaction loss to capture the spatial layout of objects with which humans interact. We empirically validate that our constraints dramatically reduce the space of likely 3D spatial configurations. We demonstrate our approach on challenging, in-the-wild images of humans interacting with large objects (such as bicycles, motorcycles, and surfboards) and handheld objects (such as laptops, tennis rackets, and skateboards). We quantify the ability of our approach to recover human-object arrangements and outline remaining challenges in this relatively domain. The project webpage can be found at https://jasonyzhang.com/phosa.

1 Introduction

PHOSA addresses the ambiguity of recovering holistic 3D human-object scenes from a single in-the-wild image by jointly reasoning about humans, objects, scale, interactions, and depth. It combines local reconstruction with globally optimized spatial arrangements without scene- or object-level 3D supervision.

  • Motivation: Independent human and object estimates can produce inconsistent 3D arrangements because multiple configurations may share the same 2D projection.Human-object context helps resolve ambiguities such as whether an object is correctly positioned relative to a person.
  • Approach: PHOSA recovers human and object spatial arrangements and shapes from a single image in uncontrolled, in-the-wild scenes.The method is demonstrated on diverse human-object interactions.
  • 3D common sense: Joint reasoning exploits physical 3D constraints and human-object interactions to reduce ambiguity without requiring complete in-the-wild 3D scene supervision.The method uses contextual cues such as plausible human-object relationships to select more reasonable spatial configurations.
  • 3D common sense: PHOSA combines intrinsic scale, human-object interaction, and depth ordering to produce globally consistent reconstructions.Its overview emphasizes realistic interaction, preserved depth ordering, and physically plausible scene layouts.
  • Approach: The framework reconstructs humans and objects locally, then uses per-instance intrinsic scale to place them in a coherent world coordinate frame.Human reconstruction uses predicted 3D pose and shape, while object pose is fitted to segmentation masks.

2 Related Work

Prior work estimates 3D humans, objects, and interactions using isolated, supervised, synthetic, temporal, or known-scene settings. PHOSA instead focuses on spatial arrangements of humans and objects from a single image without assuming an available 3D scene.

  • 3D human pose and shape: Single-image human reconstruction commonly uses statistical body models, kinematic structure, silhouettes, keypoints, or feed-forward prediction.These approaches primarily model humans in isolation, with some extensions handling multiple people through collision and depth constraints.
  • 3D objects: Single-view object reconstruction has used silhouette optimization or deep prediction, often relying on 3D supervision, multi-view cues, or synthetic datasets.Most such methods reason about object shape in isolation rather than human-object spatial arrangement.
  • 3D human-to-object interaction: Human-object interaction methods infer geometry and affordances from temporal observations, RGB-D data, known scenes, or image-based hand-object configurations.PHOSA draws inspiration from interaction constraints while not assuming that 3D scenes are available.

3 Method

The method reconstructs humans and objects from a single image, then jointly optimizes their poses, shapes, scales, and spatial arrangement in a common 3D coordinate system. It combines mask-based object fitting with human-object interaction and physical constraints to resolve ambiguities that independent reconstruction cannot.

  • 3.1 Estimating 3D Humans: The system separately estimates 3D humans and category-specific object meshes before placing them in a shared world coordinate frame.Humans use parametric 3D reconstruction, while objects are fitted to predicted instance masks and represented with category-specific mesh exemplars.
  • 3.1 Estimating 3D Humans: Per-instance intrinsic scale converts local human and object predictions into coherent metric sizes and world coordinates.The scale parameter is optimized jointly so people and objects can be compared within one globally consistent layout.
  • 3.2 Estimating 3D Objects: Multiple mesh exemplars can represent within-category shape variation, with optimization selecting the exemplar that minimizes reprojection error.The framework uses one skateboard mesh but multiple motorcycle meshes, and automatically determines the selected exemplar during optimization.
  • 3.2 Estimating 3D Objects: Independent object pose estimation fits rigid 3D mesh exemplars to 2D instance masks using differentiable rendering and an occlusion-aware silhouette objective.A no-occlusion indicator ignores pixels belonging to other instances, allowing partial occlusions even when the occluding object lacks a 3D model.
  • 3.3 Modeling Human-Object Interaction for 3D Spatial Arrangement: Joint human-object reasoning resolves scale and depth ambiguities that remain when instances are analyzed independently.Interaction cues, such as human-object proximity and corresponding contact regions, constrain relative layout; physical priors further discourage interpenetration and inconsistent arrangements.

4 Evaluation

On COCO 2017, the method reconstructs human-object arrangements across eight diverse object categories and is evaluated against independent composition and loss ablations. It outperforms independent composition overall, while qualitative and ablation results show that interaction and occlusion-aware silhouette reasoning are especially important.

  • Quantitative Analysis: Evaluation covers eight COCO 2017 object categories spanning varied sizes, shapes, and human interactions.The categories include baseball bats, benches, bicycles, laptops, motorcycles, skateboards, surfboards, and tennis rackets.
  • Quantitative Analysis: The forced-choice evaluation compares PHOSA with independently estimated human and object poses using the learned per-category mean scale.Because in-the-wild 3D ground truth is unavailable, annotators compare randomized outputs and choose better, equal, or worse.
  • Quantitative Analysis: PHOSA safely outperforms independent composition in the forced-choice evaluation.Independent composition receives the learned empirical mean category scale but does not jointly reason over all instances.
  • Quantitative Analysis: Ablations identify occlusion-aware silhouette and interaction losses as having the most significant effects on performance.The silhouette term preserves image evidence during global optimization, while the interaction term encodes object placement relative to the person.
  • Qualitative Analysis: Qualitative results reconstruct arrangements for multiple people and objects across handheld, full-sized, and large everyday objects.The examples include baseball bats, tennis rackets, laptops, skateboards, bicycles, motorcycles, surfboards, and benches.
  • Qualitative Analysis: Explicit human-object interaction reasoning produces more realistic arrangements than independently estimating human and object poses.Independent estimation often fails to resolve fundamental scale ambiguities.

5 Discussion

The discussion finds exemplar-based object modeling surprisingly effective for in-the-wild 3D understanding, while detailed object-shape recovery remains limited. It argues that human context can serve as a persistent scale cue as object categories expand.

  • Detailed 3D object-shape recovery remains challenging because the method relies on an impoverished exemplar-based shape model.The authors identify statistical object-shape models as an important next step.
  • Learning intrinsic scale distributions from data is presented as a first step toward a statistical “SMPL-for-objects.”
  • 2D instance masks, differentiable rendering, and a 3D shape library provide an effective initialization for 3D object understanding.The approach also scales relatively easily because annotations can be painted on 3D models rather than individual image instances.
  • Human contextual constraints can act as “rulers” for reasoning about increasingly diverse object categories.

6 Supplemental Material

The supplemental material details optimization, mesh processing, qualitative outputs, and failure modes across multiple COCO object categories. It shows that errors arise from unreliable human poses, object masks, and scale assumptions.

  • Implementation details: The framework optimizes object pose with an occlusion-aware silhouette loss using ADAM and multiple random rotation initializations.The implementation uses a learning rate of 1e-3 for 100 iterations and selects the lowest-loss initialization.
  • Implementation details: The global spatial-arrangement optimization jointly updates human and object intrinsic scales, object rotations, and object translations.It uses ADAM with learning rate 1e-3 for 400 iterations.
  • Mesh processing: The mesh pipeline uses multiple category models, watertight preprocessing, TSDF fusion, simplification, and vertex reduction for efficient optimization.
  • Qualitative results: Qualitative results cover baseball bats, benches, bicycles, laptops, motorcycles, skateboards, surfboards, and tennis rackets in COCO test images.
  • Failure modes: Human-pose errors can place a hand far from the interacted object, making human-object interaction reasoning difficult.The example concerns a tennis player performing a volley.
  • Failure modes: Unreliable predicted masks can prevent recovery of a plausible object pose and scene reconstruction.
  • Failure modes: Objects outside the expected scale distribution, such as a small bicycle, can challenge the method’s scale initialization.
Loading 2007.15649v2…