Source-linked AI summary

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Jiajia Lin, Mingxuan Du, Tuowen Zhou, Benfeng Xu, Hongtao Xie

arXiv:2607.27616v1cs.CV

TL;DR

Existing evaluations miss anatomical failures in multi-person contact edits, where bodies can fuse or interpenetrate despite recognizable people and requested actions. MPIE-Bench and MPIE-Eval address this with mesh-based Anatomy and Interaction axes, revealing that no single editor is strong on both.

  • Problem

    Existing evaluations often assess identity, action, and quality but overlook fused limbs, invented extremities, and interpenetrating bodies in multi-person contact edits.

  • Method

    MPIE-Bench evaluates 2,500 video-mined editing triplets, while MPIE-Eval adds mesh-based Anatomy and Interaction axes using a frozen multi-person reconstruction model.

  • Results

    Across ten editors, mesh Anatomy tops at 0.65 and mesh Interaction at 0.72 on different models, while VLM checklist scores exceed 0.95.

  • Takeaways & Limitations

    Human ratings align more closely with both mesh axes than with a zero-shot VLM judge, and editor rankings survive ablation of weights and thresholds.

  • Takeaways & Limitations

    MPIE-Eval’s Interaction metric still uses whole-body proximity, while part-conditioned contact loci remain future work.

Abstract

from arXiv · show

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

1 Introduction

MPIE-Bench targets anatomically implausible multi-person interaction edits that existing benchmarks and VLM judges miss, while MPIE-Eval measures contact-time body coherence from reconstructed human meshes. Across ten editors, mesh Anatomy reaches 0.65 and mesh Interaction 0.72 on different models, outperforming checklist evaluation in human consistency.

  • MPIE-Bench: MPIE-Bench evaluates multi-person interaction editing using 2,500 video-mined triplets spanning 405 scenes, 14 interaction categories, and four contact densities.The benchmark uses references and targets from the same real interaction and stratifies samples by interaction category and contact density.
  • MPIE-Eval: MPIE-Eval reconstructs people with a frozen public multi-person mesh model and scores Anatomy for complete-body explanation and Interaction for penetration and surface-distance agreement with requested contact.Its bands and weights are fixed before scoring, targeting fused or invented limbs and incorrect body interpenetration.
  • Results: 0.65 is the maximum mesh Anatomy score and 0.72 the maximum mesh Interaction score across ten editors, achieved on two different models.No single editor is strong on both axes, while a checklist judge rates the same images above 0.95.
  • Human consistency: Five-rater human gold shows that the mesh axes track human judgement more closely than a zero-shot VLM judge on hard contact items.The protocol’s ranking also survives ablation of every weight, gate, and threshold.

2 Related Work

Prior work advanced multi-person generation, editing, and interaction-aware geometry, but evaluation remains fragmented: contact geometry and multi-person anatomy are largely unmeasured. MPIE-Bench builds on these lines while using a frozen public RGB Multi-HMR frontend for reproducible editor rankings.

  • Multi-person generation and editing: Multi-person editing has consolidated around multi-reference in-context editors that combine several identity references with one instruction in a single pass.This is the interface evaluated by MPIE-Bench.
  • Benchmarks for generation and editing: Evaluation progressed from subject-driven generation and instruction-editing protocols to suites for multiple recognizable humans and contact-based interaction stress tests.These interaction tests target physical contact as a failure mode for identity binding, but the supplied passage truncates before detailing their diagnostics.
  • Judging bodies, and judging the judges: Existing body-judging methods cover synthetic defects, learned realism metrics, and reconstruction-assisted predictors, but remain scoped to a single isolated person.Purpose-built detectors also find general vision-language models near chance when localizing missing or redundant body parts.
  • Interaction-aware geometry and HMR: Multi-person HMR, contact fields, hand reconstruction, and contact-distance machinery make geometry-side scoring feasible for interaction evaluation.The protocol freezes a single public RGB Multi-HMR frontend so editor rankings are reproducible, while stronger contact-aware reconstructors can be used later.

3 MPIE-Bench

MPIE-Bench mines video-based multi-person interactions into editing triplets, pairing clean per-person references with harder contact-rich targets. Human review freezes 2,500 targets across 405 scenes, covering 14 interaction categories and four contact-density levels.

  • Sources: MPIE-Bench combines Pexels, Harmony4D, and CHI3D videos to provide identity and scene diversity alongside persistent labels and hard contact.Restricted-tier corpora are excluded from the redistributable split.
  • Input–output selection: Contact-sparse valley frames supply clean per-person references, while contact-dense peak frames become interaction targets from the same scene narrative.Peak targets are selected from smoothed contact-density maxima using person-box overlap, pose proximity, optical flow, and occlusion cues.
  • Prompt reverse writing: A VLM reverse-captions each sample into an instruction naming participant roles, the shared action, and contact intent without exposing target pixels at test time.Contact intent is mapped to required, forbidden, or unspecified for later Interaction mixing, while prompt text remains the editors’ only instruction.
  • Human review and distribution: 2,500 targets across 405 scenes form the reviewed benchmark, organized as a 14×4 grid of interaction categories and contact-density levels.Automatic mining produces approximately 20,000 candidates before safety, task, deduplication, and per-scene-cap review.
  • Human review and distribution: The public test set spans contact density from none (C0) through high contact (C3) across fourteen everyday-to-combat interaction categories.The four levels progress through no contact, hand-level contact, torso or point-line contact, and high contact.

4 MPIE-Eval

MPIE-Eval reports six separate axes, using a shared frozen multi-person reconstruction to measure contact-time Anatomy and Interaction with transparent geometric proxies. Anatomy evaluates unexplained human-like mass, while Interaction evaluates reconstructed pairwise contact geometry against prompt intent.

  • Evaluation design: Six axes remain separate: Count, Identity, Instruction, and Quality assess task success, while Anatomy and Interaction assess body coherence at contact.The pipeline reconstructs people once; the four established task axes stay on independent tracks.
  • Evaluation design: The evaluator uses one frozen Multi-HMR checkpoint and keeps the top-k detections, with k equal to the expected person count.Anatomy and Interaction are set to 0 when reconstruction fails or returns zero meshes, preventing unscoreable geometry from inflating rankings.
  • Anatomy: Anatomy scores whether every human-like region is explained by complete bodies, covering extra or fused limbs, floating parts, body structure, and residual mass.Diagnostics compare a classical foreground mask with rasterized kept meshes and count detached leftover blobs.
  • Interaction: Interaction measures convex-hull penetration and surface gap for every kept mesh pair, aggregates by the worst pair, and maps prompt intent to required, forbidden, or unspecified contact.The score mixes penetration, proximity, quality, or clearance terms according to the mapped intent.
  • Calibration and limits: Soft Interaction bands are calibrated on held-out CHI3D-related ground-truth reconstructions, excluding MPIE-Bench scenes and leaderboard generator outputs.A small hand–hand/gaming contact-locus check leaves Interaction top-2 rankings unchanged.

5 Experiments

Across 2,500 samples and ten editors, mesh-based Anatomy and Interaction expose substantial contact-editing failures that checklist VLM scores largely miss. The metrics track human rankings, remain stable under protocol perturbations, and reveal density- and locus-specific limitations.

  • Protocol: Ten editors were evaluated on the full N=2,500 scene-split test set using vendor-native Track A and letterbox→10242 Track B mesh scoring.Track A uses closed APIs at deployment resolution; Track B is the official resolution-matched view.
  • Main results: 0.98–0.99 closed V-Inter contrasts with 0.45–0.72 mesh M-Inter, while Gemini leads Anat at 0.65 and Seedream leads Inter at 0.72.Identity also shows a closed/open gap of Sid ≈0.49–0.58 versus 0.02–0.21.
  • Contact density: Anatomy softens from C2 to C3 for Gemini, Seedream, and FireRed, while DreamO remains nearly flat at 0.54 → 0.56.C0 is treated as a control bin because it mixes non-contact and unspecified-intent samples; high-contact comparisons use C3 Interaction.
  • Robustness: Anat rank Spearman remains ≥0.98 under weight and gate ablations, while a FireRed pad sweep changes Anat by +0.020, 0, and −0.006.Seed-level Anat/Inter standard deviations across three seeds are O(10−2) on a fixed 150-sample subset.
  • Human validation: Mesh Anatomy and Interaction are validated against human judgement, with editor-level preference correlations of ρ ≈0.87 and ρ ≈0.80, respectively.The study frames this as a ranking check on hard contact items rather than a per-image oracle because absolute agreement remains moderate.
  • Failure analysis: 55% of a hard open-source contact pool has at least one visible geometric error tag, while a hand–hand distance swap lowers overall Inter by only ≈0.005–0.011.Locus-critical prompts have pooled Inter ≈0.59 versus ≈0.56 elsewhere, indicating limited sensitivity to localized hand contact.

Ethical Statement

The benchmark uses licensed public and academic sources under attribution, access, and non-redistribution constraints. Safety measures exclude flagged or sexualized content, while the study limits use to consent-aware research diagnosis rather than non-consensual applications.

  • Data licensing: Public and academic sources are distributed as IDs, timestamps, boxes, scripts, or derived crops under their original licenses, not as unrestricted pixels.Harmony4D pixels use MIT; Pexels-derived materials follow the Pexels License without re-hosting as a stock library, and restricted-tier corpora are excluded.
  • Safety and intended use: Automated minor detection removes flagged frames, and identity-conditioned generation follows nonsexualization constraints.Combat and restraint prompts are restricted to sports- or self-defense-style scenarios.
  • Safety and intended use: The benchmark uses appearance-only public or academic likenesses and excludes placing identifiable real individuals into synthetic scenes without consent.Checklist labels were produced by adult raters only.
  • Safety and intended use: Its metrics, including ID, are intended for research diagnosis rather than rewards for non-consensual likeness use.The passage frames the metrics as diagnostic tools, not incentives for misuse.

6 Conclusion · Appendix · A Extended protocol notes

MPIE-Bench makes contact-time coherence measurable through a licensing-aware, scene-split benchmark and a six-axis evaluation protocol with mesh-based Anatomy and Interaction scores. The released artifacts support auditable rankings, alternate frontend tracks, and extended protocol diagnostics.

  • 6 Conclusion: MPIE-Bench contains 2,500 video-mined triplets across 405 scenes, stratified by 14 interaction categories and C0–C3 contact density with held-out targets.The scene split prevents editors from copying pixels.
  • 6 Conclusion: MPIE-Eval preserves Count, ID, Instruction, and Quality while adding Anatomy and Interaction from a frozen, dump-auditable Multi-HMR frontend.It supports Track A at vendor-native resolution and Track B after letterbox→10242.
  • 6 Conclusion: Closed V-Inter saturates near ceiling while M-Inter retains dynamic range, showing that mesh axes reveal distinctions missed by checklists.Identity remains the steepest closed/open gap, and Anatomy and Interaction leaders need not coincide.
  • 6 Conclusion: Density, corruption, human-checklist, and ablation checks support model-level ranking under the frozen frontend rather than a single blended score.These checks reinforce the protocol’s use for comparing editors across evaluation conditions.
  • 6 Conclusion: The release includes a redistributable subset, evaluation code, per-sample Multi-HMR dumps, and a calibration recipe for auditing rankings and adding alternate frontend tracks.These artifacts allow alternate-frontends without redefining the editing task.
  • A Extended protocol notes: Instr v2 freezes a per-sample QA bank from text-only atomic-claim extraction, while evaluation answers those questions from generated images without rewriting them.Questions cover role, asymm, and prop, with optional scene diagnostics; interaction-occurrence, count, face-identity, and anatomy questions are forbidden.
  • A Extended protocol notes: Mesh Interaction reads contact intent from prompt text and combines interpenetration volume, enclosure fraction, proximity, and contact-quality penalties with GT-calibrated soft ramps.Higher Inter than GT can occur when a generation under-penetrates relative to real contact.
  • A Extended protocol notes: Release artifacts include the redistributable scene-level test split, frozen reference face embeddings, six-axis evaluation scripts, human-checklist agreement scripts, and metric calibration materials.The extended release documents implementation components beyond the core benchmark.

B Toward anatomy-aware training (future work) … E End-to-end pipeline (Algorithm 1)

The supplement sketches anatomy-aware training built on MPIE-Bench signals, while documenting why mesh-anchored metrics remain necessary despite saturated checklist judges. It also specifies coverage requirements and freezes the constants and consistency details underlying the end-to-end pipeline.

  • B Toward anatomy-aware training (future work): The paper is evaluation-only; the proposed training recipe is exploratory and not required to use MPIE-Bench.This training direction is presented for completeness as future work.
  • B Toward anatomy-aware training (future work): A lightweight auxiliary LoRA and projection head predicts 2D multi-person skeletons during training to preserve complete per-person anatomy under occlusion.The stream is discarded at inference and targets limb fusion and limb-attribution failures.
  • B Toward anatomy-aware training (future work): The C0–C3 curriculum schedules early C0–C1, mid C2, and late C3 examples to stabilize body completeness before deep mutual occlusion.This is a dataloader-level change with zero architectural cost.
  • C VLM-as-judge saturation diagnostic: 0.98–0.99 Anatomy/Interaction means from an all-checklist visual judge coexist with visible fused limbs and pathological interpenetration in strong closed-source editors.Parametric joint/bone priors also reach ≈0.99 because hallucinated extremities can lie outside fitted topology, so primary tables use mesh-anchored metrics.
  • D Coverage rules: Every reported axis is subject to a minimum scored-generation coverage rule, and all eight main-table systems satisfy it for every reported axis.The rule is implemented in the release scripts.
  • E End-to-end pipeline (Algorithm 1): The supplement freezes the numeric constants, sensitivity checks, and human-consistency details used by the mining and mesh Anat/Inter pipeline.These specifications support the end-to-end procedure described in Algorithm 1.

F Dataset distribution

MPIE-Bench’s dataset distribution is presented through a public test-set sunburst organized by density and category, with additional source-stratified and per-class counts in the release tables. The benchmark freezes a public split of N=2,500 samples spanning contact-density labels C0–C3.

  • Dataset distribution: The public test-set distribution is visualized as a sunburst over density and interaction category.Source-stratified counts and per-class breakdowns are provided in the release tables.
  • Dataset distribution: N=2,500 samples comprise the frozen public split D.The split is frozen after human review of scenes during dataset construction.
  • Dataset distribution: Each mined triplet is assigned one of four contact-density labels: C0–C3.The pipeline emits a prompt, expected number of people, and the density label for each sample.

G Frozen metric constants … L Instruction QA protocol

MPIE-Eval freezes a public Multi-HMR-based metric protocol with explicit Anatomy and Interaction handling, while separately defining calibrated mesh-to-checklist mappings and two non-interchangeable VLM evaluation paths. The Instruction score uses a frozen text-authored question bank answered only on generated images, preventing evaluation questions from drifting with the editor.

  • G Frozen metric constants: Public rankings freeze the Multi-HMR backend multiHMR_896_L and the constants used by mining gates and mesh Anat/Inter scorers.The frozen protocol also keeps the top-k detections with k = nexp.
  • H Anatomy and Interaction formulas: Anatomy combines leftover, detached-blob, scale, ownership, part-mesh, and foreground-residual terms, with attached-signature gating for leftover and fuse penalties.Under the attached signature, ℓ≥0.62 and nb = 0; otherwise the old penalty is scaled by 0.15.
  • H Anatomy and Interaction formulas: Interaction mixes penetration, proximity, and quality scores according to intent, and aggregates multi-person diagnostics so one severely intersecting pair dominates.For at least two kept bodies, penetration/enclosure uses the maximum over pairs, while surface distance uses the minimum over pairs.
  • H Anatomy and Interaction formulas: The public protocol keeps the top k=nexp detections without identity-aware matching; failed reconstruction yields Anat and Inter of 0, while missing detections halve Interaction.Surplus detections beyond k are ignored for pairing.
  • H Anatomy and Interaction formulas: Hull-normalized Vp is treated as a ranking observable rather than metric contact volume because outstretched limbs can inflate convex hulls and overstate fusion.Soft cutoffs are calibrated on held-out ground-truth frames in the same units, with a synthetic hull-inflation stress case lowering the effective fusion signal.
  • I Mesh→checklist mapping (M): Mesh-to-checklist mapping fits ridge regressors on held-out human labels, discretizes predictions by Spearman-optimized cuts, and freezes weights, biases, and cuts before the N=170 consistency pool.Interaction items retain the same intent-conditioned dependency rules as humans.
  • J Two VLM paths: checklist vs. Instruction QA: MPIE-Eval separates the shared checklist VLM path from per-sample Instruction QA: the former supports Count and V-Anat/V-Inter, while the latter is the main-table Instruction score.Main-table Instruction-like fields from vlm_judge are ignored.
  • L Instruction QA protocol: Instruction QA freezes 2,500 text-only atomic banks, then has a VLM answer fixed items on each generation using weights of 0.50 for asymmetry, 0.35 for role, and 0.15 for props.Questions exclude contact geometry, person count, face identity, and anatomy, and the VLM cannot rewrite them.

M Metric anchors: GT reference and controlled corruption … V Contact-intent parsing

The paper validates MPIE-Eval through controlled corruption, clustered uncertainty, sensitivity analyses, and alternative scoring probes, while exposing limits in contact-region semantics and reconstruction confidence. Across these checks, rankings are generally stable, but category-specific contact geometry and frontend choices remain important qualifications.

  • M Metric anchors: GT reference and controlled corruption: GT serves as a qualitative upper anchor, while controlled corruption lowers Interaction monotonically under penetration, separation, person-drop, and duplication perturbations.A CPU synthetic two-person mesh suite passes 6/6 required direction checks.
  • N Scene-clustered uncertainty for Anat/Inter: +0.042 Seedream−Gemini Inter difference has a 95% CI of [0.019, 0.064], confirming Seedream’s clustered scene-level lead; ≈0.01 gaps should be treated as ties.The bootstrap resamples 405 scenes with replacement, using B=2000.
  • O Sensitivity analyses: ρ ≥0.98 across Anatomy weight, signature, and threshold variants preserves the leading ranks, while ROI padding leaves Interaction invariant and changes Anatomy only mildly.The AC-style 0.30/0.30/0.20/0.20 mix gives ρ=0.988 with the same top-3.
  • Q Interaction gaming proxy: ≈1.3% of required-ok units trigger the high-spen/high-sprox gaming proxy despite a VLM contact_points judgment of no, without changing the Interaction top-2.The proxy rate is ≈2–2.5% for several open editors and ≈0.5% for closed APIs.
  • R Detection-confidence proxy: Anat/Inter means are ≈0.55/0.46 in the low-detection bin versus ≈0.65/0.76 in the high-detection bin, but detection score is only a ranking diagnostic.Public scoring does not downweight unstable reconstructions or yet apply identity-aware person selection.
  • S Multi-seed generation variance: Seed-level Anat/Inter standard deviations remain O(10^-2), below the main closed/open and density gaps, while letterbox→1024 raises Anatomy by ≈0.06–0.14 without changing the Anat top-3.Track B correlates with Track A at ρ=0.89 for open editors and ρ=0.83 across all ten models; Track A remains public.
  • U Alternate mesh frontend (HMR2) / V Contact-intent parsing: A two-frontend mean agrees with Multi-HMR at rank Spearman ≈0.92 for Anatomy and ≈0.68 for Interaction, while human intent matches keyword/system intent on 145/170 cases and leaves rankings unchanged.Replacing keyword intent with human intent changes per-model Interaction means by <10^-3 and yields Spearman 1.00.

W Human-consistency statistics … Z Release diagnostics

The release documents human-consistency statistics, identity conditioning, benchmark coverage gaps, and per-sample diagnostics. Human agreement is moderate, model-level mesh preferences align strongly with checklist scores, and identity gains under full face visibility remain incomplete.

  • W Human-consistency statistics: Five annotators independently scored all N=170 consistency-pool units, with main-paper gold defined as the per-item mean after skipping undecidable marks.Per-annotator exports are included in the release.
  • W Human-consistency statistics: α=0.53 is the unweighted mean across ten items; core contact items reach I1 0.68 and Ic 0.65, while near-ceiling items have lower reliability under class imbalance.Human gold is therefore treated as a moderately reliable ranking signal rather than a near-perfect oracle.
  • W Human-consistency statistics: p<0.05 differences favor M>V on I3 and Ir, while V>M on I0; other M edges are directional but nonsignificant, including A3 at p≈0.09.These comparisons use dependent Spearman correlations ρ(H, M) and ρ(H, V) with Steiger’s test and empirical ρ(M, V).
  • W Human-consistency statistics: ρ(Qanat, Sanat)≈0.87 and ρ(Qinter, Sinter)≈0.80 across editors show strong model-level alignment between checklist and mesh preference anchors.The averages combine Qanat/Qinter and mesh Anat/Inter, although unit-level correlations remain weaker.
  • X Face-Visible Rate and ID conditioning: ≈+0.05–0.06 is the mean Sid lift after restricting to FVR ≥0.9999, yet closed-source identity remains in the mid-0.5s–low-0.6s.The conditioning tests whether the closed/open identity gap is explained solely by hidden faces; residual identity failures remain.
  • Y Extended gap list (benchmark survey): The extended survey marks benchmark support across Count, ID, Instr, Anat-under-contact, and Inter-geometry, broadening the compact main-paper gap table.It covers multi-person, anatomy, personalization, and editing suites rather than every paper in each family.
  • Y Extended gap list (benchmark survey): † denotes single-person anatomy without multi-person contact-time scoring, while “—” denotes reconstruction or ground-truth resources rather than generative editing benchmarks.The extended list is explicitly broader than the main-paper name-only gap view.
  • Z Release diagnostics: Per-sample Multi-HMR JSON exposes detection scores, pen_volume_m3, pen_inside_ratio, min_- surf_dist, leftover fractions, and contact intent for offline audit and recomposition.These fields support inspection of contact-time reconstruction and editing diagnostics.
Loading 2607.27616v1…