Source-linked AI summary

HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

Yichen Liu, Donghao Zhou, Jie Wang, Xin Gao, Guisheng Liu, Jiatong Li, Quanwei Zhang, Qiang Lyu, Lanqing Guo, Shilei Wen, Weiqiang Wang, Pheng-Ann Heng

arXiv:2603.02210v3cs.CV

TL;DR

Generating human-product images requires preserving fine-grained product details, while existing reference-based inpainting methods remain limited by training data, detail modeling, and supervision. HiFi-Inpaint combines HP-Image-40K, high-frequency-guided modeling with SEA, and DAL to address these limitations. Experiments report superior performance and effective preservation of intricate product details.

  • Problem

    Existing human-product image methods struggle to preserve fine-grained product details, while reference-based inpainting remains limited by data, detail modeling, and coarse supervision.

  • Method

    HiFi-Inpaint uses HP-Image-40K, Shared Enhancement Attention (SEA), and Detail-Aware Loss (DAL) with high-frequency guidance for reference-based inpainting.

  • Results

    HiFi-Inpaint achieves superior performance while effectively preserving intricate product details and generating visually coherent human-product images.

  • Takeaways & Limitations

    The framework provides detail-preserving reference-based inpainting for human-product image generation within the demonstrated task and setting.

Abstract

from arXiv · show

Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. The essential challenge of generating such images lies in ensuring the high-fidelity preservation of product details. Among existing paradigms, reference-based inpainting offers a targeted solution by leveraging product reference images to guide the inpainting process. However, limitations remain in three key aspects: the lack of diverse large-scale training data, the struggle of current models to focus on product detail preservation, and the inability of coarse supervision for achieving precise guidance. To address these issues, we propose HiFi-Inpaint, a novel high-fidelity reference-based inpainting framework tailored for generating human-product images. HiFi-Inpaint introduces Shared Enhancement Attention (SEA) to refine fine-grained product features and Detail-Aware Loss (DAL) to enforce precise pixel-level supervision using high-frequency maps. Additionally, we construct a new dataset, HP-Image-40K, with samples curated from self-synthesis data and processed with automatic filtering. Experimental results show that HiFi-Inpaint achieves state-of-the-art performance, delivering detail-preserving human-product images.

1 Introduction

Human-product image generation must preserve fine-grained product details, but existing approaches struggle with spatial and appearance fidelity. HiFi-Inpaint addresses this challenge with a reference-based framework, specialized detail-enhancement components, and the HP-Image-40K dataset.

  • High-fidelity product-detail preservation is critical because inaccuracies in shapes, colors, patterns, and textures can undermine consumer trust and commercial effectiveness.
  • Reference-based inpainting provides targeted product guidance, yet diffusion methods can average or hallucinate details because strict spatial and appearance alignment is insufficient.
  • HiFi-Inpaint combines Shared Enhancement Attention (SEA) for fine-grained product features with Detail-Aware Loss (DAL) for precise pixel-level supervision.
  • HP-Image-40K supplies a large-scale, diverse training foundation using self-synthesized samples and automated filtering.
  • Extensive experiments report that HiFi-Inpaint generates detail-preserving human-product images with superior performance.

2 Related Works

Text-to-image generation has progressed from GANs and autoregressive transformers to diffusion models, which have driven rapid advances in related applications. Image inpainting has likewise evolved from optimization and patch-based methods toward diffusion-based restoration with additional conditioning for improved control.

  • Early text-to-image methods primarily used GANs, while autoregressive transformers later demonstrated potential for image synthesis from text.
  • Diffusion models have revolutionized text-to-image generation and accelerated progress in related applications such as image customization.
  • Image inpainting restores missing or corrupted regions while maintaining visual coherence, evolving from optimization and patch-based approaches to diffusion-based denoising.
  • Additional conditioning enables diffusion-based inpainting models to provide better control over the restoration process.

3 Methodology

HiFi-Inpaint combines the HP-Image-40K dataset with high-frequency-guided architecture and supervision for detail-preserving human-product reference-based inpainting.

  • Framework overview: HiFi-Inpaint generates human-product images from a text prompt, a masked human image, and a product reference image.The framework aims to integrate the reference product into the masked region while following the text description.
  • Dataset construction: HP-Image-40K contains 40,000+ synthesized and automatically filtered human-product samples for model training.The pipeline uses diptych synthesis, Sobel-based segmentation, semantic filtering, and textual filtering.
  • High-frequency guidance: High-frequency extraction uses Fourier-domain high-pass filtering to target product details such as text and logos more selectively than conventional edge detection.The method applies a circular high-pass mask and inverse DFT to obtain high-frequency maps for product and masked ground-truth regions.
  • Shared Enhancement Attention: Shared Enhancement Attention inserts high-frequency visual tokens into dual-stream visual DiT blocks to refine product features within masked regions.SEA uses attention masking and a learnable weighting factor, with the latter producing more harmonious results than fixed weighting.
  • Detail-Aware Training Strategy: Detail-Aware Loss adds high-frequency pixel-level supervision to latent-level training, guiding reconstruction of intricate masked-region details.The overall loss combines latent-level MSE with DAL, balancing global consistency and local detail fidelity.

4 Experiments

HiFi-Inpaint is evaluated against four reference-based inpainting methods using fixed settings, automatic metrics, qualitative comparisons, and a user study. It achieves state-of-the-art performance, while ablations show benefits from HP-Image-40K, SEA, and DAL.

  • Experimental Setup: Four reference-based inpainting methods are compared at 1024 × 576 resolution under the same inference configurations and settings.The baselines are Paint-by-Example, ACE++, Insert Anything, and FLUX-Kontext.
  • Quantitative Comparison: HiFi-Inpaint achieves state-of-the-art performance, with top CLIP-I (0.950), DINO (0.919), SSIM (0.634), and SSIM-HF (0.429).It also achieves or matches strong results on LAION-Aes (4.40) and Q-Align-IQ (4.36).
  • Qualitative Comparison: Qualitatively, HiFi-Inpaint generates naturally composited images while preserving product text, patterns, branding, object shape, and fine details under small masks.FLUX-Kontext often produces standalone product images or loses fine details during inpainting.
  • User Study: A user study with 31 valid responses across 11 image groups finds that HiFi-Inpaint consistently outperforms existing methods on text alignment, visual consistency, and generation quality.The averaged selection rates are reported in Table 3.
  • Ablation Analysis: Adding HP-Image-40K under otherwise identical conditions produces significant gains in text alignment and visual consistency.The dataset contribution is evaluated through comparisons between Schemes A and B and between Schemes D and E.
  • Ablation Analysis: SEA improves alignment of intricate details and patterns, while removing DAL causes blurry or semantically incomplete renderings that fail to preserve text and intricate patterns.Quantitative and qualitative ablations support the contributions of both components.

5 Conclusions

HiFi-Inpaint targets high-quality human-product image generation through reference-based inpainting, preserving intricate product details while producing visually coherent images. The paper also presents HP-Image-40K and identifies future directions in diversity, realism, and video generation.

  • HiFi-Inpaint uses Shared Enhancement Attention to capture fine-grained product features and Detail-Aware Loss for precise pixel-level supervision.
  • HP-Image-40K is introduced as a high-quality dataset designed to facilitate training for human-product image generation.
  • Experiments show superior performance in preserving intricate product details while generating visually coherent images.
  • Future work will address generated-image diversity and realism and extend the method to video generation.

A HP-Image-40K Dataset Statistics

HP-Image-40K is designed with broad coverage of mask sizes, spatial distributions, and product categories. Its statistics indicate diversity in object scales, shapes, materials, and structures relevant to real-world human-product image generation.

  • Variation in mask area ratios covers object sizes and spatial distributions from small localized objects to large prominent ones.
  • The dataset includes product categories such as bottles, containers, jars, tubes, and dispensers.
  • Its category diversity exposes models to a broad spectrum of product shapes, materials, and structural characteristics.
  • HP-Image-40K covers a diverse range of mask area ratios across real-world scenarios.The mask area ratio is the proportion of mask area to total image area.

B Information of Internal Real-World Dataset

The internal real-world dataset supplements synthetic data with more complex human-product scenes and is split into training and held-out evaluation samples. Evaluation uses standardized preprocessing and input conditions across baselines for comparison.

  • The internal real-world dataset is collected from publicly available internet images and aligned with the synthetic dataset in resolution and aspect ratio.
  • Real-world samples contain varied scenes, human poses and interactions, product appearances, lighting, viewpoints, occlusions, and background clutter.
  • Approximately 14,000 samples are used for training and 2,000 separate samples for evaluation.The evaluation samples are not used during training.
  • All baselines use 1024 × 576 resolution and identical masked regions for fair comparison.
  • FLUX-Kontext receives a width-concatenated composite of the product reference and masked human images with an instruction prompt for object replacement.

D Evaluation on Real-World Data

HiFi-Inpaint and other baselines are evaluated on a 2,000-sample internal real-world test set with substantial variation in lighting, poses, and product appearance. Automatic metrics report overall state-of-the-art performance for HiFi-Inpaint.

  • The internal real-world test set contains 2,000 diverse human-product samples.
  • The test set varies substantially in lighting conditions, pose configurations, and product appearance.
  • Automatic metrics demonstrate HiFi-Inpaint’s overall state-of-the-art performance on real-world data.

D.1 Quantitative Comparison

HiFi-Inpaint remains highly competitive on challenging real-world data, combining strong text alignment, visual similarity, and structural preservation with competitive perceptual quality. Qualitative comparisons further show stronger integration and fine-detail preservation than existing methods.

  • 86.8 CLIP-I and 79.8 DINO are the highest reported visual-similarity scores, indicating strong preservation of product identity and local appearance details.These results are reported for the challenging real-world setting.
  • 60.5 SSIM and 44.1 SSIM-HF are the highest structural-similarity scores, reflecting preservation of object structure, text, logos, and fine patterns.
  • HiFi-Inpaint ranks third on both LAION-Aes (4.27) and Q-Align-IQ (3.29), while maintaining visually appealing and technically coherent outputs.The scores are slightly below the best-performing baselines.
  • Compared with existing methods, HiFi-Inpaint preserves fine-grained details in qualitative real-world comparisons and generalizes beyond synthetic data under realistic variations.
  • FLUX-Kontext often generates an isolated product instead of correctly integrating it into the masked region and can lose high-frequency details.

E Additional Results of Ablation Analysis

The ablation results show that SEA and DAL contribute to detail preservation, while the complete HiFi-Inpaint produces the strongest qualitative outputs. Additional examples emphasize seamless background integration across varied visual cases.

  • Ablation Analysis: Removing individual components causes noticeable degradation in detail preservation, whereas the complete HiFi-Inpaint consistently produces superior results.
  • Ablation Analysis: The complete model faithfully preserves critical details and achieves seamless integration with the background.
  • Ablation Analysis: HiFi-Inpaint with SEA and DAL achieves the best overall qualitative performance and superior detail preservation among the compared variants.

F Generalizability Analysis

HiFi-Inpaint is evaluated beyond standard inpainting conditions using challenging real-world cases. The evaluation spans missing humans, large pose variations, product interference, and substantial style adaptation.

  • Challenging Real-World Cases: The generalizability evaluation includes outdoor and indoor images without humans, full-body views with large pose variations, product interference, and substantial style adaptation.
  • Challenging Real-World Cases: These cases are designed to examine HiFi-Inpaint’s behavior beyond standard inpainting conditions.
Loading 2603.02210v3…