Source-linked AI summary

RewardFlow: Generate Images by Optimizing What You Reward

Onkar Susladkar, Dong-Hwan Jang, Tushar Prakash, Adheesh Juvekar, Vedant Shah, Ayush Barik, Nabeel Bashir, Muntasir Wahed, Ritish Shrirao, Ismini Lourentzou

arXiv:2604.08536v1cs.CVcs.AI

TL;DR

Existing image-editing methods face costly optimization and weaknesses in inversion-free control, including drift, leakage, and limited localization. RewardFlow uses inference-time multi-reward Langevin dynamics with adaptive scheduling and specialized VQA and object rewards; across benchmarks, it reports state-of-the-art editing fidelity and compositional alignment.

  • Problem

    Existing methods have limited generalization or suffer content drift, semantic leakage, weak localization, and inadequate coordination of heterogeneous objectives.

  • Method

    RewardFlow steers pretrained models at inference time with multi-reward Langevin dynamics, a prompt-aware adaptive policy, specialized VQA and SAM-guided rewards, and a KL tether.

  • Results

    Across multiple benchmarks, RewardFlow achieves state-of-the-art zero-shot performance in editing fidelity and compositional generation.

  • Takeaways & Limitations

    Reward-guided sampling provides a general test-time alignment strategy for fine-grained, spatially precise control while preserving identity and layout.

  • Takeaways & Limitations

    RewardFlow is limited by VQA reasoning failures, especially when counting small objects, which can make the reward signal uninformative.

Abstract

from arXiv · show

We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward Langevin dynamics. RewardFlow unifies complementary differentiable rewards for semantic alignment, perceptual fidelity, localized grounding, object consistency, and human preference, and further introduces a differentiable VQA-based reward that provides fine-grained semantic supervision through language-vision reasoning. To coordinate these heterogeneous objectives, we design a prompt-aware adaptive policy that extracts semantic primitives from the instruction, infers edit intent, and dynamically modulates reward weights and step sizes throughout sampling. Across several image editing and compositional generation benchmarks, RewardFlow delivers state-of-the-art edit fidelity and compositional alignment.

1. Introduction

RewardFlow addresses limitations of inversion and existing reward-guided methods with a zero-shot, training-free, inversion-free framework for controllable image editing and generation. It combines heterogeneous rewards with adaptive scheduling and specialized semantic and localization signals.

  • Existing fine-tuning methods require expensive optimization and generalize poorly beyond their training distribution.
  • Inversion-free methods avoid reconstruction but often suffer content drift, semantic leakage, weak localization, and limited fine-grained control.
  • RewardFlow guides pretrained flow-matching models at inference time through multi-reward Langevin dynamics without training or inversion.
  • The framework fuses global semantics, perceptual alignment, spatial grounding, aesthetic quality, and semantic faithfulness into one differentiable objective.
  • Differentiable VQA and SAM-guided rewards provide fine-grained semantic supervision and localized edits while penalizing leakage outside target regions.
  • A prompt-aware policy extracts semantic primitives and dynamically adjusts reward weights during denoising for coarse-to-fine optimization.

2. Related Work

Prior work spans expensive fine-tuning, latent optimization, reward-guided Langevin sampling, and trajectory-control formulations. These approaches can improve global alignment but generally lack localized rewards and adaptive control.

  • Training-Based Methods: Training-based approaches can achieve high fidelity but incur significant computational costs.
  • Latent-optimization and reward-guided methods address multi-objective feedback and zero-shot grounding, while newer formulations cast editing as trajectory optimal control.
  • Existing methods often lack localized reward models and adaptive control, leading to drift, over-editing, or weak spatial consistency.
  • RewardFlow unifies coarse-to-fine rewards with a prompt-aware policy for more adaptive and precise steering.

3. RewardFlow Method

RewardFlow guides pretrained flow-matching models at inference time by combining differentiable rewards with a prompt-aware policy that adapts reward priorities and update sizes throughout sampling. Its hierarchical rewards and localized gradients target semantic, perceptual, spatial, object-level, and human-preference objectives while limiting unintended edits.

  • Multi-Reward Langevin-Based Generation: RewardFlow evaluates differentiable rewards on decoded intermediate images, fuses them with adaptive weights, and maps the resulting gradient through the decoder–denoiser chain into latent-space drift.The fused drift augments the native flow-matching dynamics during test-time optimization.
  • Multi-Reward Langevin-Based Generation: The reverse-time update combines backbone drift, fused multi-reward guidance, and a clean-space KL tether that preserves identity.The method describes this update as a discretization of a Langevin SDE targeting a prompt-tilted density.
  • Prompt-Aware Adaptive Policy: The prompt-aware policy extracts semantic primitives, classifies edit intent, and dynamically adjusts reward weights, object direction, and step size during sampling.It uses prompt content, denoising time, and evolving generation state to control the closed-loop sampler.
  • Prompt-Aware Adaptive Policy: High rewards trigger smaller refinement steps, whereas low rewards trigger larger exploratory steps through a reward-aware adaptive step size.The policy therefore changes update magnitude according to proximity to the target.
  • Differentiable Rewards: Hierarchical SP-level and global prompt-level rewards provide spatial, perceptual, object-level, semantic, and human-preference control, with gradients concentrated on intended edit regions.The reward toolkit addresses semantic leakage caused by diffuse global-reward gradients.
  • Differentiable Rewards: A differentiable VQA-based reward supplies fine-grained semantic supervision for controllable image generation and editing.The paper identifies this integration as a first-of-its-kind contribution to inference-time controllable generation and editing.

4. Experiments

RewardFlow is evaluated on image editing and compositional generation benchmarks, where it improves fidelity, instruction alignment, and efficiency across diverse tasks and backbones.

  • Quantitative Results: RewardFlow achieves state-of-the-art editing performance on PIE-BENCH under a shared Flux backbone.It reduces Distance by 7.3% and improves PSNR, SSIM, Whole accuracy, and Edited accuracy relative to the strongest prior Flux-based baseline.
  • Efficiency: RewardFlow requires 43 NFEs and 20 sampling steps, roughly 60–80% fewer steps than typical gradient-based editors.In the four-step setting, it further improves Distance, LPIPS, Whole accuracy, and Edited accuracy over prior fast editors.
  • Qualitative Results: Qualitative comparisons show instruction-faithful, spatially precise edits that preserve surrounding layout and appearance.Examples include viewpoint changes, object replacements, animal attributes, and relational edits without the leakage or structural drift seen in other methods.
  • Quantitative Results: 12.5% overall improvement for Flux and 12.8% for Qwen Image on T2I-COMPBENCH, with gains across all six compositional categories.The improvements span attribute binding, object relationships, and complex multi-constraint prompts.
  • Text-to-Image Generation: Across PixArt-α, Flux, and Qwen Image, RewardFlow consistently improves compositional-generation accuracy.The benchmark covers fine-grained attribute binding, spatial and non-spatial relationships, and complex compositions.
  • Qualitative Results: Text-to-image examples show stronger prompt alignment, composition, aesthetics, color vibrancy, and local detail.RewardFlow also better realizes specified object counts and patterns in qualitative examples.

5. Ablation Studies

Ablations show that both the heterogeneous reward set and the adaptive policy contribute to spatially precise, instruction-faithful edits and visual consistency.

  • Reward Components: The full reward combination progressively concentrates gradients toward target regions and produces semantically precise edits.Global rewards support coherence, while object, region, and VQA rewards improve localization and instruction fidelity.
  • Reward Components: The VQA reward provides the strongest fine-grained supervision, reaching PSNR 32.09 and SSIM 90.21.Its gradient activations align with object contours across tasks.
  • Policy Components: The full adaptive-policy model achieves Distance 7.64, PSNR 32.09, and SSIM 90.21.Removing dynamic weighting, semantic primitives, or adaptive step sizes worsens the corresponding fidelity and consistency measures.
  • Policy Components: Excluding the KL tether causes the most severe degradation, with PSNR decreasing by 2.11 and SSIM by 1.89.The ablation is associated with structural drift and distortions in the edited content.
  • Policy Components: Removing semantic primitives raises Distance to 9.03 and lowers Whole by 2.33.The resulting interference between objectives produces inconsistent stylization.

6. Conclusion

RewardFlow steers pretrained text-guided models through multi-reward Langevin dynamics, combining adaptive control and a KL tether for precise, identity- and layout-preserving generation and editing.

  • Conclusion: RewardFlow combines global, localized, and VQA-based rewards with a prompt-aware adaptive policy and KL tether.The framework is training-free and operates during inference.
  • Conclusion: Extensive experiments demonstrate consistent improvements in edit fidelity, compositional alignment, and generation quality over strong training-free baselines.The paper presents reward-guided sampling as a general test-time alignment strategy.

Supplementary Material

RewardFlow generates high-resolution images in supplementary qualitative results.

  • Supplementary Material: Supplementary material presents high-resolution images generated by RewardFlow.

7. SDE Formulation

This section derives RewardFlow’s reverse update as an Euler–Maruyama discretization of an overdamped Langevin SDE targeting a prompt-tilted latent density. Reward and KL gradients augment the backbone drift, while a decreasing noise schedule moves sampling from exploration toward refinement.

  • 7. SDE Formulation: The prompt-tilted density combines the pretrained latent distribution with the total reward and KL potential, while z0 denotes the clean latent of an optional source image.The reward terms and KL potential are differentiated with respect to the latent variable.
  • 7. SDE Formulation: The KL drift is obtained by differentiating the KL potential and applies a restoring gradient toward the clean latent representation.Its coefficient is written as −λKL ∇z(k)K(z(k), tk; z0).
  • 7. SDE Formulation: Each image-space reward contributes a gradient that is transported into latent space through the decoder and denoiser Jacobians.Summing these contributions yields the fused reward drift gRtot,k.
  • 7.1. Langevin SDE and Discrete Update: The SDE uses an algorithmic time variable s and a monotone diffusion-time schedule t(s), with stationary distribution given by the prompt-tilted density ρt.The schedule satisfies t(0) = ¯t and t(S) = 0.
  • 7.1. Langevin SDE and Discrete Update: The reverse dynamics are an overdamped Langevin SDE driven by the score of the prompt-tilted density and Brownian noise with diffusion strength γ(s).γ(s) > 0 controls the diffusion strength.
  • 7.1. Langevin SDE and Discrete Update: Euler–Maruyama discretization uses step sizes ηk and evaluates the schedule and latent state at tk and z(k).The discretization defines tk = t(sk), γk = γ(sk), and z(k) ≈ z_sk.
  • 7.1. Langevin SDE and Discrete Update: The discrete score decomposes into the backbone drift fk, fused reward drift gRtot,k, and KL drift gKL,k.This is the stochastic update used in the main paper, with rewards providing controllability and the KL tether stabilizing identity and layout.
  • 7.2. Noise variance schedule γk: The noise variance γk decreases monotonically, using larger diffusion early for exploration and smaller diffusion late for refinement.When γmin = γmax, the schedule becomes a constant-noise Langevin sampler.

8. Datasets and Evaluation

The evaluation uses established dataset-specific protocols and metrics to assess compositional generation and fine-grained object-level alignment. T2I-COMPBENCH emphasizes attribute binding, spatial relations, and complex compositions, while GENEVAL probes object presence and related constraints.

  • T2I-COMPBENCH: T2I-COMPBENCH contains approximately 6,000 open-world prompts spanning attribute binding, object relationships, and complex compositions.Its prompts require alignment between textual semantics and visual structure across multiple objects and attributes.
  • T2I-COMPBENCH: T2I-COMPBENCH evaluates attribute binding with BLIP-based VQA, spatial relations with UniDet bounding-box analysis, and complex scenes with an aggregated consistency score.The aggregate combines CLIPScore, BLIP-VQA accuracy, and UniDet spatial-relation correctness.
  • Evaluation Protocol: The experiments adopt the evaluation protocols and metrics defined in the original papers for each dataset.This provides a consistent basis for comparison across the respective benchmarks.
  • GENEVAL: GENEVAL targets fine-grained object-level text-to-image alignment through object presence, cooccurrence, counting, spatial arrangement, and color attribution.Automated pipelines based on pretrained vision models evaluate these constraints.

9. Implementation Details

RewardFlow is implemented as inference-time optimization over semantic primitives and global rewards, with VQA and object-consistency signals supporting localized edits. The implementation uses pretrained backbones without weight updates and applies a KL tether only when a source image is available.

  • Implementation Setup: RewardFlow is implemented in PyTorch with automatic mixed precision on PixArt-α, Flux, and a Qwen-based latent diffusion backbone.Experiments use 1024 × 1024 resolution, official checkpoints and schedules, and a single node with 2× NVIDIA A100 GPUs.
  • Backbones and Resolution: Image editing encodes the source image into clean latent z0, initializes a noisy latent, and runs K = 35 reverse steps.Text-to-image generation instead samples z0 because no source image is provided.
  • Global and Perceptual Rewards: Global and perceptual rewards compute cosine similarities between the image and semantic-primitives text embeddings, with prompt-level scores aggregated over primitives.The aggregation is uniformly averaged and modulated by the adaptive policy.
  • Prompt Parsing: The prompt parser extracts semantic primitives, edit actions, visual attributes, and preservation constraints, then creates one Q&A pair focused on the final edited image.The parsing step is performed offline and cached for repeated sampling with the same prompt.
  • Region Grounding: Region grounding uses region proposals and soft attention weights to concentrate gradients on semantically and visually aligned regions.This reward is intended to match the localized behavior illustrated in Figure 9.
  • Object Consistency: Object consistency uses text-guided SAM2 to obtain soft masks and confidences for each semantic primitive.The object reward is further modulated by an add/remove intent scalar predicted by the adaptive policy.
  • Human Preference and VQA Rewards: The human-preference reward uses HPSv2 and normalizes its score with a fixed running mean and variance before combining it with other rewards.It is evaluated on the full prompt and primarily stabilizes aesthetic quality and prompt adherence.
  • Human Preference and VQA Rewards: The VQA reward obtains token-level logits from Qwen-2.5-VL 3B for generated question–answer pairs, with answer length capped at T⋆ ≤ 70 tokens.For text-to-image generation, the object-consistency reward is omitted and λKL = 0; editing uses λKL = 1.5.

10. Additional Results

Additional experiments report stronger compositional generation with RewardFlow across backbones and against ReNO and off-the-shelf models. The gains are especially pronounced on object counting, position, and color attribution, while larger VLMs improve reward estimation at added computational cost.

  • Text-to-Image Generation: RewardFlow consistently improves compositional faithfulness over backbone models and prior methods on GENEVAL.The evaluation is reported in Table 5.
  • GENEVAL Results: 0.45→0.65 and 0.64→0.81 are the mean-score gains from PixArt-α DMD and Flux, respectively, with further improvements over ReNO of +0.06 and +0.09.These values are reported as overall performance comparisons.
  • GENEVAL Results: 0.83 to 0.91 is the overall-performance improvement for the Qwen backbone, which surpasses ReNO on all metrics.Position rises from 0.27→0.47 and Color Attribution from 0.71→0.84.
  • GENEVAL Results: Qwen + RewardFlow achieves the best overall GENEVAL performance, exceeding SDXL, DALL-E 3, and SD3 (8B), whose mean scores are 0.55–0.68.This comparison is stated for overall GENEVAL performance.
  • Reward Ablation and Analysis: The reported gains are attributed to heterogeneous semantic, perceptual, regional, object-level, and QA-style rewards fused by a prompt-aware adaptive policy.The paper connects this feedback to correction of incorrect counts, swapped colors, and mislocalized objects while preserving backbone realism.
  • Qualitative Editing Results: The qualitative editing examples span global scene modifications, object-level edits, and fine-grained localized edits.Figure 12 presents source images, targeted edit instructions, and generated results.
  • VLM Comparison: Performance remains relatively stable when the VLM scales from 3B to 8B parameters, whereas Qwen3-Next-34B yields a noticeable ∼7% improvement across most metrics.The larger VLM also incurs increased computational overhead.

11. Additional Qualitative Results

RewardFlow produces diverse, fine-grained image edits across Flux and Qwen Image while preserving scene structure, identity, and instruction-relevant localization. Qualitative comparisons show more coherent localized transformations than several baselines, while the method’s VQA component can fail on fine-grained counting.

  • Qualitative editing results: Across examples, edits remain restricted to instruction-relevant regions and avoid semantic leakage into the rest of the scene.This behavior is reported for textured input images and includes preserving unrelated scene content.
  • Qualitative editing results: RewardFlow preserves camera pose and urban geometry while applying global scene changes, material transformations, and fine localized modifications with Qwen Image.Reported changes include rainy weather, fog, nighttime skies, velvet curtains, and removing embroidery.
  • Qualitative editing results: RewardFlow performs global, object-level, and pixel-level edits across diverse instructions while preserving background layout, identity, and coherent lighting.Examples include stylistic transformations, object removal and insertion, material changes, and localized attribute edits.
  • Qualitative comparisons: Compared with InfEdit, FlowEdit, FlowChef, InstantEdit, and KV-Edit, RewardFlow more coherently applies material and object transformations while preserving structure and scene context.The comparisons cover rusty bicycle frames, cat-to-labrador transformations, and cat-to-silver-sculpture edits.
  • Failure modes: A primary failure mode occurs when the VQA model cannot accurately count small objects, making its reward signal uninformative.The limitation concerns fine-grained reasoning rather than the general editing workflow.
  • Text-to-image generation: For text-to-image generation with Flux, RewardFlow is compared against vanilla Flux and Flux guided only by a global matching reward across diverse prompts.The reported prompts include portraits, nighttime street fashion, family cooking, and culturally specific festival scenes.
Loading 2604.08536v1…