Source-linked AI summary

Mi-Ripple: Restoring Images Degraded by Iterative AI Editing

Jiayin Chen, Yicheng Xu, Muting Wang

arXiv:2609.11317v1cs.CV

TL;DR

Iterative reference-conditioned editing can create grid-like and granular digital ripple, raising the challenge of removing artifacts without erasing legitimate structure. Mi-Ripple diagnoses whether degradation is spectrally separable, applies filtering or cleaned-reference regeneration accordingly, and links artifact reduction to cleaner outputs through measured and visual checks. Across fourteen notch-only executions, whole-image residual SD was 0.08–0.44 CIELAB lightness units, while the broader study identifies scope limits including single-draw comparisons and mixed-channel, mixed-canvas chains.

  • Problem

    Iterative reference-conditioned editing can propagate grid-like and granular digital ripple, creating a restoration problem involving artifact removal and preservation of legitimate detail.

  • Method

    Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture and routes candidates through selective notching, structure-aware suppression, cleaned-reference regeneration, and verification.

  • Results

    Mi-Ripple links measurable artifact reduction to visibly cleaner generated images; across fourteen notch-only executions, whole-image residual SD was 0.08–0.44 CIELAB lightness units.

  • Takeaways & Limitations

    Diagnosis-guided routing connects artifact measurements to restoration choices instead of treating a lower spectral score as sufficient evidence.

  • Takeaways & Limitations

    The study is limited by single-draw regeneration comparisons and mixed-channel, mixed-canvas chains that do not isolate causal editing effects.

Abstract

from arXiv · show

Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digital ripple. We present Mi-Ripple, a diagnosis-guided restoration workflow that suppresses this digital ripple while protecting image structure. Mi-Ripple separates periodic lattice artifacts from content-entangled granular texture, then combines selective spectral notching, structure-aware smoothing, and cleaned-reference regeneration. This separation enables low-distortion filtering when artifacts are spectrally isolated and visual reconstruction when filtering would erase legitimate detail. Across fourteen notch-only executions, whole-image residual standard deviation is 0.08--0.44 in CIELAB lightness units. In a paired regeneration example, reference cleaning reduces output debris density by 45\%. Mi-Ripple links measurable artifact reduction to visibly cleaner generated images, rather than optimizing a spectral score alone.

1 Introduction

Iterative reference-conditioned editing can propagate grid-like and granular artifacts that are inconspicuous in thumbnails, creating a restoration problem beyond ordinary denoising. Mi-Ripple addresses this by diagnosing artifact type and selecting filtering, suppression, regeneration, or review while protecting legitimate detail.

  • Motivation: Digital ripple includes grids, honeycomb-like patterns, and granular surfaces propagated by iterative reference-conditioned editing.The artifacts can remain inconspicuous in thumbnails despite being visible at native resolution.
  • Motivation: The restoration problem is determining whether an artifact is spectrally separable for safe removal or entangled with scene content and legitimate detail.
  • Contributions: The workflow evaluates restoration through visual examples and measured evidence, including restoration pairs, portrait cases, and aligned residual checks.
  • Contributions: Mi-Ripple combines isolated-peak notching, structure-aware suppression, and cleaned-reference regeneration to repair spectral contamination and degraded texture.
  • Contributions: Diagnosis-guided treatment distinguishes lattice from granular artifacts and selects filtering, regeneration, or review instead of applying one restoration operation universally.

2 Related Work

Prior work studies structured frequency discrepancies mainly for artifact detection or architectural explanation, while Mi-Ripple uses frequency analysis to guide practical restoration. It distinguishes treatable spectral contamination from content-entangled degradation without resolving artifact provenance.

  • Frequency-domain artifact research: Frequency-domain discrepancies have been linked to upsampling, aliasing, and related sampling effects in generated images.
  • Restoration perspective: Earlier work primarily detects synthetic images or explains artifact formation architecturally, whereas Mi-Ripple localizes a treatable component for restoration.
  • Iterative editing context: Self-consuming training studies concern degraded generative distributions, while this work studies fixed models whose recursion occurs through reference images at inference time.
  • Restoration perspective: Frequency-domain notch filtering is used as a classical treatment for periodic noise, with spatial distortion verification added to the restoration workflow.

3 Diagnosis-Guided Restoration

Mi-Ripple diagnoses periodic and granular artifacts with complementary spectral and spatial probes, then routes candidates to notching, masked suppression, cleaned-reference regeneration, or human review. Filtering is accepted only after image-difference checks and visual assessment.

  • Workflow: The workflow separates diagnosis, filtering or cleaned-reference regeneration, and verification while distinguishing deliverable-grade filtering from reference-grade cleaning.
  • Diagnosis: The diagnostic analysis uses CIELAB lightness, radial spectral baselines, autocorrelation, local spectral components, band-pass statistics, and whole-frame scale indices.Autocorrelation estimates periods separately from the diagnostic band-pass, while spatial probes cover dense foliage when flat windows do not qualify.
  • Treatment selection: Periodic lattice artifacts are treated with isolated-peak notching that preserves phase and normally modifies lightness alone, while compact peak selection avoids directional-content ridges.
  • Treatment selection: Granular or content-entangled regions use structure-protecting masked suppression, cleaned-reference regeneration, or human review rather than stronger filtering.Regenerated outputs are diagnosed again and require visual inspection because regeneration may invent detail and cannot use aligned pixel comparison.
  • Verification: Acceptance requires residual and retention checks, including structural residual SD at most 0.6 lightness units, high-frequency retention of at least 90%, and whole-image residual SD at most 1.0.Human review determines whether to adopt the candidate after these empirical checks.
  • Orchestration: Rule-based routing is the default, while an optional language-model layer selects only among permitted actions without changing numerical thresholds.Regeneration requires explicit permission and a bounded retry count.

4 Experimental Setup

The experiments sample two commercial editing channels through iterative chains, route comparisons, external sequences, and repeated image collections. Measurements are taken at native resolution or after specified normalization, but successive generations are not independent replicates.

  • Data and channels: The study samples two commercial editing channels without access to weights, seeds, or sampling parameters, treating them as access conditions rather than independently replicated models.
  • Editing protocol: The core experiments include same-scene and scene-change chains, prompt comparisons, and an eight-scene comparison of gpt-image-2 and gpt-image-2.5 routes.Each chain contains an initial output, gen0, and four edits through gen4.
  • External evidence: Banana100 contributes one starting photograph and 110 edited outputs from seven model families across eleven ten-step sequences.
  • Measurement protocol: Periodicity is measured at native resolution, while cross-resolution comparisons use Lanczos downsampling and retain separate native readings.The eight-scene comparison normalizes the long edge to 1280 pixels.
  • Measurement protocol: Successive generations are not independent replicates, and normalization does not remove differences in native-canvas sampling histories.

5 Restoration Evaluation

Mi-Ripple evaluates restoration by combining visual reconstruction, selective filtering, and measured distortion checks across scene types. The results show reduced artifacts, preserved structure under notch-only treatment, and scene-dependent outcomes for regeneration.

  • 5.1 Visual Restoration and Reference Cleaning: 45% lower output debris density follows radial-soft-clipped reference cleaning, decreasing from 1,842 to 1,020 components per megapixel.The comparison is a single-draw result, not an average treatment effect.
  • 5.2 Selective Suppression with Measured Distortion: Selective notching removes isolated lattice peaks while avoiding the broad tonal changes caused by radial-baseline soft clipping.On the garden-tilt example, the selected pre-feather mask covers 0.11% of frequency bins, and a background anomaly decreases from 3.49 to 1.77 with whole-image residual SD 0.18.
  • 5.2 Selective Suppression with Measured Distortion: Across fourteen notch-only executions, whole-image residual SD spans 0.08–0.44, while six additional hair-constrained candidates pass after notching with residual SD 0.25–0.44.The fourteen-execution range includes repeated use of one candidate rather than fourteen independent images.
  • 5.2 Selective Suppression with Measured Distortion: Strict masked suppression lowers sky and sea band-pass SD from 0.38 to 0.18 and from 1.21 to 0.71 while protecting hair and facial structure.All six initial portrait candidates pass distortion checks, with whole-image residual SD 0.21–0.49 and high-frequency retention 98.9–99.8%.
  • 5.3 Routing Across Scene Types: Treatment routing assigns selective notching to spectrally isolated artifacts and masked suppression, cleaned-reference regeneration, or review to content-entangled regions.The workflow therefore does not treat a lower artifact score as sufficient when filtering may erase legitimate structure.
  • 5.3 Routing Across Scene Types: Six of eight gen4 endpoints reached the none grade after restoration, with moss and wisteria reduced but remaining structured and ice cave reaching 0.0%.These single-chain observations indicate scene-dependent reduction rather than a universal restoration rate.

6 Characterization Results

The characterization results show that ripple signatures vary with access configuration, scene content, and repeated editing, while prompt constraints lack consistent control evidence.

  • 6.1 Lattice Signatures Depend on the Access Configuration: Spectral-anomaly medians were 1.73 for photographs and 1.95 for web references, versus 4.61, 4.17, and 4.13 for three generation sources.A photograph reached 3.12 because of JPEG blocking and repeated structures, so the measure is not specific to generated images.
  • 6.1 Lattice Signatures Depend on the Access Configuration: Channel B was lattice-positive in 43/43 outputs, while Channel A switched from 20/20 negative at 1280 × 720 to 6/6 positive at 1536 × 1024.These counts establish configuration dependence in the sampled conditions rather than invariance across prompts or images.
  • 6.1 Lattice Signatures Depend on the Access Configuration: Five Banana100 families showed persistent or late-emerging periods, while four showed strong positive autocorrelation-strength step correlations, with nonidentical sets.Different access paths and resolutions prevent direct vendor ranking, and changed wall texture can coexist with periodicity.
  • 6.2 Granular Texture Tracks Scene Content and Repetition: The phenomenon persisted on the gpt-image-2.5 route but redistributed across scenes, with gen4 scale percentages ranging from 0.0% to 41.1%.Coverage did not necessarily indicate visual conspicuity; the observations do not support model ranking.
  • 6.2 Granular Texture Tracks Scene Content and Repetition: Scene-change sequences returned to 1.1–2.1% in rainforest and remained at 0.0–0.5% in glacier after an intermediate 7.9% rainforest output.Channel and resolution also differed, so these observations do not isolate a causal scene-change effect.
  • 6.3 Prompt Constraints Do Not Provide a Consistent Control: Foliage-texture constraints increased the scale index in five of eight pairs and decreased it in three, with mean paired differences of +3.3, −1.3, and −0.3 percentage points.The two-sided sign-test p-values were 0.29, 0.73, and 0.73; repeated calls had a median observed range of 8.4 percentage points.

7 Discussion and Conclusion

Mi-Ripple uses diagnosis to route lattice artifacts to selective filtering and granular artifacts to reference cleaning and regeneration, then verifies candidate quality. The study concludes that the workflow connects measurable artifact reduction with cleaner reconstructed images, but remains limited by calibration, case-based comparisons, and unbenchmarked workflow advantages.

  • 7 Discussion and Conclusion: Diagnosis-guided restoration routes isolated spectral peaks to selective filtering and content-entangled texture to cleaned-reference regeneration with visual review.The workflow prioritizes restoration choices over optimizing a lower spectral score alone.
  • 7 Discussion and Conclusion: A descriptive recurrence models new artifact injection and reference carry-over, but its parameters are not fitted and do not identify an internal mechanism.The actionable implication is to preserve approved references and avoid unnecessary output-to-input chains.
  • 7 Discussion and Conclusion: The study-specific thresholds require broader calibration, and mixed-channel, mixed-canvas chains do not isolate causal editing effects.Single-draw regeneration comparisons demonstrate restoration routes rather than population-average gains.
  • 7 Discussion and Conclusion: The star-shaped workflow’s quality advantage remains unbenchmarked, and regeneration can alter details.Matched-channel studies and independently annotated images are identified as needed for broader treatment-success estimates.
  • 7 Discussion and Conclusion: The workflow combines diagnosis, compatible treatment selection, and candidate verification to produce cleaner reconstructed texture and small measured filtering residuals.Human review remains part of deciding whether to adopt a candidate.

A.2 Spatial Probes

The spatial probes distinguish structured tile-like texture from granular texture using complementary window and whole-frame measurements. Their treatment combines structure-aware suppression, reference-grade cleaning, and distortion checks with explicit limits.

  • A.2 Spatial Probes: The whole-frame scale index measures the percentage of qualifying 128-pixel tiles at stride 64 using blob coverage, area variation, circularity, and anisotropy.Scores below 2% are graded none, 2%–<6% suspected, and at least 6% structured for within-study triage.
  • A.2 Spatial Probes: Granular regions have weak autocorrelation maxima at inconsistent displacements, whereas the lattice produces repeatable period vectors.The approximately 5.17-pixel equivalent blob diameter is an instrument-preferred scale, not a physical period.
  • A.2 Spatial Probes: Strict spatial suppression subtracts diagnostic band-pass content through a structure-permission mask, while reference-grade cleaning permits broader suppression before regeneration with explicit face protection.The two routes use different edge, coherence, and texture-density bounds.

B.2 Sample Accounting and Canvas Dimensions

The study tracks channel, resolution, chain, and repair configurations across sampled outputs, while separating native and normalized measurements. Restoration comparisons combine reference cleaning, regeneration, and notching, but several endpoint values remain single-draw or composition-dependent.

  • Sample accounting: The 69-image lattice survey combines 43 channel-B images and 26 channel-A images, with channel-B chains and channel-A scene-change and portrait outputs sampled across resolutions.The survey must be distinguished from the separate 43 API calls in the prompt-decomposition experiment.
  • Trajectory illustration: The original two-chain illustration rises from spectral anomaly 3.028 to 3.904 and 4.204 after lighting and scene-action edits, while dominant periods remain approximately 4–5 pixels.Spatially separated patches show varied period vectors, supporting a widespread signal without uniquely identifying one oblique lattice.
  • Restoration controls: Cleaning the moss reference before regeneration lowers its scale index from 23.7% to 8.0% natively, while the foliage constraint yields 7.7% in a separate draw.After normalization, the corresponding values are 15.8% and 9.5%; the paired prompt study limits interpretation of the single constraint draw.
  • Restoration controls: Structure-aware reference cleaning lowers output band-pass SD for sea from 1.05 to 0.87 and railing from 2.61 to 1.77, but sky increases from 0.20 to 0.25.Hair and face measurements are treated as legitimate structure, and the cleaned-reference regeneration later reaches residual SD 0.21 with 99.7% high-frequency retention.
  • Restoration controls: Regeneration without a clean ancestor raises sharpness from 3.18 to 4.34, but the unseen ancestor scores 5.54 under different composition, so this is not a pixel-level recovery rate.The regenerated image is a plausible reconstruction rather than recovery against known ground truth.
  • Endpoint comparison: Table 4 compares eight scenes on a common 1280-pixel-long-edge canvas using one chain and one archived repair outcome per model–scene cell.Image 2 uses notch-only processing for C/G/S and regeneration for other scenes; Image 2.5 includes forced regeneration branches and separately finalized candidates.

C External Seven-Family Sequence Analysis

The external sequence analysis examines model-family periodicity and artifact accumulation under heterogeneous image supports and sampled configurations. It reports persistence separately from strength trends and avoids population-level vendor rankings or general success claims.

  • Dataset and support: The Banana100 subset uses one photograph and 110 edited outputs from seven model families, with eleven ten-step sequences and different-chat or same-chat variants.Window sizes vary across generated and original images, so absolute spectral comparisons across families have different support sizes.
  • Sequence analysis: Five families show persistent or late-emerging characteristic periods, while four show strong positive step correlations in autocorrelation strength; these groups overlap but are not identical.Qwen maintains a six-pixel horizontal period across all ten steps without increasing strength, illustrating why persistence and accumulation differ.
  • Sequence analysis: GPT output peak heights range from 0.06 to 0.16 across sequences versus approximately 0.05 for the photograph, but this weak signal is not generalized to every GPT access configuration.Flux.2 Max also changes wall content to cracked plaster, so its rising anomaly partly reflects changed scene content.
  • Scope: The sampled source comprises seven families and eleven sequences without independent replication across starting scenes, so the analysis infers neither population prevalence nor between-vendor quality ranking.The reported conclusions are bounded by the recorded snapshot and sampled starting conditions.
  • Artifact interpretation: Hair-like directional rendering changes are treated separately from lattice and granular artifacts because their energy forms directional lobes rather than isolated point signatures.Across twelve hair windows, anisotropy is 0.53–0.69; regeneration produces continuous strands, but without matched unconstrained repeats this is not a general success rate.
  • Sequence analysis: Table 5 reports dominant periods, occurrence counts, and Spearman correlations between autocorrelation strength and editing step across eleven Banana100 sequences.Persistence counts describe stability rather than significance, motivating separate reporting of periodicity and accumulation.

E.2 Implementation and Release Boundaries

The release separates modular diagnosis, processing, verification, and orchestration while bounding optional decision-making and regeneration. Reported software tests check implementation behavior, whereas restoration executions remain distinct from independent validation and pixel-level fidelity.

  • Implementation: The implementation separates image I/O, diagnosis, scale indexing, notching, masked suppression, reference cleaning, regeneration, verification, and orchestration.Rule-based routing is the default, and an optional language-model layer can select only among permitted actions.
  • Implementation: Execution records retain observations, allowed actions, decisions, and verification results, while regeneration requires explicit permission and a bounded retry count.The workflow reports suspected windows for human inspection rather than allowing unrestricted automated regeneration.
  • Validation boundary: The archived report contains nineteen passing synthetic and protocol tests covering lattice, granule, gradient, no-window, distortion, and mock-regeneration cases, but these are not independent detector-accuracy validation.One integrated attempt stopped after an HTTP 503 without delivery.
  • Fidelity boundary: Regeneration changes content and may return a different canvas, so the final aligned check measures subsequent filtering damage rather than fidelity to the original degraded image.This prevents interpreting post-regeneration alignment as recovery of the original image.
  • Release boundaries: The repository supplies processing modules and figure provenance, but the online-compilation package contains manuscript figures rather than the complete experimental archive.Source panels and experiment records are retained separately.
  • Release reporting: Table 8 reports whole-image distortion after processing, with SD measured in L* units and retention defined as the high-pass SD ratio.The table distinguishes initial channel-A candidates from constrained channel-B candidates, both at 1536 × 1024.
Loading 2609.11317v1…