Source-linked AI summary
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
Tianyu Liu, Weitao Xiong, Kunming Luo, Manyuan Zhang, Peng Li, Yuan Liu, Ping Tan
TL;DR
Rare adverse-weather synthesis requires costly data collection, while existing 3D-aware editing remains constrained by per-scene optimization and entangled geometry and illumination. AutoWeather4D uses feed-forward G-buffer dual-pass editing to separate geometric weather interactions from lighting control. It achieves comparable photorealism and structural consistency to generative baselines while offering parametric physical control as a practical autonomous-driving data engine.
Problem
Generative models require massive datasets for rare adverse-weather patterns, while 3D-aware editors face costly per-scene optimization and static-scene limitations in dynamic driving environments.
Method
AutoWeather4D uses feed-forward G-buffer extraction and dual-pass editing to decouple geometry from illumination in dynamic driving videos.
Results
AutoWeather4D achieves comparable photorealism and structural consistency to generative baselines while enabling parametric physical control.
Takeaways & Limitations
The framework serves as a complementary, practical data source for generating adverse autonomous-driving scenarios from existing footage.
Takeaways & Limitations
The method's pipeline is evaluated for sensitivity to imperfect intrinsic components and whether VidRefiner preserves or degrades physical correctness.
Abstract
from arXiv · showhide
Generative video models have significantly advanced the photorealistic synthesis of adverse weather for autonomous driving; however, they consistently demand massive datasets to learn rare weather scenarios. While 3D-aware editing methods alleviate these data constraints by augmenting existing video footage, they are fundamentally bottlenecked by costly per-scene optimization and suffer from inherent geometric and illumination entanglement. In this work, we introduce AutoWeather4D, a feed-forward 3D-aware weather editing framework designed to explicitly decouple geometry and illumination. At the core of our approach is a G-buffer Dual-pass Editing mechanism. The Geometry Pass leverages explicit structural foundations to enable surface-anchored physical interactions, while the Light Pass analytically resolves light transport, accumulating the contributions of local illuminants into the global illumination to enable dynamic 3D local relighting. Extensive experiments demonstrate that AutoWeather4D achieves comparable photorealism and structural consistency to generative baselines while enabling fine-grained parametric physical control, serving as a practical data engine for autonomous driving.
1 Introduction
AutoWeather4D addresses the data and optimization costs of adverse-weather video synthesis with a feed-forward 3D-aware pipeline that decouples geometry from illumination. Its G-buffer dual-pass design supports physically grounded weather and lighting edits for dynamic driving scenes.
- Generative video models require massive datasets to learn rare adverse-weather patterns, while real-world collection is expensive and logistically constrained.
- 3D-aware editing reduces data demands but can require up to an hour of per-scene optimization per video clip, limiting large-scale generation.
- AutoWeather4D replaces per-scene optimization with feed-forward editing that explicitly decouples geometry and illumination for rapid, high-quality, physically plausible weather and lighting edits.
- Feed-forward G-buffer prediction provides dense frame-wise geometry for dynamic scenes, bypassing static-scene assumptions and supporting spatially anchored weather effects.
- The G-buffer Dual-pass mechanism uses a Geometry Pass for surface-anchored weather interactions and a Light Pass for analytically accumulated local-illuminant contributions and 3D relighting.
- Experiments report comparable photorealism and structural consistency to generative baselines while providing parametric physical control for autonomous-driving data generation.
4 Experiment
AutoWeather4D is evaluated on dynamic Waymo driving sequences across weather and time-of-day conversions, using fidelity, structural, identity, and human-preference measures. Qualitative and quantitative results show physically plausible, spatially controlled editing, with continuous 4D reconstruction improving local illumination.
- Evaluation Setup: Evaluation uses 120 Waymo scenes and four target conditions: fog, midnight, rain, and snow.Baselines include Video-P2P, Ditto, Cosmos-Transfer2.5, and WAN-FUN 2.2.
- Evaluation Metrics: Editing Instruction Adherence, Structural Consistency, and Identity Stability quantify weather translation, geometric preservation, and foreground semantic invariance.Structural consistency uses relative bounding-box IoU, while identity stability uses patch-level CLIP similarity.
- Human Evaluation: 1,440 independent responses from 12 raters assess spatial fidelity and temporal coherence in standardized 512 × 512, 30-frame sequences.The study evaluates photorealism, instruction adherence, flicker, and motion continuity.
- Qualitative Results: AutoWeather4D enables precise control over shadows, lighting, geometry, weather conditions, and time of day in autonomous-driving videos.Its parameterization supports localized light-source and fog-density manipulation while preserving foreground geometry.
- Domain-Specific Comparisons: Compared with domain-specific architectures, the decoupled G-buffer representation preserves hard shadows, supports localized weather interactions, and anchors editing in explicit geometry.Reported interactions include snow accumulation, surface ripples, and spatial isolation of the sky region.
- Quantitative Results and Ablation: AutoWeather4D achieves performance comparable to existing baselines while adding deterministic physical control, and its continuous 4D reconstruction produces smooth, artifact-free illumination gradients.Integer-quantized depth causes spatial discretization and aliasing during local relighting, whereas the feed-forward reconstruction recovers a continuous floating-point manifold.
- Downstream Evaluation: Downstream proxy evaluations report consistent gains while absolute mIoU increments remain marginal, supporting geometric fidelity rather than advancing segmentation performance.The evaluation uses highly optimized zero-shot cross-domain baselines, which limits the observable mIoU improvement.
- Applications: Parameterized controls such as continuous fog scaling and local-light toggling support targeted challenging scenarios and future continuous perturbation sequences.Comprehensive downstream diagnostic benchmarking is left for future work.
5 Conclusion
AutoWeather4D combines deterministic graphics with video diffusion by anchoring synthesis to explicit G-buffer priors, providing physically grounded and parameterized adverse-weather simulation. The paper positions this framework as a complementary data source while identifying challenges in extreme dynamic interactions and severe environmental perturbations.
- Conclusion: Explicit G-buffer priors transform weather synthesis from entangled geometry and illumination toward physically grounded, parametric simulation.The framework is intended to complement purely data-driven generation for adverse-condition data creation.
- Limitations: Extreme long-tail dynamic interactions, such as complex vehicle-splash fluid dynamics, remain challenging for decoupled pipelines.Future work proposes localized generative priors for these unstructured phenomena.
- Limitations: Balancing severe perturbations such as heavy-fog occlusion with distant-background retention requires careful calibration.The current boundary conditioning anchors foreground geometries, while semantic-aware attention is identified as a future direction.
Supplementary Materials for AutoWeather4D
The supplementary materials describe metric calibration, sky-aware depth handling, and physical weather models implemented through G-buffer dual-pass editing. They cover snow, wet-ground, and falling-particle synthesis with explicit parameterizations.
- Metric Calibration: The supplementary pipeline calibrates reconstructed depth against LiDAR by minimizing mean squared error over matched points to solve scale s and bias b.The calibration samples N = 1000 non-sky, non-occluded LiDAR points selected via RANSAC.
- Sky-Aware Depth: Sky-aware processing assigns segmented sky pixels the 99th percentile of non-sky depth, preventing extreme sky depths from corrupting lighting calculations.Grounded-SAM extracts the sky mask using the prompt “sky”.
- Metric Calibration: In monocular fallback mode, known camera height determines scale, and metric depth is recovered as dmetric = s · d4D.The fallback assumes b ≈0 and uses the ratio of physical camera height to estimated relative height.
- Snow Synthesis: Snow accumulation uses a cascaded SPH Poly6 metaball kernel with smooth support, multi-scale buildup, and parameterized material blending.The support radius is ρ = 0.1 m, while the implementation uses three cascade levels with λ = 0.7 and k = 16 nearest metaballs.
- Wet Ground: Wet-ground synthesis darkens albedo according to porosity and water intensity while reducing roughness to reproduce water sheen.The implementation sets porosity p = 0.8, wetness intensity i = 0.5, and water roughness near zero.
- Falling Snow: Falling snow particles advance under gravity and wind using p_t+1 = p_t + (v_gravity + v_wind) · Δt.The system simulates 6,000 particles inside a view-frustum-aligned bounding box.
8 Rain Synthesis via G-Buffer Dual-Pass Editing
The rain synthesis pipeline models precipitation and wet surfaces in geometry-aware coordinates, combining procedural puddles, raindrop motion, SDF-based rendering, and ripple effects. G-buffer updates vary by surface type to preserve occlusion and wet appearance.
- Geometry-Anchored Puddle Modeling: Puddle boundaries are procedurally synthesized in explicit 3D world space because monocular reconstruction cannot observe millimeter-level micro-geometry.Fractional Brownian Motion is projected onto the reconstructed road manifold rather than applied as a 2D overlay.
- Geometry-Anchored Puddle Modeling: World-space FBM sampling anchors puddles to scene geometry, preserving perspective foreshortening, occlusion, and temporal consistency during camera motion.The resulting masks are integrated with deterministic G-buffer mechanics.
- Precipitation Dynamics: Raindrops are simulated with diameters from [0.5, 6.0] mm, physically modeled terminal velocities, wind perturbations, and collision or boundary resets.Drops initialize at heights uniformly distributed in [0, 51] m.
- Precipitation Dynamics: Each raindrop uses an uneven capsule whose streak length is 0.8 · Δt · v, producing asymmetric motion-blur geometry.The tail radius is set to rt = rh/0.7.
- Raindrop Rendering: Screen-space SDF rendering updates the G-buffer only for negative capsule distances, applying translucency blending and depth bias for occlusion.The capsule SDF subtracts an interpolated radius from the distance to its central axis.
- Surface and Atmospheric Effects: Rain effects modify atmospheric tint, surface roughness, ripple normals, and raindrop opacity according to surface-specific G-buffer rules.Drops use 40% opacity and a depth bias of 10^-4 m, while puddle ground roughness is reduced to 0.0.
9 Dense Semantic Annotations
Dense semantic annotations provide structural priors for weather synthesis and local relighting. Detection, bidirectional propagation, 3D reprojection, and spatial clustering identify roads, vehicles, and street-light sources.
- Annotation Pipeline: A detect-segment-propagate pipeline produces semantic tracks that constrain precipitation, accumulation, occlusion, and illumination rendering.Street lights localize emissive sources, roads confine accumulation, and vehicle masks identify dynamic occluders.
- Initialization: OWL-ViT proposals for street lights, roads, and cars seed precise SAM2 masks on uniformly sampled keyframes.The system retains the maximum-area road region and filters low-confidence detections.
- Temporal Propagation: Bidirectional SAM2 propagation extends masks across time while mitigating flicker and preserving object identity through occlusions.The propagated annotations support downstream depth analysis.
- 3D Reprojection: Metric-depth reprojection transforms street-light mask pixels into a unified world-space point cloud using camera poses and intrinsics.The aggregated cloud contains potential street-light structures across all frames.
- Instance Grouping: DSU clustering connects points within τdist = 0.5 m to merge fragmented observations into spatially coherent street-light instances.Connected components define distinct 3D clusters Ck.
- Light-Source Localization: Illuminant centers are estimated from the centroid of each cluster’s top 5% of points, while spotlight directions are aimed downward toward the road.This parameterization supports physically plausible angular attenuation in the Light Pass.
10 Nocturnal Local Relighting via G-Buffer Dual-Pass Editing
The nocturnal relighting pipeline combines a Cook-Torrance BRDF with multiple attenuated spotlights and global tone mapping. It models local incident radiance from street lights and headlights while separately handling sky regions.
- BRDF and Light Transport: Surface radiance is computed by integrating the Cook-Torrance BRDF against incident illumination over the hemisphere.The formulation uses world-space surface position, view and incident directions, surface normal, and material properties.
- BRDF and Light Transport: The Cook-Torrance model combines diffuse and specular reflection using metallic-roughness parameters, Fresnel reflectance, GGX distribution, and Smith visibility.The specular formulation follows the Filament/Disney PBR convention.
- Incident Radiance: Incident radiance sums RGB contributions from discrete sources with inverse-square distance falloff and combined angular-distance attenuation.Each source contributes through its radiant intensity, 3D position, and attenuation factor.
- Spotlight Modeling: Spotlight angular attenuation smoothly transitions between inner and outer cones, using 15°/35° for street lights and 10°/25° for headlights.The attenuation depends on the angle between the spotlight direction and the light-to-surface vector.
- Spotlight Modeling: Finite light influence is enforced by a clamped fourth-power falloff with typical radii of 10–20 m for street lights and 30–50 m for headlights.The falloff prevents contributions outside each light’s computational domain.
- Global Tone Mapping: A parametric 256×3 LUT darkens ambient regions and shifts color temperature while preserving visibility in areas without artificial lighting.Adaptive exposure and highlight compression are applied before LUT processing to reduce over-darkening.
- Sky Handling: Sky masks receive differential processing, including αsky = 0.6 darkening, mask dilation, and Gaussian blur for natural horizon transitions.The treatment prevents unrealistic sky brightening during nocturnal relighting.
11 Volumetric Fog Synthesis via G-Buffer Dual-Pass Editing
The framework models fog through simplified radiative transfer, combining attenuated surface radiance with accumulated in-scattered light. Density scaling and final color blending provide parametric control over fog appearance.
- Radiative Transfer: Observed radiance combines transmittance-attenuated surface radiance with accumulated in-scattered light.The transmittance uses extinction from absorption and scattering, while in-scattering accounts for atmospheric light contributions.
- Radiative Transfer: In-scattering sums contributions from all light sources using scattering strength, attenuated intensity, and a Henyey-Greenstein phase function.The phase function uses forward-scattering parameter g = 0.8, and absorption and scattering coefficients are density-scaled.
- Scattering Model: The Henyey-Greenstein phase function models forward scattering between the view direction and light directions with g = 0.8.The scattering angle is defined between the view direction and each light direction.
- Light Sources: Directional, point, and spot lights receive appropriate attenuation, including distance falloff and spot-light angular attenuation.Fog density can be adjusted in real time through the scaled density parameters without recomputation.
- Color Blending: Final color blends observed radiance with fog color according to fog opacity and blend strength β = 0.5.Fog opacity is defined as f = 1 −T(s), and β controls the artistic blend contribution.
12 Environment Harmonization
Environment harmonization combines ambient and directional or local illumination in linear light space, while text prompts specify scene content and target weather or time-of-day conditions. Adaptive weighting preserves stronger direct-light regions while retaining ambient illumination elsewhere.
- Illumination Fusion: Ambient and directional or local illumination are converted from sRGB to linear space before fusion to preserve additive light relationships.The linear-space operation avoids distortions caused by nonlinear gamma encoding.
- Illumination Fusion: Adaptive weighting uses clamped per-pixel direct-light illuminance, with a 0.05 lower threshold to suppress weak or noisy signals.The illuminance mean is computed across RGB channels and constrained to [0, 1].
- Illumination Fusion: The blended illumination is a linear combination of ambient and direct maps weighted by W_direct.Stronger direct illumination dominates locally, while regions without direct light retain the ambient background.
- Illumination Fusion: Spatially adaptive linear fusion supports physically plausible light mixing while avoiding over-saturation and unnatural transitions.The strategy respects the intensity hierarchy between global ambient illumination and directional or local sources.
- Prompt Construction: The final prompt combines a base scene description with a conversion instruction for the target scenario.The base prompt is generated from the original frame and covers layout, objects, motion, atmosphere, and traffic controls.
- Scenario Conditioning: Dedicated conversion prompts guide realistic weather and time-of-day edits with cinematic quality, physical details, and semantic consistency.Examples specify midnight lighting, fog scattering, snow accumulation, and rain-soaked surfaces with corresponding physical rationales.
- Scenario Conditioning: Snow and rain prompts encode condition-specific geometry, motion, and illumination cues such as falling particles, wet surfaces, ripples, and diffuse light.Their rationales connect these cues to snow accumulation, partial melting, hydrological detail, and atmospheric lighting.
- Scenario Conditioning: Rain conversion emphasizes wet roads, rain ripples, raindrops, and overcast diffuse illumination for a believable urban environment.The design rationale balances cinematic presentation with physically accurate rain morphology and motion.
14 VidRefiner Architecture and Configurations
The postprocessing configuration balances physical-edit preservation with artifact repair through moderate editing strength, prompt guidance, and a limited diffusion schedule. The collapsed searching-space scheme exposes these controls over the input video and latent denoising process.
- Postprocessing Configuration: Editing strength α is typically 0.4, injecting moderate noise to refine details without overwriting the weather simulation.Higher α permits stronger deviations, while the stated default targets a balance between preservation and refinement.
- Postprocessing Configuration: Classifier-free guidance uses γ = 6 to emphasize adherence to the conversion prompt.The configuration is part of the postprocessing stage alongside editing strength and diffusion steps.
- Postprocessing Configuration: The diffusion process uses T = 20 inference steps with the default scheduler in WAN-FUN 2.2-5B.The configuration is selected for the postprocessing pipeline described by the authors.
- Generation Scheme: The collapsed searching-space generation scheme operates on input video v, editing strength α, diffusion timesteps T, latent z, and a denoising pivot timestep.The pseudocode defines postprocess(v, prompt, α, T) as the entry point for this procedure.
15 Extended Quantitative Evaluations
Extended evaluation measures structural alignment, instruction adherence, perceptual quality, and temporal consistency across generated weather videos. The results favor AutoWeather4D for instruction and vehicle alignment, while several metrics require careful interpretation because physically grounded edits can diverge from unedited-video references.
- Structural Evaluation: Structural consistency is measured by 2D IoU between projected ground-truth and predicted bounding boxes.The formulation divides box intersection area by union area.
- Metric Caveats: A drop in IoU may reflect either geometric degradation or reduced OvMono3D detector performance under severe out-of-distribution weather.Thus, the metric is a proxy that conflates editing effects with detector robustness.
- Instruction Adherence: AutoWeather4D achieves the best CLIP score and vehicle geometric and perceptual alignment among the baseline approaches.The authors interpret these results as the strongest instruction adherence in the comparison.
- Depth Alignment: Depth alignment is evaluated with scale-invariant RMSE across 120 sequences covering fog, rain, snow, and night.Lower si-RMSE denotes better depth alignment, and the reported table averages results over 480 generated videos.
- Depth Alignment: The minor depth-alignment drop is attributed to explicit depth generation for falling snow and rain, which most baselines cannot render geometrically consistently.This comparison frames the metric decrease alongside the added geometry of weather particles.
- Edge Alignment: Edge F1 evaluates Canny-edge alignment, but original-video edges favor methods that preserve input geometry.AutoWeather4D modifies surfaces, normals, reflectance, and illumination, so its physically edited structures diverge from the reference edges.
- Edge Alignment: The lower Edge F1 is described as expected structural deviation from physically grounded editing rather than failure of edge preservation.The interpretation follows from using unedited-video edges as the ground truth.
16 Comprehensive ablation studies
The ablations isolate module contributions, test module synergies, and examine error tolerance under degraded intrinsic inputs. They support the roles of the reconstruction, Geometry Pass, Light Pass, and VidRefiner components.
- Ablation protocol: The ablation protocol compares full-pipeline outputs with module-ablated variants, using internal PSNR as a pixel-level spatial-divergence measure rather than perceptual quality.Lower internal PSNR indicates greater structural and photometric deviation from the full-pipeline reference.
- Module effectiveness: Single-module studies evaluate the 4D reconstruction and shadow-manipulation components, while tabulated experiments compare results with and without Geometry Pass and Light Pass editing.The study also includes rain, snow, fog, and night configurations.
- Module synergies: VidRefiner is evaluated alone and in combinations with functional modules to test whether it preserves physical edits while improving spatial fidelity.The synergy studies cover two- and three-module subsets.
- Geometry Pass Editing: The Geometry Pass ablations separately assess standing and falling water for rain, and accumulated snow, falling snowballs, and grid-based snow for snow.These experiments target the individual contributions of weather-specific geometric components.
- Light Pass Editing: The Light Pass ablation varies active illuminants from 0 to 4, 8, and all sources to examine individual and cumulative lighting effects.The progression evaluates lighting fidelity and scene photorealism.
- VidRefiner sensitivity: 0.4 VidRefiner strength is selected over 0.6 because it achieves PSNR=10.19 without the semantic car-color error observed at PSNR=10.36.The stronger setting yields higher quantitative fidelity but changes a black car to white, whereas 0.4 avoids that error.
- Error tolerance: An extreme-low-light case tests pipeline fragility under flawed depth, normal, and roughness extraction, showing that the decoupled design compensates for these errors.The case study reports that the pipeline breaks the chain of cascading errors through physical constraints.
17 Extensive Qualitative Results
Qualitative comparisons show that AutoWeather4D decouples illumination from geometry and handles dynamic objects more reliably than baselines. It also provides localized lighting and active weather interactions, with a documented limitation for self-illuminating objects.
- Local lighting: The Light Pass injects volumetric headlight cones, preserving dark unlit backgrounds while creating localized illumination, falloff, and road-surface reflections.This addresses the inability of global tone shifts to model active local light sources.
- Dynamic reconstruction: Feed-forward G-buffer extraction accommodates dynamic objects without monolithic 4D optimization, whereas 4D-Gaussian baselines exhibit motion ghosting and unstable vehicle geometry.The comparison is especially relevant for moving vehicles in monocular driving videos.
- Illumination decoupling: AutoWeather4D recalculates light transport to remove inherited directional shadows and produce diffuse illumination in rainy and snowy scenes.Baselines often retain source shadows because they lack explicit intrinsic decomposition or illumination decoupling.
- Rain editing: AutoWeather4D models ongoing rain with active precipitation and surface interactions, while several baselines mainly produce global color shifts or static wet-road appearances.Ditto attempts ripples but can distort the original scene layout.
- Limitation: Traffic lights can lose their original glow during sunny-to-rainy conversion because the current G-buffer omits a dedicated emissive channel.The system therefore represents them as high-albedo reflective surfaces dependent on external illumination.