Source-linked AI summary
Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets
Jialong Zuo, Haoyou Deng, Hanyu Zhou, Jiaxin Zhu, Yicheng Zhang, Yiwei Zhang, Yongxin Yan, Kaixing Huang, Weisen Chen, Yongtai Deng, Rui Jin, Nong Sang, Changxin Gao
TL;DR
The paper investigates whether Nano Banana Pro can serve as a generalist solver for traditional low-level vision, a capability that remains underexplored. It performs a broad zero-shot benchmark using simple prompts and finds strong perceptual quality but weaker pixel fidelity than specialist models, limiting its suitability for high-precision applications.
Problem
Whether Nano Banana Pro can generalize to traditional low-level vision tasks while meeting their strict pixel-fidelity requirements remains largely unexplored.
Method
The study systematically evaluates Nano Banana Pro zero-shot across 14 low-level vision tasks and 40 datasets using simple textual prompts, comparing it with specialist models.
Results
Nano Banana Pro shows exceptional perceptual quality and cross-task potential but trails domain-specific experts in traditional pixel-fidelity metrics.
Takeaways & Limitations
Nano Banana Pro is a capable zero-shot contender and semantic reconstructor, but generative low-level vision systems need evaluation and designs that reconcile perceptual quality with physical fidelity.
Takeaways & Limitations
The model is unsuitable for forensic, scientific, or other applications requiring outputs to correspond exactly to original scene data.
Abstract
from arXiv · showhide
The rapid evolution of text-to-image generation models has revolutionized visual content creation. While commercial products like Nano Banana Pro have garnered significant attention, their potential as generalist solvers for traditional low-level vision challenges remains largely underexplored. In this study, we investigate the critical question: Is Nano Banana Pro a Low-Level Vision All-Rounder? We conducted a comprehensive zero-shot evaluation across 14 distinct low-level tasks spanning 40 diverse datasets. By utilizing simple textual prompts without fine-tuning, we benchmarked Nano Banana Pro against state-of-the-art specialist models. Our extensive analysis reveals a distinct performance dichotomy: while \textbf{Nano Banana Pro demonstrates superior subjective visual quality}, often hallucinating plausible high-frequency details that surpass specialist models, it lags behind in traditional reference-based quantitative metrics. We attribute this discrepancy to the inherent stochasticity of generative models, which struggle to maintain the strict pixel-level consistency required by conventional metrics. This report identifies Nano Banana Pro as a capable zero-shot contender for low-level vision tasks, while highlighting that achieving the high fidelity of domain specialists remains a significant hurdle.
1 Introduction
The paper asks whether Nano Banana Pro can generalize to traditional low-level vision tasks and evaluates its zero-shot capabilities across diverse restoration, enhancement, and fusion settings. It finds a clear split: the model often produces perceptually appealing outputs, yet trails specialists on reference-based pixel-fidelity metrics.
- Evaluation scope: The study evaluates Nano Banana Pro zero-shot across 14 low-level vision tasks and 40 datasets using simple textual prompts.The benchmark covers image restoration, enhancement, and fusion without task-specific training.
- Qualitative performance: Across examples such as dehazing and deraining, the model can generate sharp edges and realistic textures that are often more aesthetically pleasing than specialist-model outputs.Figure 1 presents outputs across 14 tasks, including restoration, enhancement, and fusion, generated from simple prompts.
- Overall findings: Nano Banana Pro excels in perceptual quality but lags in metric-driven fidelity across low-level vision tasks.Its outputs can appear sharper and more realistic to human observers, while reference-based metrics such as PSNR and SSIM remain lower than those of domain-specific experts.
- Interpretation: Generative stochasticity prioritizes semantic plausibility over the strict pixel-wise alignment required by conventional reference-based metrics.This mismatch helps explain why visually convincing outputs can receive inferior quantitative scores.
- Real-world dehazing: In real-world dehazing, Nano Banana Pro can recover intricate details under severe haze but may introduce distorted colors and hallucinated weather elements.Reported failures include oversaturated hues and vivid blue skies in originally overcast or neutral scenes.
- Scope and implications: The model is better positioned as a creative enhancement tool than as a precise low-level restorer requiring color fidelity, physical realism, and consistency.The paper points toward hybrid approaches that combine generative strengths with task-specific physical constraints or refined prompting.
3 Super-Resolution
Nano Banana Pro is evaluated for Real-ISR against GAN- and diffusion-based methods using full-reference and no-reference metrics across synthetic and real-world benchmarks. Its outputs recover structures and produce natural-looking textures, but hallucinated content and altered local details reduce pixel-level fidelity.
- Quantitative Results: The evaluation compares Nano Banana Pro with GAN-based, multi-step diffusion, and accelerated diffusion methods using PSNR, SSIM, LPIPS, NIQE, MUSIQ, and CLIPIQA.Experiments use DIV2K-Val with 2,994 images alongside RealSR and DRealSR benchmarks.
- Quantitative Results: Nano Banana Pro significantly underperformed comparison methods on full-reference super-resolution metrics but achieved the best NIQE scores across DIV2K-Val, RealSR, and DRealSR.On DIV2K-Val, it achieved NIQE 3.52 versus 4.75 for BSRGAN, indicating stronger statistical naturalness despite altered pixel arrangements.
- Qualitative Results: Nano Banana Pro sharpens blurred edges and recovers linear patterns in architectural and lantern scenes while maintaining structural coherence and reducing noise.These qualitative improvements are shown in Real-ISR examples from DIV2K, RealSR, and DRealSR.
- Qualitative Results: The model expands image boundaries by hallucinating peripheral content instead of strictly preserving the input’s spatial extent.This unintended Field-of-View expansion reflects its lack of precise pixel-level alignment with the reference image.
- Qualitative Results: In complex textures, Nano Banana Pro generates sharp but altered leaf-vein and stone-surface structures, creating pixel-level discrepancies that lower fidelity scores.Its texture synthesis produces plausible details that are not faithfully restored from the reference.
- Analyses: Text reconstruction succeeds when degraded inputs retain recognizable structure but hallucinates incorrect strokes or characters when severe degradation obscures the original glyphs.The model’s dependence on semantic recognizability makes heavily degraded text particularly vulnerable to sharp but semantically incorrect outputs.
4 Deraining
Nano Banana Pro is evaluated for zero-shot single-image deraining across synthetic and real-world datasets using a fixed textual prompt. It produces visually plausible restorations but remains substantially weaker than specialist models on pixel-level fidelity, especially under heavy rain.
- 4.2 Experiment Setup: Zero-shot deraining uses one fixed textual prompt without training, fine-tuning, or adaptation across Rain200L, Rain200H, and SPA-Data.
- 4.3 Quantitative and Qualitative Results: 26.05 dB PSNR and 0.7954 SSIM on Rain200L, 21.10 dB and 0.6659 SSIM on Rain200H, and 32.25 dB and 0.9142 SSIM on SPA-Data remain below state-of-the-art deraining models.
- 4.3 Quantitative and Qualitative Results: Generative hallucination reconstructs local structures rather than restoring pixels exactly, explaining the gap between plausible visual outputs and PSNR/SSIM.
- 4.3 Quantitative and Qualitative Results: Under low-rain conditions, Nano Banana Pro better preserves color fidelity and fine details, whereas heavy rain causes pronounced color shifts and detail degradation.
- 4.3 Quantitative and Qualitative Results: The model can alter non-rain background elements and remove atmospheric haze, producing clearer-looking images that deviate from ground truth and lower quantitative scores.
5 Shadow Removal
Nano Banana Pro is assessed for single-image shadow removal against representative methods on the SRD dataset. It can produce visually convincing shadow removal, but generative alterations and limited shadow sensitivity reduce structural and pixel-level fidelity.
- 5 Shadow Removal: Nano Banana Pro effectively removes shadows while preserving original elements in some SRD examples.
- 5 Shadow Removal: Leading methods exceed 34 dB PSNR and 0.97 SSIM on SRD, while Nano Banana Pro records comparatively lower fidelity scores despite visually natural outputs.
- 5 Shadow Removal: Generative priors favor perceptual plausibility over exact reconstruction, causing altered textures, hallucinated details, or newly synthesized objects.
- 5 Shadow Removal: The model may leave faint shadows untreated or alter well-lit regions when detecting subtle, soft, or low-contrast shadows.
- 5 Shadow Removal: Resolution conversion and downsampling can smooth generated high-frequency details and further depress PSNR and SSIM.
6 Motion Deblurring
Nano Banana Pro shows strong perceptual restoration on selected synthetic and real-world motion-blur examples, including static textures, text, and difficult lighting. However, it introduces semantic and structural errors and scores well below specialist deblurring methods.
- 6 Motion Deblurring: On synthetic datasets, Nano Banana Pro can suppress severe blur and recover static architectural details, pavement textures, and text.
- 6 Motion Deblurring: Dynamic human scenes produce residual motion trajectories, duplicated clothing, ghosting, and hallucinated facial details.
- 6 Motion Deblurring: On real-world blur, it recovers text and handles low-light, overexposed, and high-dynamic-range scenes, but may alter faces and characters.
- 6 Motion Deblurring: 21.41 dB PSNR on GoPro and 21.35 dB on HIDE trail methods exceeding 33 dB on GoPro and 36 dB on RealBlur-R, while Nano Banana Pro’s SSIM ranges from 0.645 to 0.778.
- 6 Motion Deblurring: Generative high-frequency details improve visual sharpness but create pixel deviations that standard reference metrics penalize.
7 Defocus Deblurring
The zero-shot evaluation on DPDD and RealDOF shows that Nano Banana Pro is substantially weaker than specialized defocus deblurring methods. Its outputs often enhance contrast without reversing blur and can introduce structural hallucinations.
- 7.2 Quantitative Results: On DPDD, Nano Banana Pro achieves 20.180 dB PSNR and 0.635 SSIM, trailing GGKMNet by over 6 dB and showing similar weakness on RealDOF.
- 7.2 Quantitative Results: Across DPDD and RealDOF, the model consistently struggles to recover high-frequency details and often prioritizes global contrast enhancement over blur removal.
- 7.3 Qualitative Results: On DPDD, it behaves more like an enhancement filter: luminance and contrast increase while severe defocus remains, and text or geometry can be reconstructed incorrectly.
- 7.3 Qualitative Results: RealDOF cases show negligible deblurring in many spatially varying scenes, with occasional sharpness recovery accompanied by salt-and-pepper artifacts.
- 7.3 Qualitative Results: Perceptual reconstruction can hallucinate text and alter object scale because zero-shot generation lacks fidelity constraints.
9 Reflection Removal
Nano Banana Pro is evaluated for single-image reflection removal as a zero-shot, off-the-shelf model. It can produce perceptually strong results, but its high variance and generative deviations limit pixel-level fidelity and average performance.
- Quantitative Results: Nano Banana Pro lags specialist methods across all reflection-removal datasets on PSNR and SSIM because generative outputs can introduce intensity scaling and spatial shifts.Regression-based specialists optimize pixel-level reconstruction, whereas Nano Banana Pro prioritizes semantic coherence over structural fidelity.
- Quantitative Results: Despite natural-looking outputs, elevated LPIPS and sub-optimal MS-SSIM reveal semantic and stylistic deviations from reference images.On Postcard, LPIPS is 0.2513 for Nano Banana Pro versus 0.0549 for SOTA.
- Qualitative Results: Qualitative performance has a high ceiling but low stability: some samples approach ground truth quality, while others lose textures or misinterpret reflections as background.The model performs better when reflection and transmission layers are semantically distinct, but may misuse generative priors without domain-specific supervision.
- Failure Analysis: Failure cases include incomplete reflection removal, erroneous enhancement of reflections, chromatic domain shifts, texture-fidelity loss, structural hallucination, and compound degradation.Structural failures can hallucinate background structures or remove real objects mistaken for reflections.
10 Flare Removal
The flare-removal evaluation combines quantitative benchmarks with qualitative comparisons against specialist methods. Nano Banana Pro can deliver strong detail restoration and artifact removal, but stochastic behavior and perceptual-metric mismatch limit reliability.
- Qualitative Analysis: On favorable inputs, Nano Banana Pro can surpass specialists in detail restoration, but diffusion stochasticity causes variance and semantic hallucinations that undermine industrial reliability.Observed failures include unrelated content generation, valid light-source suppression, and erroneous illumination of inactive bulbs.
- Qualitative Analysis: Nano Banana Pro can preserve details near light sources and cleanly remove streak artifacts, but may introduce brightness changes.The brightness inconsistency is visible in the third row of the Flare7K++ examples.
- Metric Interpretation: Low quantitative scores can coexist with visually satisfactory flare removal when brightness or color differs from the ground truth.This divergence indicates that pixel-level metrics may not fully capture perceptual quality in generative reconstruction.
- Conclusion: Overall, Nano Banana Pro exhibits a high-ceiling, low-floor profile, trading perceptual potential for the stability and consistency required in robust restoration.The conclusion characterizes the model as potentially superior perceptually but not yet reliably consistent.
Image Enhancement
Nano Banana Pro is tested zero-shot for low-light enhancement across three benchmarks using a simple instruction and standard reference-based metrics. It produces visually reasonable, artifact-free results, but remains inconsistent and generally behind task-specific methods.
- Experimental Setup: The evaluation uses three low-light benchmarks, PSNR and SSIM, and comparisons with representative zero-reference, architecture-search, flow, transformer, and diffusion methods.Nano Banana Pro receives no fine-tuning or post-processing and is prompted to convert low-light images into normal images while preserving other elements.
- Quantitative Results: On LOLv1 and LOLv2-real, Nano Banana Pro falls considerably short of supervised state-of-the-art methods, while remaining competitive with several methods on SICE.On LOLv2-real, it achieves PSNR 15.661 dB and SSIM 0.537.
- Qualitative Results: Visual outputs often brighten dark regions and reveal scene content, but brightness control is inconsistent, with both overexposure and insufficient enhancement observed.Texture preservation is generally comparable to other approaches, and visible color, halo, noise, or structural artifacts are absent.
- Analysis: The artifact-free behavior reflects a conservative trade-off: it avoids common enhancement failures but may sacrifice peak benchmark performance.The remaining limitations include absent explicit illumination modeling, prompt sensitivity, and difficulty matching benchmark-specific ground truths.
12 Underwater Image Enhancement
Nano Banana Pro is evaluated for underwater enhancement across reference-based and reference-free settings spanning three datasets. It is strongest on severe degradations and no-reference evaluation, but less effective when degradation signals are mild.
- Quantitative Results: Nano Banana Pro achieves competitive UIQM and UCIQE results, reaching top performance on both metrics on the larger, more complex LSUI dataset.The evaluation compares against UWCNN, UIEC²-Net, U-Shape, PUGAN, DM-water, and WF-Diff.
- Qualitative Results: Qualitatively, it handles severe green or blue casts, low illumination, and high turbidity, sometimes producing visually superior results without paired ground-truth images.Its strengths include color-cast removal, blur reduction, and detail synthesis in heavily degraded scenes.
- Limitations: Performance is weaker for mild degradation, where the model may barely alter the input and produce results inferior to ground truth or competing methods.Weak degradation signals make it difficult to identify and localize features requiring fine-detail and contrast optimization.
- Conclusion: Across synthetic and real underwater benchmarks, the model trades robust perceptual recovery in complex scenes for precise pixel-level fidelity in benign conditions.The paper characterizes this as a distinct balance between strong recovery under severe degradation and weaker sensitivity to subtle defects.
13 HDR Imaging
NB Pro is evaluated for HDR reconstruction on HDR+ and MIT-FiveK using quantitative and qualitative comparisons. It produces visually plausible results in some scenes but trails mainstream methods on reference-based metrics and exhibits detail hallucination, loss, and oversharpening.
- Evaluation Setup: The HDR evaluation uses 500 aligned MIT-FiveK pairs and 250 HDR+ test pairs, combining PSNR, SSIM, LPIPS, ΔE, and subjective assessments.
- Qualitative Evaluation: In conventionally lit scenes, NB Pro restores dynamic range and color gradation, with visual quality approaching or sometimes rivaling the HDR reference.
- Qualitative Evaluation: In low-light, low-detail scenes, NB Pro may omit weak shadow textures or hallucinate redundant details absent from the original image.Examples include missing wall-tile and sofa textures, added cookie-surface texture, altered vegetable colors, and invented mountain elements.
- Qualitative Evaluation: Dense textures often trigger excessive sharpening, pronounced edges, localized artifacts, and masking of natural texture gradations.
- Quantitative Evaluation: NB Pro significantly lags mainstream HDR reconstruction methods across all reference metrics at 480p on HDR+ and MIT-FiveK, although the LPIPS gap narrows on MIT-FiveK.The evaluation uses PSNR, SSIM, LPIPS, and ΔE; the cited passage reports the overall comparison but not individual table values.
Image Fusion
The image-fusion evaluation compares NB Pro with established methods across multi-focus benchmarks. NB Pro often produces visually smooth, detailed images and strong non-reference scores, but source-reference metrics expose weakened pixel-level fidelity and occasional instability.
- Evaluation Setup: MFIF evaluation spans Lytro, MFFW, MFI-WHU, and SIMIF, using six objective metrics and comparisons against ten other methods.
- Quantitative Results: NB Pro performs strongly on non-reference MFIF metrics but poorly on source-reference metrics, trading source consistency for perceptual image quality.The comparison covers four benchmarks and ten representative methods, including supervised, unsupervised, and zero-shot approaches.
- Qualitative Results: NB Pro produces smooth transitions and sharp details in selected lock and coffee-cup examples while reducing boundary artifacts associated with defocus spread.
- Qualitative Results: The model handles complex fusion scenarios with fewer dark ghosting effects and unnatural transitions than traditional algorithms, indicating strong scene-structure understanding.
- Limitations: NB Pro sometimes blurs clear regions because focus detection and pixel localization remain unstable.
15 Infrared-Visible Image Fusion
The IVIF evaluation examines NB Pro on MSRS, RoadScene, and M3FD against ten representative methods. NB Pro achieves exceptional non-reference detail and clarity scores, but weaker source consistency and occasional hallucinations reveal a fidelity trade-off.
- Quantitative Results: NB Pro dominates non-reference IVIF metrics, ranking first on all four reported metrics for MSRS and exceeding RoadScene’s nearest SF competitor by nearly 43%.On MSRS, EN reaches 6.85 and AG 4.56; on RoadScene, SF is 21.81 versus EMMA at 15.21.
- Quantitative Results: Source-reference metrics are substantially weaker, with VIF reported as 0.58 on MSRS and 0.38 on M3FD.
- Quantitative Results: Performance is strongest on MSRS and RoadScene, while M3FD shows less pronounced overall dominance despite a second-best SD score of 43.44.
- Qualitative Results: Qualitatively, NB Pro recovers infrared targets in difficult lighting and reconstructs background texture, but may introduce unnatural hallucinations or artifacts in fine details.
- Analysis: The results expose a trade-off between exceptional visual clarity and strict source fidelity, reflecting generative enhancement rather than simple pixel preservation.
- Limitations: For safety-critical use, perceptual pleasantness may not ensure operational reliability because artifacts or oversharpening can affect downstream detection.
16 Discussion
The discussion characterizes NB Pro as a semantic generative reconstructor whose perceptual strengths conflict with traditional pixel-fidelity evaluation. It argues for broader evaluation and hybrid systems that combine generative detail synthesis with physical and structural constraints.
- Core Finding: Across 14 low-level vision tasks, NB Pro excels in perceptual quality but lags in traditional pixel-fidelity metrics.
- Evaluation Limitations: Full-reference metrics such as PSNR and SSIM assume a single pixel-perfect ground truth, which can penalize plausible reconstructions after severe information loss.
- Operational Scope and Limitations: NB Pro is positioned for creative enhancement and severely degraded photos, but not for forensic, scientific, or other applications requiring exact scene correspondence.
- Future Directions: The paper proposes hybrid generative-regression architectures with physics-informed constraints to combine structural recovery with detail enhancement.
- Future Directions: The evaluation used simple fixed prompts without meticulous tuning or multi-round inference, making the reported capability a conservative, unoptimized estimate.
- Future Directions: The authors call for evaluation frameworks that accommodate multiple plausible reconstructions and prioritize semantic consistency and naturalness alongside fidelity.
17 Appendix
The appendix lists task-specific prompts for 14 low-level vision tasks, consistently emphasizing targeted restoration while preserving scene content. Fusion prompts additionally require combining complementary information while minimizing artifacts.
- The prompts cover enhancement, restoration, deblurring, super-resolution, and image-fusion tasks with task-specific objectives.They include dehazing, denoising, reflection and flare removal, low-light and underwater enhancement, HDR imaging, fusion, super-resolution, deraining, shadow removal, and two deblurring settings.
- Motion deblurring uniquely forbids exposure correction, HDR tone mapping, lighting enhancement, and any change to original luminance levels.The prompt limits processing to removing motion streaks and recovering blurred details while leaving brightness, contrast, and exposure unchanged.
- Content preservation is a recurring requirement, with prompts prohibiting changes to objects, composition, colors, lighting, or original details beyond the targeted degradation.This constraint is explicit for underwater enhancement, shadow removal, motion deblurring, super-resolution, and defocus deblurring.
- HDR imaging targets recovery of clipped highlights and underexposed shadows while preserving midtone contrast, natural color, and the original forms of scene elements.Its instructions separately specify highlight recovery, shadow enhancement, and midtone preservation, while prohibiting added objects, artistic effects, and non-physical halos.
- Fusion prompts direct the model to combine complementary source information, selecting sharp regions or integrating thermal saliency with visible-image texture while suppressing artifacts.Multi-focus fusion uses local sharpness and boundary refinement; infrared-visible fusion combines thermal targets with visible structural detail.