Source-linked AI summary
NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)
Ya-nan Guan, Shaonan Zhang, Hang Guo, Yawen Wang, Xinying Fan, Tianqu Zhuang, Jie Liang, Hui Zeng, Guanyi Qin, Lishen Qu, Tao Dai, Shu-Tao Xia, Lei Zhang, Radu Timofte, Bin Chen, Yuanbo Zhou, Hongwei Wang, Qinquan Gao, Tong Tong, Yanxin Qian, Lizhao You, Jingru Cong, Lei Xiong, Shuyuan Zhu, Zhi-Qiang Zhong, Kan Lv, Yang Yang, Kailing Tang, Minjian Zhang, Zhipei Lei, Zhe Xu, Liwen Zhang, Dingyong Gou, Yanlin Wu, Cong Li, Xiaohui Cui, Jiajia Liu, Guoyi Xu, Yaoxin Jiang, Yaokun Shi, Jiachen Tu, Liqing Wang, Shihang Li, Bo Zhang, Biao Wang, Haiming Xu, Xiang Long, Xurui Liao, Yanqiao Zhai, Haozhe Li, Shijun Shi, Jiangning Zhang, Yong Liu, Kai Hu, Jing Xu, Xianfang Zeng, Yuyang Liu, Minchen Wei
TL;DR
AI Flash Portrait addresses the difficulty of restoring real-world low-light portraits while preserving noise control, detail, illumination, color, and aesthetic quality. The challenge establishes a real-world benchmark with paired data and evaluates solutions through region-aware objective metrics plus expert subjective assessment. Its final protocol weights reproducible objective scoring at 70% and expert blind-test scoring at 30%, reflecting a documented gap between metric performance and human perception.
Problem
Existing restoration approaches struggle to balance noise suppression, detail preservation, and faithful illumination and color reproduction in real-world low-light portraits.
Method
The challenge provides real-world paired portrait data and evaluates submissions with region-aware quantitative metrics, expert blind testing, and reproducibility verification.
Results
The final score combines reproducible objective assessment at 70% with expert subjective evaluation at 30%.
Takeaways & Limitations
Several models changed ranking after subjective testing, indicating a non-negligible gap between traditional image-quality metrics and human perception in real-world portrait generation.
Abstract
from arXiv · showhide
In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI Flash Portrait. Despite significant advancements in deep learning for image restoration, existing models still encounter substantial challenges in real-world low-light portrait scenarios. Specifically, they struggle to achieve an optimal balance among noise suppression, detail preservation, and faithful illumination and color reproduction. To bridge this gap, this challenge aims to establish a novel benchmark for real-world low-light portrait restoration. We comprehensively evaluate the proposed algorithms utilizing a hybrid evaluation system that integrates objective quantitative metrics with rigorous subjective assessment protocols. For this competition, we provide a dataset containing 800 groups of real-captured low-light portrait data. Each group consists of a 1K-resolution low-light input image, a 1K ground truth (GT), and a 1K person mask. This challenge has garnered widespread attention from both academia and industry, attracting over 100 participating teams and receiving more than 3,000 valid submissions. This report details the motivation behind the challenge, the dataset construction process, the evaluation metrics, and the various phases of the competition. The released dataset and baseline code for this track are publicly available from the same \href{https://github.com/zsn1434/AI_Flash-BaseLine/tree/main}{GitHub repository}, and the official challenge webpage is hosted on \href{https://www.codabench.org/competitions/12885/}{CodaBench}.
1. Introduction
AI Flash Portrait targets real-world low-light portraits that require simultaneous illumination enhancement, denoising, detail preservation, and aesthetic rendering. The challenge establishes a benchmark and evaluation framework intended to narrow the gap between academic methods and industrial applications.
- Low-light portraits suffer from severe noise, color distortion, and fine-detail loss because mobile cameras capture limited light.
- Existing low-light enhancement methods often distort skin tones, flatten facial lighting, and amplify background noise in portraits.
- Synthetic degradation data cannot reproduce the complex, nonlinear illumination shifts between weak-flash inputs and strong-flash references.
- The challenge provides professionally retouched real-world paired data to establish a benchmark for low-light portrait restoration and aesthetic enhancement.
- Its evaluation combines region-aware objective metrics with expert blind testing to assess portrait rendering and overall scene quality.
- The challenge seeks robust solutions suitable for deployment in practical real-world applications.
2.1. Training Data
Phase 1 supplies 600 real-world paired training groups for AI Flash Portrait development. Each group combines a 1K low-light input, a professionally retouched 1K reference, and a 1K person mask.
- 600 paired groups were released for Phase 1 model development.
- Each group contains a 1K low-light input image, a corresponding 1K ground-truth image, and a 1K person mask.
- The inputs are real photographs captured with weak flash rather than images from completely dark environments.
- Ground-truth images are professionally retouched to provide the aesthetic appearance of strong, studio-level flash.
- Participants may use public external datasets and pre-trained models, but must disclose those resources in their final reports.
2.2. Validation and Test Data
The challenge uses separate validation and hidden-test splits to assess generalization beyond paired training data. Phase 2 enables online objective feedback, while Phase 3 preserves confidentiality for final assessment.
- Phase 2 releases 100 low-light inputs and person masks while withholding ground-truth references.
- Participants submit Phase 2 outputs to CodaBench for immediate objective scores and iterative model refinement.
- Phase 3 adds a hidden test set of 100 sample groups for final expert evaluation under unseen degradation conditions.
- Phase 2 validation includes more challenging scenarios, such as distant portraits, to improve evaluation discrimination.
- The Phase 3 hidden set preserves the Phase 2 degradation difficulty distribution and remains confidential during the challenge.
- Final testing prohibits resizing and spatial padding by requiring outputs to match their input dimensions.
2.3. Evaluation Measures
The evaluation combines region-aware quantitative restoration measures with expert blind testing. Person-region fidelity, background quality, global structure, and human judgments jointly determine final rankings.
- The framework integrates region-aware quantitative measurement with expert-driven subjective evaluation.
- Person masks separate facial-detail and color evaluation from background illumination and global structural assessment.
- LPIPSperson and ∆Eperson measure masked-region perceptual similarity and color difference, while PSNRbg and SSIMglobal assess background and whole-image quality.
- The composite objective score discourages portrait oversharpening that sacrifices background cleanliness and global-PSNR optimization that flattens facial rendering.
- Experts rank outputs across 50 hidden samples and six dimensions including facial naturalness, detail preservation, and lighting realism.
- The final ranking combines subjective scores with objective scores using a 3:7 subjective-to-objective weighting ratio.
2.4. Phases
The challenge phases progressed from releasing aligned training data and a baseline, through online validation, to final reproducibility checks for top teams.
- Phase 1: Phase 1 released 600 aligned triplets containing low-light inputs, ground truths, and person masks, alongside an open-source baseline model.The resources supported degradation analysis, foundational model construction, and end-to-end pipeline validation.
- Phase 2: Phase 2 provided validation data for CodaBench submissions, automatically scoring results with objective metrics and updating the public leaderboard.The online environment supported algorithm validation and systematic hyperparameter tuning.
- Phase 3: Phase 3 required the top 12 teams to submit repositories, pretrained weights, and technical documentation for reproducibility and efficiency assessment.Documentation included training hardware configurations and inference time per image.
2.5. Awards
The challenge awarded one first-class prize, two second-class prizes, and three third-class prizes across each track.
- Each track offered one champion award worth US$1000.
- Each track offered two second-class awards worth US$500 each.
- Each track offered three third-class awards worth US$200 each.
2.6. Important Dates
The challenge schedule covered three data-release or competition phases from January through March 2026, ending with final-rank publication.
- 2026.01.23 marked Phase 1 data release and the beginning of Phase 1.
- 2026.01.28 marked Phase 2 data release and the beginning of Phase 2.
- 2026.03.05 marked the beginning of Phase 3.
- 2026.03.12 was the Phase 3 results submission deadline, followed by final-rank announcement on 2026.03.19.
3. Challenge Results
The challenge attracted broad participation and evaluated solutions through both objective validation and expert-reviewed final assessment. The final process addressed the mismatch between objective image-quality scores and human judgments of portrait aesthetics.
- Participation: 118 teams registered and submitted 3,187 valid entries during the competition.The top 12 teams advanced from Phase 2 to final evaluation with source code and pretrained weights.
- Phase 2: Quantitative Comparison: Phase 2 ranked submissions using quantitative objective metrics on validation data, with detailed results reported in Table 1.The table marks unreproducible or substantially divergent entries with hyphens.
- Phase 3: Comprehensive Evaluation: Phase 3 combined reproduced objective scores with expert blind-test assessment because high physical-signal scores do not always match human aesthetic preferences.The evaluation sampled predictions from the hidden test set and anonymized outputs from the shortlisted teams.
- Phase 3: Comprehensive Evaluation: The final score weighted reproduced objective performance at 70% and subjective UScore at 30%.The resulting scores determined the final rankings of the top 12 teams.
- Challenge Findings: Several models that excelled on Phase 2 objective metrics changed rank after expert subjective evaluation.The results indicate a domain gap between conventional image-quality metrics and human perception in real-world portrait generation.
4. Teams and Methods
The proposed methods for this track are described in the supplementary material due to space limitations.
- Track-specific team methods are deferred to the supplementary material because of space limitations.
6. Appendix: Teams and affiliations
The appendix presents a visual comparison of the top six teams and lists affiliations associated with named teams or contributors.
- Figure 4 compares restoration results from the top 6 teams against the low-light input and high-quality GT.
- The listed affiliations include Nanjing University of Science and Technology, Zhongxing Telecom Equipment, Jiangnan University, Zhejiang University, and University of Science and Technology of China.