Source-linked AI summary
NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Professional Image Quality Assessment (Track 1)
Guanyi Qin, Jie Liang, Bingbing Zhang, Lishen Qu, Ya-nan Guan, Hui Zeng, Lei Zhang, Radu Timofte, Jianhui Sun, Xinli Yue, Tao Shao, Huan Hou, Wenjie Liao, Shuhao Han, Jieyu Yuan, Chunle Guo, Chongyi Li, Zewen Chen, Yunze Liu, Jian Guo, Juan Wang, Yun Zeng, Bing Li, Weiming Hu, Hesong Li, Dehua Liu, Xinjie Zhang, Qiang Li, Li Yan, Wei Dong, Qingsen Yan, Xingcan Li, Shenglong Zhou, Manjiang Yin, Yinxiang Zhang, Hongbo Wang, Jikai Xu, Zhaohui Fan, Dandan Zhu, Wei Sun, Weixia Zhang, Kun Zhu, Nana Zhang, Kaiwei Zhang, Qianqian Zhang, Zhihan Zhang, William Gordon, Linwei Wu, Jiachen Tu, Guoyi Xu, Yaoxin Jiang, Cici Liu, Yaokun Shi
TL;DR
Conventional scalar IQA does not adequately distinguish uniformly high-quality images or explain comparative judgments. This challenge benchmarks MLLMs on expert selection and reasoning for high-quality image pairs, finding robust comparative selection and grounded interpretative reasoning among top-performing methods while exposing persistent rationale-alignment challenges.
Problem
Conventional IQA compresses complex visual characteristics into scalar scores, limiting subtle high-quality comparisons and explanations of why one image is superior.
Method
The challenge uses expert-annotated image pairs with binary superior-image choices and textual rationales, evaluated through comparative accuracy and reasoning-quality measures.
Results
Top-performing methods demonstrated robust comparative quality selection and grounded interpretative reasoning on the professional IQA benchmark.
Takeaways & Limitations
The benchmark supports MLLMs as a viable approach for connecting automated professional image assessment with human expert cognition.
Takeaways & Limitations
Additional training-data curation through manual annotation or closed-source VLMs was prohibited to support fairness and reproducibility.
Abstract
from arXiv · showhide
In this paper, we present an overview of the NTIRE 2026 challenge on the 3rd Restore Any Image Model in the Wild, specifically focusing on Track 1: Professional Image Quality Assessment. Conventional Image Quality Assessment (IQA) typically relies on scalar scores. By compressing complex visual characteristics into a single number, these methods fundamentally struggle to distinguish subtle differences among uniformly high-quality images. Furthermore, they fail to articulate why one image is superior, lacking the reasoning capabilities required to provide guidance for vision tasks. To bridge this gap, recent advancements in Multimodal Large Language Models (MLLMs) offer a promising paradigm. Inspired by this potential, our challenge establishes a novel benchmark exploring the ability of MLLMs to mimic human expert cognition in evaluating high-quality image pairs. Participants were tasked with overcoming critical bottlenecks in professional scenarios, centering on two primary objectives: (1) Comparative Quality Selection: reliably identifying the visually superior image within a high-quality pair; and (2) Interpretative Reasoning: generating grounded, expert-level explanations that detail the rationale behind the selection. In total, the challenge attracted nearly 200 registrations and over 2,500 submissions. The top-performing methods significantly advanced the state of the art in professional IQA. The challenge dataset is available at https://github.com/narthchin/RAIM-PIQA, and the official homepage is accessible at https://www.codabench.org/competitions/12789/.
1. Introduction
Conventional IQA compresses complex visual quality into scalar scores, limiting subtle comparisons and explanations. MLLMs offer a route toward expert-like, grounded comparative reasoning for professional image assessment.
- Motivation: Conventional IQA typically regresses visual characteristics into scalar Mean Opinion Scores, limiting practical interpretation.Scalar outputs struggle to distinguish subtle differences among uniformly high-quality images and do not explain why one image is superior.
- Motivation: MLLMs can combine visual perception with linguistic generation to formulate grounded comparative reasoning.This approach explicitly represents features such as subtle textures, noise, and artifacts rather than reducing them only to latent representations.
- Challenge Setup: Track 1 provides expert-annotated high-quality image pairs containing both the superior-image selection and detailed textual rationale.The dataset is designed to evaluate comparative and reasoning capabilities against nuanced feedback from professional practitioners.
- Challenge Setup: The challenge focuses on professional image quality assessment within the broader NTIRE 2026 challenge ecosystem.Earlier RAIM editions emphasized enhancement, restoration, and model efficiency, whereas this track introduces professional IQA.
2. NTIRE 2026 the 3rd RAIM Track 1 PIQA
Track 1 constructs a reasoning-based benchmark from expert comparisons of matched images and evaluates both superior-image selection and rationale quality. Its multi-phase protocol combines accuracy, text-based reasoning, and semantic judgment under stated data and participation constraints.
- Training Data: Each sample pairs images of the same scene captured by different devices, with experts selecting the superior image and explaining the choice in text.The selection is formulated as a binary Multiple Choice Question, with additional annotations supporting fine-grained observation.
- Training Data: The training set contains 100 annotated pairs: 75 portraits and 25 landscape or scenery pairs.The collection spans varied environments, subjects, and complex lighting conditions.
- Validation and Test Data: Validation and test splits introduce subjects unseen in earlier splits to assess generalization.The validation set has 102 pairs and the test set has 101 pairs, with new subjects withheld across splits.
- Ranking Metrics: Evaluation separates comparative accuracy from reasoning quality, measuring correct superior-image selection and interpretability of generated rationales.Reasoning is assessed through conditional text metrics and independent semantic evaluation.
- Ranking Metrics: Conditional reasoning quality averages BLEU-4 and ROUGE-L only when the model selects the correct image.This condition prevents elaborate justifications for incorrect choices from receiving reasoning credit.
- Ranking Metrics: An LLM-as-a-Judge evaluates semantic alignment between predicted and expert rationales across photographic aspects such as sharpness, texture, realism, and noise.The protocol supplements conventional NLG metrics, which may miss semantic equivalence and professional nuance.
- Ranking Metrics: Phase 2 weights accuracy and conditional reasoning through S_Phase2 = Accuracy × (w_base + w_reasoning × S_thinking), with weights 0.7 and 0.3.Phase 3 adds a scaled LLM-judge component, and the final ranking balances performance across both phases.
- Challenge Protocol: Participants receive a 100-pair training set, validation-server feedback, and a single final test submission under the competition protocol.Additional data and open-source pretrained models are allowed, but manual annotation and closed-source VLM curation are prohibited.
3. Challenge Results
The challenge drew substantial participation and produced a closely contested field led by IH-VQA and VCIP Pi Group. Models generally selected the better image successfully, but expert-aligned reasoning remained difficult and context-dependent.
- Participation: 192 participants registered, 85 contributed 2,530 validation submissions, and 57 teams submitted once in the final evaluation.The final phase enforced a strict single-attempt rule.
- Quantitative Comparison: IH-VQA ranked first in the Final Award Score, while VCIP Pi Group achieved superior Phase 3 performance through a notably high LLM Score.The leading teams remained narrowly separated across the competition.
- Quantitative Comparison: Overall rankings favored balanced performance across all evaluation criteria rather than excellence in only one metric.Some teams excelled on specific metrics, but the final standings emphasized multi-axis capability.
- Qualitative Comparison: Models generally handled binary selections better than generating rationales aligned with experts’ nuanced photographic criteria.Assessment criteria can vary substantially with image content, shooting conditions, and artifacts, leaving reasoning evaluation open.
4. Teams & Affiliations
Figures 4 and 5 provide visual comparisons of generated rationales and selections with ground-truth annotations and selections on curated test-set subsets.
- Figure 4 compares generated rationales and selections with ground-truth annotations and selections on the portrait test subset.
- Figure 5 compares generated rationales and selections with ground-truth annotations and selections on the scenery test subset.
5 Computer Vision Lab, University of W¨urzburg, Germany
The listed affiliations include universities, research institutes, and companies across China and the United States; participating methods widely adopted GRPO and voting ensembles.
- Affiliations: The listed affiliations also include WeChat, Tencent; the University of California, San Diego; the University of Illinois Urbana-Champaign; and Nankai University.
- Affiliations: The affiliations represent institutions including Northwestern Polytechnical University, Chongqing University, and the University of Science and Technology of China.
- Affiliations: Additional affiliations include East China Normal University, Shanghai Jiao Tong University, Tongji University, Donghua University, and Shanghai Artificial Intelligence Laboratory.
- Methods: Group Relative Policy Optimization and voting ensembles were widely adopted strategies among competing teams.
5. Conclusion
The challenge established a reasoning-based professional IQA benchmark using MLLMs to evaluate high-quality image pairs through comparative selection and interpretative reasoning.
- The benchmark uses high-quality image pairs with expert-annotated reasoning to move beyond simple numerical regression.
- Top-performing methods demonstrated robust comparative quality selection and grounded interpretative reasoning.
- The organizers aim for the datasets and multi-tiered metrics to support industrially valuable evaluation systems for photography and vision.
A. Team & Methods
This section describes the participating teams and their proposed methods for the track.
- The section describes participating teams and their proposed methods.
A.1. IH-VQA - Methodology
The team’s IH-VQA solution uses two complementary branches for professional pairwise image quality assessment: one predicts preferences, while the other generates structured explanations.
- The dual-branch framework combines an Answer Model for binary preference prediction with a Thinking Model for expert-style rationale generation.The branches separate decision making from interpretability while addressing the same pairwise quality comparison task.
A.1.1. Data Synthesis and Pre-processing
Each sample provides global and local views of both images, which are explicitly separated into four inputs for comparative quality modeling.
- Each image pair is represented by global-left, global-right, crop-left, and crop-right views.The decomposition is intended to help the model capture relative quality differences between candidates.
- The data are divided into person and scene semantic subsets for separate quality modeling.
A.1.2. Network Design
The network design combines discriminative visual classification with multimodal rationale generation, using specialized branches, ensembles, and progressively enriched prompts.
- Answer Model: The Answer Model processes four views with a shared visual backbone and fuses their features for left-right preference prediction.Multiple backbone predictions are ensembled, while separate person and scene branches model domain-specific quality cues.
- Thinking Model: The Thinking Model uses a multimodal large language model to generate expert-style reasoning from image pairs and structured prompts.Prompting incorporates domain templates, computer-vision statistics, learned IQA metrics, and answer-aware rationale refinement.
A.1.3. Training Details
Training uses distinct procedures for the two branches: ensembled binary classifiers for preference prediction and LoRA-based supervised fine-tuning for rationale generation.
- Answer Model: The Answer Model is trained for 10 epochs with AdamW, a 1 × 10−4 learning rate, 1 × 10−4 weight decay, and batch size 32.Inputs are resized to 256 × 256 and randomly cropped to 224 × 224, with a 5-fold training strategy.
- Answer Model: Fold-wise voting and a multi-backbone ensemble produce the Answer Model’s final prediction.
- Thinking Model: The Thinking Model fine-tunes Qwen3-VL-8B-Instruct with LoRA-based supervised fine-tuning for 6 epochs at 4 × 10−5 learning rate.The configuration uses LoRA rank r = 16 and DeepSpeed ZeRO-1 on 8 GPUs.
A.2. VCIP Pi Group - Methodology
The VCIP Pi Group develops a two-tier, single-agent VLM framework that uses dimension-specific tools and cyclic planning to compare paired images across expert-defined quality dimensions. Other teams adopt complementary strategies, including dual-branch prediction and rationale generation, multi-stage training with ensembles, and classifier–MLLM pipelines.
- VCIP Pi Group: Nine Qwen3-VL-2B-Instruct tools and one Qwen3-VL-8B-Instruct core agent form VCIP Pi Group’s two-tier framework for paired image comparison across nine expert-defined dimensions.The core agent coordinates the dimension tools through a unified memory bank and cyclic reasoning pipeline.
- VCIP Pi Group: The data pipeline generates expert-level analyses across nine dimensions, extracts {Tool, Region} calls, swaps local image pairs, and filters reasoning data with an LLM judge.The two-stage process produces structured supervision for the agent and its tools.
- VCIP Pi Group: VCIP Pi Group trains the tool layer with sequential SFT and GRPO, while training the core agent solely with GRPO for scheduling, reasoning, and result fusion.The implementation uses LoRA-based fine-tuning and separate hardware allocations for agent training and parallel tool inference.
- I2 Group: I2 Group progressively combines SFT and GRPO on Qwen3-VL-8B-Instruct and uses majority voting across three expert checkpoints for final predictions.On a 102-pair Phase 2 validation set, the final ensemble achieved 0.8627 accuracy and a 0.649 final score.
- Dual-branch methods: Fugui and test123djak separately train answer and thinking models so binary preference decisions and expert-style rationale generation can be handled as distinct outputs.Fugui fuses the two outputs, while test123djak applies low-rank adaptation to all linear layers and keeps the vision model and multimodal aligner unfrozen.
- Other participating methods: LZ, ongaku, dr0strange, and NTR use complementary pipelines involving checkpoint ensembles, geometric augmentation and thinking masks, LightGBM-based comparison with MLLM explanations, or SFT–GRPO fine-tuning.These methods combine multimodal fine-tuning or tabular quality aggregation with explicit explanation generation or robust inference strategies.