Source-linked AI summary
Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration
Sunwhi Kim, Sunyul Kim, Meounggun Jo, Jini Tae
TL;DR
AI-generated face portraits are photorealistic enough to challenge reliable human–machine comparisons. This study compares frontier VLMs with human performance under matched conditions, finding that July 2026 models exceeded young adults in sensitivity while retaining different calibration failures.
Problem
Photorealistic generated portraits challenge detection, while claims that machines beat humans often rely on measurements taken under conditions that cannot be directly compared.
Method
The study compares 19 frontier VLMs from nine providers on 198 portraits using a matched benchmark and measures accuracy, stability, and related decision characteristics.
Results
July 2026 models exceeded the young-adult human ceiling on single passes, draw averages, and sensitivity, while machine failures remained extreme and criteria varied.
Takeaways & Limitations
The comparison shows that frontier VLM sensitivity can surpass young-adult performance without eliminating machine decision biases or instability.
Takeaways & Limitations
Changing practice images altered roughly a quarter of per-item verdicts, and the data cannot test whether people are learning the visual tells of generated portraits.
Abstract
from arXiv · showhide
AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.
3 Hoseo University, Republic of Korea
This study compares frontier VLMs and humans on the same portrait-detection task, measuring accuracy alongside robustness, bias, calibration, and explanations. July 2026 models exceeded young-adult sensitivity, but machine decision criteria remained less human-like and example-dependent.
- Human benchmark: 85% mean human accuracy fell from about 88% in the 20s to about 66% in the 60s.The human benchmark used real FFHQ photographs and identity-matched ChatGPT-4o and Imagen 3 portraits.
- Benchmark and protocol: 19 VLMs judged 198 portraits under a protocol matching the human task, enabling age-stratified model–human comparisons.Models saw the same binary REAL/AI task and few-shot practice structure, with matched stimulus composition.
- Accuracy and sensitivity: June 2026 models did not surpass adults in their 20s–30s, whereas July 2026 leaders exceeded the young-adult mean on accuracy and sensitivity.The comparison was conducted under an identical protocol across the two model cohorts.
- Failure modes: Model failures remained non-human: response criteria were extreme, confidence was largely non-diagnostic, and rationales aligned with verdicts rather than image truth.These patterns persisted in the new leaders despite their higher detection performance.
- Robustness and calibration: Changing the labelled practice images altered roughly one quarter of per-item verdicts, revealing substantial example-driven instability.The benchmark varied few-shot example sets and measured stability, bias, calibration, rationales, and generator-specific difficulty.
Results
July-2026 models surpassed young adults in portrait detection sensitivity and balanced accuracy, but their decision criteria remained more biased and their judgments remained unstable. Confidence calibration improved among July leaders, while rationale cues often tracked verdicts rather than demonstrable evidence.
- Accuracy: 92.8% balanced accuracy put gpt-5.6-sol 4.3 percentage points above adults in their 20s, while claude-fable-5 reached 88.6% and detected every AI image.gpt-5.6-sol’s cluster confidence interval was 89.8–95.8%; claude-fable-5’s cost was 77.3% real-photo accuracy.
- Robustness: 25.7% of image judgments flipped across six passes, showing that changing labelled examples and run-to-run noise still destabilized model decisions.Strict REAL↔AI reversals occurred for 24.7% of items; the varied draws alone produced 23.3% reversals.
- Sensitivity and bias: d′ values up to 3.41 exceeded young adults’ approximately 2.4, but model criteria ranged from c = −1.10 to +1.45 while humans stayed near c ≈ 0.The leaders were themselves biased: claude-fable-5 had c = −0.97 and gpt-5.6-sol had c = +0.44.
- Rationales: 10.3% of model rationales mentioned eyes versus 42.8% of human reports, and models expressed uncertainty in 0% versus 49.1% of humans.Models most often cited over-smoothness, skin texture, uncanny perfection, lighting, and naturalistic appearance.
Discussion
Under an identical human–model protocol, frontier models crossed the young-adult human ceiling between June and July 2026, but their decisions remained less balanced and less stable. The remaining gap concerns calibration, contextual robustness, and evidence-grounded explanations rather than raw sensitivity alone.
- Crossing the human ceiling: June models were statistically indistinguishable from adults in their 20s–30s, whereas the July releases crossed that ceiling within four weeks.The comparison used an age-stratified human baseline under the same protocol and stimuli.
- Crossing the human ceiling: 92.8% single-pass and 92.1% draw-averaged balanced accuracy put gpt-5.6-sol above the 20s human ceiling, while claude-fable-5 reached 91.9% draw-averaged accuracy with perfect AI detection.The July leaders’ performance survived averaging over prompt draws and exceeded the young-adult comparison.
- Calibration remains human: About one quarter of image verdicts changed when labelled examples changed, and one June model’s single-pass score shifted across an 18-pp range.The instability included both example sensitivity and models’ own run-to-run noise, so single-pass leaderboards can overstate dependable performance.
- Calibration remains human: Model criteria ranged from c = −1.10 to +1.45, while humans held c ≈ 0 at every age; claude-fable-5 combined d′ = 3.41 with c = −0.97 and nearly one-quarter of real photographs misclassified as AI.The result separates sensitivity from decision balance: machines can detect better while applying strongly biased response criteria.
- Explanations and transfer: Models cited human-like cues such as texture and over-smoothness, but cue usefulness reversed with image truth, making rationales verdict markers rather than demonstrable evidence.Humans reported uncertainty and at least one cue carried evidential value, whereas model rationales generally tracked the answer.
Material and methods
The study reused an earlier human dataset and benchmarked 19 vision-capable models on identity-matched real and AI-generated portraits. Human and model evaluations used comparable binary real-versus-AI judgments, with signal-detection, confidence, cue, and calibration analyses.
- Task design: Human judgments were collected in an untimed web task with REAL and AI buttons, and models received the same portrait task with labelled examples.The human reference data came from an earlier study; no new human data were collected here.
- Human reference data: N = 1,667 human records remained after retaining first attempts and ages 20–69, with no further exclusions.The sample included 1,332 mobile and 335 PC participants.
- Human reference data: Human accuracy averaged 85.18%, declining from 88.5% in participants in their 20s to 65.9% in participants in their 60s.Real-photo and AI-detection accuracy were nearly equal overall, while Imagen 3 images were easier than ChatGPT-4o images.
- Analysis: Human signal-detection analysis found near-zero criterion bias at every age while sensitivity declined from d′ = 2.33 in the 20s to 0.87 in the 60s.The study used the public human data to position people against the models rather than to explain human factors in depth.
- Stimuli: The stimulus pool contained 198 main portraits: 66 identities represented by one real photograph, one ChatGPT-4o image, and one Imagen 3 image.Four labelled practice identities were excluded from the main trials.
- Model sample: Nineteen vision-capable models from nine providers were evaluated in June and July 2026 cohorts.The June cohort contained 14 models, and the July cohort added five releases.
Benchmark protocol
Models were tested under a fixed, human-matched protocol with repeated draws and shared-item comparisons. The protocol also separated sensitivity to labelled examples from intrinsic run-to-run variation where possible.
- Single-pass benchmark: Each model judged all 198 main stimuli once per pass using a fixed system prompt, four labelled examples, and a forced REAL-or-AI JSON response.The main benchmark comprised 3,762 judgments with zero API errors.
- Single-pass benchmark: Temperature 0 reduced variation but did not guarantee identical outputs because models ran on external servers.Unclear responses were excluded from accuracy denominators.
- Scoring and inference: Balanced accuracy weighted real, ChatGPT-4o, and Imagen 3 performance as 0.5, 0.25, and 0.25, respectively.Pairwise differences used exact, nominal, unadjusted McNemar tests on shared valid items.
- Few-shot robustness: Five varied practice-example draws were compared with a hand-selected baseline, reporting balanced-accuracy means, ranges, and per-item flip rates.The benchmark used 198 stimuli per draw and 18,810 additional judgments across models.
- Few-shot robustness: 24.7% of fleet-wide judgments were strict REAL↔AI reversals across the baseline and varied example draws.Strict reversals across the five varied draws alone were 23.3%, while any change across those draws was 24.2%.
Rationale cue coding and diagnosticity
The study coded model rationales into a shared cue taxonomy and evaluated whether cited cues were diagnostic of correctness. Additional analyses tested reasoning effort and prompt-based persona changes.
- Cue coding: A 15-category, multi-label taxonomy coded 3,738 rationales while remaining blind to true class and correctness.The taxonomy covered human-study checklist cues and cues salient in model rationales.
- Coding validation: The rationale coder was itself benchmarked, so the analysis mitigated circularity through ground-truth blinding and rule-based keyword cross-checking.The two coding approaches produced consistent per-model cue profiles.
- Diagnosticity: Cue frequency was the percentage of rationales mentioning a cue, while cue diagnosticity compared cue-invoking accuracy with base accuracy.Diagnosticity was computed overall and within true classes.
- Diagnosticity: Verdict-conditioned cue validity compared accuracy for cue-citing and non-citing rationales among judgments sharing the same verdict.The naturalistic-versus-realistic comparison used a small comparator cell of 62 REAL verdicts without the cue.
- Reasoning effort: Maximum reasoning effort changed 6.6% of gpt-5.6-sol answers and raised balanced accuracy from 92.8% to 94.3%.The single-model, single-condition pilot was reported descriptively.
- Persona pilot: The persona pilot did not meet expansion criteria, with maximum shifts of 3.8 percentage points in balanced accuracy, 3.3 points in said-AI rate, and 3.8 points in confidence.The four-model sample was selected for contrast rather than representative coverage, and the manipulation tested prompt sensitivity rather than human ageing.
- Signal-detection analysis: Signal-detection sensitivity used d′ = z(hit) − z(FA), whereas criterion c = −½·[z(hit) + z(FA)] quantified response bias.Negative c represented a liberal, AI-leaning bias.
Ethics
The study reused approved, de-identified human data and evaluated publicly available AI models without collecting new human-subject data.
- Human data: The human reference study had institutional review-board approval, digital informed consent, and no collection of identifying information.The present work reused those data rather than recruiting new human participants.
- Model evaluation: The present study evaluated publicly available AI models through a commercial API and required no additional ethical approval.No new human-subjects data were collected.
Data accessibility
The study releases aggregate results, code, per-item judgments, cue codes, and supporting data through public repositories. Some mappings, identifiers, rationale text, and displayed-example information remain restricted to protect the benchmark.
- Aggregate results, coding and figure-generation code, a data dictionary, and environment details are deposited.
- The stimulus pool, attribution and licence metadata, and de-identified human data are available through the prior study’s public repository.
- The released files support reproducing headline statistics, pairwise McNemar results, persona summaries, and cue profiles.
- Per-item judgments cover answers, confidence, correctness, experimental controls, pilot conditions, and rationale cue codes.
- Stimulus mappings, collection seeds, production identifiers, and verbatim rationales are withheld, while corresponding authors can provide restricted materials.
- Publication figures disclose ground truth for displayed examples, including a bounded nine-item disclosure accepted as the cost of worked examples.
Declaration of AI use
The paper discloses language-model use in manuscript preparation, rationale coding, and verification. The evaluated models were tested through the same OpenRouter protocol applied to all models, while authors checked the analyses and results.
- Language models assisted with drafting, editing, analysis and figure-generation code, formatting, and manuscript verification.
- Authors re-checked candidate issues and every number against their own analysis pipeline before revision.
- gpt-5.4-mini coded free-text model rationales as a documented analysis component rather than merely assisting manuscript preparation.
- The evaluated claude-fable-5 model belongs to the same family as the writing-assistance tool.
- Evaluation ran through the OpenRouter API under the identical protocol applied to all models.
Authors' contributions
The authors divide contributions across conceptualization, methodology, software, data collection, analysis, experiment design, interpretation, and writing. All authors approved the final publication and accepted accountability.
- S.W.K. led conceptualization, methodology, software, data collection, formal analysis, and the original draft.
- S.Y.K. contributed methodology, formal analysis, and review and editing.
- M.J. contributed selected experiment design, interpretation, and review and editing.
- J.T. contributed selected experiment design, interpretation, and review and editing.
- All authors gave final approval and agreed to accountability for the work performed.
Supplementary Information
Supplementary materials provide recomputed human values, aggregate model summaries, and figures covering accuracy, pairwise separability, prompt sensitivity, latency, and cue fingerprints.
- Human values are recomputed from the public data release, while model results are reported as aggregate summaries.
- Human accuracy is shown by age with model reference lines; July leaders sit above the 20s band, while June leaders are near it.
- Pairwise McNemar tests show each July leader beats both June leaders, but the July leaders do not differ from each other.
- The age-persona pilot found balanced-accuracy shifts of ≤3.8 pp and answer flips of ≤6.7%, without a human-like age gradient.
- Latency is reported as an API/compute proxy by stimulus type and is not comparable to human reaction time.
- LLM-based and rule-based cue coding produce consistent per-model cue profiles.