Source-linked AI summary
Vision Language Models are Biased
An Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Vy Tuong Dang, Anh Totti Nguyen, Daeyoung Kim
TL;DR
The paper asks whether VLMs’ memorized knowledge harms objective visual reasoning on counterfactual images, beyond biases induced by adversarial prompts. It introduces a human-supervised framework and evaluates VLMs across seven domains, finding strong reliance on prior knowledge and visual context. The results identify a measurable failure mode in which background masking helps, while excessive reasoning can worsen counting accuracy.
Problem
Prior bias evaluations mainly used artificial yes/no questions, leaving limited evidence about image-driven bias in neutral, objective counting and identification tasks.
Method
VLMBias automates biased-subject enumeration, counterfactual-image generation, and neutral-task construction, with human review, and evaluates VLMs across seven subjects.
Results
17.05% mean accuracy was achieved on counterfactual images, while VLMs defaulted to prior knowledge 75.70% of the time across seven domains.
Takeaways & Limitations
The findings show that VLMs often rely on prior knowledge rather than visual information, and that background context and reasoning length affect counterfactual counting performance.
Takeaways & Limitations
The study uses synthetic or programmatically generated images and publicly available models, with synthetic logos and flags included solely for non-commercial research.
Abstract
from arXiv · showhide
Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may notoriously sway their outputs towards wrong or biased answers. In this work, we test how the knowledge about popular subjects hurt the accuracy of vision language models (VLMs) on standard, objective visual tasks of counting and identification. We find that state-of-the-art VLMs are strongly biased (e.g., unable to recognize the 4th stripe has been added to a 3-stripe Adidas logo) scoring an average of 17.05% accuracy in counting (e.g., counting stripes in an Adidas-like logo) across 7 diverse domains from animals, logos, chess, board games, optical illusions, to patterned grids. Removing image backgrounds nearly doubles accuracy (21.09 percentage points), revealing that contextual visual cues trigger these biased responses. Further analysis of VLMs' reasoning patterns shows that counting accuracy initially rises with thinking tokens, reaching ~40%, before declining with excessive reasoning. Our work presents an interesting failure mode in VLMs and a human-supervised automated framework for testing VLM biases. Code and data are available at: vlmsarebiased.github.io.
1 INTRODUCTION
The paper tests whether VLMs’ memorized visual knowledge biases objective visual judgments on counterfactual images. Across seven domains, VLMs often default to familiar answers instead of detecting visual modifications, although masking backgrounds and limiting reasoning can improve performance.
- Benchmark and motivation: Prior work mainly used adversarial yes/no prompts, leaving unclear how much VLM errors arise from images rather than biased wording and how bias affects objective visual tasks.This paper therefore evaluates counting, identification, and basic geometry with neutral prompts and counterfactual images.
- Benchmark and motivation: VLMBias uses human-reviewed counterfactual images and neutral questions spanning seven diverse subjects, including animals, logos, flags, chess, game boards, optical illusions, and patterned grids.The framework automates biased-subject enumeration, counterfactual-image generation, and question construction, while humans reject low-quality or debatable images.
- Results across domains: 100% accuracy was achieved by all five VLMs on identification and counting questions for original, unmodified images.This establishes that the benchmark’s familiar subjects are recognized correctly before counterfactual modifications are introduced.
- Results across domains: VLMs consistently fail to count counterfactual elements across all seven domains, with animal accuracy dropping to 1.01% for birds and 2.50% for mammals after an extra leg is added.Comparable failures occur for modified logos, flags, chess pieces, game boards, optical illusions, and patterned grids.
- Mechanisms of bias: Adding subject names to counterfactual images further reduces counting accuracy by -2 to -6 points, indicating that textual knowledge can reinforce visual bias.The effect appears when names such as “Adidas” accompany modified images such as a four-striped logo.
- Mechanisms of bias: Masking background pixels nearly doubles VLM accuracy, producing a +21.09-point gain and suggesting that contextual background contents invite biased answers.The result links visual context outside the modified subject to the observed failure mode.
- Reasoning behavior: Mean accuracy rises to an empirical ceiling of 40% with additional reasoning tokens, then declines more steeply when models think longer.Thus, excessive reasoning does not reliably correct the initial visual bias.
2 RELATED WORK
Prior VLM-bias benchmarks often relied on biased prompts or identification questions, leaving unclear how visual counterfactuals affect objective visual reasoning. VLMBias addresses this gap with neutral prompts, objective counting, systematic topic comparisons, and counterfactual images.
- Scope: Unlike prior work, this study examines VLM bias in visual question answering when visual cues in counterfactual images favor common answers.The benchmark is designed to test visual bias without embedding a biased statement in the question.
- Counting benchmarks: Counting is a common VQA task requiring object localization and tracking, but prior work reports substantial difficulty, especially with many or spatially clustered objects.Reported counting accuracy ranges from 20–48% on MSCOCO and VCR1.0, while BlindTest reports 58.07% and better performance for spatially separated objects.
- Benchmark gaps: Existing benchmarks commonly use biased prompts, emphasize Yes/No or identification questions, lack systematic topic sampling, and omit in-image text-injection tests.These limitations restrict comparison across topics and analysis of whether bias originates from text-corpus knowledge.
- VLMBias response: VLMBias uses neutral prompts with biased counterfactual images and objective counting questions, comparing accuracy and bias rates across seven subjects of varying popularity.It also systematically tests in-image text-injection effects.
3 THE VLMBI A S BENCHMARK
The VLMBias benchmark constructs familiar subjects and patterned images with minimal counterfactual modifications, then evaluates neutral counting and identification questions across diverse visual domains. It includes controls for prompting, resolution, image-generation model family, and memorized versus image-internal patterns.
- Benchmark design: VLMBias modifies familiar signature elements, such as changing an Adidas logo from three stripes to four, and tests whether VLMs overlook the injected abnormality.The benchmark uses counterfactual images to separate recognition of common subjects from inspection of their altered visual details.
- Tasks: Counting is used as an objective visual-analysis task because it requires localizing relevant objects and maintaining a running total rather than relying on shortcuts.Counting comprises approximately 10% of questions in many VQA benchmarks and supports comparisons across topics.
- Domains: The benchmark spans seven topics ordered by decreasing popularity, from common animals, logos, and flags to optical illusions and a newly created visual pattern.It includes photorealistic and abstract images, with the novel patterned grids created from scratch.
- Controls: For each test image, the study uses three resolutions and two neutral descriptive prompts, then asks two counting questions and one Yes/No identification question.The resolutions are D ∈ {384, 768, 1152}; image resolution has a marginal impact on benchmark accuracy.
- Evaluation: Bias rate is defined as the frequency with which VLM answers match predefined responses corresponding to common knowledge, even when those responses are incorrect for the counterfactual image.The measure quantifies reliance on memorized prior answers in the benchmark.
- Animals: Animal counterfactuals add one extra leg to generated side-view images, producing 273 images from 91 animals across three resolutions.The retained examples show either three-legged birds or five-legged mammals clearly.
- Modified familiar patterns: Other domains use systematic minimal modifications to logos, flags, chess pieces, and game boards, while optical-illusion tests compare original images with modified versions whose correct answers reverse.The optical-illusion set includes six named illusion types and varying illusion strengths.
- Patterned grids: Patterned-grid tests place one anomalous non-edge cell inside otherwise symmetric dice or tally grids, with anomalies created by changing one shape or tally mark.The full design contains 168 images across grid types, modification types, anomaly scenarios, and resolutions.
4 RESULTS
Across seven domains, VLMs recognize familiar subjects in unmodified images but perform poorly on minimally modified counterfactuals, often defaulting to familiar answers. Background removal and additional reasoning can reduce bias, although excessive reasoning harms accuracy.
- 17.05% mean accuracy across seven counterfactual tasks shows widespread failure to count modified visual elements.
- 100% accuracy on unmodified images shows that all five VLMs recognize the familiar subjects and expected counts.
- Counterfactual counting: 0.44% versus 17.57% accuracy on car and shoe logos shows poorer detection of small, context-embedded modifications.Car emblems are smaller relative to the vehicle, while shoe logos occupy more image area.
- Counterfactual counting: 11.79% accuracy for stars versus 4.52% for stripes indicates greater sensitivity to discrete, spatially separate changes than adjacent structural changes.
- Bias and reasoning: Background removal improves counting accuracy by 21.09 points and reduces bias rate by 40.58 points.The authors report that backgrounds set VLMs up to be biased.
- Bias and reasoning: Reasoning-token increases initially improve accuracy but eventually hurt it, whereas longer tool use continuously improves accuracy while bias rate declines.o4-mini with tools uses tools for only 29.66% of VLMBias questions.
5 DISCUSSION AND CONCLUSION
The discussion concludes that VLMs often override visual evidence with memorized prior knowledge, while explicit counting training and tool use offer only partial remedies. The benchmark also has defined scope boundaries, including difficulty controlling generated counterfactual images.
- Limitations: Generated counterfactual images are difficult to control because image-generating VLMs retain their own biases.Gemini-2.0 Flash often generated Audi’s original four-circle logo when prompted for five circles.
- 75.70% of responses default to predefined biased answers, while SOTA VLMs achieve 17.05% mean accuracy on counterfactual images.
- +1.9 points is the gain from adding tools to o4-mini, which used tools on only 29.66% of VLMBias questions.The authors attribute the limited gain to overconfidence and answering immediately.
- Larger Pixtral and Qwen2.5-VL models tend to perform worse and show approximately 1.26× higher bias rates than smaller models.
- 36.02% mean accuracy from counting-trained VLMs exceeds the 17.05% achieved by SOTA VLMs, with bias rates 2.1× lower.
A.5 VLMS CANNOT COUNT ROWS AND COLUMNS IN SIMPLE GAME BOARDS
VLMs struggle with counting in structured game-board and visual settings, often defaulting to dominant or memorized patterns rather than analyzing local evidence. Visual encoders can detect relevant modifications, but language-model processing may override that information.
- Counting performance: 2.26% mean accuracy shows that VLMs fail at basic counting in structured settings, including 0% accuracy on Sudoku and Go.The models instead produce overconfident but incorrect guesses.
- Bias toward familiar answers: 78.02% of responses match well-known but false answers, yielding only 23.74% accuracy on counterfactual optical illusions.Four of five VLMs perform well on original illusions but poorly on counterfactual versions.
- Where bias enters: 95.26% vision-encoder accuracy contrasts with 49.71% VLM accuracy, showing that visual features can contain sufficient leg-count information.For five-legged animals, the VLM outputs the biased four-leg answer for 99.43% of images.
- Where bias enters: Processing stages reduce animal-image probing accuracy from 95.26% to 89.08%, while abstract-image accuracy remains near perfect, suggesting increasing language-model bias.The results support the interpretation that language models override visual-encoder evidence with memorized knowledge.
C.2 BENCHMARK SCALE
VLMBias contains a large, diverse collection of counterfactual images and expands its evaluation suite beyond the main dataset with additional interventions. Its images are designed to be realistic while remaining scalable through systematic generation.
- Dataset scale: 1,392 counterfactual images across 7 tasks form the main VLMBias dataset, while the full suite totals 4,176 images.The main dataset is larger than PhD-ccs, ViLP, and HallusionBench, though VLind-Bench has more main-dataset images.
- Dataset construction: VLMBias systematically generates photo-realistic, subtly modified familiar subjects, unlike related benchmarks using older generators or manual curation.The approach aims to combine realistic counterfactuals with scalable construction.
- Evaluation coverage: Tables 27–29 document model specifications and access details for commercial, open-source, and open-source counting VLMs.These tables cover the evaluated model groups rather than dataset images.
E.1 TASK DESIGN
The animal task tests whether VLMs count an added limb or identify the resulting counterfactual animal rather than recalling canonical leg counts. It uses systematically generated, manually filtered images and neutral counting and identification prompts.
- Task overview: Figures 17 and 18 illustrate the animal-leg data-generation pipeline and the tendency to answer from prior knowledge rather than image analysis.The task is presented as part of the broader seven-domain benchmark.
- Counterfactual animal design: The task adds one leg to familiar birds and mammals, testing whether VLMs answer 3 or 5 rather than canonical counts of 2 or 4.Birds have correct answer 3 and expected bias 2; mammals have correct answer 5 and expected bias 4.
- Dataset construction: 91 well-known animals are modified and rendered at 3 resolutions, producing 273 total images.The animal set contains 23 two-legged birds and 68 four-legged mammals.
- Generation and quality control: The pipeline generates candidate animals, creates full-body side-view images, adds one leg through image editing, manually filters failures, and renders approved images at three resolutions.Manual inspection removes cases with incorrect modifications such as more than one added leg.
- Prompts and ground truth: Evaluation uses two neutral counting questions and one identification question about the modified leg count.The identification ground truth is always “No,” while the expected biased answer is “Yes.”
F.1 TASK DESIGN
The logo task modifies familiar car and shoe logos while preserving realistic product contexts, then tests counting and identification with neutral prompts. Its design targets the possibility that contextual backgrounds and brand knowledge override visual counting.
- Task scope: The benchmark covers 5 familiar brands—Mercedes-Benz, Maserati, Audi, Adidas, and Nike—across car and shoe contexts.Car logos appear on cars, while shoe logos appear on footwear associated with athletes.
- Dataset construction: 207 images are generated from brand, background, modification, and resolution combinations.The total is 135 car-logo images plus 72 shoe-logo images.
- Generation and quality control: The pipeline proposes logo modifications, generates edited logos and contextual backgrounds, manually places logos, and renders each image at three resolutions.Quality control checks element counts, product visibility and orientation, and natural logo placement.
- Prompts and targets: Counting questions target stripes, curves, circles, star points, or prongs, while identification asks whether each modified logo belongs to its named brand.Examples include four Adidas stripes versus the expected biased count of three.
- Task overview: Figures 19–23 show the logo-generation pipelines and examples of background-context bias, including shoe logos such as the Nike Swoosh.The figures emphasize zoomed inspection of the modified logo within its typical context.
G.1 TASK DESIGN
The paper designs controlled counterfactual tasks that modify recognizable flags and chess boards, testing whether VLMs count visible elements or rely on standard visual knowledge. Each benchmark uses systematic variants, multiple resolutions, explicit prompts, and ground-truth answers tied to the modifications.
- Flag counting: 20 well-known flags receive one added or removed star or stripe, producing 120 images across three resolutions.The set contains 13 star-typed and 7 stripe-typed flags.
- Flag counting: Direct-counting questions require the actual modified element count, while the expected bias is the flag’s standard element count.Identification questions have “No” as the correct answer because each flag element has been modified, versus an expected “Yes” bias.
- Board-piece counting: Chess and xiangqi boards each receive one removed or replaced piece at 12 occupied target squares, yielding 144 images at three resolutions.The design tests total counts, added-piece-type counts, and whether the board remains a standard starting position.
H.3 QUALITATIVE RESULTS
The chess-piece counting task evaluates whether VLMs inspect modified boards rather than defaulting to familiar board structure. Its prompts separately query total pieces, piece-specific counts, and starting-position identification.
- Prompt design: The task asks models to count chess or xiangqi pieces and, for replacement variants, count the added piece type.Starting-position questions ask whether the board is the standard chess or xiangqi starting position.
- Evaluation views: The qualitative-results figures focus on VLM behavior when counting pieces and on the corresponding board-generation pipeline.The cited figures identify the evaluation and data-generation views for this task.
I.1 TASK DESIGN
The row-and-column task modifies familiar chess, xiangqi, Sudoku, and Go boards by adding or removing rows or columns, testing visual counting against reliance on standard layouts. It systematically varies board type, modification position, and resolution.
- Board construction: Four board types—Chess, Xiangqi, Sudoku, and Go—receive row or column additions and removals.Chess, Xiangqi, and Sudoku have eight positional variants, while Go has four uniform-grid variants.
- Board construction: 84 images are generated from the board variants at 384, 768, and 1152 pixels.The total combines eight variants for each of three boards with four Go variants, then multiplies by three resolutions.
- Generation pipeline: The generation pipeline creates standard boards, applies controlled row or column changes, preserves special visual elements, and renders each variant at three resolutions.Implementations preserve board-specific features such as chess coloring, xiangqi river and palace markings, Sudoku block boundaries, and Go star points.
- Prompt design: The task uses paired counting prompts to distinguish direct row or column counting from prior-knowledge responses.Prompt wording varies by board type, including rows, columns, and horizontal or vertical lines.
- Ground truth: Counting questions use the actual post-modification row or column count, whereas standard-layout questions have “No” as the correct answer.The expected bias is the board type’s standard dimension or a “Yes” response indicating the standard layout.
J.1 TASK DESIGN
The paper tests whether VLMs interpret optical illusions and anomalous patterned grids directly or rely on familiar visual patterns. Both task families use controlled counterfactual modifications, explicit ground truth, and multiple image resolutions.
- Optical illusions: Six classical optical illusions are generated in original and modified conditions, producing 396 images across three resolutions.Modified versions reverse the usual relationship between perceived effects and actual measurements.
- Optical illusions: The illusion task asks both whether target elements are objectively equal or parallel and whether an image is an instance of a named illusion.These questions separate direct visual measurement from identification of familiar illusion patterns.
- Optical illusions: Original illusions use equal target elements, whereas modified illusions use unequal elements that contradict the typical illusion effect.The intended correct answers are “Yes” for equal original cases and “No” for unequal modified cases.
- Evaluation focus: The figure descriptions frame optical-illusion evaluation as testing whether VLMs rely on prior illusion knowledge rather than directly interpreting images.The vertical–horizontal illusion is described as a contrasting case in which models are misled by the illusion itself.
- Patterned grids: The grid benchmark renders each anomaly image at 384, 768, and 1152 pixels to assess resolution sensitivity.Anomaly locations avoid edges and corners and are varied through 14 distinct base settings.
K.3 QUALITATIVE RESULTS
The qualitative results include a case where nearly all VLMs fail to identify an abnormal cell, alongside question and prompt materials for the evaluation.
- All VLMs except Sonnet-3.7 fail to correctly identify abnormal cell C3.
- The section includes example question sets covering multiple evaluation materials and sanity checks.
M.1 QUALITATIVE RESULTS ON THE USE OF HELPFUL PROMPTS
The qualitative results show that helpful prompts, few-shot labels, and tool use do not reliably overcome VLMs’ biases in counting tasks. Success depends on accurately attending to all visible objects, while close spacing, mislocalization, and incomplete crops produce systematic errors.
- M.1 QUALITATIVE RESULTS ON THE USE OF HELPFUL PROMPTS: Debiasing and double-checking prompts still leave VLMs biased on simple counting tasks, including chicken legs and flag stripes.The examples report complete failure on chicken-leg counting and continued bias toward 13 when one stripe is removed from a U.S. flag.
- M.2 QUALITATIVE RESULTS ON THE USE OF LOCATE-THEN-COUNT PROMPTS: Locate-then-count prompting fails when models default to canonical counts instead of the visible number of legs.All tested VLMs answer 2 for a 3-legged stork and 4 for a 5-legged lion.
- M.3 QUALITATIVE RESULTS ON POINTING VLMS: Counting is usually accurate for Moondream-2B and Molmo-72B when objects are sufficiently large and separated.Both models often fail when objects are close together, poorly understood, or incorrectly localized.
- M.4 QUALITATIVE RESULTS ON FEW-SHOT PROMPTING: With few-shot prompting, o4-mini still counts incorrectly as 6 legs despite the supplied examples and explanations.The qualitative result reports that stronger labels and examples remain insufficient to correct the count.
- M.4 QUALITATIVE RESULTS ON FEW-SHOT PROMPTING: Few-shot examples with explicit verified labels do not make o4-mini trust an unusual leg count.The model rejects a verified 5-legged label and continues reasoning from the canonical four-legged mammal prior.
- M.4 QUALITATIVE RESULTS ON FEW-SHOT PROMPTING: Adding a hint about an unusual leg count is tested as an extension of few-shot prompting with strong labels.The supplied material identifies the added hint and verified 5-legged example as the intervention setting.
- M.5 QUALITATIVE RESULTS ON O4-M I N I CHAT INTERFACE WITH TOOLS: o4-mini can overcome the canonical count when it autonomously crops the image to focus on all five legs.The successful tool-use case reports a correct {5} after cropping and attributes success to correctly focusing on the relevant region.
- M.5 QUALITATIVE RESULTS ON O4-M I N I CHAT INTERFACE WITH TOOLS: Tool use fails when o4-mini crops only the front legs, producing 4 instead of the visible 5.The incomplete crop leads the model to conclude that all four visible legs are present, showing that localization is crucial.