Source-linked AI summary
When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models
Mehak Gupta, Tanmoy Chakraborty
TL;DR
Aligned vision-language models may abstain despite identical visual inputs, raising whether safety alignment suppresses grounding or redirects generation. This paper analyzes decoding dynamics and finds that visual evidence remains available while late-stage refusal representations steer outputs toward abstention, with interventions recovering grounded answers.
Problem
It remains unclear whether safety-induced abstention reflects lost visual grounding or representational override during decoding.
Method
The paper compares default and safety-constrained inference across models, analyzing visual influence, layer-wise hidden states, and activation-level interventions.
Results
Across architectures and benchmarks, abstained generations remain visually grounded while late-stage refusal representations increasingly shape response behavior, and interventions recover grounded answers.
Takeaways & Limitations
A refusal does not necessarily indicate absent visual understanding, because safety alignment can separate internally available evidence from final grounded expression.
Takeaways & Limitations
The study’s scope is limited to comparisons between default and safety-constrained inference settings.
Abstract
from arXiv · showhide
Aligned vision-language models (VLMs) are designed to balance grounded visual reasoning with safe generation behavior. However, we observe a striking phenomenon: under safety-constrained instruction, models frequently abstain from answering questions that remain correctly answerable under default instruction despite receiving identical image-question inputs. This raises a fundamental question: does safety alignment suppress perceptual grounding itself, or does visual evidence remain internally available while generation is redirected toward abstention? In this work, we investigate the internal decoding dynamics underlying safety-induced abstention in aligned VLMs. Across multiple architectures and multimodal benchmarks, we show that abstained generations remain consistently influenced by visual evidence throughout decoding, indicating that perceptual grounding is largely preserved despite refusal behavior. We further demonstrate that, although the representational organization of refusal differs substantially across architectures, safety-constrained instruction consistently alters late-stage hidden-state dynamics toward refusal-oriented decoding. Finally, through targeted activation-level interventions, we show that suppressing refusal-related representations reliably restores grounded answering behavior across models without retraining or modifying visual inputs. Together, these findings reveal a previously underexplored failure mode in aligned VLMs: safety alignment can override grounded visual expression even when perceptual evidence remains internally preserved.
1 Introduction
Aligned VLMs can abstain under safety-constrained instructions even when identical visual evidence supports a correct answer under default instructions. The paper investigates whether this reflects lost perceptual grounding or a refusal-oriented decoding override, finding preserved visual influence and causally recoverable grounded answering.
- Motivation: Safety-constrained instructions can redirect aligned VLMs from correct, visually obvious answers to abstention despite identical image-question inputs.This challenges the assumption that abstention necessarily reflects uncertainty, ambiguity, or insufficient evidence.
- Preservation of Visual Grounding: Abstained generations remain influenced by visual evidence across architectures, indicating that safety-constrained instruction does not suppress perceptual processing.The paper distinguishes preserved internal visual evidence from the refusal behavior expressed during decoding.
- Emergence of Refusal Representations: Safety alters hidden-state trajectories in all models, while the organization of refusal representations varies across architectures.The study examines how refusal-related representations emerge during decoding and interact with grounded visual reasoning.
- Causal Role of Refusal Representations: Inference-time interventions on refusal-aligned representations reduce unnecessary abstention and recover grounded answers without changing model parameters or visual inputs.These interventions support a causal role for refusal representations in abstention behavior.
- Contribution: The findings identify a failure mode in which safety alignment suppresses grounded visual expression even when perceptual evidence remains internally preserved.This exposes a gap between what a model internally understands and what it ultimately expresses.
2 Related Work
Related work establishes VLMs’ multimodal reasoning and safety capabilities while documenting failures to use visual evidence. This paper focuses on safety-induced abstention and connects it to alignment and mechanistic-interpretability research on refusal representations.
- Large Vision-Language Models: Large VLMs combine vision-language pretraining and instruction tuning to support multimodal reasoning, dialogue, instruction following, and safety-aware interaction.Examples include GPT-4V, Gemini, Qwen-VL, and Phi-Vision.
- Visual Reasoning, Hallucination and Grounding: Prior reliability studies use VQA, GQA, and MMBench while identifying hallucination, modality conflict, language-prior reliance, and ignored visual evidence.Existing analyses primarily examine incorrect generation despite available visual information.
- Alignment, Refusal and Mechanistic Interpretability: Alignment methods such as RLHF and constitutional training encourage safer responses, while mechanistic studies identify and modify refusal-related activation patterns.These findings motivate analyzing refusal through hidden representations and representation-level interventions.
3 Key Terminologies
This section defines the default and safety-constrained inference setups, abstention and grounded answering, and a paired framework for analyzing behavioral transitions under fixed visual inputs. It also specifies the evaluated model families and controlled hidden-state extraction procedure.
- Inference setups: The default setup asks the model to answer directly from the image without added uncertainty or answerability constraints.
- Inference setups: The safety-constrained setup instructs the model to abstain when answers are uncertain, unsupported, or not explicitly observable, while changing only the instruction context.The image, question, parameters, decoding strategy, and generation hyperparameters remain identical across setups.
- Behavioral definitions: Abstention is defined by refusal or uncertainty-oriented generated behavior despite identical visual input and query, rather than by underlying question-answering correctness.Examples include “Not visible” and “I cannot determine this from the image”.
4 Safety-Constrained Instruction Induces Grounded Abstention
Safety-constrained instruction increases abstention on visually answerable samples without changing the image, question, model, or decoding setup. Abstained generations nevertheless retain visual sensitivity throughout decoding, indicating preserved perceptual grounding despite refusal behavior.
- 4 Safety-Constrained Instruction Induces Grounded Abstention: Safety-constrained instruction induces abstention even when samples remain visually answerable, motivating whether perception is lost or expression is redirected.The section frames abstention as a distinction between inaccessible perceptual information and internally available evidence that fails to shape the final response.
- 4.1 Constrained Instruction Systematically Induces Abstention: Across evaluated models and datasets, safety constraints consistently increase abstention without modifying images, questions, parameters, or decoding configuration.Paired samples isolate instruction as the changed condition behind the behavioral shift.
- 4.1 Constrained Instruction Systematically Induces Abstention: The behavioral shift suggests a mismatch between visual processing and willingness to express visually supported answers under safety-constrained instruction.The model’s ability to process visual information may remain distinct from its response-generation behavior.
- 4.2 Abstained Generations Remain Visually Grounded: Input corruption analysis tests whether abstained generations lose dependence on image features by comparing target-token logits from original and corrupted images.Visual influence is evaluated during decoding rather than only from final outputs.
- 4.2 Abstained Generations Remain Visually Grounded: Abstained generations maintain visual sensitivity throughout decoding despite shifting toward abstention, with layerwise trajectories varying across architectures.Figure 2 compares answered and abstained samples under default and safety-constrained instruction for Qwen3-VL, LLaVA-v1.6, and Phi-3-Vision on VQA-v2.
- Key Takeaway: Aligned VLMs can preserve perceptual grounding and still refuse to verbalize the visually supported answer under safety-constrained instruction.The key takeaway is that refusal does not imply that the model has stopped seeing the answer.
5 Refusal-Aligned Representations Govern Abstention
Safety-constrained instruction redirects decoding through refusal-aligned hidden-state dynamics while preserving the contribution of visual evidence. The magnitude and geometry of this representational shift vary across architectures, but refusal-related representations regulate whether internally available visual evidence is expressed as a grounded answer.
- Hidden-state analysis: The analysis extracts first-token hidden states at each decoding layer to characterize initial decoding dynamics before substantial autoregressive feedback.The first generated token typically marks the onset of abstention behavior.
- Refusal-aligned representations: Abstention is quantified by projecting hidden states onto a refusal direction defined by the representational difference between abstained and answered samples.Higher normalized projection values indicate stronger alignment with refusal-associated representations.
- Architecture-dependent dynamics: Safety-constrained instruction changes hidden-state evolution beyond early layers across all models, with Qwen3-VL showing the strongest separation and LLaVA-v1.6 a weaker analogous pattern.In Qwen3-VL, abstained generations progressively align with the refusal direction from the middle layers onward, while answered generations remain closer to the default trajectory.
- Joint visual and refusal dynamics: Safety-constrained decoding changes internal dynamics without removing visual evidence, indicating that visual influence and refusal projection evolve jointly rather than perceptual grounding disappearing.Figure 4 examines their joint distribution across representative decoding layers using independently standardized quantities within each model.
- Mechanistic interpretation: Refusal-related representations regulate the expression of internally available visual evidence, causing models to withhold grounded answers despite preserved perceptual grounding.This identifies safety-induced abstention as a failure of grounded visual expression rather than loss of perceptual access.
6 Late-Layer Refusal Representations Causally Control Abstention
Late-layer refusal representations causally control whether preserved visual evidence is expressed as an answer or overridden by abstention. Refusal-direction interventions during inference consistently suppress or amplify abstention across architectures without changing model parameters or visual inputs.
- Intervention method: The method modifies only each hidden state’s component along a layer-specific refusal direction, leaving the remaining representation unchanged.Refusal suppression projects away from the direction, while amplification enhances it.
- Intervention method: Interventions use α = 1 by default and occur exclusively during inference without modifying model parameters, alignment objectives, input images, or prompts.Thus, observed behavioral changes are attributed directly to manipulating internal refusal representations.
- Causal intervention: Suppressing refusal-related representations substantially decreases abstention and restores grounded answers, whereas amplification increases refusal behavior under identical visual inputs and decoding conditions.Recovered responses express available perceptual evidence rather than simply becoming more verbose or unsupported.
- Layerwise causal importance: Intervention effects increase with decoding depth: early-layer suppression produces modest changes, while deeper-layer interventions produce substantially larger changes in abstention.The depth-dependent pattern links the progressively stronger late-layer refusal representations to causal behavioral control.
- Causal intervention: Late-layer refusal representations determine whether perceptual grounding is ultimately expressed, despite architectural differences in refusal geometry.This causal role is consistent across diverse VLM architectures.
7 Discussion
Safety-induced abstention in aligned VLMs can occur despite internally preserved visual information, reflecting a disconnect between available understanding and final generation. The findings also indicate that late decoding stages are promising targets for behavioral modulation.
- Mechanistic interpretation: Under safety-constrained instruction, VLMs may refuse answers even when internal representations retain substantial visual information.Across behavioral, representational, and causal analyses, abstention is not primarily attributed to failure to perceive or encode visual evidence.
- Implications for Multimodal Alignment: A refusal output does not necessarily indicate absent visual understanding, but may reflect a disconnect between internal information and final generation policy.This distinction cautions against evaluating multimodal systems solely through external outputs, especially in safety-sensitive applications.
- Implications for Behavioral Control: Late-stage refusal-related representations influence whether available visual information becomes an answer or is suppressed through abstention.The concentration of intervention effects in deeper decoding layers suggests behavioral modulation can target specific generation stages while visual pathways remain active.
8 Conclusion
The paper shows that aligned VLMs can abstain under safety-constrained instruction even when identical visual inputs continue influencing internal decoding. Its analyses indicate that safety-induced abstention is not primarily caused by loss of visual information, but by changes in model behavior.
- Conclusion: Controlled default-versus-safety comparisons show that VLMs can abstain while the same visual input continues influencing internal decoding.This finding concerns abstention under safety-constrained instruction with identical visual inputs.
- Conclusion: Across multiple models and benchmarks, the analyses find that safety-induced abstention is not primarily explained by loss of visual information.The conclusion is based on analyses spanning multiple models and multimodal benchmarks.
- Conclusion: The results instead associate safety-induced abstention with changes in the models’ internal decoding process.The supplied passage identifies changes in decoding as the alternative explanation, though it does not specify their full nature.
Structure of the Appendix
The appendix details prompts and abstention, vision-sensitivity experiments, uncertainty dynamics, vision–refusal interactions, intervention strength, and qualitative abstention examples.
- Appendix structure: A.1 covers prompt details and abstention.
- Appendix structure: A.2 presents vision sensitivity experiments and corruption details, while A.3 compares uncertainty dynamics across architectures.
- Appendix structure: A.4 examines interactions between vision influence and refusal alignment.
- Appendix structure: A.5 analyzes sensitivity to intervention strength, and A.6 provides qualitative examples of abstention behavior.
A Appendix … A.6 Qualitative Examples of Abstention Behavior
Across controlled instruction, corruption, uncertainty, representation, intervention, and qualitative analyses, safety conditioning shifts aligned VLMs toward abstention while visual evidence often remains internally influential. This behavior is consistent across architectures, yet its refusal–vision relationship varies by model and can be modulated through refusal-direction interventions.
- A.1 Prompt Details and Abstention: Abstentions are detected through case-insensitive matching against predefined refusal phrases, including “cannot determine” and “not visible.”The keyword list also includes phrases such as “unable to answer,” “insufficient information,” and “image does not provide enough information.”
- A.1 Prompt Details and Abstention: Default and safety-constrained prompts keep visual inputs unchanged while isolating how safety conditioning alters abstention dynamics, internal representations, and answer generation.Safety-constrained instruction requires abstention when evidence is not clearly visible, whereas default instruction requests concise grounded answers.
- A.2 Vision Sensitivity Experiments and Corruption Details: Gaussian blur with radius r = 30 removes fine-grained visual evidence while preserving global structure for controlled analyses of logits, entropy, and top-2 probability gaps.Clean and blurred images are evaluated under identical decoding configurations in both instruction setups.
- A.3 Comparative Uncertainty Dynamics Across Architectures: Answered and abstained samples show broadly comparable top-2 probability gaps and vocabulary entropy across layers, so simple output uncertainty does not fully explain abstention.This pattern is reported across Qwen3-VL, LLaVA-v1.6, and Phi-3-Vision under default and safety-constrained instruction.
- A.4 Interaction Between Vision Influence and Refusal Alignment: Across Qwen3-VL, LLaVA-v1.6, and Phi-3-Vision, visual influence remains preserved throughout decoding even as refusal-related representations shape generation.The joint analysis examines layerwise visual influence alongside projection onto a learned refusal direction.
- A.5 Sensitivity to Intervention Strength: Positive refusal-direction scaling progressively increases abstention and reduces answer confidence, whereas negative scaling suppresses refusal-aligned features and recovers answer generation.The smooth transition across α values indicates controllable refusal-direction effects across architectures.
- A.6 Qualitative Examples of Abstention Behavior: Under identical images and questions, default instruction typically yields grounded answers, while safety-constrained instruction consistently produces abstention and explicit uncertainty across architectures.Abstentions frequently cite insufficient visual evidence or low confidence despite preserved visual influence, indicating coherent alignment-driven behavior rather than output randomness.