Source-linked AI summary

Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance

Anuridhi Gupta, Samara Mansoor, Hemant Purohit

arXiv:2608.14651v1cs.AIcs.HC

TL;DR

Disaster communication systems may not provide equivalent guidance when users interact through different modalities, especially users with access and functional needs. This paper evaluates open-weight multimodal LLMs across text and audio disaster scenarios and finds incomplete cross-modal consistency, with larger disparities for specialized access needs.

  • Problem

    Whether MM-LLMs provide equivalent, reliable guidance across text and audio modalities for users with access and functional needs remains insufficiently evaluated.

  • Method

    The study evaluates open-weight MM-LLMs using persona-based disaster scenarios paired as text and audio prompts with manual, semantic, and factual consistency analyses.

  • Results

    Open-weight MM-LLMs fail to achieve complete cross-modal consistency, with performance disparities most pronounced for personas with specialized access and functional needs.

  • Takeaways & Limitations

    Cross-modal consistency should be prioritized alongside accuracy when designing MM-LLM systems for equitable disaster risk communication.

  • Takeaways & Limitations

    The evaluation uses a relatively small scope and audio prompts from a single speaker, so findings are indicative rather than definitive.

Abstract

from arXiv · show

Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.

I. INTRODUCTION

Disaster risk communication must accommodate people with diverse access and functional needs, yet MM-LLMs can produce misaligned responses across input modalities. This paper introduces a quantitative and qualitative framework to evaluate multimodal consistency and inclusivity in real-world disaster scenarios.

  • Motivation: 403 weather and climate disasters costing over 1 billion dollars affected the U.S. from 1980 to 2024, while event frequency and intensity continue to rise.These trends motivate effective communication to help the public prepare, respond, and recover.
  • Access and functional needs: Vulnerable populations may be disproportionately affected by disasters because their characteristics and circumstances influence their ability to prepare, cope, respond, and recover.They are less likely to have access to resources and may require additional accommodations in risk communication systems.
  • Access and functional needs: Risk communication tools must provide accessible interaction methods that reflect individuals’ diverse needs and preferences.AI systems for disaster risk management therefore need to interpret and reason about diverse interaction needs.
  • Problem: MM-LLMs often show misalignment in understanding and response generation across input modalities despite claims of unified multimodal design.Such failures can produce hallucinations, adversarial fragility, and inappropriate high-stakes responses that neglect risk urgency or accessibility constraints.
  • Contributions: The paper introduces a consistency evaluation framework for open-weight MM-LLMs using quantitative and qualitative analyses in real-world disaster risk communication scenarios.It also evaluates behavioral variation across functional and access needs to measure inclusivity in model responses.

II. RELATED WORK · A. Risk Communication and Supporting Technologies

Disaster risk communication research emphasizes actionable, accessible, and tailored information supported by telecommunications and early warning infrastructure. Recent AI work extends this support, while this study examines whether open-weight MM-LLMs provide equivalent guidance across interaction modalities for people with access and functional needs.

  • A. Risk Communication and Supporting Technologies: Risk communication literature emphasizes information that is actionable, accessible, and tailored to diverse individuals during disasters.This framing positions inclusivity and usability as core requirements for emergency information.
  • A. Risk Communication and Supporting Technologies: The ITU identifies telecommunications and early warning systems as foundational infrastructure across disaster risk reduction and management.The passage also emphasizes that warning systems must reach all populations at risk.
  • A. Risk Communication and Supporting Technologies: Prior work explored ChatGPT for disaster prevention information dissemination, science education, and emergency response support.Its rapid availability and natural language reasoning were noted as advantages over static alert systems.
  • A. Risk Communication and Supporting Technologies: Recent LLM applications classify crisis information from social media, monitor infrastructure during disasters, and generate structured warnings grounded in official guidelines.These applications extend AI support across information analysis, infrastructure monitoring, and warning generation.
  • A. Risk Communication and Supporting Technologies: This work examines whether AI systems produce reliable, consistent outputs regardless of how users interact with them.The study systematically evaluates this assumption rather than treating interaction modality as inconsequential.
  • A. Risk Communication and Supporting Technologies: For people with access and functional needs, text-based interfaces may be unusable under disaster conditions, leaving equivalence of guidance across modalities unresolved.The paper addresses this gap by evaluating current open-weight MM-LLMs used as backbones for risk communication tools.

B. MM-LLMs based Systems · C. Evaluation of MM-LLMs

MM-LLM systems combine modality-specific encoders and generators with an LLM backbone, but cross-modality interference can degrade consistency, especially for audio inputs. Existing evaluation primarily uses modality-specific benchmarks emphasizing reasoning, answer correctness, and bias mitigation.

  • B. MM-LLMs based Systems: MM-LLM systems accept user queries and generate responses in the interaction modality specified by the user.The system is designed to process and generate content across diverse modalities.
  • B. MM-LLMs based Systems: Their architecture comprises a Modality Encoder, LLM Backbone, and Modality Generator.The encoder handles diverse inputs, the backbone supports zero-shot generalization, Chain-of-Thought, and instruction following, and the generator produces distinct modalities.
  • B. MM-LLMs based Systems: Cross-modality interference can cause training on one modality to degrade performance or consistency in another.Omnimodal LLMs may attend to dominant modalities such as text while producing degraded or inconsistent outputs for audio inputs.
  • B. MM-LLMs based Systems: AudioFlamingo and SALMONN illustrate cross-modal inconsistency when identical emergency queries submitted through text and audio receive different responses.The figure contrasts a detailed, actionable text response with a generic response.
  • B. MM-LLMs based Systems: Structural asymmetry motivates evaluating whether unified MM-LLMs generate inconsistent responses, particularly for underrepresented audio inputs in disaster risk communication.The concern is especially important in safety-critical contexts.
  • C. Evaluation of MM-LLMs: Current MM-LLM evaluation relies primarily on benchmarking foundational multimodal models and includes text-, vision-, and audio-based benchmarks.These benchmarks were developed to track and guide model development efficiently.
  • C. Evaluation of MM-LLMs: Text-based LLMs are evaluated for reasoning capabilities, correct answer generation, and possible bias mitigation.The passage situates these criteria within broader modality-specific evaluation practices.

III. METHODOLOGY: CONSISTENCY EVALUATION FRAMEWORK · A. Identification of Stakeholders

The framework evaluates multimodal disaster-assistance consistency through stakeholder identification, persona creation, paired text/audio prompts, and multifaceted output evaluation. Stakeholders are identified using CDC resources, focusing on individuals with access and functional needs who may require additional emergency assistance.

  • III. METHODOLOGY: CONSISTENCY EVALUATION FRAMEWORK: The framework has four components: Stakeholder Identification, Persona Creation, MM-LLM Inference, and Multifaceted Evaluation.Evaluation combines manual analysis, semantic similarity metrics, and factual overlap scoring.
  • III. METHODOLOGY: CONSISTENCY EVALUATION FRAMEWORK: Multimodal LLMs receive persona inputs as paired text and audio prompts for consistency evaluation.The framework evaluates responses across modalities using the same persona-based task inputs.
  • III. METHODOLOGY: CONSISTENCY EVALUATION FRAMEWORK: Evaluation uses manual analysis, semantic similarity metrics, and factual overlap scoring to assess multimodal outputs.These methods form the framework’s multifaceted evaluation component.
  • A. Identification of Stakeholders: Stakeholder identification draws on a critical review of U.S. Centers for Disease Control and Prevention resources.The review identifies individuals with “access and functional needs” as a formal vulnerable group.
  • A. Identification of Stakeholders: Individuals with access and functional needs may require additional assistance because temporary or permanent conditions can limit effective emergency response.This category is based on CDC guidance and does not require a formal diagnosis or medical evaluation.
  • A. Identification of Stakeholders: CDC-identified populations include children, pregnant women, older adults, and people with physical, sensory, intellectual, developmental, cognitive, or mental disabilities.The listed groups may be disproportionately affected during emergencies.
  • A. Identification of Stakeholders: Disaster impacts also include mothers with infants and preschool-aged children and hard of hearing individuals.The Great East Japan Earthquake and Tsunami affected mothers with infants and preschool-aged children, while hard of hearing individuals are directly impacted in multimodal risk communication research.

B. Persona Creation

The study creates four personas and 40 emergency-alert prompts based on real FEMA IPAWS alerts to evaluate multimodal consistency. It distinguishes perceptual accessibility from informational consistency as separate challenges for users with access and functional needs.

  • Scenario and Persona Design: Personas and communication scenarios were derived from real emergency alerts coordinated through FEMA’s Integrated Public Alert & Warning System (IPAWS).IPAWS delivers verified emergency information through multiple channels.
  • Scenario and Persona Design: 40 distinct prompts represented varied alert types and disaster scenarios using samples of alerts distributed through WEA, broadcast media, and NOAA Weather Radio.The source channels include Wireless Emergency Alerts, the Emergency Alert System, and NOAA Weather Radio.
  • Scenario and Persona Design: Prompt templates used a variable persona X to ask what someone should do after receiving warnings or alerts such as flash floods.The four personas and sample queries were summarized in Table I.
  • Accessibility Scope: The framework separates perceptual accessibility—receiving and decoding signals—from informational consistency across modalities.Examples include hearing an audio alert and retaining spoken instructions.

C. Multimodal LLMs: State-of-the-Art Open-Weight Models · 1) AudioFlamingo:

The study evaluates three state-of-the-art open-weight MM-LLMs across 40 prompts covering four personas. Audio Flamingo 3 is a 7B LLaVA-based model designed to understand sound, music, and speech.

  • C. Multimodal LLMs: State-of-the-Art Open-Weight Models: The evaluation covers three state-of-the-art open-weight MM-LLMs.
  • C. Multimodal LLMs: State-of-the-Art Open-Weight Models: The models are tested across 40 different prompts.
  • C. Multimodal LLMs: State-of-the-Art Open-Weight Models: The prompts represent four personas summarized in Table I.
  • 1) AudioFlamingo:: Audio Flamingo 3 is based on a 7B language model.
  • 1) AudioFlamingo:: Audio Flamingo 3 uses the LLaVA architecture and a unified AF-Whisper audio encoder based on Whisper.
  • 1) AudioFlamingo:: The AF-Whisper encoder handles understanding beyond speech recognition across sound, music, and speech.

2) Qwen: · 3) SALMONN: · D. Multifaceted Evaluation

Qwen3-Omni is a multilingual, end-to-end multimodal model supporting real-time text and natural-speech responses, while SALMONN is designed for diverse audio processing and reasoning. The evaluation compares text–audio response consistency using semantic similarity and factual overlap measures across four personas and 40 prompts.

  • 2) Qwen:: Qwen3-Omni processes text, images, audio, and video in a natively end-to-end architecture.It delivers real-time streaming responses in both text and natural speech.
  • 2) Qwen:: Qwen3-Omni provides real-time streaming responses in text and natural speech.The model incorporates architectural upgrades intended to improve performance and efficiency across multimodal applications.
  • 3) SALMONN:: SALMONN processes speech, environmental sounds, and music as diverse audio inputs.It integrates pretrained audio encoders with a language model through alignment layers.
  • 3) SALMONN:: SALMONN supports audio captioning, speech comprehension, and audio-based reasoning.The model is designed to generalize across heterogeneous audio domains.
  • D. Multifaceted Evaluation: Consistency is evaluated with semantic similarity and factual consistency metrics across text and audio responses.The study uses sBert, sBLEU, sRouge, and sUSE for semantic similarity, alongside factual overlap analysis.
  • D. Multifaceted Evaluation: 160 text responses and 160 audio responses were produced from 40 prompts across four personas.Responses are compared in text–audio pairs for each prompt.
  • D. Multifaceted Evaluation: Factual overlap measures how much content from the text response is preserved in the audio response.The text is treated as the source, while named entities and disaster-relevant action terms such as evacuate and shelter are extracted from the responses.

IV. RESULTS AND DISCUSSION · A. Model Performance Analysis · B. Error Analysis

The evaluation finds that open-weight MM-LLMs vary in semantic consistency between text and audio responses, while error analysis reveals systematic modality-dependent degradation in disaster guidance. Text responses are generally more detailed and actionable than corresponding audio responses, raising deployment concerns for underserved populations.

  • IV. RESULTS AND DISCUSSION: The study evaluates response-pair semantic similarity and factual consistency using four semantic metrics plus factual overlap across three models.Factual overlap combines entity-level fact overlap with lexical similarity to assess whether safety-critical facts in text are preserved in audio.
  • A. Model Performance Analysis: Qwen achieves the strongest performance across all reported semantic similarity metrics, indicating the highest consistency between text and audio responses.The metrics are Bert, BLEU, Rouge, and USE.
  • A. Model Performance Analysis: 0.896 is Qwen’s reported high S-BERT score, indicating that underlying meaning was preserved more effectively across modalities.SALMONN shows moderate performance, suggesting limited overlap in wording.
  • B. Error Analysis: Examples in Table III highlight systematic inconsistencies between text and audio responses across modalities.These inconsistencies have direct implications for deploying MM-LLMs in disaster-affected communities.
  • B. Error Analysis: Text responses provide detailed, hazard-specific actionable guidance, whereas audio responses are lower quality, generic, and sometimes fail to address the scenario meaningfully.Examples of text guidance include avoiding dust exposure, staying hydrated, and limiting physical exertion; the dust-advisory audio response omits critical protective measures.
  • B. Error Analysis: Current open-weight MM-LLMs are optimized for general capability benchmarks rather than the reliability demanded by disaster communication for underserved populations.Audio degradation is especially consequential because the audio interface is most likely to be used by people with visual impairments, low literacy, or motor disabilities.

C. Modality Type Analysis

Across open-weight MM-LLMs, cross-modal consistency preserves high-level intent but often fails in actionable detail and fine-grained semantics. This inconsistency can materially reduce the usefulness of safety-critical guidance, especially when audio responses omit specific instructions.

  • Cross-modal consistency: Current open-weight MM-LLMs preserve high-level intent while failing to maintain alignment in actionable detail and semantics.The consistency gap is particularly concerning for equitable access to emergency information.
  • Cross-modal consistency: Qwen shows near parity between S-USE (0.896) and S-BERT (0.896), indicating relatively robust semantic and factual grounding.Its lower S-BLEU and S-ROUGE scores indicate that this consistency relies on paraphrasing rather than faithful content reproduction.
  • Safety-critical implications: AudioFlamingo can produce vague audio responses that omit critical safety instructions while providing specific, actionable guidance in text.For a mother with a toddler facing a flash flood, the text response recommends staying indoors and moving to higher ground, whereas the audio response remains vague.
  • Safety-critical implications: Reformulating evacuation routes or hazard-avoidance steps across modalities can materially alter the usefulness of safety-critical responses.The concern arises because actionable guidance may not remain faithfully reproduced across modalities.
  • Model differences: SALMONN shows partial loss of fine-grained meaning across modalities, while AudioFlamingo performs worst with lexical divergence and semantic drift.These patterns indicate that cross-modal inconsistency is not limited to surface-level wording differences.

D. Persona Type Analysis · E. Implications, Limitations, and Future Work

Persona type affects cross-modal factual consistency, with hard-of-hearing users showing especially low scores in Qwen and AudioFlamingo. The findings motivate cross-modal reliability as a deployment criterion, inclusive persona-level evaluation, and broader future testing of naturalistic audio and baseline conditions.

  • D. Persona Type Analysis: 0.237 factual overlap is the lowest reported hard-of-hearing score for AudioFlamingo across modalities.The hard-of-hearing persona scores 0.397 in Qwen and 0.237 in AudioFlamingo.
  • D. Persona Type Analysis: 0.397 factual overlap is the lowest reported hard-of-hearing score for Qwen across modalities.The hard-of-hearing persona scores lowest in Qwen and AudioFlamingo, while scoring highest in SALMONN at 0.306.
  • D. Persona Type Analysis: 0.306 factual overlap is the highest reported hard-of-hearing score for SALMONN, indicating architectural differences in persona handling.The passage contrasts SALMONN’s hard-of-hearing result with Qwen and AudioFlamingo.
  • E. Implications, Limitations, and Future Work: Multimodal input capability differs from multimodal reliability when outputs are inconsistent across modalities.Such inconsistency can create unequal service for users whose access needs determine how they communicate.
  • E. Implications, Limitations, and Future Work: Inclusive evaluation protocols should assess diverse user needs because strong average performance can still fail hard-of-hearing or elderly users with dementia.The passage states that such a model is not fit for humanitarian use.
  • E. Implications, Limitations, and Future Work: Audio prompts from one speaker omit variation in accent, pace, pitch, and age-related vocal characteristics encountered in real-world use.Future work should use diverse speaker profiles and broader, more naturalistic prompt sets.
  • E. Implications, Limitations, and Future Work: Community engagement, broader interviews and surveys, and neutral baseline conditions are planned to refine personas and isolate the source of consistency gaps.The study lacks a neutral baseline for determining whether gaps are specific to vulnerable-persona conditioning or reflect general cross-modal behavior.

V. CONCLUSION

This study investigates open-weight MM-LLMs for accessible disaster risk communication by modeling stakeholder personas and scenarios. It evaluates paired audio- and text-mode outputs using semantic similarity metrics.

  • The study investigates existing open-weight MM-LLMs in accessible disaster risk communication contexts.Its primary goal is to assess these systems for disaster scenarios involving access and functional needs.
  • Stakeholders were identified through a review of existing literature and represented through personas and scenarios.The personas and scenarios model risk communication situations involving the identified disaster stakeholders.
  • The personas and scenarios served as prompts in both audio and text modes.
  • Output pairs from the two modalities were evaluated using semantic similarity metrics.
Loading 2608.14651v1…