Source-linked AI summary

Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language

Vinmay Khandode, Sai Karthik Kosuri, Neil K. R. Sehgal, Adam Greene, Elif Alpoge, Elana Duffy, Matthew Lee Smith, Thomas K. M. Cudjoe, Sharath Chandra Guntuku

arXiv:2609.02606v1cs.CL

TL;DR

Loneliness poses substantial health risks for older adults, while existing detection methods remain limited in natural conversation. The study analyzes linguistic and acoustic features from semi-structured telephone interviews with older adults to examine associations with self-reported loneliness, finding that multimodal modeling outperforms unimodal approaches.

  • Problem

    Existing loneliness assessments largely rely on self-report instruments, and few studies examine emotional loneliness in older adults using multimodal, naturalistic conversational data.

  • Method

    The study analyzes linguistic and acoustic features from semi-structured telephone interviews to predict self-reported loneliness among older adults.

  • Results

    The multimodal model outperformed text-only and audio-only models, while linguistic and vocal patterns showed associations with different loneliness levels.

  • Takeaways & Limitations

    Speech-based analysis may support psychological assessments and early indication of emotional loneliness when used alongside existing assessments rather than as a standalone diagnostic tool.

  • Takeaways & Limitations

    Interview timing and frequency were not uniform, and generalizability may be limited to populations with similar demographics, language use, and communication patterns.

Abstract

from arXiv · show

Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural conversational contexts. We analyzed speech and language markers of loneliness in 310 older adults using semi-structured telephone interviews to help understand how they process feeling lonely and how their language differs at different levels of feeling loneliness. Our multimodal framework combined linguistic features (psycholinguistic dictionaries, n-grams, and topic models) with acoustic features (pitch, tone, loudness) to examine associations with self-reported loneliness scores. Both predefined and data-driven methods captured patterns in verbal content and vocal delivery. Higher loneliness was associated with negations(r = 0.11), negative tone(r = 0.12), and conflict-related language. Lower loneliness was linked to social references(r = -0.18), motivational drives(r = -0.11), and emotional richness in speech(r = -0.12). We also found that the multimodal model (r = 0.298) outperforms the text-only and audio-only models. Findings suggest that loneliness manifests through both linguistic and acoustic cues, supporting the potential of speech-based analysis in psychological assessments and as an early indicator of emotional loneliness when used alongside existing assessments, rather than as standalone diagnostic tools.

Introduction:

Loneliness is a serious health concern for older adults, yet scalable detection in natural conversation remains limited. This study addresses that gap by examining linguistic and acoustic associations with emotional loneliness in semi-structured telephone interviews.

  • Motivation: Loneliness is linked to depression, anxiety, cardiovascular disease, cognitive decline, and premature death among older adults.One cited meta-analysis found a 26% increased mortality risk associated with loneliness.
  • Research gap: Existing loneliness assessments largely rely on self-report surveys that may overlook subtle, individualized signals in natural communication.These instruments can reduce a complex personal experience to predefined numerical scores.
  • Related evidence: Prior work links loneliness and related outcomes to both linguistic features and vocal characteristics.Reported correlates include social references, affective language, harmonic structure, resonance, and prosodic variability.
  • Related evidence: Sociocultural context, migration, hearing loss, health, and social networks shape how loneliness is experienced and expressed.These factors contribute to variation in loneliness narratives and longitudinal patterns.
  • Research gap: Few studies specifically combine linguistic and acoustic features to study emotional loneliness in older adults using longitudinal, naturalistic conversational data.This gap motivates the present study.
  • Study aim: The study tests whether linguistic and acoustic interview features predict self-reported loneliness and whether multimodal modeling outperforms unimodal approaches.The hypotheses also predict higher loneliness from negative tone and self-focus, and lower loneliness from social and affective language.
  • Scope: The analyses are exploratory and descriptive, intended to characterize associations within demographic strata rather than support causal or generalizable claims.This qualification limits how the reported patterns should be interpreted.

Linguistic Correlates of Loneliness

Across participants, lower loneliness was associated with social and affiliative language, whereas higher loneliness was associated with negative emotion and uncertainty-related cognition.

  • Across participants: r = -0.18 for social referents, the strongest reported association with low loneliness across participants.Social processes, third-person pronouns, time orientation, and drive for affiliation were also negatively associated with loneliness.
  • Across participants: r = 0.12 for negative tone and r = 0.11 for negative emotion, both associated with high loneliness.Uncertainty and discrepancy cognition were also positively associated with loneliness.
  • Across participants: r = 0.15 for uncertainty-related cognition and r = 0.13 for discrepancy-related cognition in association with high loneliness.These patterns were reported alongside negative tone and negative emotion.

Gender-Based Subgroup Patterns

Gender-stratified analyses showed distinct linguistic correlates of loneliness. Social, future-oriented, religious, and past-oriented language tracked lower loneliness, while cognition, uncertainty, discrepancy, and self-addressing tracked higher loneliness.

  • Men: Among men, social referents (r = -0.16) and future-oriented talk (r = -0.15) were associated with low loneliness.Cognition (r = 0.16) and uncertainty (r = 0.23) were associated with high loneliness.
  • Women: Among women, social processes (r = -0.17), religion (r = -0.17), and references to the past (r = -0.19) were associated with low loneliness.The largest negative association reported for women was with references to the past.

Racial Subgroup Results

Racial subgroup analyses showed different linguistic and thematic correlates of loneliness. White participants showed both social and negative-language associations, while Black participants showed past-oriented and culturally situated topic patterns.

  • White participants: Among White participants, social referents were associated with low loneliness (r = -0.18).High loneliness was associated with self-addressing speech, uncertainty, discrepancy, negative tone, and negative emotions.
  • White participants: Among White participants, negative tone showed the strongest reported high-loneliness association (r = 0.24).Negative emotions were also positively associated with high loneliness (r = 0.19).
  • Black participants: Among Black participants, past-focused language and third-person pronouns were associated with low loneliness (r = -0.19 for each).No significant LIWC correlation emerged for high loneliness in this subgroup.
  • Black participants: Among Black participants, Africa-related and lifestyle topics were associated with high loneliness (r = 0.21 and r = 0.25).The authors characterize these topics as thematic variation reflecting culturally situated experiences and conversational contexts, not loneliness indicators per se.

Acoustic Correlates of Loneliness :

Across the full sample, higher loneliness was associated with stronger and sharper vocal delivery, whereas lower loneliness was associated with flatter, calmer speech.

  • Higher loneliness correlated with strong vowel sounds (r = 0.22, p<0.001) and sharper tones (r = 0.166, p<0.001).
  • Lower loneliness correlated with low emotional peaks (r = –0.23, p<0.001) and monotones (r = -0.22, p<0.001).

Gender Subgroups Results

Acoustic correlates of loneliness varied across men and women, with higher loneliness linked to more dynamic or resonant speech and lower loneliness linked to calmer vocal patterns.

  • Men: Among men, high loneliness correlated with dynamic speech (r = 0.32, p<0.001), shift in resonance (r = 0.31, p<0.001), and expressive speech (r = 0.24, p<0.001).
  • Men: Among men, low loneliness correlated with dynamic tone (r = -0.26, p<0.001), low emotional peaks, and calmer voice (r = -0.34, p<0.001).
  • Women: Among women, high loneliness was characterized by harmonic richness (r = 0.29, p<0.001) and resonant shifts (r = 0.22, p<0.001).
  • Women: Low loneliness in women correlated with dynamic tones.

Racial Subgroups Results

Acoustic correlates of loneliness differed across White and Black participants, with each subgroup showing distinct vocal features associated with higher and lower loneliness.

  • White participants: Among White people, high loneliness correlated with dynamic speech, strong vowel sounds (r = 0.34, p<0.001), sharp tones (r = 0.27, p<0.001), and expressive speech (r = 0.26, p<0.001).
  • White participants: Among White people, low loneliness correlated with less dynamic speech, weak vowel sounds (r = -0.33, p<0.001), monotone in speech (r = -0.25, p<0.001), and change in timber (r = -0.25, p<0.001).
  • Black participants: Among Black people, high loneliness correlated with loudness (r = 0.23, p<0.001), strong vowel sounds (r = 0.22, p<0.001), and harmonic richness (r = 0.19, p<0.001).
  • Black participants: Among Black people, low loneliness correlated with less distinct harmonics (r = -0.23, p<0.001), low emotional peaks, and calmer voice (r = -0.28, p<0.001).

Full Dataset

Across the full dataset, combining linguistic and acoustic representations produced the strongest prediction of continuous emotional loneliness, while modality utility varied across demographic groups. The study characterizes complementary speech and language signals but cautions against direct numerical comparisons across differing prior-study designs and metrics.

  • Overall predictive performance: r = 0.298 was the highest Pearson correlation, achieved by the multimodal model combining text and audio for predicting CEL scores.The text-only LIWC model reached r = 0.269, while Librosa and OpenSMILE audio models reached r = 0.178 and r = 0.173, respectively.
  • Demographic subgroup patterns: For men, text features predicted loneliness better than audio features, with correlations of r = 0.234 and r = 0.141, respectively.The 1-to-3-gram model led among male text inputs at r = 0.274, while the multimodal model reached r = 0.23 with minimal incremental audio benefit.
  • Demographic subgroup patterns: For the white subgroup, the multimodal model achieved r = 0.343, exceeding audio multimodal at r = 0.286 and text multimodal at r = 0.169.OpenSMILE led the audio modality at r = 0.299, whereas LIWC led text at r = 0.245.
  • Demographic subgroup patterns: For the black subgroup, OpenSMILE audio features slightly exceeded the combined audio model, with correlations of r = 0.238 and r = 0.19.The multimodal model reached r = 0.219, providing modest gains through integration.
  • Study contribution: The analysis jointly evaluated established linguistic and acoustic feature families rather than introducing a novel feature set or state-of-the-art model.Its contribution was empirically characterizing how these representations behave when combined in naturalistic, longitudinal interviews with older adults.
  • Overall predictive performance: r = 0.283 was achieved by the combined text model, while the combined audio model reached r = 0.195.Both combined representations outperformed their individual feature classes, and the multimodal model reached r = 0.298.

Number of Interviews (Unique

The study used standardized telephone interviews with trained staff, combining natural conversation with subsequent CEL assessment. Analyses reported linguistic, topic, acoustic, and prediction results for complete-corpus and subgroup tables.

  • Data collection: Interviews followed a standardized sequence: natural conversation, CEL items, and additional engagement procedures.Participants gave verbal consent, and participation was voluntary.
  • Data collection: Interviewers avoided the term “loneliness” during conversation and assessed it afterward using validated CEL items.Open-ended prompts and neutral follow-ups were used to encourage elaboration without steering content.
  • Data collection: Staff were trained to reduce background noise and maintain recording quality during secure telephone interviews.Audio and survey data were de-identified and stored on encrypted servers with restricted access.
  • Analyses: Reported analyses included LIWC, LDA topic, audio-feature, and Extra Trees prediction tables.The tables distinguish complete-corpus analyses from subgroup analyses by gender and race.
Loading 2609.02606v1…