Source-linked AI summary

Do Depressive Facial Patterns Transfer Across Cultures and Contexts? Evidence from a German RCT and E-DAIC

Misha Sadeghi, Robert Richer, Lydia Helene Rupp, Lena Schindler-Gmelch, Marie Keinert, Farnaz Rahimi, Malin Hager, Bernhard Egger, Matthias Berking, Bjoern M. Eskofier

arXiv:2609.05543v1cs.CVcs.HC

TL;DR

Cross-corpus generalization of facial depression biomarkers remains limited, particularly because elicitation contexts differ across datasets. This study conducts bidirectional transfer between a structured RCT and a naturalistic interview corpus, finding that binary classification transfers more robustly than severity regression and that passive contexts generalize best.

  • Problem

    Cross-corpus facial biomarkers for depression are rarely tested across structured and naturalistic elicitation contexts, limiting evidence for scalable, objective monitoring.

  • Method

    The study performs bidirectional transfer between the EmpkinS RCT and E-DAIC, evaluating facial features for continuous severity regression and binary classification across contexts and label settings.

  • Results

    Binary MDD classification transfers more robustly than continuous PHQ-8 regression, achieving an AUC of 0.70 in the forward Preparatory/All configuration.

  • Takeaways & Limitations

    Functional context alignment is the primary determinant of cross-corpus generalization: passive contexts favor transferability, whereas active regulation yields stronger within-corpus signals that do not survive domain shift.

  • Takeaways & Limitations

    Reverse-transfer AUC estimates are volatile because EmpkinS condition-specific test sets are small, with n_test ≤13 in most configurations.

Abstract

from arXiv · show

Automated assessment of depression from facial dynamics holds promise for scalable mental health monitoring, yet cross-corpus generalization of learned biomarkers remains an open challenge. We present a systematic bidirectional transfer study pairing the EmpkinS-EKSpression randomized controlled trial (RCT; N = 256, SCID-5-CV diagnoses) with the Extended Distress Analysis Interview Corpus (E-DAIC; N = 275, semi-structured clinical interviews), predicting depression severity and binary diagnostic status from facial action units, head pose, and gaze. Cross-corpus binary classification proves more robust than continuous PHQ-8 severity regression, with forward transfer achieving AUC = 0.70. Regression transfer is governed by functional context alignment: passive observation phases yield the most transferable models, while active emotion regulation phases elicit stronger within-corpus signals. These findings establish functional context alignment as the primary determinant of cross-corpus generalization, with passive elicitation contexts offering the best trade-off between within-corpus sensitivity and cross-corpus robustness.

1 Introduction

The paper addresses whether facial depression biomarkers generalize across cultures and elicitation contexts, where context-dependent signals may undermine cross-corpus transfer. It introduces a bidirectional comparison of a German RCT and an independent interview corpus across severity regression and binary classification.

  • Motivation: Depression is prevalent, disabling, and substantially underdiagnosed, motivating scalable behavioral screening approaches.The paper positions facial dynamics as a non-invasive complement to limited clinical access and imperfect self-report monitoring.
  • Facial biomarkers: Facial action units, gaze, and head movement capture depression-associated patterns including blunted expressivity and restricted movement.These features provide a fine-grained representation of facial muscle activity linked to depression severity.
  • Research gap: Elicitation context is an underappreciated source of generalization gaps because structured and naturalistic paradigms are rarely compared cross-corpus.Structured tasks may elicit psychomotor dysregulation more consistently than unstructured interviews, making context alignment central to transfer.
  • Study design: The study uses 256 EmpkinS participants with SCID-5-CV diagnoses across four phases and intervention conditions, evaluating facial features against four severity scales.The within-corpus analysis examines 20 phase-condition combinations before cross-corpus transfer.
  • Study design: Bidirectional transfer compares continuous PHQ-8 regression and binary MDD classification between EmpkinS and the culturally and linguistically different E-DAIC interview corpus.The analysis tests how transfer depends on functional alignment between elicitation context and label instrument.
  • Contributions: The paper contributes a systematic cross-corpus study and practical guidance on the trade-off between passive-context robustness and structured-context sensitivity.Its comparisons address context alignment, label mismatch, and feature space in facial depression assessment.

2 Related Work

Prior work establishes facial action units and related behavioral signals as depression markers, while leaving cross-cultural, cross-context, and label-protocol transfer insufficiently resolved. This paper targets those gaps through controlled comparisons across paradigms, languages, and diagnostic definitions.

  • Established biomarkers: AU intensities, gaze, and head pose have predicted PHQ-8 or PHQ-9 scores with CCC values of 0.3–0.6 within single corpora.Key markers include AU15, AU17, and reduced AU12, while deep learning has improved within-corpus performance.
  • Generalization challenges: Facial affect varies with social display rules, cultural background, and elicitation context, complicating transfer between structured and free-conversation settings.German- and English-speaking populations may differ in baseline expressivity norms affecting AU-based indicators.
  • Research gap: The study contributes controlled evidence on paradigm contrast, language, and label protocol by comparing a structured German RCT with an interview corpus.It contrasts SCID-5-CV diagnoses with PHQ-8 self-report labels.
  • Label mismatch: Different binarization thresholds can alter class balance and apparent model performance, motivating separate label settings for contextual and label mismatch.The design uses two label settings to disentangle these sources of cross-corpus degradation.

3 Methods

The methods build a unified OpenFace-based pipeline across the EmpkinS RCT and E-DAIC interview corpus, evaluating facial features for severity regression and binary classification within and across corpora. The design spans passive and active elicitation contexts, standardized preprocessing, feature selection, multiple models, and task-specific metrics.

  • Experimental pipeline: The pipeline feeds facial behavioral features from EmpkinS and E-DAIC into unified within-corpus and cross-corpus regression and classification experiments.The targets are PHQ-8 severity and binary MDD status.
  • EmpkinS corpus: EmpkinS contains 256 RCT participants evenly split between depressive disorders and healthy controls, with SCID-5-CV diagnostic ground truth.The corpus was collected in German at the EmpkinS Lab and supports controlled phase and condition comparisons.
  • EmpkinS corpus: EmpkinS phases capture passive observation and active emotion-regulation behavior, including a preparatory baseline, mood induction, and randomized intervention conditions.The preparatory phase provides a resting-state comparison, while active conditions target cognitive and facial affect regulation.
  • E-DAIC corpus: E-DAIC provides 275 semi-clinical interview sessions conducted by a virtual agent, with official training, development, and test partitions.The corpus represents a conversational setting contrasting with EmpkinS structured emotion-regulation tasks.
  • Features: The feature set includes facial action units, gaze, and head pose extracted or distributed through OpenFace-based processing.EmpkinS processing removes low-confidence frames, while E-DAIC supplies equivalent pre-extracted features.
  • Preprocessing: Preprocessing uses shared stratified EmpkinS splits, official E-DAIC partitions, training-only fitting, and training-set median imputation.These choices preserve comparability and prevent test-set refitting.
  • Feature selection: Feature selection combines target association filtering with Benjamini–Hochberg FDR correction and recursive elimination using a Ridge surrogate.The fallback retains the top 20 features when fewer pass the association threshold.
  • Modeling: The study evaluates 12 regression models and nine classification models jointly with Standard, MinMax, and Robust scalers.This model comparison covers linear, kernel, tree-based, boosting, and neural approaches.

4 Results and Discussion

Within-corpus facial regression was strongest in selected passive or homogeneous contexts but remained modest overall, while cross-corpus performance depended sharply on transfer direction and functional context. Binary classification transferred more robustly than continuous PHQ-8 regression, with passive alignment supporting the clearest generalization.

  • Within-corpus regression: CCC peaked at 0.31 for PHQ-8 in Preparatory/AFE and 0.28 in Preparatory/SHAM, indicating reliable passive-phase signal.Minimal behavioral demands may expose trait-level psychomotor differences without structured-task confounds.
  • Within-corpus regression: Pooling all conditions reduced PHQ-8 Preparatory performance to CCC = 0.04 despite n = 251, indicating condition-specific signal dilution.Condition homogeneity appeared more important than sample size for continuous regression.
  • Within-corpus regression: Self-report scales showed non-identical performance trajectories, while HRSD-17 diverged because expert ratings integrate holistic behavioral and physical-symptom cues.Facial dynamics therefore interacted differently with questionnaire symptom weightings and clinician-rated constructs.
  • Cross-corpus regression: CCC reached 0.26 for E-DAIC → EmpkinS Prep./SHAM and 0.25 for Neg./CR+AFE, approaching the EmpkinS within-corpus peak of 0.28.The pooled Prep./All condition instead fell to 0.10, consistent with condition dilution.
  • Cross-corpus regression: Forward regression transfer peaked at CCC = 0.21 for EmpkinS → E-DAIC Neg./CR, while other configurations were near zero.Challenge-induced facial dynamics did not naturally manifest during the clinical interview, making contextual mismatch the primary explanation offered.
  • Cross-corpus classification: Binary classification remained viable across corpora, with EmpkinSall → E-DAIC reaching AUC = 0.70 in Preparatory/All under diagnostic mismatch.Reverse-transfer AUCs were higher but require caution because condition-specific EmpkinS test sets had n_test ≤ 13 and high AUC could coexist with low F1M.
  • Cross-corpus classification: EmpkinS within-corpus classification reached AUC = 0.78 in Setting 1 and 0.91 in Setting 2, whereas E-DAIC within-corpus AUC was 0.58.The contrast contextualizes the stronger cross-corpus classification signal relative to continuous regression.
  • Discussion: Classification success suggests global, context-robust facial features encode categorical MDD status, whereas continuous biomarkers remain context-dependent.Passive elicitation leaves natural facial behavior undisturbed and is presented as necessary for regression transfer.

5 Conclusion

The study finds that cross-corpus transfer differs by task and elicitation context: binary MDD classification is more robust than continuous severity regression, while passive contexts transfer best. Functional context alignment is identified as the primary determinant of generalization.

  • Continuous PHQ-8 regression transfers poorly in the forward direction, whereas binary MDD classification achieves an AUC of 0.70 in Preparatory/All.
  • Reverse regression transfer succeeds modestly only during the passive Preparatory baseline phase, which most closely resembles an unstructured interview.
  • Passive elicitation contexts produce the most transferable models, while structured active regulation phases yield stronger within-corpus signals that do not survive domain shift.
  • Functional context alignment is the primary determinant of cross-corpus generalization beyond feature extractors, label instruments, and training-set size.

A.1 Feature representation

The shared feature representation converts 49 raw OpenFace signals into 882 statistical features capturing levels, variability, distributions, dynamics, and peaks.

  • Each OpenFace signal column is summarized into 18 statistical functionals applied to the original signal and first-order frame differences.
  • The 882-feature shared set is built from 49 raw OpenFace columns covering action units, gaze, head pose, and confidence.

A.2 Feature selection

Feature selection uses a univariate statistical filter followed by recursive feature elimination, with separate procedures for regression and classification.

  • Regression ranks features by absolute Spearman correlation, while classification uses the Mann–Whitney U test before Benjamini–Hochberg FDR filtering.
  • When fewer than 20 features pass the regression or classification threshold, selection falls back to the top 20 ranked features.
  • Recursive feature elimination evaluates candidate subsets of 5, 10, 15, and 20 features using three-fold cross-validation.
  • The selected subset size minimizes cross-validation MAE for regression or maximizes AUC for classification, typically yielding 5–20 features.

A.3 Regression models and hyperparameter grids

Regression models are tuned through five-fold grid search across multiple estimators and scalers, with the best-scoring combination selected.

  • Three scalers—Standard, MinMax, and Robust—are evaluated jointly with each regression model.
  • Hyperparameter grids are searched with five-fold GridSearchCV scored by negative MAE.
  • Table 6 specifies regression hyperparameter search grids, with random_state=42 for tree ensembles and max_iter=500 for MLP.

A.4 Classification models and hyperparameter grids

Classification models are tuned through AUC-based grid search, with fixed solver and iteration settings for key estimators. The best forward-transfer classifier shown here achieves AUC = 0.70.

  • 5-fold GridSearchCV selects classification hyperparameters using AUC as the scoring metric.
  • Logistic Regression uses the lbfgs solver with max_iter = 2000, while MLP uses max_iter=500.
  • AUC = 0.70 is reported for the best forward-transfer classifier, EmpkinS → E-DAIC.

B SHAP Feature Importance

SHAP analysis identifies gaze dynamics as the dominant feature group in both transfer directions. Other influential features include facial action-unit dynamics and head-motion variability.

  • SHAP beeswarm plots rank features by mean |SHAP|, with color encoding feature value and horizontal position encoding model impact.The plots are computed using a source-partition interpretability model.
  • Gaze dynamics dominate forward transfer, followed by AU15 intensity entropy, head rotation variability, AU04, and AU25.The forward classifier is EmpkinS → E-DAIC.
  • Gaze features dominate reverse transfer even more strongly, alongside AU07 intensity entropy, head translation dynamics, and AU01.The reverse classifier is E-DAIC → EmpkinS.
Loading 2609.05543v1…