Source-linked AI summary

AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews

Tianyi Zhang, Jinwenxi Shang, Antonis Koutsoumpis, Yuan Zong, Reinout E. de Vries, Wenming Zheng

arXiv:2608.25316v1cs.HC

TL;DR

Existing personality datasets often rely on task-free videos and crowdsourced apparent-personality labels, limiting their relevance to structured interview assessment. AVI-Personality introduces a trait-activated multimodal AVI dataset with psychologically grounded annotations and evaluates its reliability, validity, fairness, and model benchmarks. Observer-rated personality traits show moderate to high reliability, while multimodal methods perform best overall but only modestly exceed strong text-based baselines.

  • Problem

    Existing datasets often use short, task-free videos and crowdsourced or simplified apparent-personality annotations, limiting construct validity and relevance to structured interview assessment.

  • Method

    AVI-Personality uses structured asynchronous interviews with generic and trait-targeted questions, combining multimodal data with self-reports, trained-rater observer ratings, and job-related competency annotations.

  • Results

    Observer-rated personality traits had reliable and meaningful annotations, while text-based methods provided strong cues and multimodal methods achieved the best overall performance.

  • Takeaways & Limitations

    AVI-Personality provides a psychometrically grounded resource for developing reliable, valid, and fair AI-based personality and competency assessment algorithms.

  • Takeaways & Limitations

    Personality-targeted questions cover only four job-related HEXACO traits, while Emotionality and Openness rely on generic-question ratings and self-reports.

Abstract

from arXiv · show

With the rapid development of AI-based personality and job-related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task-free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI-Personality, a trait-activated multimodal dataset for personality and job-related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer-reported HEXACO personality traits and job-related competency. We validate AVI-Personality through reliability, construct validity, internal nomological association, fairness, and benchmark analyses. Validation results show that the observer-rated personality traits have moderate to high reliability, especially when ratings are based on personality-targeted questions. Benchmark results show that text-based AI algorithms provide strong personality-relevant cues, while multimodal methods achieve the best overall performance but only modestly outperform text-based baselines. In general, AVI-Personality provides a psychometrically grounded dataset for developing and evaluating AI-based models for personality and competency assessment. The dataset is available are released at https://github.com/APAL-SEU/AVI6

1 Introduction

AVI-Personality addresses validity gaps in existing personality and competency datasets by using structured, trait-activated interviews and psychologically grounded annotations. It validates the dataset psychometrically and benchmarks models across behavioral modalities.

  • Existing datasets often use task-irrelevant public videos and simplified or crowdsourced personality labels, weakening trait-measurement and construct validity.
  • AVI-Personality contains 3,876 interview videos from 646 participants completing a simulated job application with two generic and four HEXACO-targeted questions.
  • The dataset combines trait-activated interview responses with self-reports, BARS-based observer ratings by trained psychologists, recruiter-rated competencies, interview performance, and cognitive ability.
  • Annotation validation examined reliability, construct validity, internal nomological association, and demographic fairness, supporting annotation consistency and practical relevance.
  • Text-based methods provide strong personality-relevant cues, while multimodal methods achieve the best overall performance but only modestly outperform strong text-based baselines.

2 Related Work

Prior work established multimodal assessment resources but often relied on unstructured scenarios, apparent personality judgments, and nonprofessional ratings. AVI-Personality links structured, theory-based interview elicitation with validated self- and observer-report annotations.

  • Personality and competency assessment matters in personnel selection, while questionnaire-based assessment is interpretable but vulnerable to social desirability bias.
  • Automatic assessment has progressed from hand-crafted multimodal features and traditional machine learning toward models requiring rich behavioral signals and reliable labels.
  • Existing datasets commonly use short social-media videos or other weakly structured scenarios and crowd ratings, limiting theoretically grounded personality and competency assessment.
  • Trait Activation Theory motivates structured questions that elicit behavioral expressions linked to theoretically relevant personality traits and workplace behaviors.
  • AVI-Personality combines validated self-report questionnaires with multi-perspective observer ratings to support personality assessment and workplace-related analysis.

3 Dataset Construction

The final AVI-Personality sample comprises 646 native English-speaking adults in the United States after prespecified data-quality and compliance exclusions, with balanced gender representation.

  • 646 participants remained after excluding incomplete responses, missing consent, failed attention checks, abnormal HEXACO patterns, low effort, corrupted audio, and rater-identified noncompliance.
  • The final sample included 309 men, 309 women, and 28 non-binary participants, with a mean age of 36.69 years and 15.96 years of work experience.
  • Participants were native English-speaking adults living in the United States, and the sample included multiple reported ethnic groups.

3.2 Interview question development

The interview questions were developed as a fixed-order, structured protocol containing generic selection questions and personality-targeted questions. Recruiters and personality experts screened and selected questions for practical use and trait relevance.

  • The structured interview used generic questions to elicit broad selection information and personality questions to target personality-relevant content.
  • Questions followed a fixed order, with generic questions presented before personality questions; Table 1 records their content, order, types, and corresponding traits.
  • An initial pool of 86 literature-based job interview questions was screened to retain 61 open-ended, broadly applicable questions conveying personality information.
  • Seventeen professional recruiters evaluated practical usage, achieving ICC(2, 17) = 0.88 inter-rater agreement before expert assessment.
  • Four personality experts selected one past-behavior question for each of Honesty-Humility, Extraversion, Agreeableness, and Conscientiousness after resolving disagreements by consensus.

3.3 Interview video collection

Participants applied for a fictitious management traineeship and completed an asynchronous video interview as part of a simulated job application. They answered two generic and four personality questions within instructed response durations.

  • Participants applied for a fictitious management traineeship and completed an AVI using a study-developed platform.
  • The AVI included two generic questions and four personality questions designed in Section 3.2.
  • Participants were instructed to answer each interview question within 1–2 minutes.

3.4 Self-report collection

Participants completed the HEXACO-100 inventory and the ICAR cognitive ability test to measure personality domains and general cognitive abilities. These assessments produced domain-level personality scores and an overall cognitive ability score.

  • The HEXACO-100 measured Honesty-Humility, Emotionality, Extraversion, Agreeableness, Conscientiousness, and Openness to Experience.Items used a 5-point Likert scale, and domain scores were calculated by averaging corresponding items.
  • The ICAR assessed general cognitive ability through verbal reasoning, series, matrix reasoning, and three-dimensional rotation tasks.Responses were aggregated into an overall cognitive ability score.

3.5 Expert annotation

Expert annotations combined trained psychologist ratings of HEXACO traits with recruiter ratings of job-related competencies and interview performance. Personality ratings contrasted generic questions with trait-activating questions and used behaviorally anchored scales.

  • Personality annotation: At least three of twelve trained personality psychologists rated each participant’s observer-reported personality traits.Raters completed nine hours of training.
  • Personality annotation: Generic questions elicited ratings for all six HEXACO traits without targeting a specific trait.
  • Personality annotation: Personality questions elicited ratings for Honesty-Humility, Extraversion, Agreeableness, and Conscientiousness through trait-relevant behaviors.These questions were designed according to Trait Activation Theory.
  • Personality annotation: Using both question types enabled comparison between non-activated and activated assessment conditions.
  • Rating procedure: Personality raters used behaviorally anchored rating scales with scores from 1, Very low, to 5, Very high.
  • Competency annotation: At least two professional recruiters rated each participant on four competencies and overall interview performance after reviewing all six questions.The competencies were Integrity, Collegiality, Social versatility, and Development orientation.
  • Competency annotation: The five competency and interview-performance variables loaded on a single principal component in exploratory PCA.

4 Dataset Analysis and Validation

AVI-Personality was evaluated through reliability, construct validity, nomological association, predictive validity, and demographic fairness analyses. Observer ratings were generally reliable and more work-related than self-reports, while targeted questions improved agreement and reduced demographic sensitivity.

  • Reliability: Mixed-effects models separated participant, rater, and residual variance to estimate observer-rating reliability.
  • Reliability: MAICC treated systematic rater differences as measurement error, whereas MCICC assessed rank-order similarity after accounting for rater differences.
  • Reliability: Observer-rated personality traits showed moderate to high reliability, with higher reliability for personality-targeted than generic questions.Extraversion had the highest reliability among the traits discussed.
  • Reliability: Competency ratings had lower reliability than personality traits, with modest absolute agreement but moderate consistency in candidate ranking.Consistency was especially evident for Development Orientation and overall interview performance.
  • Construct validity: Personality-targeted questions produced the strongest self-other agreement for Extraversion (r = 0.42) and Conscientiousness (r = 0.41).Agreeableness reached r = 0.25 and Honesty-Humility r = 0.22.
  • Nomological association: The competency regression model explained 87.1% of interview-performance variance (R2 = 0.871), with social versatility the strongest predictor (β = 0.405).All four competencies showed significant positive associations with interview performance.
  • Predictive validity: Observer-rated personality models explained 17.7%–27.1% of competency and hireability variance, compared with 2.7%–11.0% for self-reported personality models.Generic- and personality-question observer models showed similar ranges.
  • Fairness: Personality-question observer ratings showed fewer significant demographic effects than generic-question observer ratings.

5 Benchmark

The benchmark compares text, audio, visual, and multimodal methods for personality assessment using observer ratings from personality-targeted questions and MSE evaluation. Text methods provide strong personality cues, while multimodal methods perform best overall but only modestly exceed strong text baselines.

  • Benchmark design: Observer ratings from personality-targeted questions serve as benchmark labels because trait-relevant situations facilitate personality expression.The benchmark uses these ratings instead of generic-question ratings to elicit behavioral cues for AI assessment.
  • Benchmark design: The benchmark evaluates text-based, audio-based, visual-based, and multimodal methods, with subject-level splits and balanced demographic distributions.Training, validation, and test sets contain 70% (n = 452), 10% (n = 64), and 20% (n = 130) of subjects, respectively.
  • Evaluation: MSE evaluates continuous-trait prediction, following prior AVI personality-assessment studies and benchmarks.MSE measures average squared differences between predicted and actual trait values.
  • Text-based results: PersonalityLLM achieved the best average results among text-based methods, including for H and A, whereas general-purpose GPT-4 and DeepSeek-R1 performed worse.The comparison indicates that task-specific personality alignment matters beyond model scale alone.
  • Unimodal results: Whisper, Emotion2Vec, and Wav2Vec2 outperformed handcrafted acoustic features, while Swin Transformer was the strongest visual baseline.Audio and visual-only methods generally produced higher MSEs than text-based methods, indicating more limited unimodal information.
  • Multimodal results: Multimodal methods achieved the best overall performance, but their advantage over strong text-based baselines was relatively small.EMMR achieved the best result for C, while PersonalityLLM achieved the best results for H and A; fusion effectiveness depended on integrating complementary cues.

6 Limitation and future works

The paper identifies limitations in trait coverage, rating agreement, and benchmark scope. It proposes broader trait-activated questions, improved rater calibration, and more effective multimodal fusion as future directions.

  • Trait coverage: Personality-targeted questions cover four job-related HEXACO traits, while Emotionality and Openness rely on generic-question ratings and self-reports.Future studies should design and validate targeted questions for all HEXACO domains.
  • Annotation: Trained-rater BARS annotations may still reflect rater-specific preferences or demographic sensitivity.The paper calls for examining human rating decisions and whether trained models reproduce or amplify these patterns.
  • Annotation: Observer-rated personality traits showed moderate to high consistency, but job-related competencies had lower absolute agreement and may require additional calibration.The limitation concerns competency judgments rather than the reported reliability of personality ratings.
  • Benchmark scope: The benchmark does not exhaust all modeling strategies, and multimodal fusion only slightly outperformed strong text-based baselines.Future work should investigate strategies that better integrate linguistic, acoustic, and visual cues.

7 Conclusion

AVI-Personality is a psychometrically grounded dataset for personality and competency assessment from structured asynchronous video interviews. Its validation supports reliable personality annotations and strong text and multimodal benchmark performance.

  • Dataset contribution: AVI-Personality contains 3,876 videos from 646 participants answering generic and personality interview questions.The dataset includes multimodal behavior, self- and observer-rated personality, recruiter-rated competencies, interview performance, and cognitive ability.
  • Validation: Observer-rated personality traits, especially those based on personality-targeted questions, had reliable and meaningful annotations.The dataset was evaluated through reliability, construct validity, internal nomological association, and fairness analyses.
  • Validation: Job-related competencies significantly predicted interview performance.This result forms part of the dataset’s internal validation evidence.
  • Benchmark conclusion: Text-based methods provided strong personality-relevant signals, while multimodal methods achieved the best overall performance.The benchmark supports using AVI-Personality to develop and evaluate AI-based assessment algorithms.
Loading 2608.25316v1…