Source-linked AI summary
AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews
Tianyi Zhang, Jinwenxi Shang, Antonis Koutsoumpis, Yuan Zong, Reinout E. de Vries, Wenming Zheng
TL;DR
Existing personality datasets often rely on task-free videos and crowdsourced apparent-personality labels, limiting their relevance to structured interview assessment. AVI-Personality introduces a trait-activated multimodal AVI dataset with psychologically grounded annotations and evaluates its reliability, validity, fairness, and model benchmarks. Observer-rated personality traits show moderate to high reliability, while multimodal methods perform best overall but only modestly exceed strong text-based baselines.
Problem
Existing datasets often use short, task-free videos and crowdsourced or simplified apparent-personality annotations, limiting construct validity and relevance to structured interview assessment.
Method
AVI-Personality uses structured asynchronous interviews with generic and trait-targeted questions, combining multimodal data with self-reports, trained-rater observer ratings, and job-related competency annotations.
Results
Observer-rated personality traits had reliable and meaningful annotations, while text-based methods provided strong cues and multimodal methods achieved the best overall performance.
Takeaways & Limitations
AVI-Personality provides a psychometrically grounded resource for developing reliable, valid, and fair AI-based personality and competency assessment algorithms.
Takeaways & Limitations
Personality-targeted questions cover only four job-related HEXACO traits, while Emotionality and Openness rely on generic-question ratings and self-reports.
Abstract
from arXiv · showhide
With the rapid development of AI-based personality and job-related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task-free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI-Personality, a trait-activated multimodal dataset for personality and job-related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer-reported HEXACO personality traits and job-related competency. We validate AVI-Personality through reliability, construct validity, internal nomological association, fairness, and benchmark analyses. Validation results show that the observer-rated personality traits have moderate to high reliability, especially when ratings are based on personality-targeted questions. Benchmark results show that text-based AI algorithms provide strong personality-relevant cues, while multimodal methods achieve the best overall performance but only modestly outperform text-based baselines. In general, AVI-Personality provides a psychometrically grounded dataset for developing and evaluating AI-based models for personality and competency assessment. The dataset is available are released at https://github.com/APAL-SEU/AVI6
1 Introduction
AVI-Personality addresses validity gaps in existing personality and competency datasets by using structured, trait-activated interviews and psychologically grounded annotations. It validates the dataset psychometrically and benchmarks models across behavioral modalities.
- Existing datasets often use task-irrelevant public videos and simplified or crowdsourced personality labels, weakening trait-measurement and construct validity.
- AVI-Personality contains 3,876 interview videos from 646 participants completing a simulated job application with two generic and four HEXACO-targeted questions.
- The dataset combines trait-activated interview responses with self-reports, BARS-based observer ratings by trained psychologists, recruiter-rated competencies, interview performance, and cognitive ability.
- Annotation validation examined reliability, construct validity, internal nomological association, and demographic fairness, supporting annotation consistency and practical relevance.
- Text-based methods provide strong personality-relevant cues, while multimodal methods achieve the best overall performance but only modestly outperform strong text-based baselines.
2 Related Work
Prior work established multimodal assessment resources but often relied on unstructured scenarios, apparent personality judgments, and nonprofessional ratings. AVI-Personality links structured, theory-based interview elicitation with validated self- and observer-report annotations.
- Personality and competency assessment matters in personnel selection, while questionnaire-based assessment is interpretable but vulnerable to social desirability bias.
- Automatic assessment has progressed from hand-crafted multimodal features and traditional machine learning toward models requiring rich behavioral signals and reliable labels.
- Existing datasets commonly use short social-media videos or other weakly structured scenarios and crowd ratings, limiting theoretically grounded personality and competency assessment.
- Trait Activation Theory motivates structured questions that elicit behavioral expressions linked to theoretically relevant personality traits and workplace behaviors.
- AVI-Personality combines validated self-report questionnaires with multi-perspective observer ratings to support personality assessment and workplace-related analysis.
3 Dataset Construction
The final AVI-Personality sample comprises 646 native English-speaking adults in the United States after prespecified data-quality and compliance exclusions, with balanced gender representation.
- 646 participants remained after excluding incomplete responses, missing consent, failed attention checks, abnormal HEXACO patterns, low effort, corrupted audio, and rater-identified noncompliance.
- The final sample included 309 men, 309 women, and 28 non-binary participants, with a mean age of 36.69 years and 15.96 years of work experience.
- Participants were native English-speaking adults living in the United States, and the sample included multiple reported ethnic groups.
3.2 Interview question development
The interview questions were developed as a fixed-order, structured protocol containing generic selection questions and personality-targeted questions. Recruiters and personality experts screened and selected questions for practical use and trait relevance.
- The structured interview used generic questions to elicit broad selection information and personality questions to target personality-relevant content.
- Questions followed a fixed order, with generic questions presented before personality questions; Table 1 records their content, order, types, and corresponding traits.
- An initial pool of 86 literature-based job interview questions was screened to retain 61 open-ended, broadly applicable questions conveying personality information.
- Seventeen professional recruiters evaluated practical usage, achieving ICC(2, 17) = 0.88 inter-rater agreement before expert assessment.
- Four personality experts selected one past-behavior question for each of Honesty-Humility, Extraversion, Agreeableness, and Conscientiousness after resolving disagreements by consensus.
3.3 Interview video collection
Participants applied for a fictitious management traineeship and completed an asynchronous video interview as part of a simulated job application. They answered two generic and four personality questions within instructed response durations.
- Participants applied for a fictitious management traineeship and completed an AVI using a study-developed platform.
- The AVI included two generic questions and four personality questions designed in Section 3.2.
- Participants were instructed to answer each interview question within 1–2 minutes.
3.4 Self-report collection
Participants completed the HEXACO-100 inventory and the ICAR cognitive ability test to measure personality domains and general cognitive abilities. These assessments produced domain-level personality scores and an overall cognitive ability score.
- The HEXACO-100 measured Honesty-Humility, Emotionality, Extraversion, Agreeableness, Conscientiousness, and Openness to Experience.Items used a 5-point Likert scale, and domain scores were calculated by averaging corresponding items.
- The ICAR assessed general cognitive ability through verbal reasoning, series, matrix reasoning, and three-dimensional rotation tasks.Responses were aggregated into an overall cognitive ability score.
3.5 Expert annotation
Expert annotations combined trained psychologist ratings of HEXACO traits with recruiter ratings of job-related competencies and interview performance. Personality ratings contrasted generic questions with trait-activating questions and used behaviorally anchored scales.
- Personality annotation: At least three of twelve trained personality psychologists rated each participant’s observer-reported personality traits.Raters completed nine hours of training.
- Personality annotation: Generic questions elicited ratings for all six HEXACO traits without targeting a specific trait.
- Personality annotation: Personality questions elicited ratings for Honesty-Humility, Extraversion, Agreeableness, and Conscientiousness through trait-relevant behaviors.These questions were designed according to Trait Activation Theory.
- Personality annotation: Using both question types enabled comparison between non-activated and activated assessment conditions.
- Rating procedure: Personality raters used behaviorally anchored rating scales with scores from 1, Very low, to 5, Very high.
- Competency annotation: At least two professional recruiters rated each participant on four competencies and overall interview performance after reviewing all six questions.The competencies were Integrity, Collegiality, Social versatility, and Development orientation.
- Competency annotation: The five competency and interview-performance variables loaded on a single principal component in exploratory PCA.
4 Dataset Analysis and Validation
AVI-Personality was evaluated through reliability, construct validity, nomological association, predictive validity, and demographic fairness analyses. Observer ratings were generally reliable and more work-related than self-reports, while targeted questions improved agreement and reduced demographic sensitivity.
- Reliability: Mixed-effects models separated participant, rater, and residual variance to estimate observer-rating reliability.
- Reliability: MAICC treated systematic rater differences as measurement error, whereas MCICC assessed rank-order similarity after accounting for rater differences.
- Reliability: Observer-rated personality traits showed moderate to high reliability, with higher reliability for personality-targeted than generic questions.Extraversion had the highest reliability among the traits discussed.
- Reliability: Competency ratings had lower reliability than personality traits, with modest absolute agreement but moderate consistency in candidate ranking.Consistency was especially evident for Development Orientation and overall interview performance.
- Construct validity: Personality-targeted questions produced the strongest self-other agreement for Extraversion (r = 0.42) and Conscientiousness (r = 0.41).Agreeableness reached r = 0.25 and Honesty-Humility r = 0.22.
- Nomological association: The competency regression model explained 87.1% of interview-performance variance (R2 = 0.871), with social versatility the strongest predictor (β = 0.405).All four competencies showed significant positive associations with interview performance.
- Predictive validity: Observer-rated personality models explained 17.7%–27.1% of competency and hireability variance, compared with 2.7%–11.0% for self-reported personality models.Generic- and personality-question observer models showed similar ranges.
- Fairness: Personality-question observer ratings showed fewer significant demographic effects than generic-question observer ratings.
5 Benchmark
The benchmark compares text, audio, visual, and multimodal methods for personality assessment using observer ratings from personality-targeted questions and MSE evaluation. Text methods provide strong personality cues, while multimodal methods perform best overall but only modestly exceed strong text baselines.
- Benchmark design: Observer ratings from personality-targeted questions serve as benchmark labels because trait-relevant situations facilitate personality expression.The benchmark uses these ratings instead of generic-question ratings to elicit behavioral cues for AI assessment.
- Benchmark design: The benchmark evaluates text-based, audio-based, visual-based, and multimodal methods, with subject-level splits and balanced demographic distributions.Training, validation, and test sets contain 70% (n = 452), 10% (n = 64), and 20% (n = 130) of subjects, respectively.
- Evaluation: MSE evaluates continuous-trait prediction, following prior AVI personality-assessment studies and benchmarks.MSE measures average squared differences between predicted and actual trait values.
- Text-based results: PersonalityLLM achieved the best average results among text-based methods, including for H and A, whereas general-purpose GPT-4 and DeepSeek-R1 performed worse.The comparison indicates that task-specific personality alignment matters beyond model scale alone.
- Unimodal results: Whisper, Emotion2Vec, and Wav2Vec2 outperformed handcrafted acoustic features, while Swin Transformer was the strongest visual baseline.Audio and visual-only methods generally produced higher MSEs than text-based methods, indicating more limited unimodal information.
- Multimodal results: Multimodal methods achieved the best overall performance, but their advantage over strong text-based baselines was relatively small.EMMR achieved the best result for C, while PersonalityLLM achieved the best results for H and A; fusion effectiveness depended on integrating complementary cues.
6 Limitation and future works
The paper identifies limitations in trait coverage, rating agreement, and benchmark scope. It proposes broader trait-activated questions, improved rater calibration, and more effective multimodal fusion as future directions.
- Trait coverage: Personality-targeted questions cover four job-related HEXACO traits, while Emotionality and Openness rely on generic-question ratings and self-reports.Future studies should design and validate targeted questions for all HEXACO domains.
- Annotation: Trained-rater BARS annotations may still reflect rater-specific preferences or demographic sensitivity.The paper calls for examining human rating decisions and whether trained models reproduce or amplify these patterns.
- Annotation: Observer-rated personality traits showed moderate to high consistency, but job-related competencies had lower absolute agreement and may require additional calibration.The limitation concerns competency judgments rather than the reported reliability of personality ratings.
- Benchmark scope: The benchmark does not exhaust all modeling strategies, and multimodal fusion only slightly outperformed strong text-based baselines.Future work should investigate strategies that better integrate linguistic, acoustic, and visual cues.
7 Conclusion
AVI-Personality is a psychometrically grounded dataset for personality and competency assessment from structured asynchronous video interviews. Its validation supports reliable personality annotations and strong text and multimodal benchmark performance.
- Dataset contribution: AVI-Personality contains 3,876 videos from 646 participants answering generic and personality interview questions.The dataset includes multimodal behavior, self- and observer-rated personality, recruiter-rated competencies, interview performance, and cognitive ability.
- Validation: Observer-rated personality traits, especially those based on personality-targeted questions, had reliable and meaningful annotations.The dataset was evaluated through reliability, construct validity, internal nomological association, and fairness analyses.
- Validation: Job-related competencies significantly predicted interview performance.This result forms part of the dataset’s internal validation evidence.
- Benchmark conclusion: Text-based methods provided strong personality-relevant signals, while multimodal methods achieved the best overall performance.The benchmark supports using AVI-Personality to develop and evaluate AI-based assessment algorithms.