Source-linked AI summary

Different representation learning objectives recover distinct latent structures from the same psychometric data

Cong Cao, Tassos C. Kyriakides, Pambos Vrasidas

arXiv:2609.00100v1cs.AIcs.LGstat.ME

TL;DR

It remains unclear whether different representation learning objectives recover the same latent organization from psychometric data. Using matched teacher–child psychometric data, the study compares PCA-derived behavioral structure with contrastive and multi-task representations. Contrastive learning improves teacher–child correspondence but preserves behavioral phenotypes less effectively, while multi-task learning partially restores behavioral organization at the cost of retrieval performance.

  • Problem

    Whether different representation learning objectives recover the same latent structure from the same psychometric observations remains unanswered.

  • Method

    The study uses 757 matched teacher–child pairs, identifies behavioral phenotypes with PCA and clustering, and compares contrastive and multi-task representation objectives.

  • Results

    Contrastive representations recover strong teacher–child correspondence but preserve behavioral phenotype structure less effectively than PCA-based representations.

  • Takeaways & Limitations

    Teacher–child correspondence and behavioral phenotypes are distinct latent organizations, and the recovered structure depends on the representation learning objective.

  • Takeaways & Limitations

    The analyses were limited to the baseline assessment of the Cyprus cohort, so generalizability to other educational settings remains to be established.

Abstract

from arXiv · show

Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. We investigated this question using 757 matched teacher-child pairs from the baseline assessment of the Cyprus ProW preschool trial. Behavioral structure was characterized from child SDQ, ASBI, and CBRS item responses using principal component analysis and clustering, yielding four behavioral phenotypes. A contrastive objective substantially improved teacher-child retrieval relative to PCA-based representations, increasing Top-1 accuracy from 0.13% to 7.27% and Top-10 accuracy from 1.98% to 56.14%. However, contrastive representations preserved behavioral phenotype structure less effectively than PCA-based representations. A multi-task objective jointly optimizing alignment and behavioral prediction partially restored behavioral organization but reduced retrieval performance. These findings indicate that teacher-child correspondence and behavioral phenotypes represent distinct forms of latent organization and demonstrate that the latent structure recovered from linked psychometric data depends on the representation learning objective.

3. Center for the Advancement of Research & Development in Educational Technology (CARDET), Nicosia, Cyprus

The paper concerns representation learning, latent structure, contrastive learning, and psychometric data, including teacher–child correspondence and behavioral phenotypes.

  • The study addresses representation learning and latent structure in psychometric data.
  • Teacher–child correspondence and behavioral phenotypes are identified as central concepts.

1. Introduction

The introduction frames an unresolved question: whether different representation learning objectives recover the same latent structure from identical psychometric observations. It motivates a comparison of behavioral phenotypes, teacher–child correspondence, and objective-specific representations using linked questionnaire data.

  • Item-level questionnaire responses may retain psychological information that aggregated scale scores do not fully preserve.
  • Representation learning can produce informative low-dimensional representations directly from high-dimensional item-level observations.
  • Prediction-oriented representations are not necessarily optimized for scientific interpretation.
  • The study asks whether different representation learning objectives recover different latent structures from the same psychometric observations.
  • Linked teacher–child data contain two related but potentially distinct organizations: behavioral similarity among children and teacher–child correspondence.
  • Using 757 matched teacher–child pairs, the study identifies behavioral phenotypes with PCA and clustering, then compares contrastive and multi-task representations.

2. Method

The study uses linked baseline teacher–child questionnaire data, item-level preprocessing, behavioral phenotype construction, and dual-encoder representation learning. It evaluates alignment, behavioral information, and teacher–child structure under contrastive and multi-task objectives.

  • 2.1 Dataset and Preprocessing: Teacher questionnaires assess wellbeing, self-efficacy, job satisfaction, burnout, and professional climate, while child assessments use SDQ, ASBI, and CBRS.
  • 2.1 Dataset and Preprocessing: The dataset contains 757 linked teacher–child pairs from Wave 1 baseline data in the Cyprus ProW cohort.
  • 2.1 Dataset and Preprocessing: Analyses use item-level questionnaire responses, with designated missing codes median-imputed and variables standardized by z-score normalization.
  • 2.5 Evaluation: Retrieval evaluates alignment, while ridge regression with five-fold cross-validation tests whether child embeddings retain behavioral information.
  • 2.2 Behavioral Structure: Behavioral phenotypes are derived from 82 SDQ, ASBI, and CBRS indicators using PCA, retained principal components, and clustering.
  • 2.2 Behavioral Structure: Phenotype solutions are assessed for cluster separation and interpretability, with assignments compared across behavioral outcomes using one-way ANOVA.
  • 2.4 Teacher–Child Contrastive Representation Learning: A dual-encoder framework independently encodes teacher and child questionnaire items into a shared latent space and learns alignment with cosine similarity and InfoNCE.
  • 2.4 Teacher–Child Contrastive Representation Learning: Matched pairs are positives and unmatched pairs are negatives within mini-batches, encouraging matched responses to become more similar.

3. Results

Results showed that representation learning objectives recovered different aspects of structure from the psychometric data. PCA captured well-separated behavioral phenotypes, contrastive learning improved teacher–child retrieval, and multi-task learning traded some retrieval for stronger behavioral organization.

  • Behavioral latent structure: PCA explained 37.7% of variance in its first component and 64.1% across its first ten components of 82 behavioral indicators.
  • Behavioral latent structure: K-means applied to PCA embeddings identified four behavioral phenotypes containing 339, 217, 117, and 97 children.
  • Behavioral latent structure: The PCA-derived phenotypes formed a coherent gradient from adaptive functioning to elevated behavioral risk, with differences across SDQ, ASBI, and CBRS.ANOVA results were F = 161.79 for SDQ, F = 682.54 for ASBI, and F = 834.08 for CBRS, all p < 0.001.
  • Teacher characteristics: Teacher characteristics showed generally modest associations with classroom behavioral composition, and teacher questionnaires alone reached 60.5% accuracy versus a 63.5% majority-class baseline.The strongest scale-level association involved the Professional Climate Scale and the High Adaptive Functioning group.
  • Contrastive learning: Contrastive representations improved teacher–child retrieval over PCA, raising Top-1 accuracy from 0.13% to 7.27% and Top-10 accuracy from 1.98% to 56.14%.The reported Transformer estimates had narrow bootstrap confidence intervals, indicating stable retrieval performance across resamples.

4 Discussion

The latent structure recovered from linked psychometric data depended on the representation learning objective. Contrastive learning improved teacher–child correspondence, whereas PCA better preserved behavioral phenotypes; multi-task learning partially balanced these objectives.

  • Objective-dependent representations: Contrastive learning recovered strong teacher–child correspondence but preserved behavioral phenotype structure less effectively than PCA.PCA recovered clear behavioral phenotypes but provided little information about teacher–child correspondence.
  • Interpretation of retrieval: Retrieval improved substantially under contrastive learning, although Top-1 accuracy remained modest in absolute terms.The result should be interpreted in the context of a difficult retrieval task involving many possible teacher–child matches.
  • Interpretation of retrieval: The large improvement over PCA-based retrieval is likely more informative than the absolute retrieval performance.
  • Objective-dependent representations: PCA preserves dominant variation in child behavioral data, whereas contrastive learning explicitly optimizes teacher–child alignment.
  • Objective-dependent representations: Behavioral-phenotype features may contribute little to teacher–child matching, while alignment features may not strengthen behavioral separation.
  • Objective-dependent representations: Adding behavioral supervision partially restored behavioral organization but reduced retrieval performance, suggesting competition for representational capacity.
  • Multiple latent structures: Teacher–child correspondence and behavioral phenotypes are related but distinct latent structures, and no single representation recovered both equally well.
  • Multiple latent structures: The recovered representation reflects both the underlying observations and the learning objective used to analyze them.

5 Conclusion

The study identifies multiple latent structures in linked teacher–child psychometric data, with alignment and behavioral organization recovered differently by the learning objective. The findings argue against a single representation being optimal for every downstream analysis.

  • Teacher–child alignment and behavioral organization emerged as distinct latent structures.
  • Alignment-optimized representations recovered stronger teacher–child correspondence, whereas behaviorally optimized representations produced clearer behavioral phenotypes.
  • Neither objective recovered both structures equally well.
  • The latent structure recovered from psychometric data depends on the representation learning objective.
  • Learning objectives should be matched to the scientific question rather than assuming one representation is optimal for all downstream analyses.

Supplementary Material

Supplementary analyses characterize associations and Transformer-derived clusters, showing limited behavioral differentiation for standard embeddings and stronger separation for multi-task embeddings.

  • Transformer-derived clusters were relatively balanced in size but showed limited behavioral differentiation across SDQ, ASBI, and CBRS scores.
  • Cluster 2 of the multi-task Transformer had the highest adaptive classroom functioning, while Cluster 3 had the lowest adaptive profile.
  • The multi-task Transformer achieved its highest silhouette coefficient at K = 2 (0.595), indicating well-defined cluster separation.
  • Standard Transformer embeddings reached a maximum silhouette coefficient of only 0.085 at K = 8, suggesting limited intrinsic cluster structure.

N SDQ Mean (SD) ASBI Mean

The representation comparisons show that contrastive learning improves teacher–child retrieval but weakens behavioral phenotype separation, while multi-task supervision restores some behavioral organization.

  • Multi-task embeddings showed substantially stronger cluster structure than standard Transformer embeddings.
  • Multi-task embeddings reached a silhouette coefficient of 0.595 at K = 2, compared with 0.085 at K = 8 for Transformer embeddings.
  • The multi-task training loss decreased rapidly early and stabilized after approximately 60 epochs.
  • Contrastive embeddings achieved strong teacher–child retrieval performance, but behavioral phenotype separation remained limited.
  • The multi-task objective improved phenotype separation relative to the original contrastive objective, consistent with increased behavioral organization in latent space.
Loading 2609.00100v1…