Source-linked AI summary

Getting aligned on representational alignment

Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C. Love, Christopher J. Cueva, Erin Grant, Iris Groen, Jascha Achterberg, Joshua B. Tenenbaum, Katherine M. Collins, Katherine L. Hermann, Kerem Oktar, Klaus Greff, Martin N. Hebart, Nathan Cloos, Nikolaus Kriegeskorte, Nori Jacoby, Qiuyi Zhang, Raja Marjieh, Robert Geirhos, Sherol Chen, Simon Kornblith, Sunayana Rane, Talia Konkle, Thomas P. O'Connell, Thomas Unterthiner, Andrew K. Lampinen, Klaus-Robert Müller, Mariya Toneva, Thomas L. Griffiths

arXiv:2310.13018v3q-bio.NCcs.AIcs.LGcs.NE

TL;DR

Representational alignment research asks how to compare, connect, and modify the representations formed by biological and artificial systems, but its communities lack a shared language. This Perspective reviews work across cognitive science, neuroscience, and machine learning, proposes a unifying framework, and identifies open problems. It also shows that alignment and behavioral performance need not improve together and that measurement and stimulus choices can substantially affect conclusions.

  • Problem

    Representational alignment research is fragmented across fields, with limited knowledge transfer, duplicated efforts, and inconsistent terminology and implementations.

  • Method

    The Perspective conducts a broad cross-disciplinary literature review and organizes studies using a unifying framework spanning alignment measurement, shared spaces, and alignment increases.

  • Results

    The framework maps diverse studies across cognitive science, neuroscience, and machine learning, while reviewed work shows that higher model performance does not guarantee higher human alignment.

  • Takeaways & Limitations

    A common language can help synthesize methods and findings across disciplines and support progress on shared representational-alignment problems.

  • Takeaways & Limitations

    Alignment conclusions depend on the chosen measure and dataset, because different measures and restricted or confounded stimuli can produce different inferences or limit generalization.

Abstract

from arXiv · show

Biological and artificial information processing systems form representations of the world that they can use to categorize, reason, plan, navigate, and make decisions. How can we measure the similarity between the representations formed by these diverse systems? Do similarities in representations then translate into similar behavior? If so, then how can a system's representations be modified to better match those of another system? These questions pertaining to the study of representational alignment are at the heart of some of the most promising research areas in contemporary cognitive science, neuroscience, and machine learning. In this Perspective, we survey the exciting recent developments in representational alignment research in the fields of cognitive science, neuroscience, and machine learning. Despite their overlapping interests, there is limited knowledge transfer between these fields, so work in one field ends up duplicated in another, and useful innovations are not shared effectively. To improve communication, we propose a unifying framework that can serve as a common language for research on representational alignment, and map several streams of existing work across fields within our framework. We also lay out open problems in representational alignment where progress can benefit all three of these fields. We hope that this paper will catalyze cross-disciplinary collaboration and accelerate progress for all communities studying and developing information processing systems.

1 Introduction

Representational alignment concerns how similarly biological and artificial systems represent the world, and how those similarities can be measured, connected, or increased. The Perspective reviews these goals across disciplines and proposes a shared framework to improve communication and progress.

  • Representational alignment is the extent to which two or more information processing systems’ internal representations agree.
  • Limited knowledge transfer and a lack of shared language have produced duplicated efforts, motivating a unifying framework and identification of cross-disciplinary open problems.
  • The field studies three objectives: measuring alignment, bridging representations into a shared space, and increasing alignment by updating at least one system.
  • Measuring: Measuring alignment compares representational structures at an abstract information-processing level and can evaluate one system as a model of another.
  • Bridging: Bridging establishes correspondence between systems so their representations can be pooled and compared along common dimensions.
  • Increasing: Increasing alignment updates representations to make one system more like another, either to improve a model or to support downstream performance.

2 Background and review

Representational alignment research spans cognitive science, neuroscience, and machine learning, where researchers often use related techniques from different perspectives. The review uses this literature to motivate a unifying framework and identify future gaps.

  • Researchers across cognitive science, neuroscience, and machine learning study representational alignment from differing perspectives while often converging on similar techniques.
  • The cross-field review motivates a unifying framework and identifies gaps for future work.

2.1 Cognitive Science

Cognitive science uses representational alignment to study shared and differing representations across people, cultures, development, and human–machine comparisons. These studies use similarity-based methods and show that alignment can vary across domains and need not track behavioral performance.

  • Whether different people share representations of the world is a central question spanning cross-cultural and developmental cognitive science.
  • Methods: Multidimensional scaling embeds stimuli into low-dimensional spaces so distances reflect participants’ similarity judgments.
  • Methods: Representational similarity methods accommodate continuous or discrete, symmetric or asymmetric, and hierarchical or non-hierarchical systems.
  • Human-machine alignment: Human similarity judgments can correlate with final-layer convolutional-network activation inner products, but better object-recognition performance can coincide with worse human alignment.
  • Semantic representations: Representational alignment is also used to study changing semantic representations during learning and their decay under neurodegenerative disease.
  • Cross-linguistic alignment: Semantic neighborhoods across 41 languages were more aligned in highly structured domains such as number and kinship than in natural kinds and common actions.

2.2 Neuroscience

Neuroscience applies representational alignment to compare heterogeneous neural and computational systems, align data across individuals, and test hypotheses about information processing. Studies also use alignment to investigate communication and select informative stimuli.

  • Representational Similarity Analysis was motivated by the challenge of comparing heterogeneous internal activities across individuals, species, and biological and artificial systems.
  • Bridging neural representations: Neural responses can be aligned across individuals into a common space, avoiding information loss associated with standard anatomical warping caused by morphological differences.
  • Brain–model alignment: Alignment between brain regions and computational models helps study relationships between model representations and neural anatomy and function.
  • Hypothesis testing: Neuroscientists compare candidate representations against brain responses to test hypotheses about task-dependent processing in vision and language.
  • Stimulus selection: Representational alignment can guide stimulus selection, including controversial stimuli that distinguish competing accounts of neural activity or behavior.
  • Alignment as communication: Speaker–listener neural alignment emerges during successful communication, while prefrontal alignment during play emerges during joint but not independent play.

2.3 Artificial intelligence and machine learning

Machine learning research applies representational alignment to compare, bridge, interpret, and modify model representations, including for multimodal systems, distillation, human alignment, and value alignment.

  • Machine learning researchers use alignment to compare models, fuse representation spaces, improve robustness, and interpret performance.
  • Multimodality: Multimodal models jointly optimize representations from different input modalities rather than merely bridging independently learned spaces.
  • Cross-model alignment can use linear regression or similarity-based relative representation spaces to translate between latent spaces.
  • Knowledge distillation: Knowledge distillation aligns a smaller student network with a larger teacher by transferring prior knowledge through probabilistic outputs.
  • Increasing alignment between human and neural-network representations is pursued to understand similarities and improve system outputs at lower computational cost.
  • Behavioral and value alignment: Behavioral alignment can match outputs without matching internal representations, while output alignment alone may not predict continued value alignment.

3 Framework for representational alignment

Representational alignment research is fragmented across disciplines, with differing terminology and limited knowledge sharing. The paper proposes a general formalism as a shared language to accelerate progress.

  • Limited knowledge transfer has produced duplicated ideas, repeated mistakes, and underused opportunities for cross-disciplinary collaboration.
  • The proposed formalism aims to unify disparate definitions and communities studying representational alignment.

3.1 High-level overview

The framework represents alignment studies as a controlled pipeline from data presentation through inferred embeddings to an alignment score. It distinguishes measuring, bridging, and increasing alignment as three objectives.

  • Most alignment studies contain five controllable components: data, systems, measurements, embeddings, and an alignment function.
  • Data may be sensory or cognitive content, and the framework assumes static samples while allowing generalization to dynamic environments.
  • Interfaces shape how stimuli reach systems, which then form internal representations that researchers measure through potentially different procedures.
  • Embeddings transform measured outputs into spaces suitable for comparison, including lower-dimensional or continuous representations.
  • Measuring computes an alignment score, whereas bridging or increasing uses that score as feedback to update embeddings, measurements, or internal representations.
  • The framework is intended to communicate alignment methodology and results across disciplines using a simple, general language.

3.2 Formalizing representation spaces

The formalization models stimuli, systems, measurements, and embeddings as functions in a common pipeline. It accommodates varied data types, active task settings, and measurement structures.

  • The framework begins with a dataset of n trials, whose elements may be images, strings, sequences, videos, or environment states.
  • Two systems are modeled as functions mapping stimuli to internal states, with interface effects absorbed into their parameters.
  • The formalization can include active task engagement, environmental modification, stationary context, and task context represented as part of stimuli.
  • Measurements summarize the states produced across all trials through possibly parameterized measurement functions.
  • Optional embedding functions map measurements into another, potentially lower-dimensional space for comparison or denoising.
  • Measurements may be vectors, matrices, graphs, programs, or strings if an appropriate alignment measure exists.

3.3 Measuring alignment

Representational alignment is quantified by functions that map two embedded representations to scalar similarity or dissimilarity values, with choices determining whether alignment is described, increased, or interpreted directionally. The section emphasizes that implementation details and measure properties can materially affect conclusions.

  • Alignment functions: An alignment function δ maps two embedded vectors to a scalar that quantifies their degree of alignment.Under the dissimilarity convention, δ(v, w)=0 indicates fully aligned embedding vectors.
  • Alignment functions: Valid alignment functions must quantify similarity or dissimilarity, be descriptive or differentiable, and be symmetric or directional.These dimensions distinguish functions used to describe alignment from those used to increase it.
  • Similarity or dissimilarity: Similarity measures provide bounded reference points, while dissimilarity measures have a known zero point but often lack an interpretable upper bound.Dissimilarity functions can nevertheless serve as error functions minimized by gradient descent.
  • Descriptive or differentiable: Descriptive measures quantify relationships between representations, whereas differentiable measures can define losses for increasing alignment through gradient-based optimization.Rank correlations used in RSA are descriptive but not differentiable; differentiable alignment functions can be minimized as L_alignment := δ(v, w).
  • Increasing alignment: Representational transformation learns an add-on mapping while keeping a model’s parameters frozen, requiring a choice about transformation flexibility.Linear encoding models mapping neural-network representations to animal single-neuron responses provide one example.
  • Measure choice: Different alignment measures can yield different conclusions because parameter fitting may underestimate similarity or become too flexible, while symmetric measures can hide information and noise differences.The authors recommend evaluating conclusions across multiple similarity measures where possible.
  • Implementation specificity: CKA scores depend substantially on the chosen HSIC estimator and related implementation conventions, so unambiguous comparison requires specifying these choices.Variants include unbiased or low-bias estimators, metric transformations, and local nearest-neighbor versions.

4 Representational alignment in diverse communities

The framework provides a common language for comparing representational-alignment approaches across fields. The authors illustrate it with research examples and formal descriptions to expose connections and support transfer of practices between communities.

  • Framework application: The framework organizes diverse alignment research by presenting conceptual summaries and formal mathematical descriptions for highlighted projects.Additional related literature is summarized in Table 2 according to the same framework.
  • Cross-disciplinary use: The framework is intended to help researchers recognize cross-community connections and transfer best practices to their own research topics.Its examples span communities studying the alignment of intelligent systems.

4.1 Cognitive Science

Cognitive-science studies use representational alignment to compare people’s internal structures and to identify shared or differing dimensions underlying behavior. Reviewed examples measure, bridge, or increase alignment using similarity judgments, embeddings, and alignment functions.

  • Measuring representational alignment: Serial reproduction studies found categorical rhythm prototypes near simple integer ratios across cultures, although category importance varied substantially between places.The paradigm iteratively reproduces randomized rhythms and identifies high-density response regions as emergent categories.
  • Measuring representational alignment: A formalized cognitive-science example represents three-interval rhythms with group-level kernel-density embeddings and compares groups using Jensen–Shannon divergence.The systems can represent different participant groups, such as US and Bolivian Amazon participants.
  • Bridging representational spaces: SPoSE recovered 49 interpretable dimensions from 1.46 million human triplet judgments across 1854 object categories, enabling direct comparison of candidate representational dimensions.The dimensions were highly predictive of single-trial choices and support interpretable alignment across individuals or modalities.
  • Bridging representational spaces: The framework describes human triplet choices as measurements and learns a low-dimensional embedding whose pairwise inner products encode object similarities.The embedding uses learnable variables with sparsity-inducing regularization; a Bayesian variant models uncertainty with a spike-and-slab prior.
  • Increasing representational alignment: Human triplet judgments were used to increase neural-network alignment with human object-similarity spaces, while preserving local similarity structure from the original network.The resulting representation showed increased alignment with human perception and better downstream performance on various computer-vision tasks.

4.2 Neuroscience

Neuroscience examples measure representational geometry across species and modalities, bridge individual fMRI spaces into shared coordinates, and align brain activity with behavior or artificial networks. These studies use embeddings and descriptive, symmetric, or directional alignment functions suited to neural and behavioral measurements.

  • Measuring representational alignment: RSA compared monkey electrophysiology and human fMRI representational geometry to assess whether inferotemporal cortex is homologous across primate species.The comparison used a descriptive and symmetric alignment function.
  • Measuring representational alignment: Monkey and human inferior-temporal responses are embedded as flattened representational-dissimilarity vectors derived from pairwise Pearson correlations, then compared with Spearman correlation.The monkey data come from electrophysiology and the human data from fMRI.
  • Bridging representational spaces: Across-individual fMRI alignment uses PLS regression to map each participant’s responses into a shared CNN-determined representation space, including held-out data.Rows of the fMRI measurements are flattened before PLS regression.
  • Bridging representational spaces: Within individuals, group-level CNN-transformed fMRI activity was averaged across feature dimensions and layers to produce a brain-based spatial-priority map for predicting image fixations.The comparison aligns neural activity with continuous eye-movement recordings using Normalized Scanpath Salience.
  • Increasing representational alignment: CNNs optimized to predict high-level visual-brain responses were reported to recapitulate visual behaviors, illustrating alignment increases between neural-network and fMRI representations.The cited work directly optimizes CNNs against fMRI responses rather than only measuring descriptive similarity.

4.3 Artificial Intelligence and Machine Learning

Machine-learning research applies representational alignment to compare neural-network layers, bridge image and text or modality-specific spaces, and distill teacher knowledge into students. Examples use kernel-based similarity, task-specific losses, and learned critics to connect alignment with transfer or downstream performance.

  • Measuring representational alignment: Centered Kernel Alignment provides a simple, widely used measure for comparing representations from artificial neural networks.When feature mappings are expensive or infinite-dimensional, CKA can operate through kernel similarity matrices.
  • Bridging representational spaces: Multimodal models enforcing text–image representational alignment achieved better cross-task transfer than standard multitask learning.The cited example found improved transfer from visual recognition to visual question answering, with additional recognition gains for sparsely labeled categories frequent in queries.
  • Bridging representational spaces: Gupta et al. partitioned descriptions into object and attribute words and used different alignment losses for the two groups.Object losses use a margin against alternative categories, whereas attribute losses use a sigmoid-based objective with minibatch positive-sample fractions.
  • Increasing representational alignment: Knowledge distillation can transfer representations across modalities by using a pretrained teacher and a trainable student on paired inputs.The teacher and student embeddings are their activation matrices, while the student learns to align with the teacher through a critic.
  • Increasing representational alignment: The distillation critic is trained to distinguish matching from non-matching input pairs, and its objective is reported as a lower bound on mutual information between teacher and student embeddings.The relative frequency of non-matching pairs is controlled by a hyperparameter.

5 Open problems & challenges in representational alignment

Open problems span dataset selection, representation extraction, black-box access, implementation-level confounds, and the uncertain relationship between representational and behavioral alignment. Addressing these challenges may support more informative comparisons and safer alignment efforts.

  • 5.1 Selecting data and stimuli: Alignment results can depend dramatically on the dataset, with restricted datasets risking poor generalization and naturalistic stimuli introducing feature confounds.Natural photos may correlate shape and texture, obscuring differences between human and CNN feature use.
  • 5.1 Selecting data and stimuli: Overly simplistic stimuli can also invalidate inferences because naturalistic tasks and complex feature interactions may engage different processes.Examples include continuous text versus short controlled tasks and retinal responses to motion interactions absent from simple bars or gratings.
  • 5.2 Extracting representations: Researchers must decide how to present stimuli and extract representations because processing time, model components, and measurement methods can change the observed alignment.Human recurrence and indirect measures such as fMRI or EEG can affect the representations available for comparison.
  • 5.2 Extracting representations: Full pairwise comparison across system regions is often infeasible, requiring prior literature and available tools to constrain hypotheses about where alignment should be assessed.Pairwise analyses can reveal parallels in processing progression between visual cortex and artificial CNNs.
  • 5.3 Black-box systems: Black-box systems require alternatives to direct similarity experiments, including representation elicitation methods whose mathematical parallels with generative processes may improve interpretability.Such experiments can be impractical for high-dimensional, large datasets.
  • 5.4 Representation, computation, and behavior: Implementation-level factors can distort alignment estimates, while similar representations do not guarantee similar outputs and aligned outputs do not require similar representations.These factors include energetic constraints in biology and pretraining or learnability biases in deep networks.
  • 5.4 Representation, computation, and behavior: Representational alignment may help diagnose causes of output misalignment and complement direct strategies because representations and outputs constrain one another.This relationship is especially relevant when designing human-centric AI thought partners.
  • 5.5 Possible risks of representational alignment: Increasing alignment may create risks, including making model-generated outputs harder to distinguish from human outputs and introducing downstream biases through alignment targets.The paper calls for frameworks to characterize and guard against these ramifications.

6 Conclusion

Representational alignment is central across cognitive science, neuroscience, and machine learning, but these communities lack a common language. The Perspective connects terminology, methods, developments, challenges, and open questions to encourage cross-field exchange.

  • 6 Conclusion: The Perspective builds bridges across cognitive science, neuroscience, and machine learning by aligning terminology and methods for representational alignment research.It also highlights shared histories, recent developments, common challenges, and open questions.
  • 6 Conclusion: The authors hope this shared view increases cross-field exchange and inspires applications of representational alignment to understanding or building more intelligent systems.The stated aim is to improve sharing of related ideas, methods, and empirical results.
Loading 2310.13018v3…