Source-linked AI summary
Quantum Models with Multi-Stage Training for Compositional Concept Generalization
Mina Abbaszadeh, Matilda Karabina Moore, Mehrnoosh Sadrzadeh, Martha Lewis
TL;DR
Compositional generalization remains difficult for AI models, especially when relational concepts must transfer to unseen combinations. This paper evaluates quantum DisCoCat models with staged object-to-relation training and finds stronger OOD relational generalization than classical baselines with far fewer trainable parameters.
Problem
AI models still struggle to generalize relational concepts beyond combinations seen during training, motivating compositional approaches to multimodal learning.
Method
The paper translates DisCoCat representations into variational quantum circuits and uses multi-stage training that freezes learned object representations before relational learning.
Results
Multi-stage DisCoCat VQCs achieve stronger OOD relational generalization than classical baselines, using 426 versus 151 million trainable parameters.
Takeaways & Limitations
Structured quantum representations and staged learning support compositional generalization in grounded vision–language models.
Takeaways & Limitations
The evaluation is simulation-based and does not assess hardware-level performance, noise, or device-specific constraints.
Abstract
from arXiv · showhide
Compositional Concept Generalization (CoCoGen), the ability to systematically recombine learned primitives in novel contexts, is a key challenge for multimodal learning. In this work, we provide a solution using a compositional model of meaning that separates nouns from relations and uses tensors and variational quantum circuits to train them on data. This model enables us to employ a multi stage training paradigm, one that first learns object representations from single-object image-caption pairs, then subsequently transfers these to the relational stage where object parameters are frozen and optimisation is only applied to relational components. This design explicitly enforces compositional factorisation at the circuit, ensuring that relations are learned as transformations over stable primitives. The training paradigm is tested on the CLEVR dataset developed specificially for CoCoGen. For text, we work with vector representations of nouns and higher order tensor representations of relations using a set of different ansatz. For images, we work with quantum encodings of image embeddings dervied from Open AI's Vision Language tool CLIP and contrast amplitude encoding, which preserves the original embedding geometry, with angle encoding, which introduces nonlinear feature transformations. Our results show that multi-staged training combined with structured encodings significantly improves out of distribution relational generalisation, while using orders of magnitude fewer trainable parameters than classical baselines. We find that performance gains arise from the interaction between representation and encoding, with nonlinear quantum encodings enhancing the separability of compositional structure. These findings demonstrate that structured quantum representations and staged learning provide an effective framework for compositional generalisation in multimodal quantum machine learning.
I. INTRODUCTION · II. BACKGROUND · A. DisCoCat Meaning Representations in Hilbert Space
The paper addresses CoCoGen by evaluating DisCoCat-based variational quantum circuits and introducing multi-stage training that separates object grounding from relational learning. It situates this approach within DisCoCat’s categorical and Hilbert-space foundations, emphasizing stronger OOD relational generalization with substantially fewer trainable parameters.
- I. INTRODUCTION: CoCoGen requires models to recombine learned concepts in novel contexts, especially relational combinations beyond those observed during training.The paper motivates this challenge with humans recognizing an unseen yellow car as dangerous despite not having learned that exact combination.
- I. INTRODUCTION: DisCoCat and its variational quantum-circuit translation are evaluated on a controlled benchmark of geometric objects and spatial relations.The framework is presented as a compositional model of meaning whose tensor representations can be translated into trainable VQCs.
- I. INTRODUCTION: The proposed multi-stage curriculum first learns object-level representations, then disentangles relational learning from object learning.This design separates simpler object grounding from subsequent relational reasoning, rather than optimizing both stages identically.
- I. INTRODUCTION: DisCoCat-based VQCs achieve stronger OOD relational generalization than classical baselines, including OpenAI’s CLIP.The introduction attributes this improvement to separating object grounding from relational learning.
- I. INTRODUCTION: 426 vs 151 million trainable parameters accompanies the reported gains of the quantum approach over the cited classical comparison.The passage states that these improvements use very substantially fewer trainable parameters.
- II. BACKGROUND: The background frames the section around DisCoCat’s theoretical foundations, its VQC translation, and relevant multimodal baselines.This scope connects the formal representation framework to multimodal evaluation.
- A. DisCoCat Meaning Representations in Hilbert Space: DisCoCat maps grammatical structure to meaning through a functor between compact closed categories, preserving composition as morphisms over word representations.Meaning is formalized in finite-dimensional Hilbert spaces and tensor products, which form the relevant compact closed structure.
- A. DisCoCat Meaning Representations in Hilbert Space: For the dataset’s spatial captions, sentences are simplified to ‘noun {isLeftOf/isRightOf} noun’ with type n (nrsnl) n ≤1s1 = s to reduce circuit complexity.The resulting meaning representations use finite-dimensional Hilbert spaces and their tensor products.
B. From DisCoCat to Variational Quantum Circuits (VQCs)
The DisCoCat-to-quantum-circuit translation depends on the chosen ansatz. The model uses Sim4 to balance qubit and parameter counts with expressive power, while IQP encodes image embeddings.
- Ansatz selection: Quantum circuit translation depends on the ansatz, with IQP, Sim-family, and Matrix Product State constructions available.These ansatzes support quantum–classical hybrid learning.
- Ansatz selection: Sim4 is selected to control qubit and trainable-parameter counts while maintaining expressive power.It translates DisCoCat diagrams of cap...
- Image encoding: IQP encodes image embedding vectors in the quantum model.Further implementation details are provided in Section IV-A1.
C. CLIP and the Contrastive Learning Paradigm · III. METHODS · A. Task and Datasets
The study uses frozen CLIP image embeddings and a matched CLIP–Text Projection baseline, then evaluates compositional image–caption alignment on withheld combinations of known shapes and relations. The benchmark tests whether models can recombine familiar primitives into unseen configurations rather than recognize new primitives.
- C. CLIP and the Contrastive Learning Paradigm: CLIP provides separate Transformer-based image and text encoders that produce embeddings for paired image–text inputs.The image encoder produces a vector representing the image.
- C. CLIP and the Contrastive Learning Paradigm: CLIP image embeddings are frozen and used as input features for both classical and quantum models.Freezing prevents the input representations from changing during training.
- C. CLIP and the Contrastive Learning Paradigm: The CLIP–Text Projection baseline jointly fine-tunes a CLIP text encoder and lightweight projection head under the quantum models’ supervised contrastive objective.This matches training objectives so performance differences reflect representational structure rather than objective differences.
- A. Task and Datasets: The task trains models to align images of shapes in particular configurations with captions, then evaluates compositions absent from training.For example, training on “cone right cube” tests generalization to an unseen “cone left cube” composition.
- A. Task and Datasets: The CoBi2 benchmark uses single-object and relational splits covering four shapes—cube, sphere, cylinder, and cone—with varied sizes, colors, and positions.Relational examples contain two distinct shapes, a correct left/right caption, and a relation-swapped distractor.
- A. Task and Datasets: Dataset splits withhold relational triples, while training still exposes every individual shape and relation.Validation and test compositions therefore require recombining previously observed primitives into novel configurations.
B. Two-Stage Training Procedure · C. Compositional Factorisation via Multi-Stage Training
The paper uses two-stage training to learn stable object primitives before relational composition, with frozen object parameters forcing relations to operate as transformations over fixed representations. This factorisation supports systematic recombination of independently learned primitives in unseen compositions.
- B. Two-Stage Training Procedure: The two-stage regime first trains models on single-object image–caption pairs, then transfers their frozen parameters to relational captioning models.Stage 1 uses contrastive learning on “cone”, “cube”, “sphere”, and “cylinder” examples.
- C. Compositional Factorisation via Multi-Stage Training: The relational setting represents structured inputs as compositions of reusable object primitives and relation components.Objects are mapped by g(·), relations by h(·), and combined through a composition operation C(·).
- C. Compositional Factorisation via Multi-Stage Training: DisCoCat models implement composition through tensor contraction, while quantum models use circuit composition and wire contraction according to grammatical structure.This provides the compositional mechanism for combining noun and relation representations.
- C. Compositional Factorisation via Multi-Stage Training: Stage 1 learns object representations g(x; θobj) independently of relational context, establishing stable conceptual primitives for later composition.The object representations are learned from single-object image–caption pairs without introducing relational structure.
- C. Compositional Factorisation via Multi-Stage Training: In Stage 2, object parameters θobj remain frozen while only relation parameters θrel are optimised.Relational representations are constructed by composing learned object representations with a relation component.
- C. Compositional Factorisation via Multi-Stage Training: Freezing object parameters constrains relations to transform fixed object representations, reducing object-specific re-encoding and encouraging generalisation to unseen compositions.The model is therefore encouraged to generalise by recombining independently learned primitives.
D. Image Encoding … 3) Collage Encoding:
The image-encoding section presents one-hot and multi-hot quantum encodings, quantum encodings of frozen CLIP embeddings, and a Collage encoding that composes shape-specific circuits. These methods use different circuit constructions to represent individual and relational images while supporting comparable image–caption outputs.
- 1) One-Hot Encoding (OHE) and Multi-Hot Encoding (MHE):: OHE and MHE use five qubits per noun, with images encoded as 5-qubit IQP circuits.The vectors in Table III parameterize entangling gates.
- 1) One-Hot Encoding (OHE) and Multi-Hot Encoding (MHE):: MHE concatenates noun vectors to encode relational images.This follows the encoding procedure illustrated in Table II.
- 2) Quantum Encodings of CLIP Embeddings:: Frozen CLIP image embeddings have 512 dimensions and are encoded using both angle and amplitude methods in variational quantum circuits.The embeddings are kept frozen during this image-encoding procedure.
- 2) Quantum Encodings of CLIP Embeddings:: Angle encoding reduces each 512-dimensional CLIP embedding to 8 PCA dimensions and maps them to rotation angles on 9 IQP qubits.The PCA reduction precedes the quantum encoding.
- 2) Quantum Encodings of CLIP Embeddings:: Amplitude encoding normalizes each 512-dimensional CLIP embedding to unit norm before loading it as a quantum state.The passage defines this normalized representation using amplitude encoding.
- 2) Quantum Encodings of CLIP Embeddings:: Image and sentence circuits produce output vectors with matching dimensionality so image–caption cosine similarity can be measured.Matching output dimensions connect the two branches for similarity evaluation.
- 3) Collage Encoding:: Collage forms an image circuit by combining circuits for each shape, using averaged CLIP vectors from single-object images.Each shape vector is loaded onto a quantum state using either of the two quantum-encoding methods.
- 3) Collage Encoding:: The two shape circuits are combined either with an entangling gate or by juxtaposition without a connection.These are the two circuit-composition options described for Collage encoding.
E. Caption Encoding and Relation Representation · IV. EXPERIMENTAL DETAILS · A. Single-Object Training
The relational caption circuit composes transferred noun representations with a trainable relation operator, producing sentence vectors matched to image vectors through cosine similarity. Stage 1 trains noun and image representations with supervised in-batch contrastive learning before relational training.
- E. Caption Encoding and Relation Representation: The relational operator treats isLeftOf and isRightOf as one pregroup-type nrsnl operator whose noun wires contract with duals, leaving the sentence wire s.The chosen quantum ansatz is applied to the relational box in the DisCoCat diagram.
- E. Caption Encoding and Relation Representation: The sentence wire uses the same qubit count as the image representation, so caption and image circuits output vectors with matching dimensionality.This matching dimensionality enables cosine similarity between modalities.
- E. Caption Encoding and Relation Representation: During relational training, noun parameters transferred from Stage 1 are frozen, while relational parameters remain trainable.The noun wires contract with the relation and are post-selected, leaving the s-wire as the sentence representation.
- E. Caption Encoding and Relation Representation: Image and caption circuit outputs are real-valued vectors compared with cosine similarity rather than quantum state overlap, enabling consistency with CLIP-style classical comparisons.The shared similarity framework supports direct comparison between classical and quantum representations.
- IV. EXPERIMENTAL DETAILS: The single-object and relational experiments use the lambeq framework, with all variational parameters randomly initialized and optimized during training.This describes the general experimental setup before the stage-specific optimization procedure.
- A. Single-Object Training: Stage 1 single-object training uses a cosine-similarity contrastive objective that raises similarity for matching image-caption pairs and lowers it for mismatched pairs.In-batch negatives are images from different shape classes; for a cube caption, cone, sphere, and cylinder images are negatives.
- A. Single-Object Training: N = B × K = 4 × 8 = 32 samples form each batch, with K −1 = 7 positives and N −K = 24 negatives per caption.The model computes a 32×32 = 1024 = N ×N caption-image cosine-similarity matrix per batch, using d = 512 and τ = 0.07.
1) Quantum Circuit Implementation Details:
The single-object angle-encoding stage compares Sim4 and IQP ansatzes for sentence and image encoders, tracking qubits, trainable parameters, and accuracy. For 8-dimensional PCA-reduced image embeddings, IQP uses 9 qubits while Sim4 uses 3, with qubit compatibility enforced across image and noun representations.
- Ansatz and encoding choices: The single-object angle-encoding experiments compare Sim4 and IQP ansatzes for sentence and image encoders using qubit counts, trainable parameters, and accuracy.These results are reported in Table IV.
- Ansatz and encoding choices: 9 qubits are required by IQP for 8-dimensional PCA-reduced image embeddings, compared with 3 qubits for Sim4.The qubit counts are specified for the image encoder in the single-object stage.
- Representation compatibility: The image and noun representations use the same number of qubits to remain compatible when computing cosine similarity.This compatibility constraint links the image and noun encoders during single-object training.
- Circuit implementation: Figure 5 depicts the quantum circuits used for sentence and image encoding during single-object training.The sentence-encoding circuit is shown on the left and the image-encoding circuit on the right.
B. Relational Training
Relational training uses pairwise ranking between an image’s correct caption and a relation-swapped distractor, with cosine similarity and a margin-based loss. Noun parameters transferred from single-object training are frozen, while relational parameters remain trainable across quantum and classical encoding variants.
- Optimisation objective: Relational training maximizes cosine similarity for the correct caption over a relation-swapped distractor using a margin-based ranking loss.The comparison is formulated pairwise because only two captions are contrasted.
- Optimisation objective: The quantum relational experiments use a margin m of 0.5.The passage defines y as the ground-truth label and s as cosine similarity.
- Staged relational training: Caption encoders reuse frozen noun parameters from single-object training, while relation parameters remain trainable.This staged freezing scheme is applied in the relational setting, including multi-hot image encoding and angle encoding of frozen CLIP embeddings.
- Relational encoding variants: The relational setup compares multi-hot, collage, and quantum CLIP encodings with a classical baseline using fixed CLIP representations.The classical baseline computes cosine similarity directly between classical CLIP embeddings.
V. RESULTS AND ANALYSIS
The results show that quantum encodings and multi-stage training achieve strong compositional generalization with far fewer trainable parameters than classical baselines. Performance varies substantially by encoding and model configuration, while OOD validation and test accuracy can differ because their compositional splits have different difficulty.
- Single-object results: OHE achieves 93.34% accuracy on the OOD test set in the single-object task.Results are reported as average test accuracy over five random seeds, with hyperparameters selected using validation performance.
- Single-object results: Amplitude Encoding is the strongest Q-CLIP model at 80.97% accuracy, versus 91% for the classical baseline.The quantum model uses 312 trainable parameters, compared with approximately 63M in the classical model.
- Relational results: MHE reaches 85.88% average OOD test accuracy, while Collage with angle encoding reaches 72% and outperforms random guessing and Q-Single-CLIP.Q-Multi-CLIP with angle and amplitude encoding reaches 55.33% and 50%, respectively.
- Relational results: Classic-Rel-CLIP reaches 75.94% training accuracy but only 50% OOD test accuracy, indicating poor generalization to unseen relational compositions.Q-Multi-CLIP transfers noun parameters learned during Stage 1, whereas Q-Single-CLIP is trained directly on the relational task.
- Evaluation caveat: OOD validation accuracy may be lower than OOD test accuracy because held-out validation and test compositions differ in difficulty.The difference arises from how compositions are partitioned between validation and test sets.
A. Analysis
Representational similarity analysis shows that neither quantum encoding consistently produces the desired block-diagonal similarity structure across relational classes. The findings suggest angle encoding enhances relational separability, whereas amplitude encoding preserves CLIP geometry that lacks explicit relational structure.
- Representational similarity: The desired block-diagonal pattern requires high within-class similarity and low between-class similarity across quantum encodings.Figure 6 evaluates amplitude and angle encoding state fidelity on test images, using cube, cylinder, cone, sphere, left, and right labels.
- Representational similarity: Neither amplitude nor angle encoding exhibits the desired block-diagonal pattern in Figure 6.Amplitude encoding shows some block-diagonal structure but also off-diagonal blocks and overall high similarity; angle encoding shows low similarity within and outside classes.
- Underlying representations: CLIP captures strong object-level semantic similarity but does not reliably encode compositional relationships, consistent with classical baselines failing on unseen relational compositions.The quantum embedding is interpreted as enhancing or restructuring the representation space to better reflect compositional structure.
- Encoding effects: Angle encoding introduces nonlinear sinusoidal feature transformations that increase spatial-relation separability, whereas amplitude encoding preserves CLIP geometry and leaves left and right highly similar.The absence of required block-diagonal patterns indicates that CLIP image embeddings do not explicitly represent the relevant relational structure.
VI. CONCLUSION AND FUTURE WORK
The paper presents a quantum multimodal framework that separates object grounding from relational learning through multi-stage training, improving generalization and parameter efficiency. Future work will test the approach beyond controlled simulations and small compositional benchmarks.
- Conclusion: The framework combines DisCoCat, variational quantum circuits, and multi-stage training to separate object grounding from relational learning.This design improves generalization and parameter efficiency compared to classical baselines.
- Conclusion: All code for single-object and relational training, image–caption alignment, and multi-stage parameter transfer is publicly available.The repository is provided at https://github.com/Mina-Abbaszade/Quantum-Multimodality.
- Limitations and future work: Experiments were conducted in simulation to isolate representational and compositional effects independently of hardware noise.The circuits use standard variational ansatz families and operate on average qubit counts, including 9 qubits for singl…
- Limitations and future work: The evaluation uses a controlled benchmark with few objects and spatial relations, motivating analysis of behavior as compositional complexity increases.The authors specifically identify richer relational vocabularies and more complex compositional settings as important directions.