Source-linked AI summary

Imagination improves Multimodal Translation

Desmond Elliott, Ákos Kádár

arXiv:1705.04350v2cs.CLcs.CV

TL;DR

Multimodal translation has lacked convincing evidence that visual context improves quality. The paper jointly learns translation and visually grounded representations through a shared encoder, achieving state-of-the-art Multi30K results without images at translation time. The improvements remain when image prediction or translation uses external datasets.

  • Problem

    Earlier multimodal translation systems had not convincingly shown that visual context improves translation quality.

  • Method

    Imagination uses multitask learning with a shared encoder, an attention-based translation decoder, and an image-prediction decoder for grounded representations.

  • Results

    Imagination achieves state-of-the-art Multi30K results without images at translation time, with improvements retained using external MS COCO images and News Commentary parallel text.

  • Takeaways & Limitations

    Separate parallel-text and described-image datasets can support multimodal translation without hurting performance, while image prediction still improves a stronger text-only baseline.

  • Takeaways & Limitations

    The experiments do not use Multi30K’s 155K independently collected German and English descriptions.

Abstract

from arXiv · show

We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.

1 Introduction

The paper addresses whether visual context can improve multimodal translation by learning grounded representations alongside translation. Imagination achieves competitive or state-of-the-art Multi30K results without using images at translation time and can use external image and text resources.

  • Motivation: Multimodal translation remains an open problem because earlier systems did not convincingly show that visual context improves translation quality.In the First Multimodal Translation Shared Task, only three systems outperformed an off-the-shelf text-only phrase-based model, and the best system was equally effective with or without visual features.
  • Method: Imagination decomposes multimodal translation into translation learning and visually grounded representation learning through a shared encoder and task-specific decoders.The translation decoder is attention-based, while the image prediction decoder predicts a global image feature vector associated with a sentence.
  • Results: 59.3 Meteor is achieved on Multi30K by ensembling models trained on in-domain and out-of-domain resources.The approach also reports 55.8 Meteor for a single model and 57.6 Meteor for an ensemble trained on multimodal in-domain data.
  • Results: Translation improvements persist when image prediction is trained on MS COCO and when the text-only baseline uses News Commentary parallel text.The paper reports no significant difference between in-domain and out-of-domain described-image training.
  • Contribution: Multitask learning permits external parallel text and described-image datasets to supplement expensive source-target-image data.The two tasks can be trained using resources such as parallel news text or images described in a single language.
  • Contribution: The model achieves state-of-the-art results without using images as an input during translation.It learns visually grounded source representations through an auxiliary image prediction objective and requires no additional parameters for unseen sentences.

2 Problem Formulation

The formulation replaces direct image-conditioned translation with two jointly learned tasks sharing source-encoder parameters. This enables optimization using separate image-description and parallel-text datasets.

  • Problem Formulation: The task takes a source sentence x and contextual image v as inputs and produces a target translation y.Training examples are tuples containing an image description, its translation, and the associated image.
  • Problem Formulation: Multimodal translation is formulated as learning translation and visually grounded representations with shared encoder parameters and task-specific decoders.The translation decoder minimizes translation loss, while the image-prediction decoder minimizes image-prediction loss.
  • Problem Formulation: Separate described-image and parallel-text datasets can be used because the decomposition does not require paired image-text-translation tuples.External datasets may contain image descriptions (x, v) or parallel corpora (x, y).
  • Problem Formulation: The joint objective mixes translation and image-prediction losses using parameter w.The training procedure updates the primary task with probability w and the auxiliary task with probability 1−w.

3 Imagination Model

Imagination shares a bidirectional encoder between attention-based translation and image prediction tasks. The image-prediction objective trains source-language representations to be visually grounded.

  • 3.1 Shared Encoder: The shared encoder processes source sentences with forward and backward recurrent networks, concatenating both hidden states into each token representation.
  • 3.2 Neural Machine Translation Decoder: The attention-based decoder conditions each prediction on the previous token, decoder state, and an attention-weighted context over encoder states.
  • 3.2 Neural Machine Translation Decoder: The translation objective minimizes the negative log likelihood of the target-language sequence.
  • 3.3 Imaginet Decoder: The image decoder averages encoder states and feeds that sentence representation to a network predicting the associated image feature vector.
  • 3.3 Imaginet Decoder: The image prediction decoder uses a margin-based objective comparing the true image vector with contrastive vectors sampled from the minibatch.

4 Data

The experiments use Multi30K as the primary multimodal translation benchmark and evaluate external image and parallel-text resources. The study excludes Multi30K's independently collected descriptions.

  • 4 Data: Multi30K contains 31,014 image–English sentence–German translation pairs, split into 29,000 training, 1,014 development, and 1,000 evaluation instances.
  • 4 Data: English and German text is normalized, lowercased, tokenized, and German text is decompounded before training.
  • 4 Data: The experiments additionally use MS COCO English image descriptions and the English–German News Commentary parallel corpus.
  • 4 Data: The study does not use Multi30K’s 155K independently collected German and English descriptions.

5 Experiments

Experiments evaluate Imagination on Multi30K with in-domain and external image and parallel-text resources, using translation quality and ensemble comparisons. The results show competitive in-domain performance, continued gains with external resources, and a 59.3 Meteor ensemble result.

  • In-domain experiments: Multitasking improves the text-only NMT baseline by 1.8 Meteor points, although it remains 1.1 points below Moses.The model is competitive with prior approaches using visual features and with Toyama et al., which uses images only during training.
  • External described image data: Training image prediction on COCO produces no significant difference from training it on in-domain Multi30K images.This supports separating parallel-text data from described-image data sources.
  • External parallel text data: Adding News Commentary improves the text-only baseline by 2.7 Meteor points, while adding COCO to the multitask setup yields a further 0.4-point gain.The balanced multitask setting improves over the concatenated-text model, whereas combining in-domain parallel text with image prediction does not add an improvement.
  • Ensemble results: 59.3 Meteor is achieved by the best ensemble trained on Multi30K, News Commentary, and COCO data.The ensemble uses both external parallel text and described images.
  • Qualitative Examples: Qualitative examples show that multitasking corrects some pose, lexical, and grammatical errors but can also produce worse translations.Both systems miss the low-frequency word “dangling” in one example.

6 Discussion

The discussion examines whether the model learns visually grounded representations and how visual feature choices affect multitask translation. Qualitative examples show both improvements and remaining errors, while feature type strongly affects performance.

  • 6.1 Does the model learn grounded representations?: The model correctly translates “on their stomachs” as “bäuchlings” and preserves the children’s location under a swing, unlike the NMT baseline.
  • 6.1 Does the model learn grounded representations?: The model corrects the baseline’s dog description by translating the costume, dangling flowers, and purpose of reaching them.
  • 6.1 Does the model learn grounded representations?: The model changes “across the water” to “through the water,” illustrating that qualitative improvements are not uniform across examples.
  • 6.1 Does the model learn grounded representations?: Table 6 compares translations where the multitask model improves or worsens over the NMT baseline, including corrections to body-part, verb, and noun translations.The examples also show shared errors, such as omitting “pipe,” and a case where the multitask model mistranslates a preposition.
  • 6.2 The effect of visual feature vectors: The type of visual feature predicted by the IMAGINET decoder has a strong impact on multitask model performance.The study compares VGG-19, InceptionNet V3, and ResNet-50 feature representations.
  • 6.2 The effect of visual feature vectors: Image-prediction models with fewer parameters are suggested to be easier to learn, but the pronounced Inception-V3 versus ResNet-50 difference remains unexplained.The parameter counts are 8.192 million for VGG-19 versus 4.096 million for Inception-V3 and ResNet-50.

7 Related work

Related work covers image features used in multimodal translation, visual representation prediction, and multitask learning. The paper differs from prior approaches by learning grounded source representations while allowing separate source-image and source-target resources.

  • Multimodal translation: Earlier multimodal translation systems used semantic or spatially preserving image features as inputs to translation models.Semantic features typically come from a CNN’s final pooling layer, while spatial features preserve positional information deeper within the network.
  • Multimodal translation: Prior work incorporated image features into encoders, decoders, or phrase-based translation models, whereas this paper learns grounded representations through multitask learning.
  • Related multimodal translation: Related multimodal translation approaches model latent semantics, correlate image-language representations, or learn joint multimodal spaces for zero-resource translation.
  • Related multimodal translation: Unlike approaches assuming source-target-image alignment, this work can use separate source-image and source-target datasets.
  • Multitask learning: The paper’s multitask framework shares grounded learning with translation while focusing on enriching source-language representations rather than solving zero-resource translation.
  • Visual representation prediction: Visual representation prediction has been studied through unsupervised reconstruction or future prediction and supervised scene-imagination and image-vector prediction tasks.

8 Conclusion

The conclusion frames multimodal translation as learning both translation and visually grounded representations through a shared encoder. The approach reaches state-of-the-art Multi30K results and remains effective with separate external text and image resources.

  • 8 Conclusion: The paper decomposes multimodal translation into learning to translate and learning visually grounded representations.
  • 8 Conclusion: A shared encoder connects an attention-based translation model with an image prediction model in the multitask framework.
  • 8 Conclusion: The approach achieves state-of-the-art results on Multi30K without using images during translation.
  • 8 Conclusion: Training on separate parallel-text and described-image datasets does not hurt performance, and image prediction still improves a text-only baseline strengthened with out-of-domain parallel text.
Loading 1705.04350v2…