Source-linked AI summary

Including Signed Languages in Natural Language Processing

Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani

arXiv:2105.05222v2cs.CLcs.AIcs.LG

TL;DR

Most NLP models exclude signed languages, despite their linguistic completeness and importance to Deaf communities, while current SLP systems often fail to model their linguistic structure. This position paper reviews signed-language properties and SLP limitations, then advocates standardized tokenization, linguistically informed models, representative real-world data, and Deaf-community leadership. Its evidence shows that models successful on restricted data can fail on realistic open-domain data, underscoring the need for better data and more suitable approaches.

  • Problem

    Most NLP models require speech or text, excluding signed languages, while current SLP methods do not adequately capture linguistic structure, spatial relations, or simultaneous cues.

  • Method

    The paper synthesizes signed-language linguistic properties, reviews limitations and open challenges in SLP, and proposes a roadmap for NLP-oriented modeling, data collection, and community collaboration.

  • Results

    Current models can achieve 22.17 BLEU on RWTH-PHOENIX-Weather 2014T but only 3.2 BLEU on the Public DGS Corpus, showing failure on more realistic open-domain data.

  • Takeaways & Limitations

    Signed-language NLP should use linguistically informed representations and models, real-world data, and active Deaf-community collaboration.

  • Takeaways & Limitations

    Linear glosses lack a single agreed standard and lose simultaneous, non-manual, and spatial information, potentially affecting downstream processing.

Abstract

from arXiv · show

Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of Natural Language Processing (NLP) are crucial towards its modeling. However, existing research in Sign Language Processing (SLP) seldom attempt to explore and leverage the linguistic organization of signed languages. This position paper calls on the NLP community to include signed languages as a research area with high social and scientific impact. We first discuss the linguistic properties of signed languages to consider during their modeling. Then, we review the limitations of current SLP models and identify the open challenges to extend NLP to signed languages. Finally, we urge (1) the adoption of an efficient tokenization method; (2) the development of linguistically-informed models; (3) the collection of real-world signed language data; (4) the inclusion of local signed language communities as an active and leading voice in the direction of research.

1 Introduction

Signed languages remain largely excluded from language technologies despite serving Deaf communities and exhibiting the fundamental properties of natural languages. The paper argues for linguistically informed SLP and calls for standardized representations, suitable models and data, and sustained Deaf-community collaboration.

  • Most NLP models require speech or text, excluding around 200 signed languages and up to 70 million deaf people from modern language technologies.
  • Signed languages are fully fledged natural-language systems, and their exclusion from technology disregards Deaf communities’ preference for signed communication.
  • Existing SLP research has largely focused on visual processing with limited NLP involvement, while current techniques fail to leverage signed languages’ linguistic structure.
  • Signed languages pose NLP challenges through visual-gestural modality, simultaneity, spatial coherence, and lack of written form.
  • Potential applications include endangered-language documentation, educational tools, video retrieval, signed-language assistants, and real-time interpretation.
  • The paper urges standardized tokenization, linguistically informed models, representative real-world data, and collaboration with Deaf communities throughout research.

2 Background and Related Work

Signed languages are central to Deaf communities and have historically been marginalized despite being natural languages. Recent SLP work has improved through vision-based approaches, but the paper argues that NLP must address linguistic modeling challenges more directly.

  • Historical policies and educational practices favored speech over signed languages, including the 1880 ban on teaching signed languages at an international deaf-education conference.
  • Signed languages are primary communication languages for Deaf communities, and access to technologies aligned with their lived experience remains an NLP responsibility.
  • The paper distinguishes capitalized “Deaf,” referring to a language-and-culture community, from lowercase “deaf,” referring to an audiological condition.
  • Early SLP research relied mainly on sensors, fingerspelling, isolated signs, or rule-based synthesis because video-processing technology was inadequate.
  • Signed languages remained relatively overlooked in NLP, motivating an interdisciplinary roadmap focused on linguistic modeling challenges and open questions.

3 Sign Language Lingusitics

Signed languages use visual-gestural structure to encode linguistic and social meaning through manual features, non-manual cues, simultaneity, spatial reference, and language-contact processes such as fingerspelling.

  • Signed languages contain phonological, morphological, syntactic, and semantic structure serving the same social, cognitive, and communicative purposes as other natural languages.
  • Their visual-gestural modality uses the face, hands, body, and surrounding space to create meaning distinctions.
  • Phonology: Signs combine manual features such as hand configuration, orientation, placement, contact, and movement with non-manual features including eye aperture, head movement, and torso positioning.
  • Simultaneity: Simultaneity lets signed languages convey different information through multiple visual cues at the same time.
  • Referencing: Signers establish discourse referents through pointing, assigned regions in signing space, directional signs, body shift, and eye gaze.
  • Referencing: Classifiers describe referents’ characteristics, movements, and relations to other entities, while role shift marks represented speakers through spatial and embodied changes.
  • Fingerspelling: Fingerspelling links manual gestures to a surrounding spoken language’s written orthography or phonetic system and is used for names, places, and new concepts.

4 Current State of SLP

Current SLP represents signed languages through videos, poses, notations, glosses, and varied corpora, but these resources and tasks often lose linguistic structure or lack sufficient real-world data. The paper identifies representation, resource availability, evaluation, and linguistically informed modeling as central limitations across SLP.

  • Representations: Videos preserve rich signing information but are high-dimensional, costly to store and transmit, and difficult to anonymize because facial features are essential.These constraints limit public distribution of raw video data.
  • Representations: Pose representations reduce visual complexity and preserve relatively little information, but remain continuous, multidimensional, and poorly adapted to most NLP models.Pose estimation from video is currently preferred over expensive and intrusive motion-capture equipment.
  • Representations: No written notation system is widely adopted, while linear glosses omit simultaneous cues and spatial relations, causing information loss that can affect downstream SLP performance.Glosses transcribe signs individually, but they do not adequately represent body posture, eye gaze, or spatial relations.
  • Resources: Available resources range from bilingual dictionaries and fingerspelling or isolated-sign corpora to continuous corpora, but many lack contextual grammar, coarticulation, vocabulary breadth, or sufficient scale.Continuous corpora contain 4-6 orders of magnitude fewer sentence pairs than comparable spoken-language translation corpora; the largest contains 1,150 hours, only 50 publicly available.
  • Resources: Signed-language resources are scarce, frequently restricted, and difficult to anonymize without losing facial and other physical features, limiting open distribution.Developing accurate anonymous representations with minimal information loss remains a promising research problem.
  • Tasks and methods: Current SLP tasks often rely on low-level visual features or loosely mapped units and lack data or methods for linguistic phenomena such as lexical structure, sentence boundaries, and spatially grounded reference.Translation methods use glosses or pose-based representations but do not handle spatial relations and discourse grounding for ambiguous referents.
  • Tasks and methods: Sign-language production commonly uses poses as an intermediate representation, but back-translation quality does not accurately evaluate sign quality or real-world usability.The paper calls for better evaluation grounded in how distinctions in meaning are created in signed language.

5 Towards Including Signed Languages in Natural Language Processing

The paper identifies linguistic, data, modeling, and community challenges that limit signed language processing, and proposes NLP-informed, multimodal approaches grounded in real-world data and Deaf collaboration.

  • Research direction: Current SLP models often fail to explore or leverage signed languages’ linguistic structure, motivating collaboration between NLP, computer vision, signing communities, and sign linguists.The paper links model limitations to insufficient attention to signed-language linguistics and lived signer experience.
  • Building NLP pipelines: An efficient, universal, standardized tokenization method is needed because glosses miss spatial constructions, require language-specific models, and lack cross-corpus standards.The paper frames tokenization as a foundation for downstream NLP applications and raises questions about lexical units, articulators, and cross-linguistic phonology.
  • Building NLP pipelines: Signed-language NLP also requires syntactic analysis, automated named-entity recognition, and multimodal models that relate visual-gestural and linguistic information.Open questions include whether spoken-language tags generalize, how visual markers introduce entities, and how translation, alignment, and co-learning should connect modalities.
  • Collect real-world data: Deployable SLP models require real-world datasets covering broad domains, natural signing, diverse and native signers, realistic conditions, sufficient vocabulary, and dense annotations when applicable.The paper presents these as criteria for datasets that accurately represent signed-language complexity and diversity.
  • Collect real-world data: Real-world data collection is costly because annotation can take up to 600 minutes per minute of signed language video and requires scarce expertise.The paper proposes automating parts of parsing, boundary detection, articulatory-feature extraction, collection, and annotation to reduce this bottleneck.
  • Practice Deaf collaboration: Intrusive methods and oversimplified claims about sign-language translation can be rejected by signing communities, so Deaf collaboration and leadership should guide research and evaluation.The paper recommends long-term collaboration so users can identify meaningful challenges, review research, and help ensure tools address community needs.

6 Conclusions

The paper concludes that signed languages should be included in NLP, combining spoken-language methods with computer vision and linguistic insight. It also calls for more resources, tools, and sustained collaboration with signing communities.

  • The authors urge the NLP community to include signed languages as a research area.
  • NLP is positioned to contribute linguistic insight by combining successful spoken-language methods with recent computer vision tools for video.
  • The paper calls for increased efforts to collect signed-language resources, develop signed-language tools, and build strong collaboration with signing communities.
Loading 2105.05222v2…