Source-linked AI summary
Sign Language Recognition, Generation, and Translation: An Interdisciplinary Perspective
Danielle Bragg, Oscar Koller, Mary Bellard, Larwan Berke, Patrick Boudrealt, Annelies Braffort, Naomi Caselli, Matt Huenerfauth, Hernisa Kacorri, Tessa Verhoef, Christian Vogler, Meredith Ringel Morris
TL;DR
Sign language processing faces communication barriers and fragmented research across disciplinary silos. This paper synthesizes findings from a 39-expert interdisciplinary workshop to review the field, identify challenges, and formulate calls to action for future research.
Problem
Communication barriers affect deaf sign language users, while sign language processing research remains divided across disciplinary and pipeline silos.
Method
The paper synthesizes findings from an interdisciplinary workshop with 39 domain experts from diverse backgrounds.
Results
The workshop produced an interdisciplinary overview, background on Deaf culture and sign language linguistics, a state-of-the-art review, pressing challenges, and calls to action.
Takeaways & Limitations
The paper orients readers across and beyond computer science, highlights opportunities for interdisciplinary collaboration, and helps prioritize future problems, especially data challenges.
Takeaways & Limitations
Current avatar-generation pipelines are not fully automated and require human intervention to produce smooth, coherent signing avatars.
Abstract
from arXiv · showhide
Developing successful sign language recognition, generation, and translation systems requires expertise in a wide range of fields, including computer vision, computer graphics, natural language processing, human-computer interaction, linguistics, and Deaf culture. Despite the need for deep interdisciplinary knowledge, existing research occurs in separate disciplinary silos, and tackles separate portions of the sign language processing pipeline. This leads to three key questions: 1) What does an interdisciplinary view of the current landscape reveal? 2) What are the biggest challenges facing the field? and 3) What are the calls to action for people working in the field? To help answer these questions, we brought together a diverse group of experts for a two-day workshop. This paper presents the results of that interdisciplinary workshop, providing key background that is often overlooked by computer scientists, a review of the state-of-the-art, a set of pressing challenges, and a call to action for the research community.
INTRODUCTION
Sign language processing has substantial potential impact, but research remains fragmented across disciplinary silos and often fails to address sign languages comprehensively. The paper uses an interdisciplinary workshop to synthesize the field, identify challenges, and prioritize action.
- Sign language processing spans recognition, translation, and generation for more than 300 sign languages used by approximately 70 million deaf people.
- Communication technologies largely exclude sign languages, creating barriers for deaf sign language users.
- Recognition, translation, transcription, search, interpreting, and educational tools could expand access to voice-activated and text-based systems.
- Research in disciplinary silos often lacks Deaf expertise, linguistic knowledge, and datasets reflecting real-world signing.
- Successful systems require Deaf studies, linguistics, NLP and machine translation, computer vision, computer graphics, HCI, and design.
- A workshop with 39 participants synthesized domain knowledge, reviewed the state of the art, identified challenges, and formulated calls to action.
BACKGROUND AND RELATED WORK
Sign languages are natural, culturally central languages with complex linguistic structure and substantial variation across users and contexts. Building useful technology therefore requires attention to Deaf culture, language form, embodied expression, and diversity.
- Deaf Culture: Deaf culture treats sign languages as central to community identity, making technology development sensitive to cultural and linguistic representation.
- Deaf Culture: Historical suppression of sign language communication contributes to the sensitivity surrounding technology development for Deaf communities.
- Sign Language Linguistics: Signs use phonological features such as handshape, location, and movement under language-specific rules.
- Sign Language Linguistics: Classifiers combine handshapes, movements, and locations to express classes of entities and nuanced actions.
- Sign Language Linguistics: Fingerspelling spells words through letter handshapes, whose execution varies through coarticulation and must be distinguished from other functions.
- Sign Language Linguistics: Meaning can depend on eyebrows, mouth, head, shoulders, eye gaze, and bodily depiction, not only manual signs.
- Sign Language Linguistics: Execution varies by ethnicity, region, age, gender, education, proficiency, and hearing status, increasing training-data requirements.
Reviews
Earlier reviews are commonly technical, outdated, or limited to one subfield, with insufficient integration of linguistic, social, design, and user perspectives. This paper instead uses an interdisciplinary workshop to connect domains and identify shared priorities.
- Existing reviews often focus on specific technical subareas and predate deep learning, while few cover recognition, translation, and generation together.
- Prior reviews include limited linguistic, social, and design perspectives needed for sign language systems with real-world use.
- Gesture-recognition reviews risk framing sign language recognition as generic gesture recognition, overlooking linguistic complexity and social context.
- A two-day workshop convened experts in sign language processing and related fields to synthesize the workshop findings.
- The 39 attendees represented universities, schools, a technology company, and disciplines including computer science, linguistics, education, psychology, and Deaf studies.
Procedure
The workshop combined interdisciplinary background-sharing with structured problem-solving and planning. Participants learned from domain experts, incorporated Deaf users’ perspectives, and worked in topic-focused groups to map challenges and actions.
- Day 1 provided interdisciplinary domain knowledge, while Day 2 addressed the field’s landscape, challenges, and path forward.
- Domain lectures covered Deaf culture, sign language linguistics, NLP, computer vision, computer graphics, and dataset curation.
- A Deaf-moderated panel discussed Deaf users’ experiences, needs, and concerns about technology.
- Participants formed groups of 8–9 around datasets, recognition and vision, modeling and NLP, avatars and graphics, and UI/UX design.
- Each group considered the state of the art, challenges, solutions, future vision, and a community call to action.
- Breakout groups reported their topics through presentations and discussion with the larger workshop group.
Q1: WHAT IS THE CURRENT LANDSCAPE?
The landscape spans diverse sign-language datasets, recognition approaches, and continuous-signing benchmarks, with data central to progress. Realistic recognition remains difficult, especially across signers.
- Datasets: Sign-language processing research draws on video, motion capture, and depth-camera datasets, with collection methods shaping content and signer identity.Public corpora vary in recording format, signer provenance, geographic coverage, and vocabulary size.
- Datasets: Annotations can mark sign components, sign identities and order, or translations, but vary in format and temporal granularity.Producing annotations is time-intensive and expensive, and sign languages lack a standard written form.
- Recognition & Computer Vision: Non-intrusive vision-based recognition dominates current work because it reduces signer inconvenience and can incorporate non-manual signing.Depth cameras, multiple cameras, triangulation, and machine learning address aspects of the three-dimensional recognition problem.
- Recognition & Computer Vision: Continuous recognition is more challenging and realistic than isolated-sign recognition because of epenthesis, co-articulation, and spontaneous production.These effects complicate deciphering a continuous stream of signing.
- Recognition & Computer Vision: 42.8% letter accuracy is achieved on a recently released real-life fingerspelling dataset.On a real-life continuous-sign benchmark with utterance- or sentence-level segmentation, WER is 22.9% for the same signers and 39.6% for different signers.
Modeling & Natural Language Processing
Modeling and NLP for sign languages use annotations, symbolic representations, and neural methods, while output systems include avatars and interface integrations. Current avatar generation still requires human intervention.
- Modeling & Natural Language Processing: Because sign languages are minority languages lacking data, MT and NLP research has focused mainly on spoken and written languages.Recognition identifies signs from complex signals, whereas MT and NLP generally process already identified language using annotated inputs.
- Modeling & Natural Language Processing: Computational modeling uses gloss, SignWriting, si5s, HamNoSys, and other notation systems to represent sign languages.These systems differ in whether they are intended for human use or computational representation.
- Modeling & Natural Language Processing: Translation systems either use predefined intermediary representations compatible with grammatical rules or learn internal representations with neural methods.The latter representations may not be human-understandable.
- Avatars & Computer Graphics: Sign-language avatars can improve information access for DHH individuals who prefer signing or have lower literacy in written language.Avatars are especially suitable when automatically generated content must remain editable, scalable, or frequently changing.
- Avatars & Computer Graphics: Avatar pipelines assemble sign animations and non-manual signals from symbolic plans, lexicons, key frames, symbolic subsigns, or motion capture.The state of the art is not fully automated: every pipeline stage requires human intervention for smooth, coherent signing.
- UI/UX Design: Interactive recognition systems have mainly targeted applications that remain useful despite limited accuracy and coverage.Few systems robustly recognize full ASL phrases for real-world deployment or use.
Q2: WHAT ARE THE FIELD’S BIGGEST CHALLENGES?
The field’s biggest challenges involve inadequate and nonrepresentative data, sign-language-specific structure, depiction, generalization, and acceptable real-world system behavior. These constraints affect recognition, NLP, translation, graphics, and interfaces.
- Datasets: Public sign-language datasets limit the power and generalizability of systems trained on them.Key gaps include dataset size, continuous signing, native signers, signer variety, signer-independent evaluation, and coverage beyond ASL.
- Depiction: Depiction challenges recognition and translation because it requires Deaf-culture and linguistic knowledge, lacks speech-style modeling support, and is difficult to annotate.Multiple depictions can express the same concept, while annotation systems lack a standard way to encode that richness.
- Annotations: Sign-language annotations are time-consuming and error-prone, lack standardized systems and granularity, and impede combining datasets for supervised training.Annotators require extensive training, and low inter-annotator agreement remains a problem.
- Generalization: Larger, more diverse datasets are essential for generalization to unseen situations and individuals, but generating them is extremely time-consuming and expensive.Signer-independent datasets enable assessment by training and testing on different signers.
- Modeling & Natural Language Processing: MT and NLP methods cannot be straightforwardly transferred because sign languages differ structurally from spoken and written languages.Sign languages may convey multiple channels simultaneously, and spatial context can directly determine interpretation.
- Avatars & Computer Graphics: Avatar systems must avoid an uncanny valley while producing meaningful non-manual cues, varied facial expressions, and smooth transitions between signs.Creating avatars acceptable to Deaf users remains a technical challenge.
- Recognition & Computer Vision: Few sign-recognition systems are robust enough for real-world deployment, motivating near-term applications matched to current capabilities.Examples include finite-sign or phrase interactions such as meal ordering and personal-assistant commands.
Q3: WHAT ARE THE CALLS TO ACTION?
The workshop calls for an interdisciplinary research agenda spanning the complete sign-language-processing pipeline. It presents these calls as previously unarticulated and largely disregarded by the community.
- Q3: WHAT ARE THE CALLS TO ACTION?: Researchers should pursue an interdisciplinary call to action across any part of the end-to-end sign-language-processing pipeline.The workshop formulated these calls because they had not previously been articulated and were largely disregarded.
Deaf Involvement
Deaf involvement and leadership are presented as essential throughout sign language processing research and development. This involvement helps align systems with user needs, respect Deaf ownership, and support adoption.
- Deaf Involvement: Deaf community involvement is essential for designing usable systems that match users’ needs and contexts.All-hearing teams lack lived experience of Deafness and cannot speak for Deaf needs.
- Deaf Involvement: Technology development must recognize individual and community freedoms because imposed systems can provoke resentment or rejection.Disrespecting Deaf ownership of sign languages can further histories of audism and exclusion.
- Deaf Involvement: Deaf contributors should participate throughout research and development, including dataset creation, evaluation, and ownership.This helps produce representative data, address meaningful problems, and avoid cultural appropriation.
- Deaf Involvement: Computing constraints may lead signers to simplify vocabulary, restrict movement, or reduce linguistic richness to accommodate technology.These possible adaptations heighten the importance of Deaf involvement in technology design.
- Deaf Involvement: Deaf involvement and leadership are crucial for useful systems, respect for language ownership, and technology adoption.The paper summarizes this priority as involving Deaf team members throughout the process.
Application Domain
Sign language processing should target concrete real-world applications while accounting for current technical limits. Progress also depends on representative public datasets, user-interface guidance, reproducible evaluation, and careful data-collection practices.
- Application Domain: Sign language processing should focus on specific real-world domains whose requirements differ across settings.Potential applications include interpreter-free interactions, personal assistants, and everyday services.
- Application Domain: Technical limitations should shape near-term application choices and intermediary goals that also support longer-term end-to-end systems.Dictionaries and sign-language writing are examples of useful intermediate resources.
- Interface Design: The field lacks systematic user-interaction research, leaving teams to design interfaces largely from scratch.Wizard-of-Oz studies can investigate interface reactions and interaction designs before end-to-end technologies mature.
- Interface Design: User-interface guidelines and error metrics would support consistently effective sign-language systems.The field currently lacks a systematic understanding of how people interact with these technologies.
- Application Domain: Large, diverse, publicly available datasets are needed because existing corpora limit system power and generalizability.Public data supports development, competitive evaluation, and equal Deaf-community ownership.
- Data Collection: Data-collection approaches trade off cost, naturalism, quality control, privacy, scalability, and participant diversity.Scraping and controlled in-lab collection illustrate contrasting benefits and risks.
- Data Collection: Collection should record signer demographics and reproducible process metadata to assess bias, generalizability, and geographic diversity.Relevant demographics include fluency, acquisition age, education, audiological status, socioeconomic status, gender, race or ethnicity, and geography.
- Calls to Action: The field should create larger, more representative public video datasets and develop user-interface guidelines for sign-language systems.These priorities are stated as explicit calls to action.
Annotations
Annotations are a central bottleneck because they are costly, error-prone, and inconsistent across datasets. Standardization and software support could improve sharing, compatibility, quality, and training efficiency.
- Annotations: A standard annotation system would make datasets easier to combine and share while expanding compatible training data.Standardization could also reduce annotation cost and errors.
- Annotations: Current annotation practices require expensive annotator training and produce ambiguous, error-prone labels.There is no standardized system or annotation granularity, and inter-annotator agreement can be low.
- Annotations: A writing or reading-oriented sign-language standard could make everyday digital tools usable without translation into spoken or written language.User-generated writing would also create an annotated corpus for modeling language structure.
- Annotations: Computer-aided annotation could use sign-language models to assist with video segmentation and transcription.Support should leverage models of whole signs and sign subunits.
- Annotations: The paper calls for standardized annotations and annotation-support software to improve data sharing, compatibility, quality, accuracy, reliability, and cost.Annotations support recognition, NLP, machine translation, and signing-avatar generation.
CONTRIBUTIONS
The paper offers an interdisciplinary account of sign language processing, synthesizing its landscape, challenges, and priorities. It argues that progress requires collaboration across technical, linguistic, design, and Deaf-community perspectives.
- CONTRIBUTIONS: The workshop synthesis gives computer scientists background on Deaf culture and linguistics while orienting readers outside computer science to the field and its technical challenges.The paper is intended for both newcomers and researchers working within particular subdomains.
- CONTRIBUTIONS: Interdisciplinary synthesis relates technical domains to one another and shows that sign language processing depends on all of them.This extends beyond reviews focused on a single discipline.
- CONTRIBUTIONS: Many open problems span datasets, algorithms, interfaces, and annotation systems while also requiring linguistic validity and Deaf-community acceptance.Addressing them will require strong interdisciplinary teams.
- CONTRIBUTIONS: The paper identifies lack of large, annotated, representative, public datasets as arguably the field’s biggest current obstacle.Data collection is constrained by cost, recording requirements, a small contributor pool, and nonstandard annotations.
- CONTRIBUTIONS: The workshop methodology is presented as a potentially reusable model for other fields organized into disciplinary silos.Adapting it would require identifying relevant domains and experts.
- CONTRIBUTIONS: The paper provides an interdisciplinary overview of sign language recognition, generation, and translation in response to fragmented disciplinary research.It addresses the current landscape, major challenges, and calls to action.
- CONTRIBUTIONS: The workshop brought together 39 diverse domain experts to synthesize the state of the art, challenges, and research priorities.Its findings provide the paper’s interdisciplinary foundation and call to action.