Source-linked AI summary
A Survey of Code-switched Speech and Language Processing
Sunayana Sitaram, Khyathi Raghavi Chandu, Sai Krishna Rallabandi, Alan W Black
TL;DR
Code-switched language processing still lacks full end-to-end systems for interacting with multilingual humans. This survey comprehensively reviews speech and NLP techniques, datasets, progress, and open problems, reporting task-specific findings across the field.
Problem
Full end-to-end digital systems that interact with multilingual humans in code-switched language are not yet available.
Method
The survey provides a comprehensive description of code-switched Speech and NLP work, discusses open directions, and lists available code-switched datasets.
Results
The survey synthesizes progress across the field and reports task-specific findings, including challenging NLI and high perplexity reductions over the SEAME corpus.
Takeaways & Limitations
The field has made progress, but full end-to-end interaction in code-switched language remains an open problem.
Takeaways & Limitations
Most models of code-switched data generation remain an unresolved challenge.
Abstract
from arXiv · showhide
Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched Speech and Natural Language Processing. We motivate why processing code-switched text and speech is essential for building intelligent agents and systems that interact with users in multilingual communities. As code-switching data and resources are scarce, we list what is available in various code-switched language pairs with the language processing tasks they can be used for. We review code-switching research in various Speech and NLP applications, including language processing tools and end-to-end systems. We conclude with future directions and open problems in the field.
1. Introduction
Code-switching is a widespread, structured form of multilingual communication that appears across languages, dialects, registers, and settings. The survey argues that speech and language technologies must account for it while addressing scarce data, linguistic variability, and modeling complexity.
- What is code-switching?: Code-switching shifts between languages within communication and can include borrowed words, fillers, phrases, morphological mixing, and grammatical mixing.The paper uses code-switching and code-mixing interchangeably, while noting that their distinction may matter in some settings.
- What is code-switching?: Code-switching conveys group identity, societal patterning, and cultural discourse strategies, occurring in formal, informal, and semi-formal settings.Examples include social media, newspaper headlines, teaching, Spanglish, Hinglish, and dialect switching.
- Why process code-switched language?: Consumer-facing technologies need code-switched speech and text processing because ignoring one language can lead to incorrect conclusions about user sentiment.The paper connects this need to healthcare, education, entertainment, advertising, and human-machine communication.
- Why process code-switched language?: Code-switched processing is difficult because languages interact through cross-lingual transfer, lexical borrowing, speech errors, and complex grammatical structure.Tasks such as semantic role labeling require complex cross-lingual analysis, while code-switched varieties remain dynamic and diverse across speakers.
- Challenges: Code-switched data generation cannot choose languages randomly because underlying grammars and constraints make some switches unacceptable.The paper notes that code-switched varieties are dynamic, diverse, and difficult to model, although machine-learning methods can handle uncertainty.
- Survey scope: The survey reviews linguistic studies, corpora, computational techniques, shared tasks, benchmarks, remaining challenges, and future directions for code-switched speech and NLP.Its stated aim is to describe progress in the field and discuss open problems.
2. Background
Research on code-switching spans linguistic definitions, grammatical theories, computational modeling, sociolinguistic triggers, and quantitative measures. The background emphasizes that code-switching varies across language pairs and modalities, while its patterns remain constrained and measurable.
- Definitions and variation: The survey distinguishes code-switching as juxtaposing grammatical systems from code-mixing as embedding units from one language into another, while treating the terms interchangeably.The distinction between switching, mixing, and borrowing is often viewed as a continuum.
- Definitions and variation: Code-switching varies across language pairs, media, and communicative contexts, with Twitter studies finding about 3.5% of tweets were code-switched.English-Turkish tweets showed more switch points, whereas English-German tweets typically contained one switch point.
- Linguistic models: Linguistic models propose constraints on switch points, including the Free Morpheme Constraint and Equivalence Constraint, alongside constituent and grammatical relationships.Other work examines harmonization, neutralization, compromise, blocking, and variability among bilingual speakers.
- Computational approaches: Computational studies use linguistic theories to generate synthetic code-switched text and model acceptability, which depends on sociolinguistic and cognitive factors.These approaches address issues including missing literal translation pairs, alignment errors, and underspecified original models.
- Sociolinguistic factors: Empirical work links code-switching with cognate precedence, part-of-speech information, and convergence in conversational entrainment.Other studies analyze pragmatic triggers, identity, social familiarity, hierarchy, modality, discourse, and granularity.
- Measuring code-switching: Code-switching is measured with indices and information-theoretic metrics covering language proportions, switching probability, span distributions, burstiness, and memory.Examples include the Code-mixing Index, M-Index, Integration Index, Language Entropy, Span Entropy, Burstiness, and Memory.
3. Data and resources
Code-switched speech and text resources remain scarce, especially compared with the data requirements of modern speech and NLP systems. The survey catalogs available corpora across language pairs, modalities, and tasks.
- DNN-based speech and NLP systems typically require large labeled corpora, but most languages lack the necessary data.
- Code-switched resources are particularly limited because monolingual resources often exclude foreign words and code-switched speech and language data remain scarce.
- The survey lists available code-switched datasets separately by task, covering speech, text, language identification, and multiple language pairs.
- Speech data: Speech resources include corpora such as SEAME, HKUST, CECOS, BANGOR-MIAMI, and databases for Hindi-English, Malay-English, Arabic-English, and other pairs.
- Speech data: Available speech datasets vary substantially in size and design, from 3 minutes of Hindi-English speech to 1000 hours of Malay-English speech.
- Speech data: Speech synthesis lacks code-switched databases, although bilingual TTS databases from the same speakers are available for some Indian languages and English.
4. Code-switched Speech and NLP Techniques
The survey reviews code-switched speech and NLP techniques shaped by the availability of monolingual, bilingual, and code-switched data. Approaches include transfer, multilingual modeling, synthetic data, language identification, phone-set design, and data augmentation.
- When code-switched data are unavailable, systems can use monolingual resources from the two languages with domain adaptation or transfer learning.
- Code-switched resources are often unavailable in practice, motivating synthetic data generation for training code-switched embeddings and models.
- Automatic Speech Recognition: ASR approaches range from language-boundary detection with monolingual decoding to parallel recognizers, rescoring, and single-pass systems with soft language decisions.
- Automatic Speech Recognition: Phone-set design combines linguistic knowledge, IPA mappings, acoustic similarity, manual merging, or extensions from one language to another.
- Automatic Speech Recognition: The clustering approach for Mandarin-English ASR is comparable to combining phone inventories and outperforms IPA-based mapping.
- Automatic Speech Recognition: Data augmentation uses synthetic code-mixed speech, semi-supervised transcription, active learning, and selected monolingual data to improve code-switched ASR.
4.2. Language Modeling
Code-switched language modeling addresses limited and mismatched data using monolingual resources, grammatical constraints, synthetic text, multitask information, and bilingual modeling. Reported studies show improvements from curriculum training, data augmentation, and cross-lingual representations.
- Code-switched language models face scarce data, while Internet text may not follow the patterns of code-switched speech.
- Researchers use monolingual data, grammatical constraints, artificial code-switched text, and small real-data samples to build language models.
- An RNN language-model curriculum that trains first on interleaved monolingual data and then on code-switched data gives the best English-Spanish results.
- RNN and n-gram combinations or backoff models improve perplexity compared with speech-based code-switched text data.
- Synthesized isiZulu-English bigrams augment language-model training and reduce perplexity on a soap-opera speech corpus.
- A bilingual attention language model achieves high perplexity reductions on the SEAME corpus by learning cross-lingual probabilities from parallel data.
4.3. Code-switching detection from speech
Code-switching detection from speech includes language identification, switch-style classification, and acoustic or lexical modeling. The reviewed work also indicates that prosodic and acoustic cues can support detection and anticipation of switching behavior.
- Humans exploit prosodic cues to detect code-switching and can anticipate switch points even in noisy speech.
- For intra-sentential switching, detecting an utterance’s code-switching style can support adaptation through specialized language models.
- Studies classify code-switched corpora by switching style and report that acoustic features alone can distinguish different kinds of code-switching.
- Detection systems have used retrained multilingual DNNs, lexical information, HMM acoustic models, and SVM classifiers for language mixing.
4.4. Speech Synthesis
Code-switched TTS must handle mixed-language input, which monolingual systems may mispronounce or omit. Surveyed approaches include phone mapping, multilingual or polyglot synthesis, linguistic assimilation, voice conversion, and bilingual end-to-end models.
- Motivation: Monolingual TTS systems can mispronounce or omit words from another language, despite those words often carrying important message content.This motivates systems designed for mixed-language speech input.
- Approaches: Phone mapping substitutes foreign phones with similar primary-language phones, often producing strongly accented speech.Multilingual synthesis instead uses separate monolingual systems for different text portions, while polyglot synthesis trains one system on multilingual-speaker data.
- Approaches: Linguistic assimilation can improve intelligibility over an unmodified non-native monolingual system but retains primary-language accent because phone sets do not correspond exactly.The approach is reported as fairly successful for phonetically similar languages.
- Approaches: Randomized bilingual training data improved subjective metrics for Hindi-English, Tamil-English, and Hindi-Tamil code-switched speech, with marginal degradation in another setting.Other work used shared or language-specific encoders, separate decoders, phonetic posteriorgrams, and multi-head attention for end-to-end synthesis.
- Language identification: Lexical-level language identification uses dictionary lookup, supervised word-level models, sequence labeling, character n-grams, and contextual features.Character-level n-gram features and contextual information are reported as useful features.
- Language identification: Additional code-switched language-identification work models intra-word switching, dialectal variability, POS cues, and weak supervision, with some models outperforming classical baselines.A random forest using POS tags, word length, and the word itself achieved the highest performance in one study.
4.6. Named Entity Recognition
Code-switched NER research addresses noisy social-media text, limited resources, spelling variation, and high out-of-vocabulary rates. Approaches combine sequence models with character, word, contextual, multilingual, and external-resource representations.
- Representations: Embedding studies compare traditional and contextual representations, finding that their combination performs better on Spanish-English and Arabic-English tweets.Other systems use bilingual character representations, spelling normalization, self-attention, and stacked BiLSTMs to address OOV words.
- Methods: Character-level CNNs with word-level Bi-LSTMs model non-standard spelling, while gazetteer lists are important in this low-resource setting.The work frames code-switched NER as a multitask learning problem and highlights external lexical resources.
- Datasets and challenges: Code-switched NER datasets include social-media, conversational-speech, and translated data, including an Arabish benchmark with 6k sentences and 130k tokens.These resources support benchmarking across informal and multilingual settings.
- Representations: Hierarchical meta-embeddings combine word- and sub-word-level representations to achieve state-of-the-art performance.Multilingual meta-embeddings extend representation coverage to related and similar languages.
4.7. POS Tagging
Code-switched POS tagging and related structured prediction research evaluates language-specific, machine-learning, stacked, joint, and pipeline methods. Results emphasize switching-related errors, normalization, multilingual parsing, and resources for downstream applications.
- Methods: Structured prediction for code-switched text includes POS tagging, parsing, and other sequence-labeling tasks using monolingual resources, heuristics, and machine-learning models.Reported models include SVM, Logit Boost, Naive Bayes, J48, CRFs, and Random Forests.
- POS tagging: Intra-sentential switching produces many errors, establishing the complexity of code-switched structured prediction.In one comparison, Random Forests performed best, but only marginally better than combinations of individual language taggers.
- POS tagging: The best stacked model using all features outperformed joint and pipeline-based models.This comparison directly evaluates alternative architectures for code-switched structured prediction.
- POS tagging: Automatic normalization leads to a performance gain in POS tagging.A Facebook dataset was annotated for language identification, normalization, back transliteration, and POS tagging, with joint modeling proposed across tasks.
- Parsing: Multilingual BIST parsing handles code-switched data relatively well, while other parsers use transfer learning, cross-lingual embeddings, and neural transition models.A Hindi-English Universal Dependencies dataset and a neural stacking parser with a new decoding scheme are also reported.
- Question answering: Code-switched question answering research uses translated or crowd-sourced questions, lexical and transliteration features, retrieval models, and end-to-end systems such as WebShodh.Parallel English/code-switched collection can introduce lexical bias through entrainment.
4.10. Sentiment Analysis/stance detection
Sentiment and stance research in code-switched data uses benchmark datasets, n-gram and neural representations, multitask learning, and graphical models. Findings cover strong classical baselines, multilingual modeling, and joint treatment of language and emotion information.
- Model comparison: Comparisons among multilingual, separate monolingual, and language-identified monolingual models demonstrate the effectiveness of multilingual models for code-switched scenarios.The comparison directly evaluates alternative choices for handling mixed-language input.
- Sentiment analysis: A sentiment shared task released about 12k Hindi-English tweets and 2,500 Bengali-English tweets for training.These datasets support comparative evaluation of code-switched sentiment systems.
- Sentiment analysis: The best shared-task system used word- and character-level n-gram features with an SVM classifier.A similar comparison found SVM competitive with Naive Bayes for Bengali-English movie-review sentiment classification.
- Neural methods: Neural approaches model sentiment with sub-word representations, dual encoders, and multitask CNNs for sentence-level and sub-word-level information.The same line of work also applies multitask learning to stance detection on a public issue.
- Emotion detection: Emotion datasets and models represent language and sentiment information jointly using bilingual graphs, label propagation, factor graphs, and belief propagation.These approaches accommodate sentiment expressed in Chinese, English, or mixed-language text.
4.11. Hate Speech Detection
Code-switched language and speech research spans hate-speech detection, inference, translation, dialogue, speech cues, script mixing, OCR, and cross-lingual NLP. Across these applications, studies explore transfer, multilingual modeling, synthetic data, and multimodal features, while results show persistent difficulty and task-specific gains.
- Hate Speech Detection: Hate-speech detection studies use transfer learning, psycho-linguistic features, model averaging, and hierarchical phonemic or word-level models.Approaches include CNN pretraining on hateful tweets followed by fine-tuning on transliterated data.
- Natural Language Inference: Multilingual BERT for code-switched natural language inference performs only slightly better than chance, making NLI especially challenging.The task predicts whether a hypothesis entails or contradicts a premise from Hindi movie conversations.
- Machine Translation: Zero-shot neural machine translation handles code-switched inputs, but performs worse than on monolingual inputs.A separate system translates Hinglish into pure English and Hindi using cross-morphological analysis.
- Speech Processing: Speech studies find that embedded English fragments are spoken more slowly with greater vocal effort and higher pitch variation than monolingual dialogue portions.For speech segments, i-Vector features outperform spectral features, although text features produce the best system.
- Dialogue: Language accommodation across Spanish-English and Hindi-English dialogues can be delayed, with effects shaped by contextual language markedness.Code-switching research also examines lexical and prosodic features in goal-oriented dialogue settings.
- Scripts and Cross-lingual NLP: Script choice in Hindi-English can signal emphasis, disambiguation, or borrowing, while joint OCR and word-level language identification reduces historical-text errors.Other work uses code-switched text to improve cross-lingual representations and tasks such as XNLI.
5. Evaluation of Code-switched Systems
Shared tasks and benchmarks have expanded evaluation across code-switched speech and NLP, but performance varies sharply by task. Multilingual contextual models are strong on some word-level tasks yet remain substantially weaker on harder code-switched tasks than on monolingual or cross-lingual settings.
- Shared Tasks: Shared tasks cover language identification, transliterated search, entity extraction, mixed-script retrieval, POS tagging, NER, sentiment analysis, and question answering.Speech challenges include synthesis, ASR, and spoken language identification for several language settings.
- Model Comparisons: Evaluations indicate that massively multilingual contextual models such as multilingual BERT outperform cross-lingual and task-specific models.The benchmark results also report improvements from adding synthetic code-switched data during pre-training or fine-tuning.
- Task Difficulty: Language Identification and Named Entity Recognition reach high accuracy, whereas sentiment analysis, question answering, and NLI perform much worse.The gap between monolingual and code-switched performance remains large on these harder tasks.
- Task Difficulty: Massive multilingual models do not perform as well on code-switching as on monolingual or even cross-lingual tasks.Synthetic code-switched pre-training or fine-tuning is identified as a promising direction when real data are unavailable.
6. Challenges and Future Directions
Code-switched speech and NLP remain constrained by scarce, unevenly distributed data, limited cross-language generalization, and immature end-to-end systems. Future work emphasizes sociolinguistically informed synthetic data, transfer across language pairs, standardized evaluation, and systems that support multilingual interaction.
- Data Scarcity: Code-switching research is inherently data-starved because many language pairs, especially involving lower-resource languages, have difficult-to-access data.The survey concludes that models must expect to work with limited data.
- Generalization: Most studies focus on one language pair, and grammatical, social, fluency, prestige, and topical differences make general models across pairs difficult.Architectures covering multiple pairs are not yet broadly emerging.
- Data Distribution: Code-switching is more common in informal settings and social media, while formal discourse and archival resources are more often monolingual.Forum, topic, and task distributions therefore affect where code-switched data can be found.
- End-to-end Systems: Large-data end-to-end systems may work in narrowly defined tasks but are not expected to generalize across the full space of code-switching.The survey notes that less task-specific intelligent agents must be more than the sum of individually code-switching-capable components.
- Data Generation: Synthetic code-switched data is promising for exploiting unlabeled data, but current generators mainly rely on syntactic constraints and omit sociolinguistic factors.Incorporating those factors could produce more realistic data and better code-switched speech and NLP models.
- Future Directions: Transfer from resource-rich languages and techniques operating across multiple code-switched language pairs may accelerate development and generalization.The survey also recommends leveraging sociolinguistic work to understand how, when, and why speakers code-switch.
- Evaluation and Theory: Standardized benchmarks spanning speech, NLP tasks, and typologically diverse language pairs are needed for comprehensive evaluation.The survey also leaves open whether code-switching should primarily be treated as translation or as a language in its own right.