Source-linked AI summary

DravidianCodeMix: Sentiment Analysis and Offensive Language Identification Dataset for Dravidian Languages in Code-Mixed Text

Bharathi Raja Chakravarthi, Ruba Priyadharshini, Vigneshwaran Muralidaran, Navya Jose, Shardul Suryawanshi, Elizabeth Sherly, John P. McCrae

arXiv:2106.09460v1cs.CL

TL;DR

Sentiment and offensive-language research lacks large resources for under-resourced, code-mixed Dravidian languages. The paper constructs and manually annotates a multilingual YouTube-comment dataset, then evaluates machine-learning classifiers as benchmarks. The resulting corpus contains more than 60,000 comments, while experiments show sentiment analysis is generally easier than offensive-language detection and that several classes remain difficult.

  • Problem

    Large datasets for Tamil-English, Kannada-English, and Malayalam-English sentiment and offensive-language analysis were unavailable despite growing code-mixed social-media use.

  • Method

    The paper builds a manually annotated YouTube-comment corpus for three Dravidian language pairs and evaluates logistic regression, naive Bayes, decision trees, random forests, and SVMs.

  • Results

    More than 60,000 comments were collected, and sentiment analysis generally outperformed offensive-language detection; logistic regression and random forest were relatively stronger classifiers.

  • Takeaways & Limitations

    The corpus provides a benchmark and resource for further research on sentiment and offensive-language identification in under-resourced Dravidian code-mixed text.

  • Takeaways & Limitations

    Movie-trailer collection skews sentiment toward positive comments, and mixed-feelings and neutral examples are difficult for annotators to label consistently.

Abstract

from arXiv · show

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language identification for a total of more than 60,000 YouTube comments. The dataset consists of around 44,000 comments in Tamil-English, around 7,000 comments in Kannada-English, and around 20,000 comments in Malayalam-English. The data was manually annotated by volunteer annotators and has a high inter-annotator agreement in Krippendorff's alpha. The dataset contains all types of code-mixing phenomena since it comprises user-generated content from a multilingual country. We also present baseline experiments to establish benchmarks on the dataset using machine learning methods. The dataset is available on Github (https://github.com/bharathichezhiyan/DravidianCodeMix-Dataset) and Zenodo (https://zenodo.org/record/4750858\#.YJtw0SYo\_0M).

1 Introduction

Sentiment and offensive-language systems face a resource gap for informal, multilingual, code-mixed social-media text, especially in under-resourced Dravidian languages. The paper addresses this gap by presenting a dataset for Tamil-English, Kannada-English, and Malayalam-English.

  • Most NLP systems are trained on formal, grammatically proper language, creating difficulties for user-generated comments.
  • Code-mixing alternates between languages across units ranging from documents and sentences to words and morphemes.
  • Resources for sentiment analysis and offensive-language identification are primarily monolingual, while most social-media comments are code-mixed.
  • Few datasets cover Tamil, Kannada, and Malayalam code-mixed text, creating a resource bottleneck for NLP support.
  • The paper presents sentiment-analysis and offensive-language datasets for Tamil-English, Kannada-English, and Malayalam-English.
  • The dataset includes all types of code-mixing and provides machine-learning experiments as benchmarks for further research.

2 Related Work

Prior work established the importance of sentiment and offensive-language analysis but provided limited resources for code-mixed Dravidian languages. This paper extends that line of work with a large YouTube-comment corpus and benchmark experiments.

  • Sentiment analysis supports understanding audience polarity and can contribute to recommendation and hate-speech-detection tasks.
  • Automatic moderation research has expanded beyond English, motivating offensive-language resources for additional languages.
  • Social-media interaction in code-mixed native languages has increased, while freely available code-mixed datasets remain limited in size, number, and availability.
  • Earlier Kannada-English work applied sentence embeddings and neural networks, but its dataset was not readily available for research.
  • The paper creates manually annotated YouTube corpora exceeding 60,000 comments across Tamil-English, Kannada-English, and Malayalam-English.

3 Raw Data

The raw data comprises YouTube comments from Tamil, Kannada, and Malayalam film trailers, intentionally preserving diverse real-world code-mixing for annotation.

  • Data source: YouTube comments were collected from Tamil, Kannada, and Malayalam film trailers because movies are popular among speakers of these languages.The comments were gathered in 2019 using a YouTube Comment Scraper.
  • Collection goals: The collection targeted code-mixed comments with representation for each sentiment and offensive-language class.The dataset was assembled for manual annotation in both tasks.
  • Code-mixing coverage: The corpora include monolingual native-language texts, script mixing, word and morphological mixing, and inter- and intra-sentential switching.These forms reflect code-mixing observed in real-world social-media data.
  • Code-mixing coverage: All code-mixing instances were retained to preserve real-world usage.The raw-data process did not filter out mixed-language examples.

4 Methodology of Annotation

The annotation methodology combines bilingual volunteer labeling, task-specific schemas, quality control, and Krippendorff’s alpha to assess agreement across sentiment and offensive-language annotations.

  • Annotation setup: The data were anonymized, and bilingual student volunteers annotated sentiment and offensive language through Google Forms.Volunteers were recruited from institutions in Kerala, Tamil Nadu, and Karnataka.
  • Sentiment analysis: Sentiment annotations used at least three annotators per sentence and labels for positive, negative, mixed, neutral, and not-in-intended-language states.The annotation schema was provided in English and the Dravidian languages.
  • Offensive language identification: Offensive-language annotation used a three-level hierarchical schema with six labels, including untargeted, targeted, not offensive, and not-in-intended-language categories.Targeted categories distinguish individuals, groups, and other targets.
  • Quality control: Annotation quality was controlled with expert-produced gold standards and removal of volunteers whose initial submissions showed poor labeling quality.The process involved 22 sentiment annotators and 23 offensive-language annotators.
  • Inter-annotator agreement: Krippendorff’s alpha was selected because it handles incomplete annotation and differing disagreement levels across nominal and ordinal data.The reported agreement results cover multiple languages and both annotation tasks.

5 Corpus Statistics

The corpus spans sentiment and offensive-language annotations across Tamil, Malayalam, and Kannada, with class distributions varying substantially by language and task. Most comments are positive in sentiment and non-offensive, while offensive comments are usually targeted.

  • Corpus size and structure: The Tamil dataset had the most samples, Kannada the fewest, and comments averaged one sentence in both tasks.Corpus statistics cover words, vocabulary, comments, sentences, and average words per sentence.
  • Sentiment distribution: Positive comments were the largest sentiment class in all three languages, with the strongest imbalance in Tamil.Malayalam had 6,502 neutral-state comments, while Kannada had the fewest neutral comments.
  • Offensive-language distribution: Non-offensive comments formed the majority in every language: 71% in Tamil and 85% in Malayalam.The passage states that Kannada also had a majority non-offensive class but does not provide its percentage here.
  • Offensive-language distribution: Among offensive comments, targeted comments comprised 60% in Tamil, 66% in Malayalam, and 79% in Kannada.Targeting referred to individuals or groups; targeted-other comments were absent or scarce.
  • Data format: The dataset files store each YouTube comment in the first TSV column and its final annotation in the second.The corpus statistics are organized separately for sentiment analysis and offensive-language identification.

6 Difficult Examples

Annotation was difficult because code-mixed, non-standardized, and context-dependent comments often supported multiple sentiment or offensiveness interpretations. Ambiguity arose from comparisons, multiple movie aspects, dialect variation, sarcasm, and implicit social references.

  • Sources of difficulty: Code-mixed comments combine languages and scripts, while non-standard spelling and language-specific meanings complicate annotation by bilingual volunteers.Annotators needed familiarity with multiple scripts, phonological adaptations, and local meanings of English words.
  • Ambiguous sentiment: Movie comparisons could be interpreted as positive, negative, neutral, or mixed depending on the comparison target and surrounding context.Comparisons involved other films, industries, or broader standards, and annotators sometimes disagreed about the intended sentiment.
  • Ambiguous sentiment: Comments mentioning different aspects of a movie in one sentence created further ambiguity about the viewer’s overall sentiment.One example praised a trailer while expressing doubt about the story, yet was labelled positive because the optimism was considered strong.
  • Ambiguous offensiveness: Implicit offensiveness could depend on sarcasm, social context, or whether a seemingly ordinary remark targeted a group or individual.A comment about dislikes was interpreted as offensively implying that scheduled-caste people disliked a film.
  • Language and cultural context: Dialect and cultural context produced disagreement when words had regional meanings or when references required knowledge of songs, actors, and films.A Malayalam word translated as “stupid” was offensive in some regions but generally meant “bad,” while another comment was resolved as non-offensive after contextual interpretation.

7 Benchmark Systems

The benchmark treats both tasks as text classification and evaluates several traditional machine-learning models under a controlled train-development-test split. Logistic regression and SVM use TF-IDF features, while the other baselines apply task-specific conventional classifiers.

  • Benchmark models: The study evaluates Logistic Regression, Support Vector Machine, Multinomial Naive Bayes, K-Nearest Neighbours, Decision Trees, and Random Forests.These traditional algorithms are applied separately to sentiment analysis and offensive-language identification.
  • Evaluation setup: The experiments use a 90%-5%-5% random split for training, development, and testing after duplicate removal.Models are tuned on the development set and evaluated on the test set.
  • Feature and model settings: Logistic Regression uses L2 regularization and TF-IDF features containing up to 3 grams, without pretrained embeddings.Its real-valued features are weighted and passed through a sigmoid function to obtain class probabilities.
  • Feature and model settings: SVM uses the same TF-IDF features as Logistic Regression and applies L2 regularization.The classifier seeks a decision boundary that separates vector-encoded data points.
  • Other baseline settings: Multinomial Naive Bayes models word counts in bag-of-words representations, applies Laplace smoothing with α=1, and uses TF-IDF vectors.KNN is tested with 3, 4, 5, and 9 neighbours using uniform weights, while decision trees use Gini and entropy criteria.

8 Results and Discussion

The code-mixed datasets support baseline classification experiments, but performance is generally limited, especially for offensive language detection. Class imbalance and difficult sentiment categories constrain results, while the movie-trailer source skews sentiment distributions.

  • Evaluation: Macro-averaged precision, recall, and F1-score were reported for sentiment analysis and offensive language detection across the evaluated classifiers.Macro-averaging computes each class metric independently before averaging, giving all classes equal weight.
  • Sentiment analysis: Logistic regression, random forests, and decision trees performed comparatively better for sentiment analysis, although overall performance ranged from inadequate to average.SVM performed poorly, while Positive and Negative classes received higher scores than the remaining sentiment classes.
  • Sentiment analysis: Mixed feelings and Neutral state were difficult for annotators to label because of problematic examples.These challenging classes contributed to their poor classification performance.
  • Offensive language detection: All classifiers performed poorly on offensive language detection, with logistic regression and random forests relatively better than the other methods.The Not Offensive class had higher precision, recall, and F1-score than the offensive categories.
  • Class distributions: 72.4% of Tamil comments and 88.44% of Malayalam comments were non-offensive, compared with 55.79% in Kannada.The stronger non-offensive-class scores in Tamil and Malayalam were associated with their larger non-offensive class proportions.
  • Dataset effects: Movie-trailer comments produced more positive sentiment than other classes, skewing the overall distribution.The authors identify the trailer-based collection source as a dataset characteristic affecting sentiment balance.

9 Conclusion

The paper contributes a large annotated corpus for under-resourced Dravidian code-mixed languages and establishes machine-learning baselines. It reports annotation agreement and positions the resource for further code-mixed research and extension to other Dravidian languages.

  • Corpus contribution: The dataset contains more than 60,000 comments from under-resourced Dravidian languages annotated for sentiment analysis and offensive language identification.The corpus is presented as a resource for research on code-mixed Dravidian languages.
  • Annotation: The authors created an annotation scheme and achieved high inter-annotator agreement using Krippendorff’s alpha with volunteer annotators.The annotations were collected through contributions submitted using Google Forms.
  • Baselines: The paper presents gold-standard baselines with precision, recall, and F-score results for each class.These baselines use the manually annotated data.
  • Future directions: The authors expect the resource to support new code-mixed research problems and plan to investigate its application to other under-resourced Dravidian languages.The stated future direction is extending the corpora beyond the languages covered here.
Loading 2106.09460v1…