Source-linked AI summary

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Sahil Deepak Gawande, Mayank Singh

arXiv:2607.23242v1cs.CLcs.LG

TL;DR

IndicTalk addresses the scarcity of large-scale, high-quality multilingual code-mixed conversational resources for Indic languages. It introduces an automated, event-grounded corpus-generation pipeline and produces fluent, coherent, naturally code-mixed conversations across supported language varieties and script variants.

  • Problem

    Large-scale, multi-turn conversational resources for Indic code-mixed language varieties remain scarce, especially across native-script and Romanized forms.

  • Method

    IndicTalk combines real-world news grounding, persona-conditioned multilingual LLM dialogue generation, and automatic quality validation to construct its corpus.

  • Results

    Evaluations indicate fluent, coherent, and naturally code-mixed conversations across all supported language varieties and both script variants.

  • Takeaways & Limitations

    IndicTalk provides a resource for developing and evaluating multilingual conversational AI for underrepresented Indic languages.

  • Takeaways & Limitations

    As a synthetic, news- and blog-derived corpus, IndicTalk may not capture natural bilingual pragmatic nuances, dialectal variation, or disfluencies and has formal, event-centric topical bias.

Abstract

from arXiv · show

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .

1 Introduction

IndicTalk addresses the scarcity of large-scale, high-quality multilingual Indic code-mixed conversational resources by providing event-grounded, multi-turn conversations across language varieties and script variants. Its fully automated generation pipeline and evaluations demonstrate fluent, coherent, and naturally code-mixed dialogue.

  • Corpus and motivation: Existing Indic datasets mainly target discriminative tasks, while conversational resources are comparatively rare, Hinglish-limited, small-scale, or manually curated.The introduction identifies a lack of publicly available large-scale, event-grounded, multi-turn conversations across multiple Indic varieties.
  • Corpus and motivation: IndicTalk comprises over 13,28,604 event-grounded multi-turn conversations spanning 18 language varieties across 9 Indic languages.The corpus covers native-script and Romanized code-mixed variants paired with English.
  • Generation pipeline: Each IndicTalk conversation contains 6–8 dialogue turns generated through a fully automated pipeline combining real-world news grounding, persona-conditioned dialogue generation, multilingual LLMs, and automatic quality validation.This design is presented as a response to the need for corpora that capture natural code-mixing across multiple turns.
  • Evaluation: Evaluation consistently demonstrates fluent, coherent, and naturally code-mixed conversations across all supported language varieties.The evaluation includes code-mixing metrics, native-script fluency analysis, LLM-as-a-Judge assessment, and human evaluation by native or highly proficient speakers.

2 Related Work

Prior Indic code-mixed resources mainly support discriminative tasks, while conversational datasets remain few, narrow in language or domain, and often translation-based. IndicTalk addresses these limitations with a large multilingual corpus spanning 18 language varieties across 9 Indic languages and a fully automated generation pipeline.

  • Code-mixed conversational datasets: Most Indic code-mixed datasets target discriminative tasks and contain isolated sentences or social media posts, limiting their suitability for conversational LLM training.Examples include sentiment analysis, hate speech detection, named entity recognition, and natural language inference.
  • Code-mixed conversational datasets: Only a handful of datasets address code-mixed dialogue, including translated task-oriented, conversational summarization, and knowledge-grounded resources.Prior examples include DSTC2 extensions, GupShup, and KCM.
  • Large-scale Indic conversational resources: Existing large-scale Indic conversational corpora are predominantly monolingual and do not model natural code-mixing.IndicDialogue covers subtitle-based conversations across ten Indic languages, while IndicLLMSuite provides multilingual pre-training and instruction-tuning resources.
  • How our work differs: 18 language varieties across 9 Indic languages distinguish IndicTalk from resources restricted to Hinglish or a small number of language pairs.The corpus includes both native-script and Romanized forms.
  • How our work differs: IndicTalk uses a fully automated pipeline to generate event-grounded, persona-conditioned conversations rather than translating or manually curating them.This design is identified as a key difference from existing efforts.

3 Methodology

IndicTalk’s methodology converts cleaned, event-grounded web documents into a shared semantic representation, then generates persona-conditioned multilingual code-mixed dialogues for 18 Indic-English varieties. Automatic validation filters outputs using script, code-mixing, and length requirements before retention.

  • Persona-Conditioned Generation: Dialogues are generated independently for 18 language varieties using multilingual LLM self-play conditioned on fixed persona specifications and dialogue history.The five persona categories are Friends, Family Members, Colleagues, Experts, and Student–Teacher.
  • Source Corpus: The pipeline extracts primary content from publicly available news articles and blog posts spanning current affairs, technology, business, sports, health, and other domains.Boilerplate elements such as navigation menus, advertisements, scripts, and duplicated templates are removed.
  • Semantic Representation: Each document is transformed into a concise English summary of 6–8 factual sentences shared across all target languages.This representation reduces document-length and stylistic variation while keeping conversations grounded in identical factual content.
  • Code-Mixed Variants: Each Indic language receives Native-Script and Romanized code-mixed variants, differing in whether Indic text uses native script or Roman script.Native-Script variants combine Indic-script tokens with Roman-script English, whereas Romanized variants use Roman script for both languages.
  • Automatic Validation: Generated conversations are retained only after automatic validation checks script usage, CMI ≥τCMI for native-script utterances, and a minimum length of τlen tokens.Romanized variants are checked for Roman characters; the validator addresses monolingual or weakly code-mixed outputs.
  • Generation Model: GPT-OSS-120B was selected for dialogue generation because it more reliably satisfied persona consistency, multilingual fidelity, and structural validation criteria than the other tested models.The comparison included Qwen2.5-72B, Gemma 4-31B, and smaller variants.

4 Dataset Analysis

IndicTalk is a large, multilingual code-mixed corpus spanning 18 language varieties, with analyses showing frequent, balanced language alternation across native-script and Romanized forms. Automatic, LLM-based, and human evaluations indicate fluent, coherent, and high-quality conversations under both writing conventions.

  • Corpus composition and scale: IndicTalk contains over 13,28,604 conversations from approximately 1,42,053 articles and blog posts, covering 9 Indic languages across 18 native-script and Romanized varieties.Each conversation averages 7.55 dialogue turns and uses the same source document across language variants for controlled comparisons.
  • Code-mixing properties: Native-script variants average CMI 37.81, SPF 0.438, I-index 0.437, and M-index 0.865, indicating frequent switching and balanced bilingual contribution.Telugu and Kannada exhibit the strongest bilingual mixing.
  • Code-mixing properties: Romanized variants retain positive average SPF (0.311), I-index (0.311), and M-index (0.508), showing meaningful code-mixing despite shared orthography.Lower scores reflect less reliable automatic language identification in Roman script.
  • Automatic fluency evaluation: Native-script variants achieve PPPL scores of 9.14–18.52, indicating fluent conversations despite frequent code-switching.PPPL is not reported for Romanized variants because transliteration variability makes cross-language comparisons less reliable.
  • LLM-as-a-Judge evaluation: Across 9,000 held-out conversations, both LLM judges assign high scores across all 18 varieties, with rankings highly correlated (ρ = 0.91, p < 0.001).The majority of varieties exceed an overall quality score of 4.0, and native-script and Romanized variants show comparable quality.
  • Human evaluation: Human evaluation finds high overall quality across languages, with native-script variants generally scoring slightly higher than Romanized variants.Marathi and Kannada receive the highest native-script ratings, while Hindi receives the highest Romanized rating.

5 Conclusion

IndicTalk is a large-scale, fully automated corpus of event-grounded, persona-based multilingual code-mixed conversations spanning 18 language varieties across 9 Indic languages in native-script and Romanized forms. Automatic and human evaluations show fluent, coherent, naturally code-mixed conversations, supporting multilingual conversational AI development and evaluation for Indic languages.

  • Corpus: 13,28,604 event-grounded multi-turn conversations span 18 language varieties covering 9 Indic languages in native-script and Romanized forms.IndicTalk is presented as a large-scale multilingual Indic code-mixed conversational corpus.
  • Corpus: The fully automated pipeline combines real-world news grounding, persona-conditioned dialogue generation, multilingual LLMs, and automatic quality validation.These components define the corpus construction process.
  • Evaluation: Automatic and human evaluations demonstrate fluent, coherent, and naturally code-mixed conversations across all supported languages.The evaluations cover the corpus’s conversational quality across its language varieties.
  • Impact: IndicTalk supports developing and evaluating multilingual conversational AI for Indic languages.The corpus is positioned as a resource for multilingual conversational AI research.

6 Future Work

Future work will broaden IndicTalk’s linguistic and interaction coverage while establishing standardized benchmarks for evaluating multilingual conversational models on downstream Indic code-mixed tasks.

  • 6 Future Work: Future work will extend IndicTalk to additional Indic languages, dialects, and script variants while adding richer personas and interaction scenarios.Proposed scenarios include customer support and multilingual debates.
  • 6 Future Work: Future work will establish standardized benchmarks for conversational summarization, sentiment analysis, and intent detection.These benchmarks would enable systematic evaluation and comparison of multilingual conversational models for Indic code-mixed natural language.

7 Limitations

IndicTalk remains limited by its synthetic nature and news/blog-based sourcing, which may reduce coverage of natural bilingual conversational phenomena and bias topics toward formal, event-centric discourse.

  • As a synthetic corpus, IndicTalk may not fully capture pragmatic nuances, dialectal variation, or disfluencies in naturally occurring bilingual conversations.The limitation persists despite the corpus’s carefully designed generation pipeline and multiple quality-control stages.
  • Because its source documents primarily come from news and blogs, IndicTalk exhibits a topical bias toward formal, event-centric discourse.

8 Ethical Considerations · Appendix

IndicTalk is produced through an automated, license-conscious process using public news and blogs, with source documents withheld and detected PII removed. Human evaluation involved informed-consent procedures, and the work aims to support more inclusive research on multilingual code-mixed conversational AI for underrepresented Indic languages.

  • 8 Ethical Considerations: IndicTalk is generated from publicly available news articles and blogs through a fully automated pipeline.The corpus generation process does not rely on releasing the underlying source documents.
  • 8 Ethical Considerations: Source documents are not released, and personally identifiable information is removed where detected during preprocessing.These measures apply to the publicly sourced news articles and blogs used for generation.
  • 8 Ethical Considerations: Human evaluation was conducted with informed consent from volunteer graduate student annotators.The annotators participated voluntarily and provided consent before evaluation.
  • 8 Ethical Considerations: Datasets and models used in the work will be publicly available or used according to their respective licenses.The stated availability and usage plans are conditioned on applicable licensing requirements.
  • 8 Ethical Considerations: IndicTalk is intended to facilitate more inclusive research on multilingual code-mixed conversational AI.The stated goal concerns broader inclusion in research on multilingual code-mixed conversational systems.
  • 8 Ethical Considerations: The corpus targets conversational AI research for underrepresented Indic languages.This motivation is stated in the context of multilingual, code-mixed conversational AI.

A Source Data Extraction and Cleaning · B Summarization

The pipeline extracts source content from HTML pages and PDFs, removes unusable documents, and supports parallel, fault-tolerant processing.

  • A Source Data Extraction and Cleaning: HTML pages are fetched with browser-mimicking User-Agent headers and parsed using BeautifulSoup.The extraction process begins with standard HTTP GET requests.
  • A Source Data Extraction and Cleaning: BeautifulSoup removes navigation menus, headers, footers, sidebars, scripts, and advertisements before text extraction.These elements are treated as boilerplate.
  • A Source Data Extraction and Cleaning: PDF documents are processed with pdfplumber, extracting text page by page before concatenation.This provides a separate extraction path for PDF sources.
  • A Source Data Extraction and Cleaning: 100 words is the minimum extracted-text threshold; documents below it are discarded.Documents are also discarded when their HTTP requests fail.
  • A Source Data Extraction and Cleaning: Sequential batches use nonoverlapping index ranges, allowing multiple pipeline instances to run in parallel.Batching supports concurrent processing without overlapping article ranges.
  • A Source Data Extraction and Cleaning: Fault-tolerant error handling isolates individual article failures so they do not interrupt the broader pipeline run.The pipeline continues processing despite errors affecting particular articles.

B.1 Model Selection and Evaluation … D Human Evaluation Prompt

The paper selects Sarvam-M through blind pairwise human evaluation and uses it for factual article summarization, while human dialogue evaluation scores fluency, coherence, and code-mixing naturalness on a 1–5 scale.

  • B.1 Model Selection and Evaluation: The model-selection study compared Sarvam-M, Llama 7B, and Gemma 7B using arena-style blind human evaluation on documents from four language tracks.Annotators compared model outputs with identities hidden before voting.
  • B.1 Model Selection and Evaluation: Annotators assessed summaries for factual accuracy, content coverage, and named-entity preservation using the original article and source text.Model identities were revealed only after votes were cast to support unbiased preference judgments.
  • B.1 Model Selection and Evaluation: 509 pairwise preference votes yielded Sarvam-M’s highest Elo rating of 1861.5, outperforming Llama 7B and Gemma 7B across all four language tracks.The tracks were English, Gujarati, Hindi, and Tamil; Sarvam-M was consequently used throughout the summarization pipeline.
  • B.2 Article Summarization Prompt: The summarization prompt requests 6–8 concise, factual, neutral English sentences forming a coherent paragraph without bullet points.It instructs the model to include only information explicitly supported by the article and avoid opinions, speculation, and external knowledge.
  • B.2 Article Summarization Prompt: The prompt preserves named entities, numerical values, dates, temporal expressions, and locations to create a language-independent semantic representation of each source document.The representation is generated from the source document using the summarization model.
  • D Human Evaluation Prompt: The evaluation prompt presents a synthetic dialogue between Speaker_1 and Speaker_2 simulating a conversation about a topic.For Romanized variants, it specifies Roman script without native-script characters and transliterated native-language words mixed with English.
  • D Human Evaluation Prompt: Human evaluators rate synthetic dialogues on fluency, coherence, and code-mixing naturalness using integer scores from 1 to 5 and brief justifications.They are instructed to penalize grammatical errors, spelling mistakes, unnatural expressions, and implausible code-switching patterns.

E Qwen2.5-72B Vs Gemma 4-31B Vs GPT-OSS-120B

GPT-OSS-120B generates the most coherent, contextually grounded, and naturally code-mixed conversations among the three evaluated models, while Qwen2.5-72B and Gemma 4-31B show weaker grounding or consistency.

  • Qualitative comparison: GPT-OSS-120B consistently produces the most coherent, contextually grounded, and naturally code-mixed conversations among the three models.Qwen2.5-72B and Gemma 4-31B exhibit relatively weaker grounding or conversational consistency.
  • LLM-as-Judge evaluation: The same LLM-as-Judge prompt was used for Gemini-2.5-Flash and GPT-OSS-120B across all language varieties.Romanized variants additionally included the italicized instruction; both judges used greedy decoding with T = 0.0, and invalid outputs were retried up to three times.
  • Representative examples: Representative Hindi-English code-mixed conversations are presented for Qwen2.5-72B, Gemma 4-31B, and GPT-OSS-120B.The examples appear in Tables 10, 11, and 12, respectively.
Loading 2607.23242v1…