Source-linked AI summary

Creating a Live, Public Short Message Service Corpus: The NUS SMS Corpus

Tao Chen, Min-Yen Kan

arXiv:1112.2468v1cs.CL

TL;DR

Public SMS research lacks shared raw data because messages are private and existing corpora are scarce. The paper builds a live, publicly released English–Mandarin corpus using multiple collection methods, privacy safeguards, and metadata. By October 2011, it had collected 28,724 English and 29,100 Chinese messages, while the authors note continuing privacy and anonymization challenges.

  • Problem

    Private SMS data and the scarcity of public corpora limit comparative research using shared raw messages.

  • Method

    The project uses multiple collection methodologies, anonymization, metadata gathering, and monthly public releases for a live English–Mandarin corpus.

  • Results

    By October 2011, the corpus contained 28,724 English and 29,100 Chinese SMS messages and was described as the largest public corpus in each language.

  • Takeaways & Limitations

    The corpus is intended to support comparative SMS studies by providing a growing public reference dataset.

  • Takeaways & Limitations

    Privacy remains a constraint because manual review and possible contributor identification discouraged participation despite safeguards.

Abstract

from arXiv · show

Short Message Service (SMS) messages are largely sent directly from one person to another from their mobile phones. They represent a means of personal communication that is an important communicative artifact in our current digital era. As most existing studies have used private access to SMS corpora, comparative studies using the same raw SMS data has not been possible up to now. We describe our efforts to collect a public SMS corpus to address this problem. We use a battery of methodologies to collect the corpus, paying particular attention to privacy issues to address contributors' concerns. Our live project collects new SMS message submissions, checks their quality and adds the valid messages, releasing the resultant corpus as XML and as SQL dumps, along with corpus statistics, every month. We opportunistically collect as much metadata about the messages and their sender as possible, so as to enable different types of analyses. To date, we have collected about 60,000 messages, focusing on English and Mandarin Chinese.

1 Introduction

SMS is important but difficult to study publicly because messages are private, scarce in accessible corpora, and unlike broadcast short-message media. The NUS project addresses this gap with a continuously expanding, anonymized corpus for English and Mandarin Chinese.

  • Motivation: Public SMS research is constrained by private messages, limited corpus availability, and contributors’ reluctance to share personal data.Researchers often require many phone owners’ cooperation and software spanning mobile platforms to collect large datasets.
  • Motivation: Unlike tweets, SMS is private interpersonal communication, can contain sensitive information, and differs in typical message length despite a shared 140-character restriction.These differences limit how directly social-media corpora can substitute for SMS data.
  • Project goals: The revived project collects English and Mandarin Chinese SMS through multiple methodologies to improve representativeness beyond a single local community.The project also targets fewer transcription errors and public, copyright-free release.
  • Project outcomes: By October 2011, the corpus contained 28,724 English and 29,100 Chinese messages, described as the largest public corpora for either language.The project releases corpus data and statistics through a public website in XML and SQL formats on a regular schedule.
  • Article scope: The article reports collection methods, privacy practices, public dissemination, and comparisons involving Chinese and U.S. crowdsourcing platforms.Its contributions include a website for direct browsing and recurring corpus updates.

2 Related Work

Prior SMS corpora are generally small, linguistically uneven, and often inaccessible because privacy and consent constrain publication. The paper reviews collection approaches and motivates distributed methods for building a large comparative resource.

  • Corpus availability: Publicly available SMS corpora are scarce, creating a cycle in which researchers repeatedly collect private, project-specific datasets.The review inventories existing corpora and their contributors and collection methods.
  • Corpus characteristics: Most SMS corpora are small: 50% contain fewer than 1,000 messages, only five exceed 10,000, and the largest listed corpus has 85,870 messages.The review links small scale to collection difficulty and notes that small samples can weaken statistical significance.
  • Corpus characteristics: Existing corpora span multiple languages, but European languages dominate, only two Chinese and two African corpora are noted, and most collections are monolingual.Only five existing corpora were identified as multilingual.
  • Collection methods: Collection methods include manual transcription, later transcription from phone photographs, and software-based export or upload of messages.The reviewed methods vary with contributor access and collection design.
  • Corpus availability: Privacy, non-disclosure, contributor consent, and institutional review rules explain why most existing corpora are private or only partially public.Researchers may protect contributors by withholding data or may be unable to obtain permission for public release.
  • Collection methods: Large authoritative corpora require distributing collection across many contributors, making crowdsourcing a relevant strategy for SMS data gathering.The paper focuses on mechanized labor, exemplified by Amazon Mechanical Turk, as one crowdsourcing form used in language-data collection.

3 Methodology

The corpus combines diverse contributor sources and collection methods to build representative, accurate English and Chinese SMS data while addressing privacy and replicability.

  • The project targets representative, accurate, free public English and Chinese SMS data and regular versioned releases for reference use.Its goals include expanding the earlier English corpus and addressing the scarcity of public Chinese SMS data.
  • 3.2 Source of Contributors: The collection accepted personal sent messages across unrestricted topics and gathered contributor demographics, texting habits, and phone information.Received messages were excluded because sender consent was not guaranteed and complete sender demographics could not be obtained.
  • 3.2 Source of Contributors: Contributors came from MTurk, Chinese crowdsourcing, Singapore recruitment, and unpaid online communities to diversify backgrounds.MTurk was suitable for English collection, while its shortage of Chinese workers motivated use of a more suitable Chinese platform.
  • 3.3 Technical methods: Three technical methods were used: web transcription, SMS exporting, and direct smartphone-app contribution, with exporting improving accuracy and batch submission.The methods were designed to be simple, convenient, and less prone to transcription errors.
  • A live corpus grows, is maintained, and is released at regular short intervals rather than changing immediately with every new submission.Regular releases reduce replicability problems caused by researchers using different corpus versions.

4 Properties and Statistics

By October 2011, the corpus had nearly 60,000 English and Chinese messages gathered through multiple collection methods, with contributor demographics, metadata, and costs documented. The statistics show substantial language-specific differences in contributor counts, collection sources, profile coverage, and texting habits.

  • 28,724 English and 29,100 Chinese messages had been collected by October 2011, totaling nearly 60,000 messages.
  • Chinese collection involved 515 contributors averaging 56.5 messages each, compared with 116 English contributors averaging 247.6 messages each, because most Chinese contributors submitted fewer than 30 messages.
  • 98.3% of English messages and 45.9% of Chinese messages came from SMS Export or SMS Upload, which provide typo-free text and sender, receiver, and timestamp metadata.
  • 99.5% of English messages and 93.8% of Chinese messages had associated user profiles, although missing surveys were more common in the Chinese Zhubajie collection.
  • Contributor demographics were skewed toward young adults, and English and Chinese contributors differed in daily texting frequency and SMS-use experience.
  • Crowdsourcing was economical relative to local collection, while collection costs and participation patterns varied across MTurk, ShortTask, Zhubajie, and local contributors.

5 Discussion

The discussion examines privacy, contributor reactions, altruistic recruitment, and the relative performance of Zhubajie and MTurk for SMS collection. Zhubajie was more economical for the reported SMS collection, while privacy concerns constrained recruitment and platform access.

  • Reactions to our Collection Efforts: Privacy concerns limited participation and ultimately led Amazon to suspend the project’s MTurk account without further detail.Potential contributors also worried that manual review or identifiable names could expose their messages.
  • Zhubajie compared to MTurk: The discussion frames Zhubajie as a Chinese crowdsourcing platform whose task categories, payment structure, and reputation system differ from MTurk.The authors compare the platforms across conceptualization, task characteristics, cost, completion time, and result quality.
  • Zhubajie compared to MTurk: Zhubajie collected SMS at $0.0057 per message versus $0.00815 in MTurk, making it 30.1% more economical for the reported collection.The authors compare the two platforms’ per-message costs for their SMS tasks.
  • Zhubajie compared to MTurk: Zhubajie completed a Chinese transcription task in less than 30 minutes, whereas comparable MTurk tasks took 2 days for English and 20 days for Chinese.Under SMS Export, Zhubajie collected 16 Chinese submissions in 40 days while MTurk collected 27 English submissions in 50 days.
  • Zhubajie compared to MTurk: The authors judged Zhubajie work quality higher than MTurk work, associating this difference with Zhubajie’s open worker reputation system.SMS Upload’s unique application code could discourage cheating, whereas SMS Export permitted arbitrary file uploads; the comparison also reports a platform-level quality judgment.
  • Altruism as a Possible Motivator: Unpaid community recruitment produced only 5 anonymous contributions totaling 149 English and 236 Chinese SMS, which the authors judged unsuccessful.The authors do not recommend this collection method in its current form.

6 Conclusion

The project revives the SMS corpus as a live English-and-Mandarin resource designed for representative, accurate, openly released, and reusable research data. By October 2011 it contained nearly 58,000 messages, while the authors continued expanding languages, users, and downstream annotations.

  • 6 Conclusion: The revised project aims to make the corpus representative, accurate, copyright-free for unlimited use, and useful as a reference dataset.These goals motivated the project’s resurrection in October 2010 for English and Mandarin Chinese SMS.
  • 6 Conclusion: By October 2011, the corpus contained 28,724 English and 29,100 Chinese SMS, collected for $497 and approximately 300 human hours.Because the project is live, these totals and resource costs were continuing to grow.
  • 6 Conclusion: The collection combines mobile applications, Chinese crowdsourcing, platform comparison, and an unsuccessful test of uncompensated contribution.The authors report that large-scale publicity, rather than altruistic motivation alone, is key to successful collection.
  • 6 Conclusion: The authors plan to extend collection to more languages and a wider user population, test additional methods such as an iOS application, and benchmark their efficacy.They also identify possible future annotation for part of speech, translation, and semantic markup.

7 Data

The paper makes the NUS SMS corpus publicly available through its corpus website. This provides access to the dataset described in the study.

  • 7 Data: The corpus described in the paper is publicly available at the NUS SMS corpus website.The paper identifies the corpus website as the access point for the resource.
Loading 1112.2468v1…