Source-linked AI summary
JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
Ryosuke Sonobe, Shinnosuke Takamichi, Hiroshi Saruwatari
TL;DR
Japanese end-to-end speech synthesis lacked a freely available large-scale corpus covering the pronunciation variation of daily-use characters. The paper constructs and analyzes JSUT, a 10-hour Japanese reading-speech corpus with transcription, pronunciation-focused coverage, and multiple domains, and releases it online for research use.
Problem
Freely available large-scale Japanese speech corpora for end-to-end speech synthesis were absent, despite the expected value of such resources for related research.
Method
The paper constructs and analyzes JSUT from nine pronunciation- and domain-oriented sub-corpora, including coverage of daily-use kanji readings and varied utterances.
Results
JSUT contains 10 hours of speech read by a native Japanese speaker, with Japanese text and speech data freely available online.
Takeaways & Limitations
JSUT is intended for end-to-end speech synthesis research by academic institutions and non-commercial research, including research within commercial organizations.
Abstract
from arXiv · showhide
Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, such a corpus for Japanese speech synthesis does not exist. In this paper, we designed a novel Japanese speech corpus, named the "JSUT corpus," that is aimed at achieving end-to-end speech synthesis. The corpus consists of 10 hours of reading-style speech data and its transcription and covers all of the main pronunciations of daily-use Japanese characters. In this paper, we describe how we designed and analyzed the corpus. The corpus is freely available online.
1. INTRODUCTION
Deep learning has accelerated speech research, but freely available large-scale Japanese corpora for end-to-end speech synthesis were lacking. The JSUT corpus addresses this gap with pronunciation-focused coverage, diverse utterances, and 10 hours of freely available speech data.
- Freely available Japanese speech corpora for end-to-end speech synthesis were not available despite the expected value for accelerating related research.
- The JSUT corpus was constructed as a free, large-scale Japanese speech corpus designed to cover pronunciations of daily-use characters and individual readings.
- 10 hours of speech data were recorded by a native Japanese speaker, analyzed linguistically and acoustically, and released online with Japanese text.
2.1. Structures
JSUT is organized into nine sub-corpora whose names encode their utterance counts. Together, they target pronunciation coverage and varied linguistic domains, styles, and phenomena.
- The corpus contains nine sub-corpora, with each name formatted as a category followed by its number of utterances.
- basic5000 covers the main pronunciations of daily-use Japanese characters, while countersuffix26 covers individual readings of counter suffixes.
- loanword128 contains loanword utterances, and onomatopee300 contains utterances featuring famous Japanese onomatopoeia.
- utparaphrase512 contains paraphrase substitutions, repeat500 contains repeatedly spoken utterances, and voiceactress100 provides para-speech for a Japanese voice-actress corpus.
- travel1000 and precedent130 provide travel-domain and precedent-domain utterances, respectively.
2.2. Components
The corpus components combine a pronunciation-focused main subset with additional linguistic and domain-oriented material. The design uses collected and manually constructed sentences to cover readings and spoken-language phenomena.
- The nine sub-corpora are described as the structural components of JSUT.
- basic5000: basic5000 is the main sub-corpus and targets coverage of individual pronunciations of Japan’s 2136 officially defined daily-use kanji characters.
- basic5000: 5000 sentences from Wikipedia and the TANAKA corpus were selected, with manually created additions covering pronunciations absent from those sources.
2.2.2. countersuffix26
The component design addresses Japanese pronunciation variation through counter suffixes and loanwords. These materials target context-dependent numeral readings and spoken forms outside the modern Japanese system.
- countersuffix26: Japanese numeral pronunciations change with counter suffixes, such as “ni” with ko and “futa” with tsu.
- countersuffix26: countersuffix26 was built from 26 crowdsourced sentences containing such counter-suffix constructions.
- loanword128: Loanword material includes everyday verbs and nouns, whose pronunciations and accents are relevant to spoken-language processing.
2.2.4. utparaphrase512
The utparaphrase512 sub-corpus pairs original and paraphrased Japanese sentences to support reading comprehension through lexical substitution.
- 256 sentences were selected with one paraphrased word per sentence, producing 512 original and paraphrased sentences.Each constructed sentence is paired with its paraphrased version.
- Paraphrasing substitutes a word or phrase into another sentence and can support reading comprehension in speech communication.
- The corpus uses paraphrased sentences from the SNOW E4 corpus as its source material.
2.2.6. Onomatopee300
The Onomatopee300 sub-corpus targets Japanese onomatopoeia, which connects speech with non-speech sounds in nature, by crowdsourcing 300 sentences.
- Onomatopoeia has an important role in connecting speech and non-speech sounds in nature.
- 300 sentences containing individual onomatopoeia words were crowdsourced.
- The sub-corpus focuses on Japanese onomatopoeia, which is rich in the language.
2.2.7. repeat500
The repeat500 sub-corpus captures within-context speech variability by repeatedly recording the same sentences and adding sentences from distinct travel and precedent domains.
- 500 repeated utterances were recorded by having one speaker produce each of 100 sentences five times.The repetitions were used to quantify speech randomness.
- The repeated recordings use sentences from the Voice Actress Corpus.
- The sub-corpus adds 1000 travel-domain sentences and 138 copyright-free precedent sentences.Some precedent sentences were manually removed or modified to make them easier to read.
3. RESULTS OF DATA COLLECTION
The corpus contains 10 hours of recorded Japanese speech with accompanying recording information and varied utterance lengths. Analysis found broad length variation and higher mean log F0 in the later recording days.
- Data collection: 10 hours of speech were recorded from a female native Japanese speaker in an anechoic room and sampled at 48 kHz.Recording took place over several months in 2017.
- Data collection: Recording information identifies the recording day, while speech power was normalized and commas were manually annotated between breath groups.
- Linguistic analysis: The mora histogram had minimum, mean, and maximum values of 7, 37.14, and 133, respectively.
- Linguistic analysis: The word histogram had minimum, mean, and maximum values of 2, 18.03, and 70, respectively.
- Linguistic analysis: Utterances ranged from a few words and moras to 70 words and 140 moras.Mora and word counts were computed with MeCab and NEologd.
- Speech analysis: Mean log F0 increased during the second half of the recording days, with no special tendency in the first half.F0 was extracted using the WORLD analysis-synthesis system.
4. CONCLUSION
The paper presents JSUT as a free, large-scale Japanese speech corpus for end-to-end speech synthesis, designed to cover daily-use kanji pronunciations and sentences from several domains.
- JSUT is freely available for academic and non-commercial research, including research conducted within commercial organizations.