paper-with-me

Papers

Building Tailored Speech Recognizers for Japanese Speaking Assessment

2025-09-25 · Yotaro Kubo, Richard Sproat, Chihiro Taguchi, Llion Jones arxiv

This paper presents methods for building speech recognizers tailored for Japanese speaking assessment tasks. Specifically, we build a speech recognizer that outputs phonemic labels with accent markers. Although Japanese is resource-rich, there is only a small amount of data for training models to produce accurate phonemic transcriptions that include accent marks. We propose two methods to mitigate data sparsity. First, a multitask training scheme introduces auxiliary loss functions to estimate orthographic text labels and pitch patterns of the input signal, so that utterances with only orthographic annotations can be leveraged in training. The second fuses two estimators, one over phonetic alphabet strings, and the other over text token sequences. To combine these estimates we develop an algorithm based on the finite-state transducer framework. Our results indicate that the use of multitask learning and fusion is effective for building an accurate phonemic recognizer. We show that this approach is advantageous compared to the use of generic multilingual recognizers. The relative advantages of the proposed methods were also compared. Our proposed methods reduced the average of mora-label error rates from 12.3% to 7.1% over the CSJ core evaluation sets.

📄 PDF Abstract BibTeX arXiv:2509.20655

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sentence Suggestion of Japanese Functional Expressions for Chinese-speaking Learners

2018-07-01 · ACL 2018 7 · Jun Liu, Hiroyuki Shindo, Yuji Matsumoto

We present a computer-assisted learning system, Jastudy, which is particularly designed for Chinese-speaking learners of Japanese as a second language (JSL) to learn Japanese functional expressions with suggestion of app…

ClusteringSentence

JSSS: free Japanese speech corpus for summarization and simplification

2020-10-05

In this paper, we construct a new Japanese speech corpus for speech-based summarization and simplification, "JSSS" (pronounced "j-triple-s"). Given the success of reading-style speech synthesis from short-form sentences,…

FormSpeech Synthesistext-to-speechText to Speech

Item Development and Scoring for Japanese Oral Proficiency Testing

2012-05-01 · LREC 2012 5 · Hitokazu Matsushita, Deryle Lonsdale

This study introduces and evaluates a computerized approach to measuring Japanese L2 oral proficiency. We present a testing and scoring method that uses a type of structured speech called elicited imitation (EI) to evalu…

Language ModelingLanguage ModellingSpeech Recognition

The Amplitude Modulation Structure of Japanese Infant- and Child-Directed Speech: Longitudinal Data Reveal Universal Acoustic Physical Structures Underpinning Moraic Timing

2025-03-07 · Tatsuya Daikoku, Usha Goswami

Infant-directed speech (IDS) is highly rhythmic, and in European languages IDS is dominated by patterns of amplitude modulation (AM) at ~2Hz (reflecting prosody) and ~5Hz (reflecting individual syllables). The rhythm str…

Rhythm

Simplification of Example Sentences for Learners of Japanese Functional Expressions

2016-12-01 · WS 2016 12 · Jun Liu, Yuji Matsumoto

Learning functional expressions is one of the difficulties for language learners, since functional expressions tend to have multiple meanings and complicated usages in various situations. In this paper, we report an expe…