Extracting linguistic speech patterns of Japanese fictional characters using subword units
This study extracted and analyzed the linguistic speech patterns that characterize Japanese anime or game characters. Conventional morphological analyzers, such as MeCab, segment words with high performance, but they are unable to segment broken expressions or utterance endings that are not listed in the dictionary, which often appears in lines of anime or game characters. To overcome this challenge, we propose segmenting lines of Japanese anime or game characters using subword units that were proposed mainly for deep learning, and extracting frequently occurring strings to obtain expressions that characterize their utterances. We analyzed the subword units weighted by TF/IDF according to gender, age, and each anime character and show that they are linguistic speech patterns that are specific for each feature. Additionally, a classification experiment shows that the model with subword units outperformed that with the conventional method.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Towards an Entertaining Natural Language Generation System: Linguistic Peculiarities of Japanese Fictional Characters
Statistical laws and linguistics differ in naturalistic video and fictional conversations
Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations…
The Amplitude Modulation Structure of Japanese Infant- and Child-Directed Speech: Longitudinal Data Reveal Universal Acoustic Physical Structures Underpinning Moraic Timing
Infant-directed speech (IDS) is highly rhythmic, and in European languages IDS is dominated by patterns of amplitude modulation (AM) at ~2Hz (reflecting prosody) and ~5Hz (reflecting individual syllables). The rhythm str…
RhythmApplying Syntax$\unicode{x2013}$Prosody Mapping Hypothesis and Prosodic Well-Formedness Constraints to Neural Sequence-to-Sequence Speech Synthesis
End-to-end text-to-speech synthesis (TTS), which generates speech sounds directly from strings of texts or phonemes, has improved the quality of speech synthesis over the conventional TTS. However, most previous studies …
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisSentiment Analysis with R: Natural Language Processing for Semi-Automated Assessments of Qualitative Data
Sentiment analysis is a sub-discipline in the field of natural language processing and computational linguistics and can be used for automated or semi-automated analyses of text documents. One of the aims of these analys…
Sentiment Analysis