paper-with-me

Papers

A Comparative Study on Neural Architectures and Training Methods for Japanese Speech Recognition

2021-06-09 · Shigeki Karita, Yotaro Kubo, Michiel Adriaan Unico Bacchiani, Llion Jones

End-to-end (E2E) modeling is advantageous for automatic speech recognition (ASR) especially for Japanese since word-based tokenization of Japanese is not trivial, and E2E modeling is able to model character sequences directly. This paper focuses on the latest E2E modeling techniques, and investigates their performances on character-based Japanese ASR by conducting comparative experiments. The results are analyzed and discussed in order to understand the relative advantages of long short-term memory (LSTM), and Conformer models in combination with connectionist temporal classification, transducer, and attention-based loss functions. Furthermore, the paper investigates on effectivity of the recent training techniques such as data augmentation (SpecAugment), variational noise injection, and exponential moving average. The best configuration found in the paper achieved the state-of-the-art character error rates of 4.1%, 3.2%, and 3.5% for Corpus of Spontaneous Japanese (CSJ) eval1, eval2, and eval3 tasks, respectively. The system is also shown to be computationally efficient thanks to the efficiency of Conformer transducers.

📄 PDF Abstract BibTeX arXiv:2106.05111

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Implementing a Logical Inference System for Japanese Comparatives

2025-09-17 · Yosuke Mikami, Daiki Matsuoka, Hitomi Yanaka arxiv

Natural Language Inference (NLI) involving comparatives is challenging because it requires understanding quantities and comparative relations expressed by sentences. While some approaches leverage Large Language Models (…

Natural Language Inference

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

2025-10-25 · Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara 외 arxiv

Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for impro…

Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?

2025-09-17 · Yosuke Mikami, Daiki Matsuoka, Hitomi Yanaka arxiv

Large Language Models (LLMs) perform remarkably well in Natural Language Inference (NLI). However, NLI involving numerical and logical expressions remains challenging. Comparatives are a key linguistic phenomenon related…

Natural Language Inference

Adapting Methods for Domain-Specific Japanese Small LMs: Scale, Architecture, and Quantization

2026-03-12 · Takato Yasuno arxiv

This paper presents a systematic methodology for building domain-specific Japanese small language models using QLoRA fine-tuning. We address three core questions: optimal training scale, base-model selection, and archite…

Bridging National and International Legal Data: Two Projects Based on the Japanese Legal Standard XML Schema for Comparative Law Studies

2026-03-16 · Makoto Nakamura arxiv

This paper presents an integrated framework for computational comparative law by connecting two consecutive research projects based on the Japanese Legal Standard (JLS) XML schema. The first project establishes structura…

Semantic Textual Similarity