paper-with-me

홈 › Papers

Interpretable Mnemonic Generation for Kanji Learning via Expectation-Maximization

2025-07-07 · Jaewook Lee, Alexander Scarlatos, Andrew Lan arxiv

Learning Japanese vocabulary is a challenge for learners from Roman alphabet backgrounds due to script differences. Japanese combines syllabaries like hiragana with kanji, which are logographic characters of Chinese origin. Kanji are also complicated due to their complexity and volume. Keyword mnemonics are a common strategy to aid memorization, often using the compositional structure of kanji to form vivid associations. Despite recent efforts to use large language models (LLMs) to assist learners, existing methods for LLM-based keyword mnemonic generation function as a black box, offering limited interpretability. We propose a generative framework that explicitly models the mnemonic construction process as driven by a set of common rules, and learn them using a novel Expectation-Maximization-type algorithm. Trained on learner-authored mnemonics from an online platform, our method learns latent structures and compositional rules, enabling interpretable and systematic mnemonics generation. Experiments show that our method performs well in the cold-start setting for new learners while providing insight into the mechanisms behind effective mnemonic creation.

📄 PDF Abstract BibTeX arXiv:2507.05137

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeepMnemonic: Password Mnemonic Generation via Deep Attentive Encoder-Decoder Model

2020-06-24 · Yao Cheng, Chang Xu, Zhen Hai, Yingjiu Li

Strong passwords are fundamental to the security of password-based user authentication systems. In recent years, much effort has been made to evaluate password strength or to generate strong passwords. Unfortunately, the…

DecoderSentence

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

2026-07-01 · Ryota Mibayashi, Hiroya Takamura, Hitomi Yanaka arxiv

We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character often has multiple possible readings, mak…

PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs

2025-07-07 · Sana Kang, Myeongseok Gwon, Su Young Kwon, Jaewook Lee 외 arxiv

Vocabulary acquisition poses a significant challenge for second-language (L2) learners, especially when learning typologically distant languages such as English and Korean, where phonological and structural mismatches co…

A Character-based Approach to Distributional Semantic Models: Exploiting Kanji Characters for Constructing JapaneseWord Vectors

2014-05-01 · LREC 2014 5 · Akira Utsumi

Many Japanese words are made of kanji characters, which themselves represent meanings. However traditional word-based distributional semantic models (DSMs) do not benefit from the useful semantic information of kanji cha…

Collaborative FilteringInformation Retrieval

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

2026-06-24 · Lianbo Liu, Shiao Zhu, Kai Washizaki, Reo Yoneyama 외 arxiv

While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on English and Chinese. Japanese, however, remains under-explored, and its uniq…

Data AugmentationSpeech Synthesis