paper-with-me

Papers

Fast, Not Fancy: Rethinking G2P with Rich Data and Rule-Based Models

2025-05-19 · Mahta Fetrat Qharabagh, Zahra Dehghanian, Hamid R. Rabiee

Homograph disambiguation remains a significant challenge in grapheme-to-phoneme (G2P) conversion, especially for low-resource languages. This challenge is twofold: (1) creating balanced and comprehensive homograph datasets is labor-intensive and costly, and (2) specific disambiguation strategies introduce additional latency, making them unsuitable for real-time applications such as screen readers and other accessibility tools. In this paper, we address both issues. First, we propose a semi-automated pipeline for constructing homograph-focused datasets, introduce the HomoRich dataset generated through this pipeline, and demonstrate its effectiveness by applying it to enhance a state-of-the-art deep learning-based G2P system for Persian. Second, we advocate for a paradigm shift - utilizing rich offline datasets to inform the development of fast, rule-based methods suitable for latency-sensitive accessibility applications like screen readers. To this end, we improve one of the most well-known rule-based G2P systems, eSpeak, into a fast homograph-aware version, HomoFast eSpeak. Our results show an approximate 30% improvement in homograph disambiguation accuracy for the deep learning-based and eSpeak systems.

📄 PDF Abstract BibTeX arXiv:2505.12973

Code (3)

MahtaFetrat/Homo-GE2PE-Persian 공식 구현 pytorch
MahtaFetrat/HomoRich-G2P-Persian 공식 구현
MahtaFetrat/Persian-G2P-Tools-Benchmark 공식 구현

Similar Papers 제목 키워드 기반

ALFA: A Safe-by-Design Approach to Mitigate Quishing Attacks Launched via Fancy QR Codes

2026-01-11 · Muhammad Wahid Akram, Keshav Sood, Muneeb Ul Hassan, Dhananjay Thiruvady arxiv

Phishing with Quick Response (QR) codes is termed as Quishing. The attackers exploit this method to manipulate individuals into revealing their confidential data. Recently, we see the colorful and fancy representations o…

Learning Structured Declarative Rule Sets -- A Challenge for Deep Discrete Learning

2020-12-08 · Johannes Fürnkranz, Eyke Hüllermeier, Eneldo Loza Mencía, Michael Rapp

Arguably the key reason for the success of deep neural networks is their ability to autonomously form non-linear combinations of the input features, which can be used in subsequent layers of the network. The analogon to …

Position

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

2024-08-15 · Jiasong Feng, Ao Ma, Jing Wang, Bo Cheng 외

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross…

TARVideo Generation

Modular TTT: Rethinking Test-Time Training as Composable Modules

2026-08-07 · Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li 외 hf

Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically har…

Rethinking Performance Gains in Image Dehazing Networks

2022-09-23 · Yuda Song, Yang Zhou, Hui Qian, Xin Du

Image dehazing is an active topic in low-level vision, and many image dehazing networks have been proposed with the rapid development of deep learning. Although these networks' pipelines work fine, the key mechanism to i…

Image DehazingSingle Image Dehazing