paper-with-me

홈 › Papers

On the Compression of Lexicon Transducers

2019-09-01 · WS 2019 9 · Marco Cognetta, Cyril Allauzen, Michael Riley

In finite-state language processing pipelines, a lexicon is often a key component. It needs to be comprehensive to ensure accuracy, reducing out-of-vocabulary misses. However, in memory-constrained environments (e.g., mobile phones), the size of the component automata must be kept small. Indeed, a delicate balance between comprehensiveness, speed, and memory must be struck to conform to device requirements while providing a good user experience.In this paper, we describe a compression scheme for lexicons when represented as finite-state transducers. We efficiently encode the graph of the transducer while storing transition labels separately. The graph encoding scheme is based on the LOUDS (Level Order Unary Degree Sequence) tree representation, which has constant time tree traversal for queries while being information-theoretically optimal in space. We find that our encoding is near the theoretical lower bound for such graphs and substantially outperforms more traditional representations in space while remaining competitive in latency benchmarks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JCLext: A Java Tool for Compiling Finite-State Transducers from Full-Form Lexicons

2015-11-01 · WS 2015 11 · Leonel F. de Alencar, Philipp B. Costa, Mardonio J. C. Fran{\c{c}}a, Alex Ewart 외
Form

Mlphon: A Multifunctional Grapheme-Phoneme Conversion Tool Using Finite State Transducers

2022-09-05 · IEEE Access 2022 9 · Kavya Manohar, A R jayan, Rajeev Rajan

In this article we present the design and the development of a knowledge based computational linguistic tool, Mlphon for Malayalam language. Mlphon computationally models linguistic rules using finite state transducers a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityGrapheme-to-Phoneme Conversion+10

A Syntactically Expressive Morphological Analyzer for Turkish

2019-09-01 · WS 2019 9 · Adnan Ozturel, Tolga Kayadelen, Isin Demirsahin

We present a broad coverage model of Turkish morphology and an open-source morphological analyzer that implements it. The model captures intricacies of Turkish morphology-syntax interface, thus could be used as a baselin…

Language ModelingLanguage Modelling

Reinforcement learning of minimalist grammars

2020-04-30 · Peter beim Graben, Ronald Römer, Werner Meyer, Markus Huber 외

Speech-controlled user interfaces facilitate the operation of devices and household functions to laymen. State-of-the-art language technology scans the acoustically analyzed speech signal for relevant keywords that are s…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)slot-filling+1

PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers

2023-03-30 · Rahul Pandey, Roger Ren, Qi Luo, Jing Liu 외

End-to-End (E2E) automatic speech recognition (ASR) systems used in voice assistants often have difficulties recognizing infrequent words personalized to the user, such as names and places. Rare words often have non-triv…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition