paper-with-me

Papers

Rule based Approach for Word Normalization by resolving Transcription Ambiguity in Transliterated Search Queries

2019-10-16 · Varsha Pathak, Manish Joshi

Query term matching with document term matching is the basic function of any best effort Information Retrieval models like Vector Space Model. In our problem of SMS based Information Systems we expect common people to participate in information search. Our system allows mobile users to formulate their queries in their own words, own transliteration style and spelling formation. To achieve this flexibility we have resolved the term level ambiguity due to inherent transcription noise in user query terms. We have developed a rule based approach to select most relevantly close standard term for each noisy term in the user query. We have used four different versions of the rule based algorithm with variation in the rule set. We have formulated this rule set including the basic Levenshtein minimum edit distance algorithm for term matching. This paper presents the experiments and corresponding results of Marathi and Hindi language literature information system. We have experimented on Marathi and Hindi literature which include songs, gazals, powadas, bharud and other types in a standard transliteration form like ITRANS.

📄 PDF Abstract BibTeX arXiv:1910.07233

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrievalTransliteration

Similar Papers 제목 키워드 기반

Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration

2024-02-05 · Xiliang Zhu, Chia-Tien Chang, Shayna Gardiner, David Rossouw 외

Punctuation restoration is a crucial step after Automatic Speech Recognition (ASR) systems to enhance transcript readability and facilitate subsequent NLP tasks. Nevertheless, conventional lexical-based approaches are in…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+4

Towards Resolving Word Ambiguity with Word Embeddings

2023-07-25 · Matthias Thurnbauer, Johannes Reisinger, Christoph Goller, Andreas Fischer

Ambiguity is ubiquitous in natural language. Resolving ambiguous meanings is especially important in information retrieval tasks. While word embeddings carry semantic information, they fail to handle ambiguity well. Tran…

Information RetrievalRetrievalWord Embeddings

VDN-NeRF: Resolving Shape-Radiance Ambiguity via View-Dependence Normalization

2023-03-31 · CVPR 2023 1 · Bingfan Zhu, Yanchao Yang, Xulong Wang, Youyi Zheng 외

We propose VDN-NeRF, a method to train neural radiance fields (NeRFs) for better geometry under non-Lambertian surface and dynamic lighting conditions that cause significant variation in the radiance of a point when view…

NeRF

Investigating Transcription Normalization in the Faetar ASR Benchmark

2025-08-15 · Leo Peckham, Michael Ong, Naomi Nagy, Ewan Dunbar arxiv

We examine the role of transcription inconsistencies in the Faetar Automatic Speech Recognition benchmark, a challenging low-resource ASR benchmark. With the help of a small, hand-constructed lexicon, we conclude that fi…

Speech RecognitionLanguage Modelling

TS-Net: OCR Trained to Switch Between Text Transcription Styles

2021-03-09 · Jan Kohút, Michal Hradiš

Users of OCR systems, from different institutions and scientific disciplines, prefer and produce different transcription styles. This presents a problem for training of consistent text recognition neural networks on real…

Optical Character Recognition (OCR)