paper-with-me

홈 › Papers

Bilingual Document Alignment with Latent Semantic Indexing

2017-07-29 · WS 2016 8 · Ulrich Germann

We apply cross-lingual Latent Semantic Indexing to the Bilingual Document Alignment Task at WMT16. Reduced-rank singular value decomposition of a bilingual term-document matrix derived from known English/French page pairs in the training data allows us to map monolingual documents into a joint semantic space. Two variants of cosine similarity between the vectors that place each document into the joint semantic space are combined with a measure of string similarity between corresponding URLs to produce 1:1 alignments of English/French web pages in a variety of domains. The system achieves a recall of ca. 88% if no in-domain data is used for building the latent semantic model, and 93% if such data is included. Analysing the system's errors on the training data, we argue that evaluating aligner performance based on exact URL matches under-estimates their true performance and propose an alternative that is able to account for duplicates and near-duplicates in the underlying data.

📄 PDF Abstract BibTeX arXiv:1707.09443

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Novel Two-Step Method for Cross Language Representation Learning

2013-12-01 · NeurIPS 2013 12 · Min Xiao, Yuhong Guo

Cross language text classification is an important learning task in natural language processing. A critical challenge of cross language learning lies in that words of different languages are in disjoint feature spaces. In…

Matrix CompletionRepresentation LearningVocal Bursts Valence Prediction

Building and Aligning Comparable Corpora

2025-08-04 · Motaz Saad, David Langlois, Kamel Smaili arxiv

Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no …

HM-BiTAM: Bilingual Topic Exploration, Word Alignment, and Translation

2007-12-01 · NeurIPS 2007 12 · Bing Zhao, Eric P. Xing

We present a novel paradigm for statistical machine translation (SMT), based on joint modeling of word alignment and the topical aspects underlying bilingual document pairs via a hidden Markov Bilingual Topic AdMixture (…

Machine TranslationSentenceTranslationWord Alignment

Probabilistic Latent Semantic Analysis (PLSA) untuk Klasifikasi Dokumen Teks Berbahasa Indonesia

2015-12-02 · Derwin Suhartono

One task that is included in managing documents is how to find substantial information inside. Topic modeling is a technique that has been developed to produce document representation in form of keywords. The keywords wi…

Retrieval

Indexation et appariement de documents cliniques avec le mod\`ele vectoriel (Indexing and matching clinical documents using the vector space model)

2019-07-01 · JEPTALNRECITAL 2019 7 · Khadim Dram{\'e}, Ibrahima Diop, Lamine Faty, Birame Ndoye

Dans ce papier, nous pr{\'e}sentons les m{\'e}thodes que nous avons d{\'e}velopp{\'e}es pour participer aux t{\^a}ches 1 et 2 de l{'}{\'e}dition 2019 du d{\'e}fi fouille de textes (DEFT 2019). Pour la premi{\`e}re t{\^a}…