Bilingual Document Alignment with Latent Semantic Indexing
We apply cross-lingual Latent Semantic Indexing to the Bilingual Document Alignment Task at WMT16. Reduced-rank singular value decomposition of a bilingual term-document matrix derived from known English/French page pairs in the training data allows us to map monolingual documents into a joint semantic space. Two variants of cosine similarity between the vectors that place each document into the joint semantic space are combined with a measure of string similarity between corresponding URLs to produce 1:1 alignments of English/French web pages in a variety of domains. The system achieves a recall of ca. 88% if no in-domain data is used for building the latent semantic model, and 93% if such data is included. Analysing the system's errors on the training data, we argue that evaluating aligner performance based on exact URL matches under-estimates their true performance and propose an alternative that is able to account for duplicates and near-duplicates in the underlying data.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Novel Two-Step Method for Cross Language Representation Learning
Cross language text classification is an important learning task in natural language processing. A critical challenge of cross language learning lies in that words of different languages are in disjoint feature spaces. In…
Matrix CompletionRepresentation LearningVocal Bursts Valence PredictionBuilding and Aligning Comparable Corpora
Comparable corpus is a set of topic aligned documents in multiple languages, which are not necessarily translations of each other. These documents are useful for multilingual natural language processing when there is no …
HM-BiTAM: Bilingual Topic Exploration, Word Alignment, and Translation
We present a novel paradigm for statistical machine translation (SMT), based on joint modeling of word alignment and the topical aspects underlying bilingual document pairs via a hidden Markov Bilingual Topic AdMixture (…
Machine TranslationSentenceTranslationWord AlignmentProbabilistic Latent Semantic Analysis (PLSA) untuk Klasifikasi Dokumen Teks Berbahasa Indonesia
One task that is included in managing documents is how to find substantial information inside. Topic modeling is a technique that has been developed to produce document representation in form of keywords. The keywords wi…
RetrievalIndexation et appariement de documents cliniques avec le mod\`ele vectoriel (Indexing and matching clinical documents using the vector space model)
Dans ce papier, nous pr{\'e}sentons les m{\'e}thodes que nous avons d{\'e}velopp{\'e}es pour participer aux t{\^a}ches 1 et 2 de l{'}{\'e}dition 2019 du d{\'e}fi fouille de textes (DEFT 2019). Pour la premi{\`e}re t{\^a}…