paper-with-me

Papers

D-Terminer: Online Demo for Monolingual and Bilingual Automatic Term Extraction

2022-06-01 · TERM (LREC) 2022 6 · Ayla Rigouts Terryn, Veronique Hoste, Els Lefever

This contribution presents D-Terminer: an open access, online demo for monolingual and multilingual automatic term extraction from parallel corpora. The monolingual term extraction is based on a recurrent neural network, with a supervised methodology that relies on pretrained embeddings. Candidate terms can be tagged in their original context and there is no need for a large corpus, as the methodology will work even for single sentences. With the bilingual term extraction from parallel corpora, potentially equivalent candidate term pairs are extracted from translation memories and manual annotation of the results shows that good equivalents are found for most candidate terms. Accompanying the release of the demo is an updated version of the ACTER Annotated Corpora for Term Extraction Research (version 1.5).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Term ExtractionTranslation

Similar Papers 제목 키워드 기반

Bilingual Streaming ASR with Grapheme units and Auxiliary Monolingual Loss

2023-08-11 · Mohammad Soleymanpour, Mahmoud Al Ismail, Fahimeh Bahmaninezhad, Kshitiz Kumar 외

We introduce a bilingual solution to support English as secondary locale for most primary locales in hybrid automatic speech recognition (ASR) settings. Our key developments constitute: (a) pronunciation lexicon with gra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+1

Unsupervised Word Mapping Using Structural Similarities in Monolingual Embeddings

2017-12-19 · TACL 2018 1 · Hanan Aldarmaki, Mahesh Mohan, Mona Diab

Most existing methods for automatic bilingual dictionary induction rely on prior alignments between the source and target languages, such as parallel corpora or seed dictionaries. For many language pairs, such supervised…

Word Embeddings

Anchor-based Bilingual Word Embeddings for Low-Resource Languages

2020-10-23 · ACL 2021 5 · Tobias Eder, Viktor Hangya, Alexander Fraser

Good quality monolingual word embeddings (MWEs) can be built for languages which have large amounts of unlabeled text. MWEs can be aligned to bilingual spaces using only a few thousand word translation pairs. For low res…

Bilingual Lexicon InductionCross-Lingual TransferTransfer LearningTranslation+3

Multi-Graph Decoding for Code-Switching ASR

2019-06-18 · Emre Yilmaz, Samuel Cohen, Xianghu Yue, David van Leeuwen 외

In the FAME! Project, a code-switching (CS) automatic speech recognition (ASR) system for Frisian-Dutch speech is developed that can accurately transcribe the local broadcaster's bilingual archives with CS speech. This a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognition+1

Towards producing bilingual lexica from monolingual corpora

2016-05-01 · LREC 2016 5 · Jingyi Han, N{\'u}ria Bel

Bilingual lexica are the basis for many cross-lingual natural language processing tasks. Recent works have shown success in learning bilingual dictionary by taking advantages of comparable corpora and a diverse set of si…

Machine TranslationTranslation