paper-with-me

Papers

Line-a-line: A Tool for Annotating Word-Alignments

2020-05-01 · LREC 2020 5 · Maria Skeppstedt, Magnus Ahltorp, Gunnar Eriksson, Rickard Domeij

We here describe line-a-line, a web-based tool for manual annotation of word-alignments in sentence-aligned parallel corpora. The graphical user interface, which builds on a design template from the Jigsaw system for investigative analysis, displays the words from each sentence pair that is to be annotated as elements in two vertical lists. An alignment between two words is annotated by drag-and-drop, i.e. by dragging an element from the left-hand list and dropping it on an element in the right-hand list. The tool indicates that two words are aligned by lines that connect them and by highlighting associated words when the mouse is hovered over them. Line-a-line uses the efmaral library for producing pre-annotated alignments, on which the user can base the manual annotation. The tool is mainly planned to be used on moderately under-resourced languages, for which resources in the form of parallel corpora are scarce. The automatic word-alignment functionality therefore also incorporates information derived from non-parallel resources, in the form of pre-trained multilingual word embeddings from the MUSE library.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Multilingual Word EmbeddingsSentenceWord AlignmentWord Embeddings

Methods 이 논문이 사용한 방법론

Jigsaw Jigsaw is a self-supervision approach that relies on jigsaw-like puzzles as the pretext task in order to learn image representations.

Similar Papers 제목 키워드 기반

A Web Tool for Building Parallel Corpora of Spoken and Sign Languages

2016-05-01 · LREC 2016 5 · Alex Becker, Fabio Kepler, C, Sara eias

In this paper we describe our work in building an online tool for manually annotating texts in any spoken language with SignWriting in any sign language. The existence of such tool will allow the creation of parallel cor…

Translation

A Study of Word-Classing for MT Reordering

2012-05-01 · LREC 2012 5 · Ananthakrishnan Ramanathan, Karthik Visweswariah

MT systems typically use parsers to help reorder constituents. However most languages do not have adequate treebank data to learn good parsers, and such training data is extremely time-consuming to annotate. Our earlier …

Dependency ParsingLanguage ModellingMachine TranslationPOS+1

Novel elicitation and annotation schemes for sentential and sub-sentential alignments of bitexts

2016-05-01 · LREC 2016 5 · Yong Xu, Fran{\c{c}}ois Yvon

Resources for evaluating sentence-level and word-level alignment algorithms are unsatisfactory. Regarding sentence alignments, the existing data is too scarce, especially when it comes to difficult bitexts, containing in…

Sentence

Saliency-driven Word Alignment Interpretation for Neural Machine Translation

2019-06-25 · WS 2019 8 · Shuoyang Ding, Hainan Xu, Philipp Koehn

Despite their original goal to jointly learn to align and translate, Neural Machine Translation (NMT) models, especially Transformer, are often perceived as not learning interpretable word alignments. In this paper, we s…

Machine TranslationNMTTranslationWord Alignment

A Corpus of Word-Aligned Asked and Anticipated Questions in a Virtual Patient Dialogue System

2016-05-01 · LREC 2016 5 · Ajda Gokcen, Evan Jaffe, Johnsey Erdmann, Michael White 외

We present a corpus of virtual patient dialogues to which we have added manually annotated gold standard word alignments. Since each question asked by a medical student in the dialogues is mapped to a canonical, anticipa…