paper-with-me

홈 › Papers

TagNText: A parallel corpus for the induction of resource-specific non-taxonomical relations from tagged images

2014-05-01 · LREC 2014 5 · Theodosia Togia, Ann Copestake

When producing textual descriptions, humans express propositions regarding an object; but what do they express when annotating a document with simple tags? To answer this question, we have studied what users of tagging systems would have said if they were to describe a resource with fully fledged text. In particular, our work attempts to answer the following questions: if users were to use full descriptions, would their current tags be words present in these hypothetical sentences? If yes, what kind of language would connect these words? Such questions, although central to the problem of extracting binary relations between tags, have been sidestepped in the existing literature, which has focused on a small subset of possible inter-tag relations, namely hierarchical ones (e.g. {`}car{''} --is-a-- {}vehicle{''}), as opposed to non-taxonomical relations (e.g. {}woman{''} --wears-- {`}hat{''}). TagNText is the first attempt to construct a parallel corpus of tags and textual descriptions with respect to particular resources. The corpus provides enough data for the researcher to gain an insight into the nature of underlying relations, as well as the tools and methodology for constructing larger-scale parallel corpora that can aid non-taxonomical relation extraction.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Relation ExtractionTAG

Similar Papers 제목 키워드 기반

JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

2019-07-01 · ACL 2019 7 · {\v{Z}}eljko Agi{\'c}, Ivan Vuli{\'c}

Viable cross-lingual transfer critically depends on the availability of parallel texts. Shortage of such resources imposes a development and evaluation bottleneck in multilingual processing. We introduce JW300, a paralle…

Cross-Lingual Transfer

Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages

2022-01-28 · LREC 2022 6 · Silvia Severini, Ayyoob Imani, Philipp Dufter, Hinrich Schütze

Parallel corpora are ideal for extracting a multilingual named entity (MNE) resource, i.e., a dataset of names translated into multiple languages. Prior work on extracting MNE datasets from parallel corpora required reso…

Bilingual Lexicon InductionTransliteration

The Trilingual ALLEGRA Corpus: Presentation and Possible Use for Lexicon Induction

2012-05-01 · LREC 2012 5 · Yves Scherrer, Bruno Cartoni

In this paper, we present a trilingual parallel corpus for German, Italian and Romansh, a Swiss minority language spoken in the canton of Grisons. The corpus called ALLEGRA contains press releases automatically gathered …

Sentence

Improving Translation of Out Of Vocabulary Words using Bilingual Lexicon Induction in Low-Resource Machine Translation

2022-09-01 · AMTA 2022 9 · Jonas Waldendorf, Alexandra Birch, Barry Hadow, Antonio Valerio Micele Barone

Dictionary-based data augmentation techniques have been used in the field of domain adaptation to learn words that do not appear in the parallel training data of a machine translation model. These techniques strive to le…

Bilingual Lexicon InductionData AugmentationDomain AdaptationMachine Translation+3

Multilingual Lexicon Bootstrapping - Improving a Lexicon Induction System Using a Parallel Corpus

2013-10-01 · IJCNLP 2013 10 · Patrick Ziering, Lonneke van der Plas, Hinrich Sch{\"u}tze
Coreference ResolutionWord AlignmentWord Sense Disambiguation