paper-with-me

홈 › Papers

Understood in Translation, Transformers for Domain Understanding

2020-12-18 · Dimitrios Christofidellis, Matteo Manica, Leonidas Georgopoulos, Hans Vandierendonck

Knowledge acquisition is the essential first step of any Knowledge Graph (KG) application. This knowledge can be extracted from a given corpus (KG generation process) or specified from an existing KG (KG specification process). Focusing on domain specific solutions, knowledge acquisition is a labor intensive task usually orchestrated and supervised by subject matter experts. Specifically, the domain of interest is usually manually defined and then the needed generation or extraction tools are utilized to produce the KG. Herein, we propose a supervised machine learning method, based on Transformers, for domain definition of a corpus. We argue why such automated definition of the domain's structure is beneficial both in terms of construction time and quality of the generated graph. The proposed method is extensively validated on three public datasets (WebNLG, NYT and DocRED) by comparing it with two reference methods based on CNNs and RNNs models. The evaluation shows the efficiency of our model in this task. Focusing on scientific document understanding, we present a new health domain dataset based on publications extracted from PubMed and we successfully utilize our method on this. Lastly, we demonstrate how this work lays the foundation for fully automated and unsupervised KG generation.

📄 PDF Abstract BibTeX arXiv:2012.10271

Code (1)

christofid/DomainUnderstanding 공식 구현 pytorch

Tasks

document understandingTranslation

Similar Papers 제목 키워드 기반

Transformers are Deep Infinite-Dimensional Non-Mercer Binary Kernel Machines

2021-06-02 · Matthew A. Wright, Joseph E. Gonzalez

Despite their ubiquity in core AI fields like natural language processing, the mechanics of deep attention-based neural networks like the Transformer model are not fully understood. In this article, we present a new pers…

Deep Attention

BERTuit: Understanding Spanish language in Twitter through a native transformer

2022-04-07 · Javier Huertas-Tato, Alejandro Martin, David Camacho

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models enc…

Misinformation

How Do Transformers Learn Topic Structure: Towards a Mechanistic Understanding

2023-03-07 · Yuchen Li, Yuanzhi Li, Andrej Risteski

While the successes of transformers across many domains are indisputable, accurate understanding of the learning mechanics is still largely lacking. Their capabilities have been probed on benchmarks which include a varie…

Translation Transformers Rediscover Inherent Data Domains

2021-09-16 · WMT (EMNLP) 2021 11 · Maksym Del, Elizaveta Korotkova, Mark Fishel

Many works proposed methods to improve the performance of Neural Machine Translation (NMT) models in a domain/multi-domain adaptation scenario. However, an understanding of how NMT baselines represent text domain informa…

ClusteringDomain AdaptationMachine TranslationNMT+3

Distill, Adapt, Distill: Training Small, In-Domain Models for Neural Machine Translation

2020-03-05 · WS 2020 7 · Mitchell A. Gordon, Kevin Duh

We explore best practices for training small, memory efficient machine translation models with sequence-level knowledge distillation in the domain adaptation setting. While both domain adaptation and knowledge distillati…

Domain AdaptationKnowledge DistillationMachine TranslationTranslation