paper-with-me

Papers

MuLMS: A Multi-Layer Annotated Text Corpus for Information Extraction in the Materials Science Domain

2023-10-24 · Timo Pierre Schrader, Matteo Finco, Stefan Grünewald, Felix Hildebrand, Annemarie Friedrich

Keeping track of all relevant recent publications and experimental results for a research area is a challenging task. Prior work has demonstrated the efficacy of information extraction models in various scientific areas. Recently, several datasets have been released for the yet understudied materials science domain. However, these datasets focus on sub-problems such as parsing synthesis procedures or on sub-domains, e.g., solid oxide fuel cells. In this resource paper, we present MuLMS, a new dataset of 50 open-access articles, spanning seven sub-domains of materials science. The corpus has been annotated by domain experts with several layers ranging from named entities over relations to frame structures. We present competitive neural models for all tasks and demonstrate that multi-task training with existing related resources leads to benefits.

📄 PDF Abstract BibTeX arXiv:2310.15569

Code (1)

boschresearch/mulms-wiesp2023 공식 구현 pytorch

Tasks

Articles

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MuLMS-AZ: An Argumentative Zoning Dataset for the Materials Science Domain

2023-07-05 · Timo Pierre Schrader, Teresa Bürkle, Sophie Henning, Sherry Tan 외

Scientific publications follow conventionalized rhetorical structures. Classifying the Argumentative Zone (AZ), e.g., identifying whether a sentence states a Motivation, a Result or Background information, has been propo…

ArticlesSentence

The N2 corpus: A semantically annotated collection of Islamist extremist stories

2014-05-01 · LREC 2014 5 · Mark Finlayson, Jeffry Halverson, Steven Corman

We describe the N2 (Narrative Networks) Corpus, a new language resource. The corpus is unique in three important ways. First, every text in the corpus is a story, which is in contrast to other language resources that may…

Translation

A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations

2018-05-01 · LREC 2018 5 · Elena Musi, Manfred Stede, Leonard Kriese, Smar Muresan 외
Text Generation

Deriving a PropBank Corpus from Parallel FrameNet and UD Corpora

2020-05-01 · LREC 2020 5 · Normunds Gruzitis, Roberts Dar{\c{g}}is, Laura Rituma, Gunta Ne{\v{s}}pore-B{\=e}rzkalne 외

We propose an approach for generating an accurate and consistent PropBank-annotated corpus, given a FrameNet-annotated corpus which has an underlying dependency annotation layer, namely, a parallel Universal Dependencies…

FRACCO: A gold-standard annotated corpus of oncological entities with ICD-O-3.1 normalisation

2025-10-13 · Johann Pignat, Milena Vucetic, Christophe Gaudet-Blavignac, Jamil Zaghir 외 arxiv

Developing natural language processing tools for clinical text requires annotated datasets, yet French oncology resources remain scarce. We present FRACCO (FRench Annotated Corpus for Clinical Oncology) an expert-annotat…