paper-with-me

홈 › Papers

RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced Labour

2022-05-05 · LREC 2022 6 · Erick Mendez Guzman, Viktor Schlegel, Riza Batista-Navarro

Forced labour is the most common type of modern slavery, and it is increasingly gaining the attention of the research and social community. Recent studies suggest that artificial intelligence (AI) holds immense potential for augmenting anti-slavery action. However, AI tools need to be developed transparently in cooperation with different stakeholders. Such tools are contingent on the availability and access to domain-specific data, which are scarce due to the near-invisible nature of forced labour. To the best of our knowledge, this paper presents the first openly accessible English corpus annotated for multi-class and multi-label forced labour detection. The corpus consists of 989 news articles retrieved from specialised data sources and annotated according to risk indicators defined by the International Labour Organization (ILO). Each news article was annotated for two aspects: (1) indicators of forced labour as classification labels and (2) snippets of the text that justify labelling decisions. We hope that our data set can help promote research on explainability for multi-class and multi-label text classification. In this work, we explain our process for collecting the data underpinning the proposed corpus, describe our annotation guidelines and present some statistical analysis of its content. Finally, we summarise the results of baseline experiments based on different variants of the Bidirectional Encoder Representation from Transformer (BERT) model.

📄 PDF Abstract BibTeX arXiv:2205.02684

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesMulti Label Text ClassificationMulti-Label Text Classificationtext-classificationText Classification

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Towards Constructing a Corpus for Studying the Effects of Treatments and Substances Reported in PubMed Abstracts

2019-12-04 · Evgeni Stefchov, Galia Angelova, Preslav Nakov

We present the construction of an annotated corpus of PubMed abstracts reporting about positive, negative or neutral effects of treatments or substances. Our ultimate goal is to annotate one sentence (rationale) for each…

Sentencetext-classificationText Classification

BiasLab: Toward Explainable Political Bias Detection with Dual-Axis Annotations and Rationale Indicators

2025-05-21 · KMA Solaiman

We present BiasLab, a dataset of 300 political news articles annotated for perceived ideological bias. These articles were selected from a curated 900-document pool covering diverse political events and source biases. Ea…

ArticlesBias Detection

Semi-Supervised Iterative Approach for Domain-Specific Complaint Detection in Social Media

2020-07-01 · WS 2020 7 · Akash Gautam, Debanjan Mahata, Rakesh Gosangi, Rajiv Ratn Shah

In this paper, we present a semi-supervised bootstrapping approach to detect product or service related complaints in social media. Our approach begins with a small collection of annotated samples which are used to ident…

The N2 corpus: A semantically annotated collection of Islamist extremist stories

2014-05-01 · LREC 2014 5 · Mark Finlayson, Jeffry Halverson, Steven Corman

We describe the N2 (Narrative Networks) Corpus, a new language resource. The corpus is unique in three important ways. First, every text in the corpus is a story, which is in contrast to other language resources that may…

Translation

POMELO: Medline corpus with manually annotated food-drug interactions

2017-09-01 · RANLP 2017 9 · Thierry Hamon, Vincent Tabanou, Fleur Mougin, Natalia Grabar 외

When patients take more than one medication, they may be at risk of drug interactions, which means that a given drug can cause unexpected effects when taken in combination with other drugs. Similar effects may occur when…