paper-with-me

Papers

I Wish I Would Have Loved This One, But I Didn’t – A Multilingual Dataset for Counterfactual Detection in Product Review

2021-11-01 · EMNLP 2021 11 · James O’Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, Danushka Bollegala

Counterfactual statements describe events that did not or cannot take place. We consider the problem of counterfactual detection (CFD) in product reviews. For this purpose, we annotate a multilingual CFD dataset from Amazon product reviews covering counterfactual statements written in English, German, and Japanese languages. The dataset is unique as it contains counterfactuals in multiple languages, covers a new application area of e-commerce reviews, and provides high quality professional annotations. We train CFD models using different text representation methods and classifiers. We find that these models are robust against the selectional biases introduced due to cue phrase-based sentence selection. Moreover, our CFD dataset is compatible with prior datasets and can be merged to learn accurate CFD models. Applying machine translation on English counterfactual examples to create multilingual data performs poorly, demonstrating the language-specificity of this problem, which has been ignored so far.

📄 PDF Abstract BibTeX

Code (1)

amazon-research/amazon-multilingual-counterfactual-dataset 공식 구현

Tasks

counterfactualCounterfactual DetectionMachine TranslationSentenceSpecificityTranslation

Methods 이 논문이 사용한 방법론

Counterfactuals 설명 없음

Similar Papers 제목 키워드 기반

I Wish I Would Have Loved This One, But I Didn't -- A Multilingual Dataset for Counterfactual Detection in Product Reviews

2021-04-14 · James O'Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota 외

Counterfactual statements describe events that did not or cannot take place. We consider the problem of counterfactual detection (CFD) in product reviews. For this purpose, we annotate a multilingual CFD dataset from Ama…

counterfactualCounterfactual DetectionMachine TranslationSentence+2

I Wish I Didn't Say That! Analyzing and Predicting Deleted Messages in Twitter

2013-05-14 · Sasa Petrovic, Miles Osborne, Victor Lavrenko

Twitter has become a major source of data for social media researchers. One important aspect of Twitter not previously considered are {\em deletions} -- removal of tweets from the stream. Deletions can be due to a multit…

LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation

2021-10-17 · Junjue Wang, Zhuo Zheng, Ailong Ma, Xiaoyan Lu 외

Deep learning approaches have shown promising results in remote sensing high spatial resolution (HSR) land-cover mapping. However, urban and rural scenes can show completely different geographical landscapes, and the ina…

Domain AdaptationPseudo LabelSegmentationSemantic Segmentation+1

Task Proposal: The TL;DR Challenge

2018-11-01 · WS 2018 11 · Shahbaz Syed, Michael V{\"o}lske, Martin Potthast, Nedim Lipka 외

The TL;DR challenge fosters research in abstractive summarization of informal text, the largest and fastest-growing source of textual data on the web, which has been overlooked by summarization research so far. The chall…

Abstractive Text SummarizationInformation RetrievalText GenerationText Summarization

A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic

2025-07-07 · Juan Moreno Gonzalez, Bashar Alhafni, Nizar Habash arxiv

Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for J…

Machine Translation