ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue Summarization
Dialogue agents have been receiving increasing attention for years, and this trend has been further boosted by the recent progress of large language models (LLMs). Stance detection and dialogue summarization are two core tasks of dialogue agents in application scenarios that involve argumentative dialogues. However, research on these tasks is limited by the insufficiency of public datasets, especially for non-English languages. To address this language resource gap in Chinese, we present ORCHID (Oral Chinese Debate), the first Chinese dataset for benchmarking target-independent stance detection and debate summarization. Our dataset consists of 1,218 real-world debates that were conducted in Chinese on 476 unique topics, containing 2,436 stance-specific summaries and 14,133 fully annotated utterances. Besides providing a versatile testbed for future research, we also conduct an empirical study on the dataset and propose an integrated task. The results show the challenging nature of the dataset and suggest a potential of incorporating stance detection in summarization for argumentative dialogue.
Code (1)
Tasks
BenchmarkingStance DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Orchid2024: A cultivar-level dataset and methodology for fine-grained classification of Chinese Cymbidium Orchids
The authors dedicated over a year to collecting a cultivar image dataset for Chinese Cymbidium orchids named Orchid2024. This dataset contains over 150,000 images spanning 1,275 different categories, involving visits to …
Fine-Grained Image ClassificationImage Classificationparameter-efficient fine-tuningR-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory
We present R-Debater, an agentic framework for generating multi-turn debates built on argumentative memory. Grounded in rhetoric and memory studies, the system views debate as a process of recalling and adapting prior ar…
Constructing a Chinese---Japanese Parallel Corpus from Wikipedia
Parallel corpora are crucial for statistical machine translation (SMT). However, they are quite scarce for most language pairs, such as Chinese―Japanese. As comparable corpora are far more available, many studies have be…
Machine TranslationSentenceTranslationTarget-based Sentiment Annotation in Chinese Financial News
This paper presents the design and construction of a large-scale target-based sentiment annotation corpus on Chinese financial news text. Different from the most existing paragraph/document-based annotation corpus, in th…
Sentiment AnalysisIdentification of Orchid Species Using Content-Based Flower Image Retrieval
In this paper, we developed the system for recognizing the orchid species by using the images of flower. We used MSRM (Maximal Similarity based on Region Merging) method for segmenting the flower object from the backgrou…
Image RetrievalRetrieval