The Kyutech corpus and topic segmentation using a combined method
Summarization of multi-party conversation is one of the important tasks in natural language processing. In this paper, we explain a Japanese corpus and a topic segmentation task. To the best of our knowledge, the corpus is the first Japanese corpus annotated for summarization tasks and freely available to anyone. We call it {``}the Kyutech corpus.{''} The task of the corpus is a decision-making task with four participants and it contains utterances with time information, topic segmentation and reference summaries. As a case study for the corpus, we describe a method combined with LCSeg and TopicTiling for a topic segmentation task. We discuss the effectiveness and the problems of the combined method through the experiment with the Kyutech corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingSegmentationSimilar Papers 제목 키워드 기반
Annotation and Analysis of Extractive Summaries for the Kyutech Corpus
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
Topic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-lev…
Discourse ParsingInformation RetrievalRetrievalSentenceSubtopic Annotation in a Corpus of News Texts: Steps Towards Automatic Subtopic Segmentation
Improving Unsupervised Dialogue Topic Segmentation with Utterance-Pair Coherence Scoring
Dialogue topic segmentation is critical in several dialogue modeling problems. However, popular unsupervised approaches only exploit surface features in assessing topical coherence among utterances. In this work, we addr…
SegmentationText Segmentation using Named Entity Recognition and Co-reference Resolution in English and Greek Texts
In this paper we examine the benefit of performing named entity recognition (NER) and co-reference resolution to an English and a Greek corpus used for text segmentation. The aim here is to examine whether the combinatio…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2