Assessing Discourse Relations in Language Generation from GPT-2
Recent advances in NLP have been attributed to the emergence of large-scale pre-trained language models. GPT-2, in particular, is suited for generation tasks given its left-to-right language modeling objective, yet the linguistic quality of its generated text has largely remain unexplored. Our work takes a step in understanding GPT-2's outputs in terms of discourse coherence. We perform a comprehensive study on the validity of explicit discourse relations in GPT-2's outputs under both organic generation and fine-tuned scenarios. Results show GPT-2 does not always generate text containing valid discourse relations; nevertheless, its text is more aligned with human expectation in the fine-tuned scenario. We propose a decoupled strategy to mitigate these problems and highlight the importance of explicitly modeling discourse information.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingText GenerationvalidMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Graph-based Argument Quality Assessment
The paper presents a novel discourse-based approach to argument quality assessment defined as a graph classification task, where the depth of reasoning (argumentation) is evident from the number and type of detected disc…
Graph ClassificationTowards Automatic Detection of Narrative Structure
We present novel computational experiments using William Labov{'}s theory of narrative analysis. We describe his six elements of narrative structure and construct a new corpus based on his most recent work on narrative. …
Assessing Crosslingual Discourse Relations in Machine Translation
In an attempt to improve overall translation quality, there has been an increasing focus on integrating more linguistic elements into Machine Translation (MT). While significant progress has been achieved, especially rec…
Machine TranslationTranslationCORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
Multimodal Large Language Models (MLLMs) are renowned for their superior instruction-following and reasoning capabilities across diverse problem domains. However, existing benchmarks primarily focus on assessing factual …
Instruction FollowingUnlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence
Recent large language models (LLMs) have shown remarkable performance in aligning generated text with user intentions across various tasks. When it comes to long-form text generation, there has been a growing interest in…
ArticlesCoherence EvaluationFormText Generation