Rethinking Coherence Modeling: Synthetic vs. Downstream Tasks
Although coherence modeling has come a long way in developing novel models, their evaluation on downstream applications for which they are purportedly developed has largely been neglected. With the advancements made by neural approaches in applications such as machine translation (MT), summarization and dialog systems, the need for coherence evaluation of these tasks is now more crucial than ever. However, coherence models are typically evaluated only on synthetic tasks, which may not be representative of their performance in downstream applications. To investigate how representative the synthetic tasks are of downstream use cases, we conduct experiments on benchmarking well-known traditional and neural coherence models on synthetic sentence ordering tasks, and contrast this with their performance on three downstream applications: coherence evaluation for MT and summarization, and next utterance prediction in retrieval-based dialog. Our results demonstrate a weak correlation between the model performances in the synthetic tasks and the downstream applications, {motivating alternate training and evaluation methods for coherence models.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingCoherence EvaluationLanguage ModellingMachine TranslationRetrievalSentenceSentence OrderingText SummarizationTranslationSimilar Papers 제목 키워드 기반
Rethinking Self-Supervision Objectives for Generalizable Coherence Modeling
Given the claims of improved text generation quality across various pre-trained neural models, we consider the coherence evaluation of machine generated text to be one of the principal applications of coherence models th…
Coherence EvaluationContrastive LearningText GenerationGenerative Modeling of Complex-Valued Brain MRI Data
Objective. Standard Magnetic Resonance Imaging (MRI) reconstruction pipelines discard phase information captured during acquisition, despite evidence that it encodes tissue properties relevant to tumor diagnosis. Current…
Synthetic Data GenerationIncremental Neural Lexical Coherence Modeling
Pretrained language models, neural models pretrained on massive amounts of data, have established the state of the art in a range of NLP tasks. They are based on a modern machine-learning technique, the Transformer which…
Language ModelingLanguage ModellingBBScore: A Brownian Bridge Based Metric for Assessing Text Coherence
Measuring the coherence of text is a vital aspect of evaluating the quality of written content. Recent advancements in neural coherence modeling have demonstrated their efficacy in capturing entity coreference and discou…
Coherence EvaluationRethinking Driving World Model as Synthetic Data Generator for Perception Tasks
Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. Existing methods primarily focus on metrics related to generation quality and controllability. H…
Synthetic Data GenerationAutonomous Driving