paper-with-me

홈 › Papers

What am I missing here?: Evaluating Large Language Models for Masked Sentence Prediction

2025-08-11 · Charlie Wyatt, Aditya Joshi, Flora Salim arxiv

Transformer-based models primarily rely on Next Token Prediction (NTP), which predicts the next token in a sequence based on the preceding context. However, NTP's focus on single-token prediction often limits a model's ability to plan ahead or maintain long-range coherence, raising questions about how well LLMs can predict longer contexts, such as full sentences within structured documents. While NTP encourages local fluency, it provides no explicit incentive to ensure global coherence across sentence boundaries-an essential skill for reconstructive or discursive tasks. To investigate this, we evaluate three commercial LLMs (GPT-4o, Claude 3.5 Sonnet, and Gemini 2.0 Flash) on Masked Sentence Prediction (MSP) - the task of infilling a randomly removed sentence - from three domains: ROCStories (narrative), Recipe1M (procedural), and Wikipedia (expository). We assess both fidelity (similarity to the original sentence) and cohesiveness (fit within the surrounding context). Our key finding reveals that commercial LLMs, despite their superlative performance in other tasks, are poor at predicting masked sentences in low-structured domains, highlighting a gap in current model capabilities.

📄 PDF Abstract BibTeX arXiv:2508.07702

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ontology Completion with Natural Language Inference and Concept Embeddings: An Analysis

2024-03-25 · Na Li, Thomas Bailleux, Zied Bouraoui, Steven Schockaert

We consider the problem of finding plausible knowledge that is missing from a given ontology, as a generalisation of the well-studied taxonomy expansion task. One line of work treats this task as a Natural Language Infer…

Natural Language InferenceTaxonomy Expansion

What We are Missing in Multimodal LLM Evaluation?

2026-06-24 · Po-han Li, Shenghui Chen, Sandeep Chinchali, Ufuk Topcu arxiv

Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced rapidly, evaluation of such models has not…

What is Missing from AI Post-Training AI: An Empirical Analysis

2026-08-19 · Joy Jia Yin Lim, Xin Huang, Hao Peng, Yaxi Lu 외 arxiv

Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that thi…

On the Winograd Schema: Situating Language Understanding in the Data-Information-Knowledge Continuum

2018-09-30 · Walid S. Saba

The Winograd Schema (WS) challenge, proposed as an al-ternative to the Turing Test, has become the new standard for evaluating progress in natural language understanding (NLU). In this paper we will not however be concer…

Natural Language Understanding

What Is Missing: Interpretable Ratings for Large Language Model Outputs

2026-02-17 · Nicholas Stranges, Yimin Yang arxiv

Current Large Language Model (LLM) preference learning methods such as Proximal Policy Optimization and Direct Preference Optimization learn from direct rankings or numerical ratings of model outputs, these rankings are …