paper-with-me

Sentence Completion

1개 벤치마크 · 논문 100편 · 이 태스크의 논문 보기 →

Benchmarks

HellaSwag

결과 89개

Most implemented

Language Models are Few-Shot Learners

2020-05-28 · 구현 67개

Papers

Evaluating Developmental Cognition Capabilities of LLMs

2026-05-08 · Xiao Xiao, Hayoun Noh, Mar Gonzalez-Franco arxiv

Conversational AI is increasingly personalized around users' preferences, histories, goals, and knowledge, but much less around how users interpret and take up model outputs to construct and understand their reality. We …

Sentence Completion

Stroke Lesions as a Rosetta Stone for Language Model Interpretability

2026-02-03 · Julius Fridriksson, Roger D. Newman-Norlund, Saeed Ahmadi, Regan Willis 외 arxiv

Large language models (LLMs) have achieved remarkable capabilities, yet methods to verify which model components are truly necessary for language function remain limited. Current interpretability approaches rely on inter…

Sentence Completion

QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks

2026-01-28 · Mae Sosto, Delfina Sol Martinez Pandiani, Laura Hollink arxiv

This paper examines how Large Language Models (LLMs) reproduce societal norms, particularly heterocisnormativity, and how these norms translate into measurable biases in their text generations. We investigate whether exp…

Sentence Completion

Decoding Workload and Agreement From EEG During Spoken Dialogue With Conversational AI

2026-01-09 · Lucija Mihić Zidar, Philipp Wicke, Praneel Bhatia, Rosa Lutz 외 arxiv

Passive brain-computer interfaces offer a potential source of implicit feedback for alignment of large language models, but most mental state decoding has been done in controlled tasks. This paper investigates whether es…

Sentence Completion

FIBER: A Multilingual Evaluation Resource for Factual Inference Bias

2025-12-11 · Evren Ayberk Munis, Deniz Yılmaz, Arianna Muti, Çağrı Toraman arxiv

Large language models are widely used across domains, yet there are concerns about their factual reliability and biases. Factual knowledge probing offers a systematic means to evaluate these aspects. Most existing benchm…

Sentence Completion

Building Domain-Specific Small Language Models via Guided Data Generation

2025-11-23 · Aman Kumar, Ekant Muljibhai Amin, Xian Yeow Lee, Lasitha Vidyaratne 외 arxiv

Large Language Models (LLMs) have shown remarkable success in supporting a wide range of knowledge-intensive tasks. In specialized domains, there is growing interest in leveraging LLMs to assist subject matter experts wi…

Synthetic Data GenerationSentence CompletionQuestion AnsweringDomain Adaptation

전체 100편 보기 →