Sentence Completion
1개 벤치마크 · 논문 100편 · 이 태스크의 논문 보기 →
Benchmarks
HellaSwag
Most implemented
Language Models are Few-Shot Learners
RoBERTa: A Robustly Optimized BERT Pretraining Approach
LLaMA: Open and Efficient Foundation Language Models
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Llama 2: Open Foundation and Fine-Tuned Chat Models
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Papers
Evaluating Developmental Cognition Capabilities of LLMs
Conversational AI is increasingly personalized around users' preferences, histories, goals, and knowledge, but much less around how users interpret and take up model outputs to construct and understand their reality. We …
Sentence CompletionStroke Lesions as a Rosetta Stone for Language Model Interpretability
Large language models (LLMs) have achieved remarkable capabilities, yet methods to verify which model components are truly necessary for language function remain limited. Current interpretability approaches rely on inter…
Sentence CompletionQueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks
This paper examines how Large Language Models (LLMs) reproduce societal norms, particularly heterocisnormativity, and how these norms translate into measurable biases in their text generations. We investigate whether exp…
Sentence CompletionDecoding Workload and Agreement From EEG During Spoken Dialogue With Conversational AI
Passive brain-computer interfaces offer a potential source of implicit feedback for alignment of large language models, but most mental state decoding has been done in controlled tasks. This paper investigates whether es…
Sentence CompletionFIBER: A Multilingual Evaluation Resource for Factual Inference Bias
Large language models are widely used across domains, yet there are concerns about their factual reliability and biases. Factual knowledge probing offers a systematic means to evaluate these aspects. Most existing benchm…
Sentence CompletionBuilding Domain-Specific Small Language Models via Guided Data Generation
Large Language Models (LLMs) have shown remarkable success in supporting a wide range of knowledge-intensive tasks. In specialized domains, there is growing interest in leveraging LLMs to assist subject matter experts wi…
Synthetic Data GenerationSentence CompletionQuestion AnsweringDomain Adaptation