paper-with-me

Papers

STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs

2026-04-20 · Sungeun An, Swanand Ravindra Kadhe, Shailja Thakur, Chad DeLuca, Hima Patel arxiv

Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.

📄 PDF Abstract BibTeX arXiv:2604.18177

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Developmentally motivated emergence of compositional communication via template transfer

2019-10-04 · Tomasz Korbak, Julian Zubek, Łukasz Kuciński, Piotr Miłoś 외

This paper explores a novel approach to achieving emergent compositional communication in multi-agent systems. We propose a training regime implementing template transfer, the idea of carrying over learned biases across …

Zero-shot Generalization

Unfinished Architectures: A Perspective from Artificial Intelligence

2023-03-03 · Elena Merino-Gómez, Pedro Reviriego, Fernando Moral

Unfinished buildings are a constant throughout the history of architecture and have given rise to intense debates on the opportuneness of their completion, in addition to offering alibis for theorizing about the composit…

CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation

2026-05-18 · Rajeev Goel, Jason Ding, Phani Harish Wajjala, Pavan Turaga 외 arxiv

Metaverse platforms rely on creator-driven marketplaces where avatars are assembled from discrete, taxonomy-labeled 3D assets (e.g., tops, bottoms, shoes, accessories) under strict category and topology constraints. Whil…

FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model

2024-10-03 · Yichen Lu, Jiaqi Song, Chao-Han Huck Yang, Shinji Watanabe

In this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction. Unlike other modalities such as vision or text, speech has unique temporal dependencies, making previou…

Emotion RecognitionLanguage ModelingLanguage ModellingQuestion Answering+1

BeSTAD: Behavior-Aware Spatio-Temporal Anomaly Detection for Human Mobility Data

2025-10-14 · Junyi Xie, Jina Kim, Yao-Yi Chiang, Lingyi Zhao 외 arxiv

Traditional anomaly detection in human mobility has primarily focused on trajectory-level analysis, identifying statistical outliers or spatiotemporal inconsistencies across aggregated movement traces. However, detecting…

Anomaly Detection