STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs
Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Developmentally motivated emergence of compositional communication via template transfer
This paper explores a novel approach to achieving emergent compositional communication in multi-agent systems. We propose a training regime implementing template transfer, the idea of carrying over learned biases across …
Zero-shot GeneralizationUnfinished Architectures: A Perspective from Artificial Intelligence
Unfinished buildings are a constant throughout the history of architecture and have given rise to intense debates on the opportuneness of their completion, in addition to offering alibis for theorizing about the composit…
CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation
Metaverse platforms rely on creator-driven marketplaces where avatars are assembled from discrete, taxonomy-labeled 3D assets (e.g., tops, bottoms, shoes, accessories) under strict category and topology constraints. Whil…
FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model
In this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction. Unlike other modalities such as vision or text, speech has unique temporal dependencies, making previou…
Emotion RecognitionLanguage ModelingLanguage ModellingQuestion Answering+1BeSTAD: Behavior-Aware Spatio-Temporal Anomaly Detection for Human Mobility Data
Traditional anomaly detection in human mobility has primarily focused on trajectory-level analysis, identifying statistical outliers or spatiotemporal inconsistencies across aggregated movement traces. However, detecting…
Anomaly Detection