paper-with-me

홈 › Papers

From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning

2025-05-20 · Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen

Instruction-tuned large language models (LLMs) have shown strong performance on a variety of tasks; however, generalizing from synthetic to human-authored instructions in grounded environments remains a challenge for them. In this work, we study generalization challenges in spatial grounding tasks where models interpret and translate instructions for building object arrangements on a $2.5$D grid. We fine-tune LLMs using only synthetic instructions and evaluate their performance on a benchmark dataset containing both synthetic and human-written instructions. Our results reveal that while models generalize well on simple tasks, their performance degrades significantly on more complex tasks. We present a detailed error analysis of the gaps in instruction generalization.

📄 PDF Abstract BibTeX arXiv:2505.14425

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Automatic Slide Updating with User-Defined Dynamic Templates and Natural Language Instructions

2026-04-20 · Kun Zhou, Jiakai He, Wenmian Yang, Zhensheng Wang 외 arxiv

Presentation slides are a primary medium for data-driven reporting, yet keeping complex, analytics-style decks up to date remains labor-intensive. Existing automation methods mostly follow fixed template filling and cann…

Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning

2024-06-13 · Janghoon Han, Changho Lee, Joongbo Shin, Stanley Jungkyu Choi 외

Instruction tuning has emerged as a powerful technique, significantly boosting zero-shot performance on unseen tasks. While recent work has explored cross-lingual generalization by applying instruction tuning to multilin…

Zero-shot Generalization

Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates

2024-08-22 · Yusuke Sakai, Adam Nohejl, Jiangnan Hang, Hidetaka Kamigaito 외

The natural language understanding (NLU) performance of large language models (LLMs) has been evaluated across various tasks and datasets. The existing evaluation methods, however, do not take into account the variance i…

Natural Language Understanding

Generalization in Online Reinforcement Learning for Mobile Agents

2026-03-08 · Li Gu, Zihuan Jiang, Zhixiang Chi, Huan Liu 외 arxiv

Graphical user interface (GUI)-based mobile agents automate digital tasks on mobile devices by interpreting natural-language instructions and interacting with the screen. While recent methods apply reinforcement learning…

Zero-shot GeneralizationReinforcement Learning

Finetuned Language Models Are Zero-Shot Learners

2021-09-03 · ICLR 2022 4 · Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 외

This paper explores a simple method for improving the zero-shot learning abilities of language models. We show that instruction tuning -- finetuning language models on a collection of tasks described via instructions -- …

ARCCommon Sense ReasoningCoreference ResolutionLanguage Modeling+8