paper-with-me

홈 › Papers

We're Calling an Intervention: Exploring Fundamental Hurdles in Adapting Language Models to Nonstandard Text

2024-04-10 · Aarohi Srivastava, David Chiang

We present a suite of experiments that allow us to understand the underlying challenges of language model adaptation to nonstandard text. We do so by designing interventions that approximate core features of user-generated text and their interactions with existing biases of language models. Applying our interventions during language model adaptation to nonstandard text variations, we gain important insights into when such adaptation is successful, as well as the aspects of text variation and noise that are particularly difficult for language models to handle. For instance, on text with character-level variation, out-of-the-box performance improves even with a few additional training examples but approaches a plateau, suggesting that more data is not the solution. In contrast, on text with variation involving new words or meanings, far more data is needed, but it leads to a massive breakthrough in performance. Our findings reveal that existing models lack the necessary infrastructure to handle diverse forms of nonstandard text, guiding the development of more resilient language modeling techniques. We make the code for our interventions, which can be applied to any English text data, publicly available.

📄 PDF Abstract BibTeX arXiv:2404.07304

Code (1)

aarsri/interventions-linguistic-variation 공식 구현

Tasks

Language ModelingLanguage ModellingText-VariationTransfer Learning

Similar Papers 제목 키워드 기반

Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training

2025-02-10 · Yuchen Zhuang, Jingfeng Yang, Haoming Jiang, Xin Liu 외

Due to the scarcity of agent-oriented pre-training data, LLM-based autonomous agents typically rely on complex prompting or extensive fine-tuning, which often fails to introduce new capabilities while preserving strong g…

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

2026-02-04 · Youjin Wang, Run Zhou, Yingjie Ma, Rong Fu 외 arxiv

Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy to deploy but often fragile under distribution shift and strict parsers, while …

parameter-efficient fine-tuning

Querying Databases with Function Calling

2025-01-23 · Connor Shorten, Charles Pierse, Thomas Benjamin Smith, Karel D'Oosterlinck 외

The capabilities of Large Language Models (LLMs) are rapidly accelerating largely thanks to their integration with external tools. Querying databases is among the most effective of these integrations, enabling LLMs to ac…

Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning

2025-09-25 · Asim Ersoy, Enes Altinisik, Husrev Taha Sencar, Kareem Darwish arxiv

Tool calling is a critical capability that allows Large Language Models (LLMs) to interact with external systems, significantly expanding their utility. However, research and resources for tool calling are predominantly …

Cross-Lingual Transfer

ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario

2025-01-17 · Lucen Zhong, Zhengxiao Du, Xiaohan Zhang, Haiyi Hu 외

Enhancing large language models (LLMs) with real-time APIs can help generate more accurate and up-to-date responses. However, evaluating the function calling abilities of LLMs in real-world scenarios remains under-explor…