paper-with-me

홈 › Papers

When More Data Hurts: A Troubling Quirk in Developing Broad-Coverage Natural Language Understanding Systems

2022-05-24 · Elias Stengel-Eskin, Emmanouil Antonios Platanios, Adam Pauls, Sam Thomson, Hao Fang, Benjamin Van Durme, Jason Eisner, Yu Su

In natural language understanding (NLU) production systems, users' evolving needs necessitate the addition of new features over time, indexed by new symbols added to the meaning representation space. This requires additional training data and results in ever-growing datasets. We present the first systematic investigation of this incremental symbol learning scenario. Our analysis reveals a troubling quirk in building broad-coverage NLU systems: as the training dataset grows, performance on the new symbol often decreases if we do not accordingly increase its training data. This suggests that it becomes more difficult to learn new symbols with a larger training dataset. We show that this trend holds for multiple mainstream models on two common NLU tasks: intent recognition and semantic parsing. Rejecting class imbalance as the sole culprit, we reveal that the trend is closely associated with an effect we call source signal dilution, where strong lexical cues for the new symbol become diluted as the training dataset grows. Selectively dropping training examples to prevent dilution often reverses the trend, showing the over-reliance of mainstream neural NLU models on simple lexical cues. Code, models, and data are available at https://aka.ms/nlu-incremental-symbol-learning

📄 PDF Abstract BibTeX arXiv:2205.12228

Code (1)

esteng/calibration_miso pytorch

Tasks

Intent RecognitionNatural Language UnderstandingSemantic Parsing

Similar Papers 제목 키워드 기반

When More Data Hurts: A Troubling Quirk in Developing Broad-Coverage Natural Language Understanding Systems

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In natural language understanding (NLU) production systems, the end users' evolving needs necessitate the addition of new abilities, indexed by discrete symbols, requiring additional training data and resulting in dynami…

Intent RecognitionNatural Language UnderstandingSemantic Parsing

Quirk or Palmer: A Comparative Study of Modal Verb Frameworks with Annotated Datasets

2022-12-20 · Risako Owan, Maria Gini, Dongyeop Kang

Modal verbs, such as "can", "may", and "must", are commonly used in daily communication to convey the speaker's perspective related to the likelihood and/or mode of the proposition. They can differ greatly in meaning dep…

Natural Language UnderstandingSentence

QuIRK: Quantum-Inspired Re-uploading KAN

2025-10-09 · Vinayak Sharma, Ashish Padhy, Lord Sen, Vijay Jagdish Karanjkar 외 arxiv

Kolmogorov-Arnold Networks or KANs have shown the ability to outperform classical Deep Neural Networks, while using far fewer trainable parameters for regression problems on scientific domains. Even more powerful has bee…

When Covariate-shifted Data Augmentation Increases Test Error And How to Fix It

2019-09-25 · Sang Michael Xie*, Aditi Raghunathan*, Fanny Yang, John C. Duchi 외

Empirically, data augmentation sometimes improves and sometimes hurts test error, even when only adding points with labels from the true conditional distribution that the hypothesis class is expressive enough to fit. In…

Data Augmentationregression

Eliciting Latent Knowledge from Quirky Language Models

2023-12-02 · Alex Mallen, Madeline Brumley, Julia Kharchenko, Nora Belrose

Eliciting Latent Knowledge (ELK) aims to find patterns in a capable neural network's activations that robustly track the true state of the world, especially in hard-to-verify cases where the model's output is untrusted. …

Anomaly DetectionMath