Testing the Generalization Power of Neural Network Models Across NLI Benchmarks
Neural network models have been very successful in natural language inference, with the best models reaching 90% accuracy in some benchmarks. However, the success of these models turns out to be largely benchmark specific. We show that models trained on a natural language inference dataset drawn from one benchmark fail to perform well in others, even if the notion of inference assumed in these benchmarks is the same or similar. We train six high performing neural network models on different datasets and show that each one of these has problems of generalizing when we replace the original test set with a test set taken from another corpus designed for the same task. In light of these results, we argue that most of the current neural network models are not able to generalize well in the task of natural language inference. We find that using large pre-trained language models helps with transfer learning when the datasets are similar enough. Our results also highlight that the current NLI datasets do not cover the different nuances of inference extensively enough.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language InferenceTransfer LearningSimilar Papers 제목 키워드 기반
Towards Effective Semantic OOD Detection in Unseen Domains: A Domain Generalization Perspective
Two prevalent types of distributional shifts in machine learning are the covariate shift (as observed across different domains) and the semantic shift (as seen across different classes). Traditional OOD detection techniq…
Domain GeneralizationAnyBody: A Benchmark Suite for Cross-Embodiment Manipulation
Generalizing control policies to novel embodiments remains a fundamental challenge in enabling scalable and transferable learning in robotics. While prior works have explored this in locomotion, a systematic study in the…
Zero-shot GeneralizationGeneralization Ability of Feature-based Performance Prediction Models: A Statistical Analysis across Benchmarks
This study examines the generalization ability of algorithm performance prediction models across various benchmark suites. Comparing the statistical similarity between the problem collections with the accuracy of perform…
Best sources forward: domain generalization through source-specific nets
A long standing problem in visual object categorization is the ability of algorithms to generalize across different testing conditions. The problem has been formalized as a covariate shift among the probability distribut…
Domain AdaptationDomain GeneralizationObject CategorizationStress-Testing Multimodal Foundation Models for Crystallographic Reasoning
Evaluating foundation models for crystallographic reasoning requires benchmarks that isolate generalization behavior while enforcing physical constraints. This work introduces a multiscale multicrystal dataset with two p…
HallucinationSpatial Interpolation