PhysNLU: A Language Resource for Evaluating Natural Language Understanding and Explanation Coherence in Physics
In order for language models to aid physics research, they must first encode representations of mathematical and natural language discourse which lead to coherent explanations, with correct ordering and relevance of statements. We present a collection of datasets developed to evaluate the performance of language models in this regard, which measure capabilities with respect to sentence ordering, position, section prediction, and discourse coherence. Analysis of the data reveals equations and sub-disciplines which are most common in physics discourse, as well as the sentence-level frequency of equations and expressions. We present baselines that demonstrate how contemporary language models are challenged by coherence related tasks in physics, even when trained on mathematical natural language objectives.
Code (1)
Tasks
PositionSentenceSentence OrderingSimilar Papers 제목 키워드 기반
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding
Although Indonesian is known to be the fourth most frequently used language over the internet, the research progress on this language in the natural language processing (NLP) is slow-moving due to a lack of available res…
BenchmarkingDiversityNatural Language UnderstandingSentence+1BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla
This work presents BanglaNLG, a comprehensive benchmark for evaluating natural language generation (NLG) models in Bangla, a widely spoken yet low-resource language. We aggregate six challenging conditional text generati…
Conditional Text GenerationDialogue GenerationLanguage ModelingLanguage Modelling+1Consolidating and Developing Benchmarking Datasets for the Nepali Natural Language Understanding Tasks
The Nepali language has distinct linguistic features, especially its complex script (Devanagari script), morphology, and various dialects, which pose a unique challenge for natural language processing (NLP) evaluation. W…
BenchmarkingNatural Language InferenceNatural Language UnderstandingSentence+1Bridging Language Gaps: Enhancing Few-Shot Language Adaptation
The disparity in language resources poses a challenge in multilingual NLP, with high-resource languages benefiting from extensive data, while low-resource languages lack sufficient data for effective training. Our Contra…
Natural Language UnderstandingNatural Language InferenceCross-Lingual TransferContrastive LearningOn Evaluating and Mitigating Gender Biases in Multilingual Settings
While understanding and removing gender biases in language models has been a long-standing problem in Natural Language Processing, prior research work has primarily been limited to English. In this work, we investigate s…