Since the Scientific Literature Is Multilingual, Our Models Should Be Too
English has long been assumed the $\textit{lingua franca}$ of scientific research, and this notion is reflected in the natural language processing (NLP) research involving scientific document representation. In this position piece, we quantitatively show that the literature is largely multilingual and argue that current models and benchmarks should reflect this linguistic diversity. We provide evidence that text-based models fail to create meaningful representations for non-English papers and highlight the negative user-facing impacts of using English-only models non-discriminately across a multilingual domain. We end with suggestions for the NLP community on how to improve performance on non-English documents.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityPositionSimilar Papers 제목 키워드 기반
Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model
The rapid proliferation of interdisciplinary and multilingual scientific literature has left traditional manual analysis and single-algorithm methods plagued by low efficiency, poor scalability, and insufficient domain a…
Information ExtractionSciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini
We present SciDraw-6K, a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models, each paired with prompts in eleven languages (English, Simplified Chinese, Traditional Chin…
Automatic semantic classification of scientific literature according to the hallmarks of cancer
The hallmarks of cancer have become highly influential in cancer research. They reduce the complexity of cancer into 10 principles (e.g. resisting cell death and sustaining proliferative signaling) that explain the biolo…
General ClassificationAutomatic Aspect Extraction from Scientific Texts
Being able to extract from scientific papers their main points, key insights, and other important information, referred to here as aspects, might facilitate the process of conducting a scientific literature review. There…
Aspect ExtractionMassively Multilingual Corpus of Sentiment Datasets and Multi-faceted Sentiment Classification Benchmark
Despite impressive advancements in multilingual corpora collection and model training, developing large-scale deployments of multilingual models still presents a significant challenge. This is particularly true for langu…