Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models
Multilingual language models (MLLMs) are crucial for handling text across various languages, yet they often show performance disparities due to differences in resource availability and linguistic characteristics. While the impact of pre-train data percentage and model size on performance is well-known, our study reveals additional critical factors that significantly influence MLLM effectiveness. Analyzing a wide range of features, including geographical, linguistic, and resource-related aspects, we focus on the SIB-200 dataset for classification and the Flores-200 dataset for machine translation, using regression models and SHAP values across 204 languages. Our findings identify token similarity and country similarity as pivotal factors, alongside pre-train data and model size, in enhancing model performance. Token similarity facilitates cross-lingual transfer, while country similarity highlights the importance of shared cultural and linguistic contexts. These insights offer valuable guidance for developing more equitable and effective multilingual language models, particularly for underrepresented languages.
Code (1)
Tasks
Cross-Lingual TransferMachine TranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improving Recommender Systems Beyond the Algorithm
Recommender systems rely heavily on the predictive accuracy of the learning algorithm. Most work on improving accuracy has focused on the learning algorithm itself. We argue that this algorithmic focus is myopic. In part…
Recommendation SystemsExploring Adversarial Robustness of LiDAR-Camera Fusion Model in Autonomous Driving
Our study assesses the adversarial robustness of LiDAR-camera fusion models in 3D object detection. We introduce an attack technique that, by simply adding a limited number of physically constrained adversarial points ab…
3D Object DetectionAdversarial RobustnessAutonomous Drivingobject-detection+1Are All Marine Species Created Equal? Performance Disparities in Underwater Object Detection
Underwater object detection is critical for monitoring marine ecosystems but poses unique challenges, including degraded image quality, imbalanced class distribution, and distinct visual characteristics. Not every specie…
Object DetectionBeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
Recent advances in large language model (LLM) pretraining have shown that simply scaling data quantity eventually leads to diminishing returns, hitting a data wall. In response, the use of synthetic data for pretraining …
Synthetic Data GenerationA Logarithmic Mean Divisia Index Decomposition of CO$_2$ Emissions from Energy Use in Romania
Carbon emissions have become a specific alarming indicators and intricate challenges that lead an extended argue about climate change. The growing trend in the utilization of fossil fuels for the economic progress and si…