paper-with-me

홈 › Papers

MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine Translation

2024-11-03 · Langlin Huang, Mengyu Bu, Yang Feng

Byte-based machine translation systems have shown significant potential in massively multilingual settings. Unicode encoding, which maps each character to specific byte(s), eliminates the emergence of unknown words, even in new languages, enabling broad language scalability. However, byte-level tokenization results in sequences that are hard to interpret due to limited semantic information per byte. Local contextualization has proven effective in assigning initial semantics to tokens, improving sentence comprehension. Nevertheless, variations in encoding rules across languages necessitate an adaptive approach for effective contextualization. To this end, we propose Adaptive MultiScale-Headed Attention (Ada-MSHA), adaptively selecting and mixing attention heads, which are treated as contextualization experts. This enhances the flexibility of contextualization scales and improves the potential to discover a better strategy than previous methods. Experiment results show that our method outperforms existing methods without extensive manual adjustment of hyper-parameters and surpasses subword-based models with fewer parameters in Ted-59 dataset. Our code is available at https://github.com/ictnlp/MoCE.

📄 PDF Abstract BibTeX arXiv:2411.01474

Code (1)

ictnlp/moce 공식 구현 pytorch

Tasks

Machine TranslationSentence

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

EEG-Based Multimodal Learning via Hyperbolic Mixture-of-Curvature Experts

2026-04-14 · Runhe Zhou, Shanglin Li, Guanxiang Huang, Xinliang Zhou 외 arxiv

Electroencephalography (EEG)-based multimodal learning integrates brain signals with complementary modalities to improve mental state assessment, providing great clinical potential. The effectiveness of such paradigms la…

Representation LearningEmotion Recognition

Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts

2024-02-08 · Zhili Liu, Kai Chen, Jianhua Han, Lanqing Hong 외

Masked Autoencoder~(MAE) is a prevailing self-supervised learning method that achieves promising results in model pre-training. However, when the various downstream tasks have data distributions different from the pre-tr…

Mixture-of-ExpertsSelf-Supervised Learning

Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning

2025-09-03 · Sugyeong Eo, Jungjun Lee, Chanjun Park, Heuiseok Lim arxiv

A sparse Mixture-of-Experts (MoE) architecture has emerged as a highly scalable solution by conditionally activating sub-modules without a proportional increase in computational costs. However, improving expert specializ…

Enhancing Molecular Property Prediction via Mixture of Collaborative Experts

2023-12-06 · Xu Yao, Shuang Liang, Songqiao Han, Hailiang Huang

Molecular Property Prediction (MPP) task involves predicting biochemical properties based on molecular features, such as molecular graph structures, contributing to the discovery of lead compounds in drug development. To…

Decision MakingDiversityMolecular Property PredictionPrediction+1

Complexity Experts are Task-Discriminative Learners for Any Image Restoration

2024-11-27 · CVPR 2025 1 · Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan 외

Recent advancements in all-in-one image restoration models have revolutionized the ability to address diverse degradations through a unified framework. However, parameters tied to specific tasks often remain inactive for…

AttributeBlind All-in-One Image RestorationImage RestorationMixture-of-Experts