paper-with-me

홈 › Papers

Two Heads are Better than One? Verification of Ensemble Effect in Neural Machine Translation

2021-11-01 · EMNLP (insights) 2021 11 · Chanjun Park, Sungjin Park, Seolhwa Lee, Taesun Whang, Heuiseok Lim

In the field of natural language processing, ensembles are broadly known to be effective in improving performance. This paper analyzes how ensemble of neural machine translation (NMT) models affect performance improvement by designing various experimental setups (i.e., intra-, inter-ensemble, and non-convergence ensemble). To an in-depth examination, we analyze each ensemble method with respect to several aspects such as different attention models and vocab strategies. Experimental results show that ensembling is not always resulting in performance increases and give noteworthy negative findings.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

ESDF: Ensemble Selection using Diversity and Frequency

2015-08-18 · Shouvick Mondal, Arko Banerjee

Recently ensemble selection for consensus clustering has emerged as a research problem in Machine Intelligence. Normally consensus clustering algorithms take into account the entire ensemble of clustering, where there is…

ClusteringDiversity

Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks

2015-11-19 · Stefan Lee, Senthil Purushwalkam, Michael Cogswell, David Crandall 외

Convolutional Neural Networks have achieved state-of-the-art performance on a wide range of tasks. Most benchmarks are led by ensembles of these powerful learners, but ensembling is typically treated as a post-hoc proced…

Diversity

On Multi-head Ensemble of Smoothed Classifiers for Certified Robustness

2022-11-20 · Kun Fang, Qinghua Tao, Yingwen Wu, Tao Li 외

Randomized Smoothing (RS) is a promising technique for certified robustness, and recently in RS the ensemble of multiple deep neural networks (DNNs) has shown state-of-the-art performances. However, such an ensemble brin…

Boosting of Head Pose Estimation by Knowledge Distillation

2021-08-20 · Andrey Sheka, Victor Samun

We propose a response-based method of knowledge distillation (KD) for the head pose estimation problem. A student model trained by the proposed KD achieves results better than a teacher model, which is atypical for the r…

Head Pose EstimationKnowledge DistillationPose Estimationregression

Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality

2026-04-08 · Kai Qian, Weijie Shi, Jiaqi Wang, Mengze Li 외 arxiv

Multimodal fake news detection (MFND) aims to verify news credibility by jointly exploiting textual and visual evidence. However, real-world news dissemination frequently suffers from missing modality due to deleted imag…

Fake News Detection