paper-with-me

홈 › Papers

DAMEX: Dataset-aware Mixture-of-Experts for visual understanding of mixture-of-datasets

2023-11-08 · NeurIPS 2023 11 · Yash Jain, Harkirat Behl, Zsolt Kira, Vibhav Vineet

Construction of a universal detector poses a crucial question: How can we most effectively train a model on a large mixture of datasets? The answer lies in learning dataset-specific features and ensembling their knowledge but do all this in a single model. Previous methods achieve this by having separate detection heads on a common backbone but that results in a significant increase in parameters. In this work, we present Mixture-of-Experts as a solution, highlighting that MoEs are much more than a scalability tool. We propose Dataset-Aware Mixture-of-Experts, DAMEX where we train the experts to become an `expert' of a dataset by learning to route each dataset tokens to its mapped expert. Experiments on Universal Object-Detection Benchmark show that we outperform the existing state-of-the-art by average +10.2 AP score and improve over our non-MoE baseline by average +2.0 AP score. We also observe consistent gains while mixing datasets with (1) limited availability, (2) disparate domains and (3) divergent label sets. Further, we qualitatively show that DAMEX is robust against expert representation collapse.

📄 PDF Abstract BibTeX arXiv:2311.04894

Code (1)

jinga-lala/damex 공식 구현 pytorch

Tasks

Mixture-of-Expertsobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Fine-Grained Zero-Shot Learning with Attribute-Centric Representations

2025-12-13 · Zhi Chen, Jingcai Guo, Taotao Cai, Yuxiang Cai arxiv

Recognizing unseen fine-grained categories demands a model that can distinguish subtle visual differences. This is typically achieved by transferring visual-attribute relationships from seen classes to unseen classes. Th…

Representation LearningZero-Shot Learning

Dynamic Mixture-of-Experts for Visual Autoregressive Model

2025-10-08 · Jort Vincenti, Metod Jazbec, Guoxuan Xia arxiv

Visual Autoregressive Models (VAR) offer efficient and high-quality image generation but suffer from computational redundancy due to repeated Transformer calls at increasing resolutions. We introduce a dynamic Mixture-of…

Image Generation

Learning to Ground VLMs without Forgetting

2024-10-14 · Aritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis Koelma, Martin R. Oswald 외

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Visual Language Models (VLMs) struggle at this task. In this paper, we introduce LynX, a framew…

DecoderLanguage ModellingMixture-of-ExpertsMultimodal Reasoning+3

Modular Approach to Machine Reading Comprehension: Mixture of Task-Aware Experts

2022-10-04 · Anirudha Rayasam, Anusha Kamath, Gabriel Bayomi Tinoco Kalejaiye

In this work we present a Mixture of Task-Aware Experts Network for Machine Reading Comprehension on a relatively small dataset. We particularly focus on the issue of common-sense learning, enforcing the common ground kn…

Common Sense ReasoningMachine Reading ComprehensionReading ComprehensionTransfer Learning+1

Understanding Structured Health Data through Interaction-Aware Mixture-of-Experts

2026-07-14 · Ji Hwan Park, Ying Ding, Tianjin Guo arxiv

We study interaction-aware mixture-of-experts for post-stroke rigidity prediction using multi-level views of structured health records. Despite minimal performance gains, routing attribution reveals systematic importance…