paper-with-me

홈 › Papers

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

2025-09-17 · Elena Camuffo, Francesco Barbato, Mete Ozay, Simone Milani, Umberto Michieli arxiv

Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Lightweight models often struggle in this setting due to their weak semantic priors, while large vision-language models (VLMs) offer strong object-level understanding but are too computationally demanding for real-time or on-device applications. We introduce MOCHA (Multi-modal Objects-aware Cross-arcHitecture Alignment), a distillation framework that transfers multimodal region-level knowledge from a frozen VLM teacher into a lightweight vision-only detector. MOCHA extracts fused visual and textual teacher's embeddings and uses them to guide student training through a dual-objective loss that enforces accurate local alignment and global relational consistency across regions. This process enables efficient transfer of semantics without the need for teacher modifications or textual input at inference. MOCHA consistently outperforms prior baselines across four personalized detection benchmarks under strict few-shot regimes, yielding a +10.1 average improvement, with minimal inference cost.

📄 PDF Abstract BibTeX arXiv:2509.14001

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description

2024-10-15 · Jiawei Mo, Yixuan Chen, Rifen Lin, Yongkang Ni 외

Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interac…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention

2025-07-30 · Yuqi Pang, Bowen Yang, Yun Cao, Rong Fan 외 arxiv

Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual models. However, these approaches face h…

MoCha:End-to-End Video Character Replacement without Structural Guidance

2026-01-13 · Zhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng 외 arxiv

Controllable video character replacement with a user-provided identity remains a challenging problem due to the lack of paired video data. Prior works have predominantly relied on a reconstruction-based paradigm that req…

Multi-head Monotonic Chunkwise Attention For Online Speech Recognition

2020-05-01 · Baiji Liu, Songjun Cao, Sining Sun, Weibin Zhang 외

The attention mechanism of the Listen, Attend and Spell (LAS) model requires the whole input sequence to calculate the attention context and thus is not suitable for online speech recognition. To deal with this problem, …

speech-recognitionSpeech Recognition

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

2026-07-17 · Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu 외 arxiv

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions…

Speech Synthesis