paper-with-me

홈 › Papers

Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip Reading

2021-06-19 · CVPR 2021 1 · Sucheng Ren, Yong Du, Jianming Lv, Guoqiang Han, Shengfeng He

Lip reading aims to predict the spoken sentences from silent lip videos. Due to the fact that such a vision task usually performs worse than its counterpart speech recognition, one potential scheme is to distill knowledge from a teacher pretrained by audio signals. However, the latent domain gap between the cross-modal data could lead to an learning ambiguity and thus limits the performance of lip reading. In this paper, we propose a novel collaborative framework for lip reading, and two aspects of issues are considered: 1) the teacher should understand bi-modal knowledge to possibly bridge the inherent cross-modal gap; 2) the teacher should adjust teaching contents adaptively with the evolution of the student. To these ends, we introduce a trainable "master" network which ingests both audio signals and silent lip videos instead of a pretrained teacher. The master produces logits from three modalities of features: audio modality, video modality, and their combination. To further provide an interactive strategy to fuse these knowledge organically, we regularize the master with the task-specific feedback from the student, in which the requirement of the student is implicitly embedded. Meanwhile we involve a couple of "tutor" networks into our system as guidance for emphasizing the fruitful knowledge flexibly. In addition, we incorporate a curriculum learning design to ensure a better convergence. Extensive experiments demonstrate that the proposed network outperforms the state-of-the-art methods on several benchmarks, including in both word-level and sentence-level scenarios.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Lip ReadingSentencespeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution

2026-05-09 · Feng Xiong, Zengbin Wang, Yong Wang, Xuecai Hu 외 arxiv

Self-evolving agents present a promising path toward continual adaptation by distilling task interactions into reusable knowledge artifacts. In practice, this paradigm remains hindered by two coupled bottlenecks: data in…

Masking Teacher and Reinforcing Student for Distilling Vision-Language Models

2025-12-23 · Byung-Kwan Lee, Yu-Chiang Frank Wang, Ryo Hachiuma arxiv

Large-scale vision-language models (VLMs) have recently achieved remarkable multimodal understanding, but their massive size makes them impractical for deployment on mobile or edge devices. This raises the need for compa…

Reinforcement LearningOffline RL

Advanced Mathematics Learning Behavior Prediction and Academic Early Warning Model Based on Multimodal Data Analysis

2026-05-31 · Liu Qiong, Li Zhengbo arxiv

Early detection of at-risk students and timely academic intervention pose major challenges in advanced mathematics education, where complex conceptual hierarchies and nonlinear learning trajectories often hold back stude…

From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion

2025-06-08 · Zheng Wang, Kai Ying, Bin Xu, Chunjiao Wang 외

Accurate near-real-time precipitation retrieval has been enhanced by satellite-based technologies. However, infrared-based algorithms have low accuracy due to weak relations with surface precipitation, whereas passive mi…

Data IntegrationRetrieval

SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model

2025-10-14 · Lin Lin, Jiefeng Long, Zhihe Wan, Yuchi Wang 외 arxiv

Multimodal embedding models aim to yield informative unified representations that empower diverse cross-modal tasks. Despite promising developments in the evolution from CLIP-based dual-tower architectures to large visio…

Representation Learning