paper-with-me

홈 › Papers

VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

2023-05-31 · Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng, Jing Shi, Pin Lv, Bo Xu

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues can also benefit in many scenarios. In this paper, we first propose ViLaS (Vision and Language into Automatic Speech Recognition), a novel multimodal ASR model based on the continuous integrate-and-fire (CIF) mechanism, which can integrate visual and textual context simultaneously or separately, to facilitate speech recognition. Next, we introduce an effective training strategy that improves performance in modal-incomplete test scenarios. Then, to explore the effects of integrating vision and language, we create VSDial, a multimodal ASR dataset with multimodal context cues in both Chinese and English versions. Finally, empirical results are reported on the public Flickr8K and self-constructed VSDial datasets. We explore various cross-modal fusion schemes, analyze fine-grained crossmodal alignment on VSDial, and provide insights into the effects of integrating multimodal information on speech recognition.

📄 PDF Abstract BibTeX arXiv:2305.19972

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

VILAS: A VLA-Integrated Low-cost Architecture with Soft Grasping for Robotic Manipulation

2026-05-03 · Zijian An, Hadi Khezam, Bill Cai, Ran Yang 외 arxiv

We present VILAS, a fully low-cost, modular robotic manipulation platform designed to support end-to-end vision-language-action (VLA) policy learning and deployment on accessible hardware. The system integrates a Fairino…

Exploring Diverse In-Context Configurations for Image Captioning

2023-05-24 · NeurIPS 2023 11 · Xu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen 외

After discovering that Language Models (LMs) can be good in-context few-shot learners, numerous strategies have been proposed to optimize in-context sequence configurations. Recently, researchers in Vision-Language (VL) …

Image CaptioningIn-Context Learning

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

2025-06-11 · Junfei Wu, Jian Guan, Kaituo Feng, Qiang Liu 외

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, exis…

Multimodal ReasoningSpatial Reasoning

COLA: Context-aware Language-driven Test-time Adaptation

2025-09-22 · Aiming Zhang, Tianyuan Yu, Liang Bai, Jun Tang 외 arxiv

Test-time adaptation (TTA) has gained increasing popularity due to its efficacy in addressing ``distribution shift'' issue while simultaneously protecting data privacy. However, most prior methods assume that a paired so…

Test-time Adaptation

What Happens When Small Is Made Smaller? Exploring the Impact of Compression on Small Data Pretrained Language Models

2024-04-06 · Busayo Awobade, Mardiyyah Oduwole, Steven Kolawole

Compression techniques have been crucial in advancing machine learning by enabling efficient training and deployment of large-scale language models. However, these techniques have received limited attention in the contex…

Knowledge DistillationLanguage ModelingLanguage ModellingQuantization