paper-with-me

Papers

A Novel Frame Structure for Cloud-Based Audio-Visual Speech Enhancement in Multimodal Hearing-aids

2022-10-24 · Abhijeet Bishnu, Ankit Gupta, Mandar Gogate, Kia Dashtipour, Ahsan Adeel, Amir Hussain, Mathini Sellathurai, Tharmalingam Ratnarajah

In this paper, we design a first of its kind transceiver (PHY layer) prototype for cloud-based audio-visual (AV) speech enhancement (SE) complying with high data rate and low latency requirements of future multimodal hearing assistive technology. The innovative design needs to meet multiple challenging constraints including up/down link communications, delay of transmission and signal processing, and real-time AV SE models processing. The transceiver includes device detection, frame detection, frequency offset estimation, and channel estimation capabilities. We develop both uplink (hearing aid to the cloud) and downlink (cloud to hearing aid) frame structures based on the data rate and latency requirements. Due to the varying nature of uplink information (audio and lip-reading), the uplink channel supports multiple data rate frame structure, while the downlink channel has a fixed data rate frame structure. In addition, we evaluate the latency of different PHY layer blocks of the transceiver for developed frame structures using LabVIEW NXG. This can be used with software defined radio (such as Universal Software Radio Peripheral) for real-time demonstration scenarios.

📄 PDF Abstract BibTeX arXiv:2210.13127

Code (0)

등록된 구현이 없습니다.

Tasks

Lip ReadingSpeech Enhancement

Similar Papers 제목 키워드 기반

Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding

2025-07-01 · Duc Cao-Dinh, Khai Le-Duc, Anh Dao, Bach Phan Tat 외 arxiv

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D…

Multi-Label ClassificationRepresentation LearningSpeech RecognitionVisual Grounding

Audio-Visual Speech Inpainting with Deep Learning

2020-10-09 · Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan, Jesper Jensen

In this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliable audio context and uncorrupted visual …

Deep LearningMulti-Task Learning

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

2024-06-22 · Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen 외

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end a…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

On the Utility of Audiovisual Dialog Technologies and Signal Analytics for Real-time Remote Monitoring of Depression Biomarkers

2020-07-01 · WS 2020 7 · Michael Neumann, Oliver Roessler, David Suendermann-Oeft, Vikram Ramanarayanan

We investigate the utility of audiovisual dialog systems combined with speech and video analytics for real-time remote monitoring of depression at scale in uncontrolled environment settings. We collected audiovisual conv…

Encrypted Speech Recognition using Deep Polynomial Networks

2019-05-11 · Shi-Xiong Zhang, Yifan Gong, Dong Yu

The cloud-based speech recognition/API provides developers or enterprises an easy way to create speech-enabled features in their applications. However, sending audios about personal or company internal information to the…

speech-recognitionSpeech Recognition