paper-with-me

홈 › Papers

LSTM-CNN Network for Audio Signature Analysis in Noisy Environments

2023-12-12 · Praveen Damacharla, Hamid Rajabalipanah, Mohammad Hosein Fakheri

There are multiple applications to automatically count people and specify their gender at work, exhibitions, malls, sales, and industrial usage. Although current speech detection methods are supposed to operate well, in most situations, in addition to genders, the number of current speakers is unknown and the classification methods are not suitable due to many possible classes. In this study, we focus on a long-short-term memory convolutional neural network (LSTM-CNN) to extract time and / or frequency-dependent features of the sound data to estimate the number / gender of simultaneous active speakers at each frame in noisy environments. Considering the maximum number of speakers as 10, we have utilized 19000 audio samples with diverse combinations of males, females, and background noise in public cities, industrial situations, malls, exhibitions, workplaces, and nature for learning purposes. This proof of concept shows promising performance with training/validation MSE values of about 0.019/0.017 in detecting count and gender.

📄 PDF Abstract BibTeX arXiv:2312.07059

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Resource aware design of a deep convolutional-recurrent neural network for speech recognition through audio-visual sensor fusion

2018-03-13 · Matthijs Van keirsbilck, Bert Moons, Marian Verhelst

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-readi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip ReadingPhoneme Recognition+3

Deep Learning for Lip Reading using Audio-Visual Information for Urdu Language

2018-02-15 · M Faisal, Sanaullah Manzoor

Human lip-reading is a challenging task. It requires not only knowledge of underlying language but also visual clues to predict spoken words. Experts need certain level of experience and understanding of visual expressio…

Lip Readingspeech-recognitionSpeech Recognition

SigGate: Enhancing Recurrent Neural Networks with Signature-Based Gating Mechanisms

2025-02-13 · Rémi Genet, Hugo Inzirillo

In this paper, we propose a novel approach that enhances recurrent neural networks (RNNs) by incorporating path signatures into their gating mechanisms. Our method modifies both Long Short-Term Memory (LSTM) and Gated Re…

Time Series Analysis

Audio-Based Pedestrian Detection in the Presence of Vehicular Noise

2025-09-23 · Yonghyun Kim, Chaeyeon Han, Akash Sarode, Noah Posner 외 arxiv

Audio-based pedestrian detection is a challenging task and has, thus far, only been explored in noise-limited environments. We present a new dataset, results, and a detailed analysis of the state-of-the-art in audio-base…

Pedestrian Detection

Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras

2019-12-05 · Ander Arriandiaga, Giovanni Morrone, Luca Pasa, Leonardo Badino 외

We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is th…

Optical Flow EstimationSpeech Separation