paper-with-me

Papers

Cocktail Party Processing via Structured Prediction

2012-12-01 · NeurIPS 2012 12 · Yuxuan Wang, DeLiang Wang

While human listeners excel at selectively attending to a conversation in a cocktail party, machine performance is still far inferior by comparison. We show that the cocktail party problem, or the speech separation problem, can be effectively approached via structured prediction. To account for temporal dynamics in speech, we employ conditional random fields (CRFs) to classify speech dominance within each time-frequency unit for a sound mixture. To capture complex, nonlinear relationship between input and output, both state and transition feature functions in CRFs are learned by deep neural networks. The formulation of the problem as classification allows us to directly optimize a measure that is well correlated with human speech intelligibility. The proposed system substantially outperforms existing ones in a variety of noises.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationPredictionSpeech SeparationStructured Prediction

Similar Papers 제목 키워드 기반

Deep Transform: Cocktail Party Source Separation via Probabilistic Re-Synthesis

2015-03-20 · Andrew J. R. Simpson

In cocktail party listening scenarios, the human brain is able to separate competing speech signals. However, the signal processing implemented by the brain to perform cocktail party listening is not well understood. Her…

Probabilistic Binary-Mask Cocktail-Party Source Separation in a Convolutional Deep Neural Network

2015-03-24 · Andrew J. R. Simpson

Separation of competing speech is a key challenge in signal processing and a feat routinely performed by the human auditory brain. A long standing benchmark of the spectrogram approach to source separation is known as th…

Prediction

L'amor\ccage s\'emantique masqu\'e en situation de cocktail party (Masked semantic priming in cocktail party situation) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Marie Dekerle, V{\'e}ronique Boulenger, Michel Hoen, Fanny Meunier

Cocktail-Party Audio-Visual Speech Recognition

2025-06-02 · Thai-Binh Nguyen, Ngoc-Quan Pham, Alexander Waibel

Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AV…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Speaker-Targeted Audio-Visual Models for Speech Recognition in Cocktail-Party Environments

2019-06-13 · Guan-Lin Chao, William Chan, Ian Lane

Speech recognition in cocktail-party environments remains a significant challenge for state-of-the-art speech recognition systems, as it is extremely difficult to extract an acoustic signal of an individual speaker from …

speech-recognitionSpeech Recognition