paper-with-me

홈 › Papers

Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification

2024-04-29 · Artem Abzaliev, Humberto Pérez Espinosa, Rada Mihalcea

Similar to humans, animals make extensive use of verbal and non-verbal forms of communication, including a large range of audio signals. In this paper, we address dog vocalizations and explore the use of self-supervised speech representation models pre-trained on human speech to address dog bark classification tasks that find parallels in human-centered tasks in speech recognition. We specifically address four tasks: dog recognition, breed identification, gender classification, and context grounding. We show that using speech embedding representations significantly improves over simpler classification baselines. Further, we also find that models pre-trained on large human speech acoustics can provide additional performance boosts on several tasks.

📄 PDF Abstract BibTeX arXiv:2404.18739

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGender Classificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM

2025-08-19 · Dariia Puhach, Amir H. Payberah, Éva Székely arxiv

Similar to text-based Large Language Models (LLMs), Speech-LLMs exhibit emergent abilities and context awareness. However, whether these similarities extend to gender bias remains an open question. This study proposes a …

Enhancing Suno's Bark Text-to-Speech Model: Addressing Limitations Through Meta's Encodec and Pre-Trained Hubert

2023-04-18 · Social Science Research Network (SSRN) 2023 4 · Devin Schumacher, Francis LaBounty Jr.

Bark, a transformer-based text-to-audio model by Suno, generates highly realistic, multilingual speech as well as other audio, including music, background noise, and simple sound effects. While this model has shown promi…

Audio GenerationExpressive Speech SynthesisSpeech Synthesistext-to-speech+3

WPS-Dataset: A benchmark for wood plate segmentation in bark removal processing

2024-04-17 · Rijun Wang, Guanghao Zhang, Fulong Liang, Bo wang 외

Using deep learning methods is a promising approach to improving bark removal efficiency and enhancing the quality of wood products. However, the lack of publicly available datasets for wood plate segmentation in bark re…

Segmentation

Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI

2024-05-29 · Che Liu, Changde Du, Xiaoyu Chen, Huiguang He

Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel coarse-to-fine audio r…

FAD

Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech

2026-07-06 · Benjamin Ballyk, Teyun Kwon, Miran Özdogan, Oiwi Parker Jones arxiv

Non-invasive brain-to-speech decoding aims to restore communication to patients suffering from neurodegenerative disease, without the risks of neurosurgery. Existing MEG- and EEG-based methods, while scalable, continue t…

Speech RecognitionData Augmentation