paper-with-me

홈 › Papers

Unsupervised Structure Discovery for Semantic Analysis of Audio

2012-12-01 · NeurIPS 2012 12 · Sourish Chaudhuri, Bhiksha Raj

Approaches to audio classification and retrieval tasks largely rely on detection-based discriminative models. We submit that such models make a simplistic assumption in mapping acoustics directly to semantics, whereas the actual process is likely more complex. We present a generative model that maps acoustics in a hierarchical manner to increasingly higher-level semantics. Our model has 2 layers with the first being generic sound units with no clear semantic associations, while the second layer attempts to find patterns over the generic sound units. We evaluate our model on a large-scale retrieval task from TRECVID 2011, and report significant improvements over standard baselines.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Audio ClassificationGeneral ClassificationRetrieval

Similar Papers 제목 키워드 기반

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

2026-05-08 · Hamze Hammami, Nidhal Abdulaziz arxiv

Discovering structure in biological signals without supervision is a fundamental problem in computational intelligence, yet existing bioacoustic methods assume vocal production models or predefined semantic units, leavin…

Unsupervised Musical Object Discovery from Audio

2023-11-13 · Joonsu Gha, Vincent Herrmann, Benjamin Grewe, Jürgen Schmidhuber 외

Current object-centric learning models such as the popular SlotAttention architecture allow for unsupervised visual scene decomposition. Our novel MusicSlots method adapts SlotAttention to the audio domain, to achieve un…

ObjectObject DiscoveryProperty Prediction

Learning Transposition-Invariant Interval Features from Symbolic Music and Audio

2018-06-21 · Stefan Lattner, Maarten Grachten, Gerhard Widmer

Many music theoretical constructs (such as scale types, modes, cadences, and chord types) are defined in terms of pitch intervals---relative distances between pitches. Therefore, when computer models are employed in musi…

Unsupervised Cross-Modal Audio Representation Learning from Unstructured Multilingual Text

2020-03-27 · Alexander Schindler, Sergiu Gordea, Peter Knees

We present an approach to unsupervised audio representation learning. Based on a triplet neural network architecture, we harnesses semantically related cross-modal information to estimate audio track-relatedness. By appl…

Representation LearningRetrievalTriplet

COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space

2026-05-31 · Yao Luan, Ni Mu, Hanfei Ge, Yiqin Yang 외 arxiv

Unsupervised skill discovery (USD) aims to learn diverse behaviors without reward functions, but often results in task-irrelevant or hazardous behaviors due to uniform exploration. Guided skill discovery (GSD) addresses …