The bag-of-frames approach: a not so sufficient model for urban soundscapes
The "bag-of-frames" approach (BOF), which encodes audio signals as the long-term statistical distribution of short-term spectral features, is commonly regarded as an effective and sufficient way to represent environmental sound recordings (soundscapes) since its introduction in an influential 2007 article. The present paper describes a concep-tual replication of this seminal article using several new soundscape datasets, with results strongly questioning the adequacy of the BOF approach for the task. We show that the good accuracy originally re-ported with BOF likely result from a particularly thankful dataset with low within-class variability, and that for more realistic datasets, BOF in fact does not perform significantly better than a mere one-point av-erage of the signal's features. Soundscape modeling, therefore, may not be the closed case it was once thought to be. Progress, we ar-gue, could lie in reconsidering the problem of considering individual acoustical events within each soundscape.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Effects of a Hovering Unmanned Aerial Vehicle on Urban Soundscapes Perception
Several industry leaders and governmental agencies are currently investigating the use of Unmanned Aerial Vehicles (UAVs), or drones as commonly known, for an ever-growing number of applications from blue light services …
NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating
Audio provides critical situational cues, yet current Audio Language Models (ALMs) face an attention bottleneck in long-form recordings where dominant background patterns can dilute rare, salient events. We introduce NAA…
Urban Rhapsody: Large-scale exploration of urban soundscapes
Noise is one of the primary quality-of-life issues in urban environments. In addition to annoyance, noise negatively impacts public health and educational performance. While low-cost sensors can be deployed to monitor am…
SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
We present a novel and practically significant problem-Geo-Contextual Soundscape-to-Landscape (GeoS2L) generation-which aims to synthesize geographically realistic landscape images from environmental soundscapes. Prior a…
Image GenerationUSM-SED - A Dataset for Polyphonic Sound Event Detection in Urban Sound Monitoring Scenarios
This paper introduces a novel dataset for polyphonic sound event detection in urban sound monitoring use-cases. Based on isolated sounds taken from the FSD50k dataset, 20,000 polyphonic soundscapes are synthesized with s…
Dataset GenerationEvent DetectionSound Event Detection