paper-with-me

홈 › Papers

Interactive Dual-Conformer with Scene-Inspired Mask for Soft Sound Event Detection

2023-11-23 · Han Yin, Jisheng Bai, Mou Wang, Dongyuan Shi, Woon-Seng Gan, Jianfeng Chen

Traditional binary hard labels for sound event detection (SED) lack details about the complexity and variability of sound event distributions. Recently, a novel annotation workflow is proposed to generate fine-grained non-binary soft labels, resulting in a new real-life dataset named MAESTRO Real for SED. In this paper, we first propose an interactive dual-conformer (IDC) module, in which a cross-interaction mechanism is applied to effectively exploit the information from soft labels. In addition, a novel scene-inspired mask (SIM) based on soft labels is incorporated for more precise SED predictions. The SIM is initially generated through a statistical approach, referred as SIM-V1. However, the fixed artificial mask may mismatch the SED model, resulting in limited effectiveness. Therefore, we further propose SIM-V2, which employs a word embedding model for adaptive SIM estimation. Experimental results show that the proposed IDC module can effectively utilize the information from soft labels, and the integration of SIM-V1 can further improve the accuracy. In addition, the impact of different word embedding dimensions on SIM-V2 is explored, and the results show that the appropriate dimension can enable SIM-V2 achieve superior performance than SIM-V1. In DCASE 2023 Challenge Task4B, the proposed system achieved the top ranking performance on the evaluation dataset of MAESTRO Real.

📄 PDF Abstract BibTeX arXiv:2311.14068

Code (0)

등록된 구현이 없습니다.

Tasks

Event DetectionSound Event Detection

Similar Papers 제목 키워드 기반

Point'n Move: Interactive Scene Object Manipulation on Gaussian Splatting Radiance Fields

2023-11-28 · Jiajun Huang, Hongchuan Yu

We propose Point'n Move, a method that achieves interactive scene object manipulation with exposed region inpainting. Interactivity here further comes from intuitive object selection and real-time editing. To achieve thi…

Object

DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement

2021-06-30 · Yuma Koizumi, Shigeki Karita, Scott Wisdom, Hakan Erdogan 외

Single-channel speech enhancement (SE) is an important task in speech processing. A widely used framework combines an analysis/synthesis filterbank with a mask prediction network, such as the Conv-TasNet architecture. In…

Computational EfficiencyDenoisingPredictionSpeech Enhancement

SelectAnyTree: A Promptable Instance Segmentation Model for 3D Forest LiDAR Point Clouds

2026-06-25 · Trung Thanh Nguyen, Daniel Lusk, Kilian Gerberding, Janusch Vajna-Jehle 외 arxiv

Instance segmentation of trees in forest LiDAR point clouds is constrained by label scarcity: A single hectare holds millions of points and hundreds of overlapping tree crowns, making manual annotation laborious, while a…

Instance SegmentationPoint Clouds

LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control

2024-06-23 · Delin Qu, Qizhi Chen, Pingrui Zhang, Xianqiang Gao 외

This paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes with multiple interactive objects. To tac…

Novel View SynthesisObjectObject Reconstruction

ABConformer: Physics-inspired Sliding Attention for Antibody-Antigen Interface Prediction

2025-09-27 · Zhang-Yu You, Jiahao Ma, Hongzong Li, Ye-Fan Hu 외 arxiv

Accurate prediction of antibody-antigen (Ab-Ag) interfaces is critical for vaccine design, immunodiagnostics, and therapeutic antibody development. However, achieving reliable predictions from sequences alone remains a c…