Papers FAD
“FAD” 태그가 달린 논문 62편 · 필터 해제
Detecting immune cells with label-free two-photon autofluorescence and deep learning
Label-free imaging has gained broad interest because of its potential to omit elaborate staining procedures which is especially relevant for in vivo use. Label-free multiphoton microscopy (MPM), for instance, exploits tw…
Binary ClassificationClassificationFADMulti-class Classification+1FAD-Net: Frequency-Domain Attention-Guided Diffusion Network for Coronary Artery Segmentation using Invasive Coronary Angiography
Background: Coronary artery disease (CAD) remains one of the leading causes of mortality worldwide. Precise segmentation of coronary arteries from invasive coronary angiography (ICA) is critical for effective clinical de…
Coronary Artery SegmentationFADSegmentationBemaGANv2: A Tutorial and Comparative Survey of GAN-based Vocoders for Long-Term Audio Generation
This paper presents a tutorial-style survey and implementation guide of BemaGANv2, an advanced GAN-based vocoder designed for high-fidelity and long-term audio generation. Built upon the original BemaGAN architecture, Be…
Audio GenerationFADSSIMIMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio gen…
AudioCapsAudio GenerationFADFAD: Frequency Adaptation and Diversion for Cross-domain Few-shot Learning
Cross-domain few-shot learning (CD-FSL) requires models to generalize from limited labeled samples under significant distribution shifts. While recent methods enhance adaptability through lightweight task-specific module…
Cross-Domain Few-Shotcross-domain few-shot learningFADFew-Shot LearningEfficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion
Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing approaches usually consider extracting 3D M…
Action GenerationFADImage GenerationMotion SynthesisDOSE : Drum One-Shot Extraction from Music Mixture
Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are pr…
FADDRAGON: Distributional Rewards Optimize Diffusion Generative Models
We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with traditional reinforcement learning with human…
FADEnhancing U.S. swine farm preparedness for infectious foreign animal diseases with rapid access to biosecurity information
The U.S. launched the Secure Pork Supply (SPS) Plan for Continuity of Business, a voluntary program providing foreign animal disease (FAD) guidance and setting biosecurity standards to maintain business continuity amid F…
FADTARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transform…
Audio SynthesisFADEnhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization
Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio gene…
Audio GenerationFADAligning Text-to-Music Evaluation with Human Preferences
Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, w…
FADFlowDec: A flow-based full-band general audio codec with high perceptual quality
We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared…
FADKAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
Although being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and h…
Audio GenerationFADGPURenderBox: Expressive Performance Rendering with Text Control
Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emoti…
DiversityFADMusic Performance RenderingDiffusion based Text-to-Music Generation with Global and Local Text based Conditioning
Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a…
FADLanguage ModelingLanguage ModellingLarge Language Model+2Sound Scene Synthesis at the DCASE 2024 Challenge
This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a sta…
FADMarket Making with Fads, Informed, and Uninformed Traders
We characterise the solutions to a continuous-time optimal liquidity provision problem in a market populated by informed and uninformed traders. In our model, the asset price exhibits fads -- these are short-term deviati…
FADFrechet Music Distance: A Metric For Generative Symbolic Music Evaluation
In this paper we introduce the Frechet Music Distance (FMD), a novel evaluation metric for generative symbolic music models, inspired by the Frechet Inception Distance (FID) in computer vision and Frechet Audio Distance …
FADMusic GenerationMusic ModelingMIMII-Gen: Generative Modeling Approach for Simulated Evaluation of Anomalous Sound Detection System
Insufficient recordings and the scarcity of anomalies present significant challenges in developing and validating robust anomaly detection systems for machine sounds. To address these limitations, we propose a novel appr…
Anomaly DetectionFAD