paper-with-me

Papers FAD

“FAD” 태그가 달린 논문 62편 · 필터 해제

Detecting immune cells with label-free two-photon autofluorescence and deep learning

2025-06-17 · Lucas Kreiss, Amey Chaware, Maryam Roohian, Sarah Lemire 외

Label-free imaging has gained broad interest because of its potential to omit elaborate staining procedures which is especially relevant for in vivo use. Label-free multiphoton microscopy (MPM), for instance, exploits tw…

Binary ClassificationClassificationFADMulti-class Classification+1

FAD-Net: Frequency-Domain Attention-Guided Diffusion Network for Coronary Artery Segmentation using Invasive Coronary Angiography

2025-06-13 · Nan Mu, Ruiqi Song, Xiaoning Li, Zhihui Xu 외

Background: Coronary artery disease (CAD) remains one of the leading causes of mortality worldwide. Precise segmentation of coronary arteries from invasive coronary angiography (ICA) is critical for effective clinical de…

Coronary Artery SegmentationFADSegmentation

BemaGANv2: A Tutorial and Comparative Survey of GAN-based Vocoders for Long-Term Audio Generation

2025-06-11 · Taesoo Park, Mungwi Jeong, MinGyu Park, Narae Kim 외

This paper presents a tutorial-style survey and implementation guide of BemaGANv2, an advanced GAN-based vocoder designed for high-fidelity and long-term audio generation. Built upon the original BemaGAN architecture, Be…

Audio GenerationFADSSIM

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling

2025-05-31 · Kuan-Po Huang, Shu-wen Yang, Huy Phan, Bo-Ru Lu 외

Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio gen…

AudioCapsAudio GenerationFAD

FAD: Frequency Adaptation and Diversion for Cross-domain Few-shot Learning

2025-05-13 · Ruixiao Shi, Fu Feng, Yucheng Xie, Jing Wang 외

Cross-domain few-shot learning (CD-FSL) requires models to generalize from limited labeled samples under significant distribution shifts. While recent methods enhance adaptability through lightweight task-specific module…

Cross-Domain Few-Shotcross-domain few-shot learningFADFew-Shot Learning

Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion

2025-04-29 · Zesheng Wang, Alexandre Bruckert, Patrick Le Callet, Guangtao Zhai

Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing approaches usually consider extracting 3D M…

Action GenerationFADImage GenerationMotion Synthesis

DOSE : Drum One-Shot Extraction from Music Mixture

2025-04-25 · Suntae Hwang, SeongHyeon Kang, KyungSu Kim, Semin Ahn 외

Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are pr…

FAD

DRAGON: Distributional Rewards Optimize Diffusion Generative Models

2025-04-21 · Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, Nicholas J. Bryan

We present Distributional RewArds for Generative OptimizatioN (DRAGON), a versatile framework for fine-tuning media generation models towards a desired outcome. Compared with traditional reinforcement learning with human…

FAD

Enhancing U.S. swine farm preparedness for infectious foreign animal diseases with rapid access to biosecurity information

2025-04-12 · Christian Fleming, Kelsey Mills, Nicolas Cardenas, Jason A. Galvis 외

The U.S. launched the Secure Pork Supply (SPS) Plan for Continuity of Business, a voluntary program providing foreign animal disease (FAD) guidance and setting biosecurity standards to maintain business continuity amid F…

FAD

TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis

2025-04-08 · Tri Ton, Ji Woo Hong, Chang D. Yoo

This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transform…

Audio SynthesisFAD

Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization

2025-03-28 · Haomin Zhang, Sizhe Shan, Haoyu Wang, Zihao Chen 외

Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio gene…

Audio GenerationFAD

Aligning Text-to-Music Evaluation with Human Preferences

2025-03-20 · Yichen Huang, Zachary Novack, Koichi Saito, Jiatong Shi 외

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, w…

FAD

FlowDec: A flow-based full-band general audio codec with high perceptual quality

2025-03-03 · Simon Welker, Matthew Le, Ricky T. Q. Chen, Wei-Ning Hsu 외

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared…

FAD

KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

2025-02-21 · Yoonjin Chung, Pilsun Eu, Junwon Lee, Keunwoo Choi 외

Although being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and h…

Audio GenerationFADGPU

RenderBox: Expressive Performance Rendering with Text Control

2025-02-11 · huan zhang, Akira Maezawa, Simon Dixon

Expressive music performance rendering involves interpreting symbolic scores with variations in timing, dynamics, articulation, and instrument-specific techniques, resulting in performances that capture musical can emoti…

DiversityFADMusic Performance Rendering

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

2025-01-24 · Jisi Zhang, Pablo Peso Parada, Md Asif Jalal, Karthikeyan Saravanan

Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a…

FADLanguage ModelingLanguage ModellingLarge Language Model+2

Sound Scene Synthesis at the DCASE 2024 Challenge

2025-01-15 · Mathieu Lagrange, Junwon Lee, Modan Tailleur, Laurie M. Heller 외

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a sta…

FAD

Market Making with Fads, Informed, and Uninformed Traders

2025-01-07 · Emilio Barucci, Adrien Mathieu, Leandro Sánchez-Betancourt

We characterise the solutions to a continuous-time optimal liquidity provision problem in a market populated by informed and uninformed traders. In our model, the asset price exhibits fads -- these are short-term deviati…

FAD

Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation

2024-12-10 · Jan Retkowski, Jakub Stępniak, Mateusz Modrzejewski

In this paper we introduce the Frechet Music Distance (FMD), a novel evaluation metric for generative symbolic music models, inspired by the Frechet Inception Distance (FID) in computer vision and Frechet Audio Distance …

FADMusic GenerationMusic Modeling

MIMII-Gen: Generative Modeling Approach for Simulated Evaluation of Anomalous Sound Detection System

2024-09-27 · Harsh Purohit, Tomoya Nishida, Kota Dohi, Takashi Endo 외

Insufficient recordings and the scarcity of anomalies present significant challenges in developing and validating robust anomaly detection systems for machine sounds. To address these limitations, we propose a novel appr…

Anomaly DetectionFAD
1–20 / 62 다음 →