paper-with-me

Papers

Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks

2025-03-07 · Youness Atif

This paper introduces a novel audio-to-image encoding framework that integrates multiple dimensions of voice characteristics into a single RGB image for speaker recognition. In this method, the green channel encodes raw audio data, the red channel embeds statistical descriptors of the voice signal (including key metrics such as median and mean values for fundamental frequency, spectral centroid, bandwidth, rolloff, zero-crossing rate, MFCCs, RMS energy, spectral flatness, spectral contrast, chroma, and harmonic-to-noise ratio), and the blue channel comprises subframes representing these features in a spatially organized format. A deep convolutional neural network trained on these composite images achieves 98% accuracy in speaker classification across two speakers, suggesting that this integrated multi-channel representation can provide a more discriminative input for voice recognition tasks.

📄 PDF Abstract BibTeX arXiv:2503.05929

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

Securing Voice-driven Interfaces against Fake (Cloned) Audio Attacks

2019-02-18

Voice cloning technologies have found applications in a variety of areas ranging from personalized speech interfaces to advertisement, robotics, and so on. Existing voice cloning systems are capable of learning speaker c…

Speech SynthesisVoice Cloning

Self Voice Conversion as an Attack against Neural Audio Watermarking

2026-01-28 · Yigitcan Özer, Wanying Ge, Zhe Zhang, Xin Wang 외 arxiv

Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and digital signal processing-based watermark…

Voice Conversion

ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA

2026-03-10 · Aviad Dahan, Moran Yanuka, Noa Kraicer, Lior Wolf 외 arxiv

Existing video personalization methods preserve visual likeness but treat video and audio separately. Without access to the visual scene, audio models cannot synchronize sounds with on-screen actions; and because classic…

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis

2025-05-25 · Minsu Kim, Pingchuan Ma, Honglie Chen, Stavros Petridis 외

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language

2024-08-19 · Manjil Karki, Pratik Shakya, Sandesh Acharya, Ravi Pandit 외

Voice cloning is a prominent feature in personalized speech interfaces. A neural vocal cloning system can mimic someone's voice using just a few audio samples. Both speaker encoding and speaker adaptation are topics of r…

Transfer LearningVoice Cloning