paper-with-me

홈 › Papers

Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering

2024-10-31 · Aris J. Aristorenas

This study presents a machine learning framework for assessing similarity between audio content and predicting sentiment score. We construct a dataset containing audio samples from music covers on YouTube along with the audio of the original song, and sentiment scores derived from user comments, serving as proxy labels for content quality. Our approach involves extensive pre-processing, segmenting audio signals into 30-second windows, and extracting high-dimensional feature representations through Mel-Frequency Cepstral Coefficients (MFCC), Chroma, Spectral Contrast, and Temporal characteristics. Leveraging these features, we train regression models to predict sentiment scores on a 0-100 scale, achieving root mean square error (RMSE) values of 3.420, 5.482, 2.783, and 4.212, respectively. Improvements over a baseline model based on absolute difference metrics are observed. These results demonstrate the potential of machine learning to capture sentiment and similarity in audio, offering an adaptable framework for AI applications in media analysis.

📄 PDF Abstract BibTeX arXiv:2411.00195

Code (0)

등록된 구현이 없습니다.

Tasks

Feature Engineering

Similar Papers 제목 키워드 기반

Music Genre Classification: Training an AI model

2024-05-23 · Keoikantse Mogonediwa

Music genre classification is an area that utilizes machine learning models and techniques for the processing of audio signals, in which applications range from content recommendation systems to music recommendation syst…

ClassificationGenre classificationmodelMusic Genre Classification+2

Transformer-based Sequence Labeling for Audio Classification based on MFCCs

2023-04-30 · C. S. Sonali, Chinmayi B S, Ahana Balasubramanian

Audio classification is vital in areas such as speech and music recognition. Feature extraction from the audio signal, such as Mel-Spectrograms and MFCCs, is a critical step in audio classification. These features are tr…

Audio ClassificationClassification

MFAAN: Unveiling Audio Deepfakes with a Multi-Feature Authenticity Network

2023-11-06 · Karthik Sivarama Krishnan, Koushik Sivarama Krishnan

In the contemporary digital age, the proliferation of deepfakes presents a formidable challenge to the sanctity of information dissemination. Audio deepfakes, in particular, can be deceptively realistic, posing significa…

Face SwappingMisinformation

FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection

2021-10-18 · Zhenyu Zhang, Yewei Gu, Xiaowei Yi, Xianfeng Zhao

As increasing development of text-to-speech (TTS) and voice conversion (VC) technologies, the detection of synthetic speech has been suffered dramatically. In order to promote the development of synthetic speech detectio…

Speech SynthesisSynthetic Speech Detectiontext-to-speechText to Speech+1

DDFAD: Dataset Distillation Framework for Audio Data

2024-07-15 · Wenbo Jiang, Rui Zhang, Hongwei Li, Xiaoyuan Liu 외

Deep neural networks (DNNs) have achieved significant success in numerous applications. The remarkable performance of DNNs is largely attributed to the availability of massive, high-quality training datasets. However, pr…

Continual LearningDataset DistillationNeural Architecture Search