paper-with-me

Papers

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

2025-05-21 · Cheng Yifan, Zhang Ruoyi, Shi Jiatong

Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional speech from unlabeled video data. Leveraging face detection and tracking algorithms, we developed an automatic emotion analysis system using a multimodal large language model (MLLM). Our results demonstrate that MIKU-PAL can achieve human-level accuracy (68.5% on MELD) and superior consistency (0.93 Fleiss kappa score) while being much cheaper and faster than human annotation. With the high-quality, flexible, and consistent annotation from MIKU-PAL, we can annotate fine-grained speech emotion categories of up to 26 types, validated by human annotators with 83% rationality ratings. Based on our proposed system, we further released a fine-grained emotional speech dataset MIKU-EmoBench(131.2 hours) as a new benchmark for emotional text-to-speech and visual voice cloning.

📄 PDF Abstract BibTeX arXiv:2505.15772

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionFace DetectionLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelSpeech Synthesistext-to-speechText to SpeechVoice Cloning

Similar Papers 제목 키워드 기반

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

2025-08-19 · Xiaolu Hou, Bing Ma, Jiaxiang Cheng, Xuhua Ren 외 arxiv

With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods mostly rely on predefined scripts, lack…

MikuDance: Animating Character Art with Mixed Motion Dynamics

2024-11-13 · Jiaxu Zhang, Xianfang Zeng, Xin Chen, Wei Zuo 외

We propose MikuDance, a diffusion-based pipeline incorporating mixed motion dynamics to animate stylized character art. MikuDance consists of two key techniques: Mixed Motion Modeling and Mixed-Control Diffusion, to addr…

An automated medical scribe for documenting clinical encounters

2018-06-01 · NAACL 2018 6 · Gregory Finley, Erik Edwards, Am Robinson, a 외

A medical scribe is a clinical professional who charts patient{--}physician encounters in real time, relieving physicians of most of their administrative burden and substantially increasing productivity and job satisfact…

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition+1

Enhancing Computer Vision with Knowledge: a Rummikub Case Study

2024-11-27 · Simon Vandevelde, Laurent Mertens, Sverre Lauwers, Joost Vennekens

Artificial Neural Networks excel at identifying individual components in an image. However, out-of-the-box, they do not manage to correctly integrate and interpret these components as a whole. One way to alleviate this w…

MultiZoo & MultiBench: A Standardized Toolkit for Multimodal Deep Learning

2023-06-28 · Paul Pu Liang, Yiwei Lyu, Xiang Fan, Arav Agarwal 외

Learning multimodal representations involves integrating information from multiple heterogeneous sources of data. In order to accelerate progress towards understudied modalities and tasks while ensuring real-world robust…

Deep LearningMultimodal Deep Learning