paper-with-me

홈 › Papers

ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024

2024-07-01

The Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC 2024) is part of the ISCSLP 2024 Competitions and Challenges track. While current text-to-speech (TTS) technology can generate high-quality audio, its ability to convey complex emotions and controlled detail content remains limited. This constraint leads to a discrepancy between the generated audio and human subjective perception in practical applications like companion robots for children and marketing bots. The core issue lies in the inconsistency between high-quality audio generation and the ultimate human subjective experience. Therefore, this challenge aims to enhance the persuasiveness and acceptability of synthesized audio, focusing on human alignment convincing and inspirational audio generation. A total of 19 teams have registered for the challenge, and the results of the competition and the competition are described in this paper.

📄 PDF Abstract BibTeX arXiv:2407.12038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge

2024-10-31 · Dake Guo, Jixun Yao, Xinfa Zhu, Kangxiang Xia 외

This paper presents the NPU-HWC system submitted to the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC). Our system consists of two modules: a speech generator for Track 1 and a backgroun…

Audio GenerationLanguage ModelingLanguage Modelling

Inspirational Adversarial Image Generation

2019-06-17 · Baptiste Rozière, Morgane Riviere, Olivier Teytaud, Jérémy Rapin 외

The task of image generation started to receive some attention from artists and designers to inspire them in new creations. However, exploiting the results of deep generative models such as Generative Adversarial Network…

Image Generation

VABench: A Comprehensive Benchmark for Audio-Video Generation

2025-12-10 · Daili Hua, Xizhi Wang, Bohan Zeng, Xinyi Huang 외 arxiv

Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual…

Video Generation

Leveraging Multimodal LLM for Inspirational User Interface Search

2025-01-29 · SeokHyeon Park, Yumin Song, Soohyun Lee, Jaeyoung Kim 외

Inspirational search, the process of exploring designs to inform and inspire new creative work, is pivotal in mobile user interface (UI) design. However, exploring the vast space of UI references remains a challenge. Exi…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

2024-02-27 · Linrui Tian, Qi Wang, Bang Zhang, Liefeng Bo

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify …

Video Generation