paper-with-me

홈 › Papers

JoyHallo: Digital human model for Mandarin

2024-09-20 · Sheng Shi, Xuyang Cao, Jun Zhao, Guoxin Wang

In audio-driven video generation, creating Mandarin videos presents significant challenges. Collecting comprehensive Mandarin datasets is difficult, and the complex lip movements in Mandarin further complicate model training compared to English. In this study, we collected 29 hours of Mandarin speech video from JD Health International Inc. employees, resulting in the jdh-Hallo dataset. This dataset includes a diverse range of ages and speaking styles, encompassing both conversational and specialized medical topics. To adapt the JoyHallo model for Mandarin, we employed the Chinese wav2vec2 model for audio feature embedding. A semi-decoupled structure is proposed to capture inter-feature relationships among lip, expression, and pose features. This integration not only improves information utilization efficiency but also accelerates inference speed by 14.3%. Notably, JoyHallo maintains its strong ability to generate English videos, demonstrating excellent cross-language generation capabilities. The code and models are available at https://jdh-algo.github.io/JoyHallo.

📄 PDF Abstract BibTeX arXiv:2409.13268

Code (0)

등록된 구현이 없습니다.

Tasks

modelText GenerationVideo Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

A Quantitative Analysis of Comparison of Emoji Sentiment: Taiwan Mandarin Users and English Users

2022-11-01 · ROCLING 2022 11 · Fang-Yu Chang

Emojis have become essential components in our digital communication. Emojis, especially smiley face emojis and heart emojis, are considered the ones conveying more emotions. In this paper, two functions of emoji usages …

Language ModelingLanguage Modelling

Unsupervised Learning and Representation of Mandarin Tonal Categories by a Generative CNN

2025-09-22 · Kai Schenck, Gašper Beguš arxiv

This paper outlines the methodology for modeling tonal learning in fully unsupervised models of human language acquisition. Tonal patterns are among the computationally most complex learning objectives in language. We ar…

Language Acquisition

A unified sequence-to-sequence front-end model for Mandarin text-to-speech synthesis

2019-11-11 · Junjie Pan, Xiang Yin, Zhiling Zhang, Shichao Liu 외

In Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech. Building a typical pipeline-based front-end which consists…

Polyphone disambiguationSpeech Synthesistext-to-speechText to Speech+1

Annotating a corpus of human interaction with prosodic profiles --- focusing on Mandarin repair/disfluency

2012-05-01 · LREC 2012 5 · Helen Kai-yun Chen

This study describes the construction of a manually annotated speech corpus that focuses on the sound profiles of repair/disfluency in Mandarin conversational interaction. Specifically, the paper focuses on how the tag s…

TAG

SimpleNLG-ZH: a Linguistic Realisation Engine for Mandarin

2018-11-01 · WS 2018 11 · Guanyi Chen, Kees Van Deemter, Chenghua Lin

We introduce SimpleNLG-ZH, a realisation engine for Mandarin that follows the software design paradigm of SimpleNLG (Gatt and Reiter, 2009). We explain the core grammar (morphology and syntax) and the lexicon of SimpleNL…

Morphological InflectionText Generation