paper-with-me

Papers

Large Generative Model-assisted Talking-face Semantic Communication System

2024-11-06 · Feibo Jiang, Siwei Tu, Li Dong, Cunhua Pan, Jiangzhou Wang, Xiaohu You

The rapid development of generative Artificial Intelligence (AI) continually unveils the potential of Semantic Communication (SemCom). However, current talking-face SemCom systems still encounter challenges such as low bandwidth utilization, semantic ambiguity, and diminished Quality of Experience (QoE). This study introduces a Large Generative Model-assisted Talking-face Semantic Communication (LGM-TSC) System tailored for the talking-face video communication. Firstly, we introduce a Generative Semantic Extractor (GSE) at the transmitter based on the FunASR model to convert semantically sparse talking-face videos into texts with high information density. Secondly, we establish a private Knowledge Base (KB) based on the Large Language Model (LLM) for semantic disambiguation and correction, complemented by a joint knowledge base-semantic-channel coding scheme. Finally, at the receiver, we propose a Generative Semantic Reconstructor (GSR) that utilizes BERT-VITS2 and SadTalker models to transform text back into a high-QoE talking-face video matching the user's timbre. Simulation results demonstrate the feasibility and effectiveness of the proposed LGM-TSC system.

📄 PDF Abstract BibTeX arXiv:2411.03876

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelSemantic Communication

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

2026-04-09 · Chanhyuk Choi, Taesoo Kim, Donggyu Lee, Siyeol Jung 외 arxiv

Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial …

Talking Face Generation

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

2024-02-25 · Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang 외

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's sp…

Face GenerationHallucinationTalking Face Generation

SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing

2024-09-05 · Lingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu 외

Audio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth)…

Face GenerationFacial EditingSegmentationTalking Face Generation

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

2024-01-16 · Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang 외

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simulta…

3D ReconstructionFace GenerationSuper-ResolutionTalking Face Generation

Speech4Mesh: Speech-Assisted Monocular 3D Facial Reconstruction for Speech-Driven 3D Facial Animation

2023-01-01 · ICCV 2023 1 · Shan He, Haonan He, Shuo Yang, Xiaoyan Wu 외

Recent audio2mesh-based methods have shown promising prospects for speech-driven 3D facial animation tasks. However, some intractable challenges are urgent to be settled. For example, the data-scarcity problem is int…

Contrastive Learning