paper-with-me

Papers

STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing

2025-03-29 · Zijun Ding, Mingdie Xiong, Congcong Zhu, Jingrun Chen

Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We argue that aligning the semantic features from spatial and temporal domains is a promising approach to stabilizing facial motion. To achieve this, we propose a Spatial-Temporal Semantic Alignment (STSA) method, which introduces a dual-path alignment mechanism and a differentiable semantic representation. The former leverages a Consistent Information Learning (CIL) module to maximize the mutual information at multiple scales, thereby reducing the manifold differences between spatial and temporal domains. The latter utilizes probabilistic heatmap as ambiguity-tolerant guidance to avoid the abnormal dynamics of the synthesized faces caused by slight semantic jittering. Extensive experimental results demonstrate the superiority of the proposed STSA, especially in terms of image quality and synthesis stability. Pre-trained weights and inference code are available at https://github.com/SCAILab-USTC/STSA.

📄 PDF Abstract BibTeX arXiv:2503.23039

Code (1)

scailab-ustc/stsa 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Heatmap 설명 없음

Similar Papers 제목 키워드 기반

Spatial-Temporal Transformer for 3D Point Cloud Sequences

2021-10-19 · Yimin Wei, Hao liu, TingTing Xie, Qiuhong Ke 외

Effective learning of spatial-temporal information within a point cloud sequence is highly important for many down-stream tasks such as 4D semantic segmentation and 3D action recognition. In this paper, we propose a nove…

3D Action RecognitionAction RecognitionSegmentationSemantic Segmentation

DSTSA-GCN: Advancing Skeleton-Based Gesture Recognition with Semantic-Aware Spatio-Temporal Topology Modeling

2025-01-21 · Hu Cui, Renjing Huang, Ruoyu Zhang, Tessai Hayama

Graph convolutional networks (GCNs) have emerged as a powerful tool for skeleton-based action and gesture recognition, thanks to their ability to model spatial and temporal dependencies in skeleton data. However, existin…

Action RecognitionGesture RecognitionHand Gesture RecognitionSkeleton Based Action Recognition

Spatio-Temporal Self-Attention Network for Video Saliency Prediction

2021-08-24 · Ziqiang Wang, Zhi Liu, Gongyang Li, Yang Wang 외

3D convolutional neural networks have achieved promising results for video tasks in computer vision, including video saliency prediction that is explored in this paper. However, 3D convolution encodes visual representati…

PredictionSaliency PredictionVideo Saliency Prediction

Adapting Segment Anything Model for Change Detection in HR Remote Sensing Images

2023-09-04 · Lei Ding, Kun Zhu, Daifeng Peng, Hao Tang 외

Vision Foundation Models (VFMs) such as the Segment Anything Model (SAM) allow zero-shot or interactive segmentation of visual contents, thus they are quickly applied in a variety of visual scenes. However, their direct …

Change DetectionInteractive Segmentation

CAD -- Contextual Multi-modal Alignment for Dynamic AVQA

2023-10-25 · Asmar Nadeem, Adrian Hilton, Robert Dawes, Graham Thomas 외

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA methods suffer from two major shortcomi…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringVisual Question Answering