paper-with-me

홈 › Papers

Pose-Guided Sign Language Video GAN with Dynamic Lambda

2021-05-06 · Christopher Kissel, Christopher Kümmel, Dennis Ritter, Kristian Hildebrand

We propose a novel approach for the synthesis of sign language videos using GANs. We extend the previous work of Stoll et al. by using the human semantic parser of the Soft-Gated Warping-GAN from to produce photorealistic videos guided by region-level spatial layouts. Synthesizing target poses improves performance on independent and contrasting signers. Therefore, we have evaluated our system with the highly heterogeneous MS-ASL dataset with over 200 signers resulting in a SSIM of 0.893. Furthermore, we introduce a periodic weighting approach to the generator that reactivates the training and leads to quantitatively better results.

📄 PDF Abstract BibTeX arXiv:2105.02742

Code (0)

등록된 구현이 없습니다.

Tasks

SSIM

Similar Papers 제목 키워드 기반

Learning Question-Guided Video Representation for Multi-Turn Video Question Answering

2019-07-31 · WS 2019 9 · Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz, Dilek Hakkani-Tür 외

Understanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans. Video question answering is a specific scenario of such…

NavigateQuestion AnsweringText GenerationVideo Question Answering

Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query

2019-10-01 · ICCV 2019 10 · Hao Wang, Cheng Deng, Junchi Yan, Dacheng Tao

Actor and action video segmentation from natural language query aims to selectively segment the actor and its action in a video based on an input textual description. Previous works mostly focus on learning simple correl…

Referring Expression SegmentationSegmentationVideo SegmentationVideo Semantic Segmentation

PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement

2025-12-04 · Yu-Wei Zhan, Xin Wang, Hong Chen, Tongtong Feng 외 arxiv

Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper understanding of physical dynamics. This li…

Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval

2025-10-21 · Jiaao Yu, Mingjie Han, Tao Gong, Jian Zhang 외 arxiv

With the rapid growth of video data, text-video retrieval technology has become increasingly important in numerous application scenarios such as recommendation and search. Early text-video retrieval methods suffer from t…

Video AlignmentVideo Retrieval

Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning

2026-01-05 · Sungjune Park, Hongda Mao, Qingshuang Chen, Yong Man Ro 외 arxiv

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the i…