paper-with-me

Papers

SignDiff: Diffusion Model for American Sign Language Production

2023-08-30 · Sen Fang, Chunyu Sui, Yanghao Zhou, Xuedong Zhang, Hongbin Zhong, Yapeng Tian, Chen Chen

In this paper, we propose a dual-condition diffusion pre-training model named SignDiff that can generate human sign language speakers from a skeleton pose. SignDiff has a novel Frame Reinforcement Network called FR-Net, similar to dense human pose estimation work, which enhances the correspondence between text lexical symbols and sign language dense pose frames, reduces the occurrence of multiple fingers in the diffusion model. In addition, we propose a new method for American Sign Language Production (ASLP), which can generate ASL skeletal pose videos from text input, integrating two new improved modules and a new loss function to improve the accuracy and quality of sign language skeletal posture and enhance the ability of the model to train on large-scale data. We propose the first baseline for ASL production and report the scores of 17.19 and 12.85 on BLEU-4 on the How2Sign dev/test sets. We evaluated our model on the previous mainstream dataset PHOENIX14T, and the experiments achieved the SOTA results. In addition, our image quality far exceeds all previous results by 10 percentage points in terms of SSIM.

📄 PDF Abstract BibTeX arXiv:2308.16082

Code (0)

등록된 구현이 없습니다.

Tasks

Pose EstimationSign Language ProductionSSIM

Methods 이 논문이 사용한 방법론

American 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production

2025-09-13 · Liqian Feng, Lintao Wang, Kun Hu, Dehui Kong 외 arxiv

Sign language production (SLP) aims to translate spoken language sentences into a sequence of pose frames in a sign language, bridging the communication gap and promoting digital inclusion for deaf and hard-of-hearing co…

DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models

2025-03-03 · CVPR 2025 1 · Zhendong Wang, Jianmin Bao, Shuyang Gu, Dong Chen 외

In this paper, we present DesignDiffusion, a simple yet effective framework for the novel task of synthesizing design images from textual descriptions. A primary challenge lies in generating accurate and style-consistent…

Image GenerationText Generation

My LLM might Mimic AAE -- But When Should it?

2025-02-06 · Sandra C. Sandoval, Christabel Acquaye, Kwesi Cobbina, Mohammad Nayeem Teli 외

We examine the representation of African American English (AAE) in large language models (LLMs), exploring (a) the perceptions Black Americans have of how effective these technologies are at producing authentic AAE, and …

SDW-ASL: A Dynamic System to Generate Large Scale Dataset for Continuous American Sign Language

2022-10-13 · Yehong Jiang

Despite tremendous progress in natural language processing using deep learning techniques in recent years, sign language production and comprehension has advanced very little. One critical barrier is the lack of largesca…

Dataset GenerationDeep LearningSign Language Production

Challenges for Linguistically-Driven Computer-Based Sign Recognition from Continuous Signing for American Sign Language

2023-11-01 · Carol Neidle

There have been recent advances in computer-based recognition of isolated, citation-form signs from video. There are many challenges for such a task, not least the naturally occurring inter- and intra- signer synchronic …