paper-with-me

홈 › Papers

TransFusion: A Practical and Effective Transformer-based Diffusion Model for 3D Human Motion Prediction

2023-07-30 · Sibo Tian, Minghui Zheng, Xiao Liang

Predicting human motion plays a crucial role in ensuring a safe and effective human-robot close collaboration in intelligent remanufacturing systems of the future. Existing works can be categorized into two groups: those focusing on accuracy, predicting a single future motion, and those generating diverse predictions based on observations. The former group fails to address the uncertainty and multi-modal nature of human motion, while the latter group often produces motion sequences that deviate too far from the ground truth or become unrealistic within historical contexts. To tackle these issues, we propose TransFusion, an innovative and practical diffusion-based model for 3D human motion prediction which can generate samples that are more likely to happen while maintaining a certain level of diversity. Our model leverages Transformer as the backbone with long skip connections between shallow and deep layers. Additionally, we employ the discrete cosine transform to model motion sequences in the frequency space, thereby improving performance. In contrast to prior diffusion-based models that utilize extra modules like cross-attention and adaptive layer normalization to condition the prediction on past observed motion, we treat all inputs, including conditions, as tokens to create a more lightweight model compared to existing approaches. Extensive experimental studies are conducted on benchmark datasets to validate the effectiveness of our human motion prediction model.

📄 PDF Abstract BibTeX arXiv:2307.16106

Code (1)

sibotian96/TransFusion 공식 구현 pytorch

Tasks

Human motion predictionHuman Pose Forecastingmotion prediction

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

2024-08-20 · Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala 외

We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single tra…

Language ModelingLanguage Modelling

TransFusion: Generating Long, High Fidelity Time Series using Diffusion Models with Transformers

2023-07-24 · Md Fahim Sikder, Resmi Ramachandranpillai, Fredrik Heintz

The generation of high-quality, long-sequenced time-series data is essential due to its wide range of applications. In the past, standalone Recurrent and Convolutional Neural Network-based Generative Adversarial Networks…

Time Series

TransFusion: Transcribing Speech with Multinomial Diffusion

2022-10-14 · Matthew Baas, Kevin Eloff, Herman Kamper

Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional text synthesis. Denoising diffusion model…

DenoisingImage GenerationSentencespeech-recognition+1

TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation

2021-10-18 · Haoyu Ma, Liangjian Chen, Deying Kong, Zhe Wang 외

Estimating the 2D human poses in each view is typically the first step in calibrated multi-view 3D pose estimation. But the performance of 2D pose detectors suffers from challenging situations such as occlusions and obli…

3D Human Pose Estimation3D Pose EstimationPose Estimation

TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers

2022-03-22 · CVPR 2022 1 · Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 외

LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illu…

3D Object DetectionAutonomous DrivingDecoderobject-detection+2