paper-with-me

홈 › Papers

LARNet: Latent Action Representation for Human Action Synthesis

2021-10-21 · Naman Biyani, Aayush J Rana, Shruti Vyas, Yogesh S Rawat

We present LARNet, a novel end-to-end approach for generating human action videos. A joint generative modeling of appearance and dynamics to synthesize a video is very challenging and therefore recent works in video synthesis have proposed to decompose these two factors. However, these methods require a driving video to model the video dynamics. In this work, we propose a generative approach instead, which explicitly learns action dynamics in latent space avoiding the need of a driving video during inference. The generated action dynamics is integrated with the appearance using a recurrent hierarchical structure which induces motion at different scales to focus on both coarse as well as fine level action details. In addition, we propose a novel mix-adversarial loss function which aims at improving the temporal coherency of synthesized videos. We evaluate the proposed approach on four real-world human action datasets demonstrating the effectiveness of the proposed approach in generating human actions. Code available at https://github.com/aayushjr/larnet.

📄 PDF Abstract BibTeX arXiv:2110.10899

Code (1)

aayushjr/larnet 공식 구현 pytorch

Similar Papers 제목 키워드 기반

PolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation

2023-09-27 · ShiZhe Chen, Ricardo Garcia, Cordelia Schmid, Ivan Laptev

The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics. The dominant approaches for language-guided manipulation use 2D image representa…

Multi-Task LearningRobot ManipulationRobot Manipulation Generalization

PolarNet: Accelerated Deep Open Space Segmentation Using Automotive Radar in Polar Domain

2021-03-04 · Farzan Erlik Nowruzi, Dhanvin Kolhatkar, Prince Kapoor, Elnaz Jahani Heravi 외

Camera and Lidar processing have been revolutionized with the rapid development of deep learning model architectures. Automotive radar is one of the crucial elements of automated driver assistance and autonomous driving …

Autonomous DrivingDecision Making

Panoptic-PolarNet: Proposal-free LiDAR Point Cloud Panoptic Segmentation

2021-03-27 · CVPR 2021 1 · Zixiang Zhou, Yang Zhang, Hassan Foroosh

Panoptic segmentation presents a new challenge in exploiting the merits of both detection and segmentation, with the aim of unifying instance segmentation and semantic segmentation in a single framework. However, an effi…

ClusteringInstance SegmentationPanoptic SegmentationSegmentation+1

PillarNet: Real-Time and High-Performance Pillar-based 3D Object Detection

2022-05-16 · Guangsheng Shi, Ruifeng Li, Chao Ma

Real-time and high-performance 3D object detection is of critical importance for autonomous driving. Recent top-performing 3D object detectors mainly rely on point-based or 3D voxel-based convolutions, which are both com…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection+1

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

2026-05-29 · Xiang Zhu, Puzhen Yuan, Yichen Liu, Jianyu Chen arxiv

Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions. While latent…

Representation Learning