paper-with-me

Papers

Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation

2024-11-06 · Ke Fan, Jiangning Zhang, Ran Yi, Jingyu Gong, Yabiao Wang, Yating Wang, Xin Tan, Chengjie Wang, Lizhuang Ma

Text-to-motion generation is a crucial task in computer vision, which generates the target 3D motion by the given text. The existing annotated datasets are limited in scale, resulting in most existing methods overfitting to the small datasets and unable to generalize to the motions of the open domain. Some methods attempt to solve the open-vocabulary motion generation problem by aligning to the CLIP space or using the Pretrain-then-Finetuning paradigm. However, the current annotated dataset's limited scale only allows them to achieve mapping from sub-text-space to sub-motion-space, instead of mapping between full-text-space and full-motion-space (full mapping), which is the key to attaining open-vocabulary motion generation. To this end, this paper proposes to leverage the atomic motion (simple body part motions over a short time period) as an intermediate representation, and leverage two orderly coupled steps, i.e., Textual Decomposition and Sub-motion-space Scattering, to address the full mapping problem. For Textual Decomposition, we design a fine-grained description conversion algorithm, and combine it with the generalization ability of a large language model to convert any given motion text into atomic texts. Sub-motion-space Scattering learns the compositional process from atomic motions to the target motions, to make the learned sub-motion-space scattered to form the full-motion-space. For a given motion of the open domain, it transforms the extrapolation into interpolation and thereby significantly improves generalization. Our network, $DSO$-Net, combines textual $d$ecomposition and sub-motion-space $s$cattering to solve the $o$pen-vocabulary motion generation. Extensive experiments demonstrate that our DSO-Net achieves significant improvements over the state-of-the-art methods on open-vocabulary motion generation. Code is available at https://vankouf.github.io/DSONet/.

📄 PDF Abstract BibTeX arXiv:2411.04079

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelMotion Generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Machine olfaction using time scattering of sensor multiresolution graphs

2016-02-13 · Leonid Gugel, Yoel Shkolnisky, Shai Dekel

In this paper we construct a learning architecture for high dimensional time series sampled by sensor arrangements. Using a redundant wavelet decomposition on a graph constructed over the sensor locations, our algorithm …

BIG-bench Machine LearningTime SeriesTime Series Analysis

CNN-Based Invertible Wavelet Scattering for the Investigation of Diffusion Properties of the In Vivo Human Heart in Diffusion Tensor Imaging

2019-12-17 · Zeyu Deng, Lihui Wang, Zixiang Kuai, Qijian Chen 외

In vivo diffusion tensor imaging (DTI) is a promising technique to investigate noninvasively the fiber structures of the in vivo human heart. However, signal loss due to motions remains a persistent problem in in vivo ca…

Motion CompensationTranslation

Unsupervised Classification of PolSAR Data Using a Scattering Similarity Measure Derived from a Geodesic Distance

2017-12-01 · Debanshu Ratha, Avik Bhattacharya, Alejandro C. Frery

In this letter, we propose a novel technique for obtaining scattering components from Polarimetric Synthetic Aperture Radar (PolSAR) data using the geodesic distance on the unit sphere. This geodesic distance is obtained…

General Classification

On Level Crossings and Fade Durations in von Mises-Fisher Scattering Channels

2025-06-06 · Kenan Turbic, Slawomir Stanczak

This paper investigates the second-order statistics of multipath fading channels with von Mises-Fisher (vMF) distributed scatters. Simple closed-form expressions for the mean Doppler shift and Doppler spread are derived …

Deep scattering network for speech emotion recognition

2021-05-11 · Premjeet Singh, Goutam Saha, Md Sahidullah

This paper introduces scattering transform for speech emotion recognition (SER). Scattering transform generates feature representations which remain stable to deformations and shifting in time and frequency without much …

Emotion RecognitionSpeech Emotion Recognition