paper-with-me

홈 › Papers

How Physics and Background Attributes Impact Video Transformers in Robotic Manipulation: A Case Study on Planar Pushing

2023-10-03 · Shutong Jin, Ruiyu Wang, Muhammad Zahid, Florian T. Pokorny

As model and dataset sizes continue to scale in robot learning, the need to understand how the composition and properties of a dataset affect model performance becomes increasingly urgent to ensure cost-effective data collection and model performance. In this work, we empirically investigate how physics attributes (color, friction coefficient, shape) and scene background characteristics, such as the complexity and dynamics of interactions with background objects, influence the performance of Video Transformers in predicting planar pushing trajectories. We investigate three primary questions: How do physics attributes and background scene characteristics influence model performance? What kind of changes in attributes are most detrimental to model generalization? What proportion of fine-tuning data is required to adapt models to novel scenarios? To facilitate this research, we present CloudGripper-Push-1K, a large real-world vision-based robot pushing dataset comprising 1278 hours and 460,000 videos of planar pushing interactions with objects with different physics and background attributes. We also propose Video Occlusion Transformer (VOT), a generic modular video-transformer-based trajectory prediction framework which features 3 choices of 2D-spatial encoders as the subject of our case study. The dataset and source code are available at https://cloudgripper.org.

📄 PDF Abstract BibTeX arXiv:2310.02044

Code (0)

등록된 구현이 없습니다.

Tasks

FrictionObjectTrajectory Prediction

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Attention 설명 없음
Adam 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

A Comprehensive Study of Image Classification Model Sensitivity to Foregrounds, Backgrounds, and Visual Attributes

2022-01-26 · CVPR 2022 1 · Mazda Moayeri, Phillip Pope, Yogesh Balaji, Soheil Feizi

While datasets with single-label supervision have propelled rapid advances in image classification, additional annotations are necessary in order to quantitatively assess how models make predictions. To this end, for a s…

image-classificationImage ClassificationSensitivity

PAVAS: Physics-Aware Video-to-Audio Synthesis

2025-12-09 · Oh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka, Tae-Hyun Oh 외 arxiv

Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing visual-acoustic correlations without consid…

3D Reconstruction

vid-TLDR: Training Free Token merging for Light-weight Video Transformer

2024-03-20 · CVPR 2024 1 · Joonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi 외

Video Transformers have become the prevalent solution for various video downstream tasks with superior expressive power and flexibility. However, these video transformers suffer from heavy computational costs induced by …

Action RecognitionComputational EfficiencyText RetrievalVideo Question Answering+5

Physics-Driven Diffusion Models for Impact Sound Synthesis from Videos

2023-03-29 · CVPR 2023 1 · Kun Su, Kaizhi Qian, Eli Shlizerman, Antonio Torralba 외

Modeling sounds emitted from physical object interactions is critical for immersive perceptual experiences in real and virtual worlds. Traditional methods of impact sound synthesis use physics simulation to obtain a set …

Extracting Latent Attributes from Video Scenes Using Text as Background Knowledge

2014-08-01 · SEMEVAL 2014 8 · Anh Tran, Mihai Surdeanu, Paul Cohen
Coreference ResolutionInformation Retrieval