Towards Generalizing Temporal Action Segmentation to Unseen Views
While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen view action segmentation where camera views for evaluating the model are unavailable during training. This includes changing from top-frontal views to a side view or even more challenging from exocentric to egocentric views. Furthermore, we present an approach for temporal action segmentation that tackles this challenge. Our approach leverages a shared representation at both the sequence and segment levels to reduce the impact of view differences during training. We achieve this by introducing a sequence loss and an action loss, which together facilitate consistent video and action representations across different views. The evaluation on the Assembly101, IkeaASM, and EgoExoLearn datasets demonstrate significant improvements, with a 12.8% increase in F1@50 for unseen exocentric views and a substantial 54% improvement for unseen egocentric views.
Code (0)
등록된 구현이 없습니다.
Tasks
Action SegmentationSegmentationTemporal Action SegmentationSimilar Papers 제목 키워드 기반
Cascaded and Generalizable Neural Radiance Fields for Fast View Synthesis
We present CG-NeRF, a cascade and generalizable neural radiance fields method for view synthesis. Recent generalizing view synthesis methods can render high-quality novel views using a set of nearby input views. However,…
GPUNeRFNeural RenderingNovel View SynthesisKeypointNeRF: Generalizing Image-based Volumetric Avatars using Relative Spatial Encoding of Keypoints
Image-based volumetric humans using pixel-aligned features promise generalization to unseen poses and identities. Prior work leverages global spatial encodings and multi-view geometric consistency to reduce spatial ambig…
3D Face Reconstruction3D Human ReconstructionGeneralizable Novel View SynthesisNovel View SynthesisLatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion.
Current research on generating 3D hand-object interaction motion primarily focuses on in-domain objects. Generalization to unseen objects is essential for practical applications, yet it remains both challenging and l…
Motion GenerationObjectGeneralizing Word Embeddings using Bag of Subwords
We approach the problem of generalizing pre-trained word embeddings beyond fixed-size vocabularies without using additional contextual information. We propose a subword-level word vector generation model that views words…
TAGWord EmbeddingsWord SimilarityOG-VLA: 3D-Aware Vision Language Action Model via Orthographic Image Generation
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of ma…
Image GenerationLarge Language ModelRobot ManipulationVision-Language-Action