paper-with-me

Papers

Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding

2025-05-12 · Oriol Barbany, Adrià Colomé, Carme Torras

Manipulating clothes is challenging due to their complex dynamics, high deformability, and frequent self-occlusions. Garments exhibit a nearly infinite number of configurations, making explicit state representations difficult to define. In this paper, we analyze BiFold, a model that predicts language-conditioned pick-and-place actions from visual observations, while implicitly encoding garment state through end-to-end learning. To address scenarios such as crumpled garments or recovery from failed manipulations, BiFold leverages temporal context to improve state estimation. We examine the internal representations of the model and present evidence that its fine-tuning and temporal context enable effective alignment between text and image regions, as well as temporal consistency.

📄 PDF Abstract BibTeX arXiv:2505.07600

Code (0)

등록된 구현이 없습니다.

Tasks

State Estimation

Similar Papers 제목 키워드 기반

Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning

2025-10-14 · Xingang Guo, Utkarsh Tyagi, Advait Gosai, Paula Vergara 외 arxiv

Multimodal Large Language Models (MLLMs) are increasingly applied in real-world scenarios where user-provided images are often imperfect, requiring active image manipulations such as cropping, editing, or enhancement to …

Graphormer-Guided Task Planning: Beyond Static Rules with LLM Safety Perception

2025-03-10 · Wanjing Huang, Tongjie Pan, Yalan Ye

Recent advancements in large language models (LLMs) have expanded their role in robotic task planning. However, while LLMs have been explored for generating feasible task sequences, their ability to ensure safe task exec…

Task Planning

Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

2024-04-04 · CVPR 2024 1 · Shuting He, Henghui Ding

Referring video segmentation relies on natural language expressions to identify and segment objects, often emphasizing motion clues. Previous works treat a sentence as a whole and directly perform identification at the v…

Contrastive LearningReferring ExpressionReferring Expression SegmentationReferring Video Object Segmentation+4

Describing Videos by Exploiting Temporal Structure

2015-02-27 · ICCV 2015 12 · Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas 외

Recent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description. However, while images are static, working with videos requires mod…

Action RecognitionImage DescriptionTemporal Action LocalizationVideo Description

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

2026-06-14 · Yanxin Xi, Xiang Su, Jie Feng, Yu Liu 외 arxiv

Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs). We introduce UrbanWe…