paper-with-me

Papers

MResT: Multi-Resolution Sensing for Real-Time Control with Vision-Language Models

2024-01-25 · Saumya Saxena, Mohit Sharma, Oliver Kroemer

Leveraging sensing modalities across diverse spatial and temporal resolutions can improve performance of robotic manipulation tasks. Multi-spatial resolution sensing provides hierarchical information captured at different spatial scales and enables both coarse and precise motions. Simultaneously multi-temporal resolution sensing enables the agent to exhibit high reactivity and real-time control. In this work, we propose a framework, MResT (Multi-Resolution Transformer), for learning generalizable language-conditioned multi-task policies that utilize sensing at different spatial and temporal resolutions using networks of varying capacities to effectively perform real time control of precise and reactive tasks. We leverage off-the-shelf pretrained vision-language models to operate on low-frequency global features along with small non-pretrained models to adapt to high frequency local feedback. Through extensive experiments in 3 domains (coarse, precise and dynamic manipulation tasks), we show that our approach significantly improves (2X on average) over recent multi-task baselines. Further, our approach generalizes well to visual and geometric variations in target objects and to varying interaction forces.

📄 PDF Abstract BibTeX arXiv:2401.14502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FiMReSt: Finite Mixture of Multivariate Regulated Skew-t Kernels -- A Flexible Probabilistic Model for Multi-Clustered Data with Asymmetrically-Scattered Non-Gaussian Kernels

2023-05-15 · Sarmad Mehrdad, S. Farokh Atashzar

Recently skew-t mixture models have been introduced as a flexible probabilistic modeling technique taking into account both skewness in data clusters and the statistical degree of freedom (S-DoF) to improve modeling gene…

Tags2Parts: Discovering Semantic Regions from Shape Tags

2017-08-22 · CVPR 2018 6 · Sanjeev Muralikrishnan, Vladimir G. Kim, Siddhartha Chaudhuri

We propose a novel method for discovering shape regions that strongly correlate with user-prescribed tags. For example, given a collection of chairs tagged as either "has armrest" or "lacks armrest", our system correctly…

TAG

GECOR: An End-to-End Generative Ellipsis and Co-reference Resolution Model for Task-Oriented Dialogue

2019-09-26 · IJCNLP 2019 11 · Jun Quan, Deyi Xiong, Bonnie Webber, Changjian Hu

Ellipsis and co-reference are common and ubiquitous especially in multi-turn dialogues. In this paper, we treat the resolution of ellipsis and co-reference in dialogue as a problem of generating omitted or referred expre…

Multi-Task Learning

The STONE Transform: Multi-Resolution Image Enhancement and Real-Time Compressive Video

2013-11-14 · Tom Goldstein, Lina Xu, Kevin F. Kelly, Richard Baraniuk

Compressed sensing enables the reconstruction of high-resolution signals from under-sampled data. While compressive methods simplify data acquisition, they require the solution of difficult recovery problems to make use …

compressed sensingCompressive SensingImage Enhancement

HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention

2025-11-29 · Yi Liu, Yi Wan, Xinyi Liu, Qiong Wu 외 arxiv

In remote sensing applications, such as disaster detection and response, real-time efficiency and model lightweighting are of critical importance. Consequently, existing remote sensing image super-resolution methods ofte…

Computational EfficiencyImage Super-Resolution