Papers Robot Manipulation Generalization
“Robot Manipulation Generalization” 태그가 달린 논문 19편 · 필터 해제
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel…
Motion PlanningRobot ManipulationRobot Manipulation GeneralizationTask PlanningSam2Rad: A Segmentation Model for Medical Images with Learnable Prompts
Foundation models like the segment anything model require high-quality manual prompts for medical image segmentation, which is time-consuming and requires expertise. SAM and its variants often fail to segment structures …
Image SegmentationMedical Image Segmentationparameter-efficient fine-tuningPrompt Learning+3SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
The advent of large models, also known as foundation models, has significantly transformed the AI research landscape, with models like Segment Anything (SAM) achieving notable success in diverse image segmentation scenar…
Image SegmentationMedical Image Segmentationobject-detectionObject Detection+4SAM 2: Segment Anything in Images and Videos
We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect …
Image SegmentationRobot Manipulation GeneralizationSegmentationSemantic Segmentation+5Segment Anything for Videos: A Systematic Survey
The recent wave of foundation models has witnessed tremendous success in computer vision (CV) and beyond, with the segment anything model (SAM) having sparked a passion for exploring task-agnostic visual foundation model…
Image SegmentationRobot Manipulation GeneralizationSemantic SegmentationSurvey+4Generative Image as Action Models
Image-generation diffusion models have been fine-tuned to unlock new capabilities such as image-editing and novel view synthesis. Can we similarly unlock image-generation models for visuomotor control? We present GENIMA,…
Image GenerationRobot ManipulationRobot Manipulation GeneralizationRVT-2: Learning Precise Manipulation from Few Demonstrations
In this work, we study how to build a robotic system that can solve multiple 3D manipulation tasks given language instructions. To be useful in industrial and household domains, such a system should be capable of learnin…
Robot ManipulationRobot Manipulation GeneralizationEvaluating Real-World Robot Manipulation Policies in Simulation
The field of robotics has made significant advances towards generalist robot manipulation policies. However, real-world evaluation of such policies is not scalable and faces reproducibility challenges, which are likely t…
Robotic GraspingRobot ManipulationRobot Manipulation Generalization3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
We marry diffusion policies and 3D scene representations for robot manipulation. Diffusion policies learn the action distribution conditioned on the robot and environment state using conditional diffusion models. They ha…
DenoisingRobot ManipulationRobot Manipulation GeneralizationZero-shot GeneralizationTHE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions. Unfortunately, a majority of studies evaluate robot performanc…
Robot Manipulation GeneralizationPolarNet: 3D Point Clouds for Language-Guided Robotic Manipulation
The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics. The dominant approaches for language-guided manipulation use 2D image representa…
Multi-Task LearningRobot ManipulationRobot Manipulation GeneralizationRVT: Robotic View Transformer for 3D Object Manipulation
For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adver…
ObjectRobot ManipulationRobot Manipulation GeneralizationLearning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Fine manipulation tasks, such as threading cable ties or slotting a battery, are notoriously difficult for robots because they require precision, careful coordination of contact forces, and closed-loop visual feedback. P…
ChunkingImitation LearningRobot ManipulationRobot Manipulation GeneralizationSegment Anything
We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with…
Event-based Object SegmentationImage SegmentationRobot Manipulation GeneralizationSegmentation+3Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
Transformers have revolutionized vision and natural language processing with their ability to scale with large datasets. But in robotic manipulation, data is both limited and expensive. Can manipulation still benefit fro…
Robot ManipulationRobot Manipulation GeneralizationInstruction-driven history-aware policies for robotic manipulations
In human environments, robots are expected to accomplish a variety of manipulation tasks given simple natural language instructions. Yet, robotic manipulation is extremely challenging as it requires fine-grained motor co…
Robot ManipulationRobot Manipulation Generalization0/1 Deep Neural Networks via Block Coordinate Descent
The step function is one of the simplest and most natural activation functions for deep neural networks (DNNs). As it counts 1 for positive variables and 0 for others, its intrinsic characteristics (e.g., discontinuity a…
10-shot image generation16k2D Object Detection+92R3M: A Universal Visual Representation for Robot Manipulation
We study how visual representations pre-trained on diverse human video data can enable data-efficient learning of downstream robotic manipulation tasks. Concretely, we pre-train a visual representation using the Ego4D hu…
Contrastive LearningRobot ManipulationRobot Manipulation GeneralizationMasked Visual Pre-training for Motor Control
This paper shows that self-supervised visual pre-training from real-world images is effective for learning motor control tasks from pixels. We first train the visual representations by masked modeling of natural images. …
Robot Manipulation GeneralizationState Estimation