paper-with-me

Papers

Graph-Structured Visual Imitation

2019-07-11 · Maximilian Sieb, Zhou Xian, Audrey Huang, Oliver Kroemer, Katerina Fragkiadaki

We cast visual imitation as a visual correspondence problem. Our robotic agent is rewarded when its actions result in better matching of relative spatial configurations for corresponding visual entities detected in its workspace and teacher's demonstration. We build upon recent advances in Computer Vision,such as human finger keypoint detectors, object detectors trained on-the-fly with synthetic augmentations, and point detectors supervised by viewpoint changes and learn multiple visual entity detectors for each demonstration without human annotations or robot interactions. We empirically show the proposed factorized visual representations of entities and their spatial arrangements drive successful imitation of a variety of manipulation skills within minutes, using a single demonstration and without any environment instrumentation. It is robust to background clutter and can effectively generalize across environment variations between demonstrator and imitator, greatly outperforming unstructured non-factorized full-frame CNN encodings of previous works.

📄 PDF Abstract BibTeX arXiv:1907.05518

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DocTr: Document Transformer for Structured Information Extraction in Documents

2023-07-16 · ICCV 2023 1 · Haofu Liao, Aruni RoyChowdhury, Weijian Li, Ankan Bansal 외

We present a new formulation for structured information extraction (SIE) from visually rich documents. It aims to address the limitations of existing IOB tagging or graph-based formulations, which are either overly relia…

Entity LinkingSemantic entity labeling

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering

2026-01-14 · Zhiyang Li, Ao Ke, Yukun Cao, Xike Xie arxiv

Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perception. Crucially, we identify that commo…

Visual Question Answering

Modular Diffusion Models for Structured Visual Recognition

2026-06-21 · Siddhesh Khandelwal, Björn Ommer, Leonid Sigal arxiv

Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often produce deterministic, fixed outputs, limiting their ability to captur…

Scene Graph GenerationInstance SegmentationObject Detection

GraphTSNE: A Visualization Technique for Graph-Structured Data

2019-04-15 · Yao Yang Leow, Thomas Laurent, Xavier Bresson

We present GraphTSNE, a novel visualization technique for graph-structured data based on t-SNE. The growing interest in graph-structured data increases the importance of gaining human insight into such datasets by means …

Dimensionality Reduction

Generating Animated Layouts as Structured Text Representations

2025-05-02 · Yeonsang Shin, JiHwan Kim, Yumin Song, Kyungseung Lee 외

Despite the remarkable progress in text-to-video models, achieving precise control over text elements and animated graphics remains a significant challenge, especially in applications such as video advertisements. To add…

Layout Generation