paper-with-me

홈 › Papers

A Joint Model of Language and Perception for Grounded Attribute Learning

2012-06-27 · Cynthia Matuszek, Nicholas FitzGerald, Luke Zettlemoyer, Liefeng Bo, Dieter Fox

As robots become more ubiquitous and capable, it becomes ever more important to enable untrained users to easily interact with them. Recently, this has led to study of the language grounding problem, where the goal is to extract representations of the meanings of natural language tied to perception and actuation in the physical world. In this paper, we present an approach for joint learning of language and perception models for grounded attribute induction. Our perception model includes attribute classifiers, for example to detect object color and shape, and the language model is based on a probabilistic categorial grammar that enables the construction of rich, compositional meaning representations. The approach is evaluated on the task of interpreting sentences that describe sets of objects in a physical workspace. We demonstrate accurate task performance and effective latent-variable concept induction in physical grounded scenes.

📄 PDF Abstract BibTeX arXiv:1206.6423

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

2026-08-13 · Yifan Mei, Qingling Shi, Changli Wu, Jiayuan Rao 외 arxiv

Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling…

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

2025-07-23 · Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li 외 arxiv

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challeng…

Relational Reasoning

Semantically Grounded Visual Embeddings for Zero-Shot Learning

2022-01-03 · Shah Nawaz, Jacopo Cavazza, Alessio Del Bue

Zero-shot learning methods rely on fixed visual and semantic embeddings, extracted from independent vision and language models, both pre-trained for other large-scale tasks. This is a weakness of current zero-shot learni…

Zero-Shot Learning

Representations of language in a model of visually grounded speech signal

2017-02-07 · ACL 2017 7 · Grzegorz Chrupała, Lieke Gelderloos, Afra Alishahi

We present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space. We use a multi-layer recurrent highway network to model the temporal nature of spoken speec…

Form

EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery

2025-09-19 · Gui Wang, Yang Wennuo, Xusen Ma, Zehao Zhong 외 arxiv

MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap…

Clinical Knowledge