paper-with-me

Papers

Exploring Affordance and Situated Meaning in Image Captions: A Multimodal Analysis

2023-05-24 · Pin-Er Chen, Po-Ya Angela Wang, Hsin-Yu Chou, Yu-Hsiang Tseng, Shu-Kai Hsieh

This paper explores the grounding issue regarding multimodal semantic representation from a computational cognitive-linguistic view. We annotate images from the Flickr30k dataset with five perceptual properties: Affordance, Perceptual Salience, Object Number, Gaze Cueing, and Ecological Niche Association (ENA), and examine their association with textual elements in the image captions. Our findings reveal that images with Gibsonian affordance show a higher frequency of captions containing 'holding-verbs' and 'container-nouns' compared to images displaying telic affordance. Perceptual Salience, Object Number, and ENA are also associated with the choice of linguistic expressions. Our study demonstrates that comprehensive understanding of objects or events requires cognitive attention, semantic nuances in language, and integration across multiple modalities. We highlight the vital importance of situated meaning and affordance grounding in natural language understanding, with the potential to advance human-like interpretation in various scenarios.

📄 PDF Abstract BibTeX arXiv:2305.14616

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningNatural Language Understanding

Similar Papers 제목 키워드 기반

Exploring Situated Stabilities of a Rhythm Generation System through Variational Cross-Examination

2025-09-05 · Błażej Kotowski, Nicholas Evans, Behzad Haki, Frederic Font 외 arxiv

This paper investigates GrooveTransformer, a real-time rhythm generation system, through the postphenomenological framework of Variational Cross-Examination (VCE). By reflecting on its deployment across three distinct ar…

Exploring the Functional and Geometric Bias of Spatial Relations Using Neural Language Models

2018-06-01 · WS 2018 6 · Simon Dobnik, Mehdi Ghanimifard, John Kelleher

The challenge for computational models of spatial descriptions for situated dialogue systems is the integration of information from different modalities. The semantics of spatial descriptions are grounded in at least two…

Image Captioning

Accelerating the Development of Multimodal, Integrative-AI Systems with Platform for Situated Intelligence

2020-10-12 · Sean Andrist, Dan Bohus

We describe Platform for Situated Intelligence, an open-source framework for multimodal, integrative-AI systems. The framework provides infrastructure, tools, and components that enable and accelerate the development of …

What is it Like to Be a Bot: Simulated, Situated, Structurally Coherent Qualia (S3Q) Theory of Consciousness

2021-03-13 · K. Schmidt, J. Culbertson, C. Cox, H. S. Clouse 외

A novel representationalist theory of consciousness is presented that is grounded in neuroscience and provides a path to artificially conscious computing. Central to the theory are representational affordances of the con…

2D Human Pose Estimation2D Object Detection2D Semantic Segmentation2D Semantic Segmentation task 1 (8 classes)+1

Self-Explainable Affordance Learning with Embodied Caption

2024-04-08 · Zhipeng Zhang, Zhimin Wei, Guolei Sun, Peng Wang 외

In the field of visual affordance learning, previous methods mainly used abundant images or videos that delineate human behavior patterns to identify action possibility regions for object manipulation, with a variety of …