paper-with-me

Papers

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

2025-03-10 · Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi, Sung-Eun Lee

This study investigated the multimodal perception of large language models (LLMs), focusing on their ability to capture human-like perceptual strength ratings across sensory modalities. Utilizing perceptual strength ratings as a benchmark, the research compared GPT-3.5, GPT-4, GPT-4o, and GPT-4o-mini, highlighting the influence of multimodal inputs on grounding and linguistic reasoning. While GPT-4 and GPT-4o demonstrated strong alignment with human evaluations and significant advancements over smaller models, qualitative analyses revealed distinct differences in processing patterns, such as multisensory overrating and reliance on loose semantic associations. Despite integrating multimodal capabilities, GPT-4o did not exhibit superior grounding compared to GPT-4, raising questions about their role in improving human-like grounding. These findings underscore how LLMs' reliance on linguistic patterns can both approximate and diverge from human embodied cognition, revealing limitations in replicating sensory experiences.

📄 PDF Abstract BibTeX arXiv:2503.06980

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Attention 설명 없음

Similar Papers 제목 키워드 기반

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

2024-06-14 · Yiqi Wu, Xiaodan Hu, Ziming Fu, Siling Zhou 외

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a …

Activity RecognitionMMR totalSemantic correspondenceVideo Understanding+1

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

2025-07-07 · Zhicheng Zhang, Wuyou Xia, Chenxi Zhao, Zhou Yan 외 arxiv

Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by a generalizable attention architecture. Advanced methods predominantly focus on languag…

TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

2024-04-14 · Ya-Qi Yu, Minghui Liao, Jihao Wu, Yongxin Liao 외

Multimodal Large Language Models (MLLMs) have shown impressive results on various multimodal tasks. However, most existing MLLMs are not well suited for document-oriented tasks, which require fine-grained image perceptio…

Exploring Perceptual Limitation of Multimodal Large Language Models

2024-02-12 · Jiarui Zhang, Jinyi Hu, Mahyar Khayatkhoei, Filip Ilievski 외

Multimodal Large Language Models (MLLMs) have recently shown remarkable perceptual capability in answering visual questions, however, little is known about the limits of their perception. In particular, while prior works…

ObjectQuestion Answering

Exploring Embodied Multimodal Large Models: Development, Datasets, and Future Directions

2025-02-21 · Shoubin Chen, Zehao Wu, Kai Zhang, Chunyu Li 외

Embodied multimodal large models (EMLMs) have gained significant attention in recent years due to their potential to bridge the gap between perception, cognition, and action in complex, real-world environments. This comp…

Decision Making