paper-with-me

홈 › Papers

Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication

2026-06-16 · Peter Zeng, Amie J. Paige, Weiling Li, Susan E. Brennan, Owen Rambow, Cameron R. Jones arxiv

Two recent studies (Jones et al. (2026); Zeng et al. (2026)) reach apparently contradictory conclusions about whether LVLMs can coordinate on efficient referring expressions. We control for task differences between the studies while directly comparing their prompting styles. We replicate the finding that models can coordinate efficient referring expressions when explicitly prompted to do so, suggesting that other task differences are not responsible for divergent results. However, we also find that the same models fail to infer the need for communicative efficiency from a more implicit prompt, highlighting critical differences between how humans and AI systems communicate.

📄 PDF Abstract BibTeX arXiv:2606.17372

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering

2024-07-31 · Danfeng Guo, Demetri Terzopoulos

Large Vision-Language Models (LVLMs) have achieved significant success in recent years, and they have been extended to the medical domain. Although demonstrating satisfactory performance on medical Visual Question Answer…

DiagnosticHallucinationMedical Visual Question AnsweringQuestion Answering+2

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

2023-09-28 · Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero 외

Popular prompt strategies like Chain-of-Thought Prompting can dramatically improve the reasoning abilities of Large Language Models (LLMs) in various domains. However, such hand-crafted prompt-strategies are often sub-op…

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

2025-07-24 · Tianheng Qiu, Jingchun Gao, Jingyu Li, Huiyi Leong 외 arxiv

Intent-oriented controlled video captioning aims to generate targeted descriptions for specific targets in a video based on customized user intent. Current Large Visual Language Models (LVLMs) have gained strong instruct…

Instruction FollowingVideo Captioning

DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models

2026-01-02 · Yue Zhou, Jue Chen, Zilun Zhang, Penghui Huang 외 arxiv

Remote sensing (RS) large vision-language models (LVLMs) have shown strong promise across visual grounding (VG) tasks. However, existing RS VG datasets predominantly rely on explicit referring expressions-such as relativ…

Reinforcement LearningVisual Grounding

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

2026-04-02 · Bhaskara Hanuma Vedula, Darshan Anghan, Ishita Goyal, Ponnurangam Kumaraguru 외 arxiv

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based pr…