paper-with-me

홈 › Papers

MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes

2025-09-16 · Liu Liu, Alexandra Kudaeva, Marco Cipriano, Fatimeh Al Ghannam, Freya Tan, Gerard de Melo, Andres Sevtsuk arxiv

Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual cues such as relations, proximity, and co-movement - semantically complex signals that go beyond traditional object detection. To address this challenge, we introduce a social group region detection task, which requires inferring and spatially grounding visual regions defined by abstract interpersonal relations. We propose MINGLE (Modeling INterpersonal Group-Level Engagement), a modular three-stage pipeline that integrates: (1) off-the-shelf human detection and depth estimation, (2) VLM-based reasoning to classify pairwise social affiliation, and (3) a lightweight spatial aggregation algorithm to localize socially connected groups. To support this task and encourage future research, we present a new dataset of 100K urban street-view images annotated with bounding boxes and labels for both individuals and socially interacting groups. The annotations combine human-created labels and outputs from the MINGLE pipeline, ensuring semantic richness and broad coverage of real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2509.13484

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationObject Detection

Similar Papers 제목 키워드 기반

Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs

2026-02-17 · Guangtao Lyu, Qi Liu, Chenghao Xu, Jiexi Yan 외 arxiv

LVLMs have achieved strong multimodal reasoning capabilities but remain prone to hallucinations, producing outputs inconsistent with visual inputs or user instructions. Existing training-free methods, including contrasti…

Multimodal ReasoningVisual Grounding

ViLLA: Fine-Grained Vision-Language Representation Learning from Real-World Data

2023-08-22 · ICCV 2023 1 · Maya Varma, Jean-Benoit Delbrouck, Sarah Hooper, Akshay Chaudhari 외

Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are sign…

Attributeobject-detectionObject DetectionRepresentation Learning+2

VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation

2024-11-18 · Ruiyang Zhang, Hu Zhang, Zhedong Zheng

Given the higher information load processed by large vision-language models (LVLMs) compared to single-modal LLMs, detecting LVLM hallucinations requires more human and time expense, and thus rise a wider safety concerns…

HallucinationLanguage ModelingLanguage Modelling

QuISE: Defense against Typographic Attacks on VLMs via Query-Irrelevant Semantic Editing

2026-08-13 · Shubin Lu, Jiaqi Yin, Yihao Huang arxiv

Typographic attacks pose a critical threat to vision-language models (VLMs) by injecting misleading text into images and causing models to rely on adversarial textual cues rather than visual evidence. Existing defenses o…

VModA: An Effective Framework for Adaptive NSFW Image Moderation

2025-05-29 · Han Bao, Qinying Wang, Zhi Chen, Qingming Li 외

Not Safe/Suitable for Work (NSFW) content is rampant on social networks and poses serious harm to citizens, especially minors. Current detection methods mainly rely on deep learning-based image recognition and classifica…