paper-with-me

Papers

IoT-Brain: Grounding LLMs for Semantic-Spatial Sensor Scheduling

2026-04-09 · Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Junda Lin, Jinke Song arxiv

Intelligent systems powered by large-scale sensor networks are shifting from predefined monitoring to intent-driven operation, revealing a critical Semantic-to-Physical Mapping Gap. While large language models (LLMs) excel at semantic understanding, existing perception-centric pipelines operate retrospectively, overlooking the fundamental decision of what to sense and when. We formalize this proactive decision as Semantic-Spatial Sensor Scheduling (S3) and demonstrate that direct LLM planning is unreliable due to inherent gaps in representation, reasoning, and optimization. To bridge these gaps, we introduce the Spatial Trajectory Graph (STG), a neuro-symbolic paradigm governed by a verify-before-commit discipline that transforms open-ended planning into a verifiable graph optimization problem. Based on STG, we implement IoT-Brain, a concrete system embodiment, and construct TopoSense-Bench, a campus-scale benchmark with 5,250 natural-language queries across 2,510 cameras. Evaluations show that IoT-Brain boosts task success rate by 37.6% over the strongest search-intensive methods while running nearly 2 times faster and using 6.6 times fewer prompt tokens. In real-world deployment, it approaches the reliability upper bound while reducing 4.1 times network bandwidth, providing a foundational framework for LLMs to interact with the physical world with unprecedented reliability and efficiency.

📄 PDF Abstract BibTeX arXiv:2604.08033

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans

2026-05-09 · Odysseas S. Chlapanis, Orfeas Menis Mastromichalakis, Christos H. Papadimitriou arxiv

Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social context. Do large language models (LLMs) gr…

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

2026-03-12 · Mohamad Alansari, Naufal Suryanto, Divya Velayudhan, Sajid Javed 외 arxiv

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temp…

Visual Grounding

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

2025-03-10 · Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi 외

This study investigated the multimodal perception of large language models (LLMs), focusing on their ability to capture human-like perceptual strength ratings across sensory modalities. Utilizing perceptual strength rati…

Faithful Grounded Visual Reasoning via Learned Proxy-Tokens

2026-06-22 · Tom Hodemon, Mohamed Chaouch, Aboubacar Tuo, Angelique Loesch arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. Grounded Visual Reasoning (GVR) approache…

Visual Question AnsweringVisual GroundingVisual Reasoning

Brain Mapping with Dense Features: Grounding Cortical Semantic Selectivity in Natural Images With Vision Transformers

2024-10-07 · Andrew F. Luo, Jacob Yeung, Rushikesh Zawar, Shaurya Dewan 외

We introduce BrainSAIL, a method for linking neural selectivity with spatially distributed semantic visual concepts in natural scenes. BrainSAIL leverages recent advances in large-scale artificial neural networks, using …

Denoising