paper-with-me

Papers

TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References

2025-12-25 · Jiahong Yu, Ziqi Wang, Hailiang Zhao, Wei Zhai, Xueqiang Yan, Shuiguang Deng arxiv

Understanding natural-language references to objects in dynamic 3D driving scenes is essential for interactive autonomous systems. In practice, many referring expressions describe targets through recent motion or short-term interactions, which cannot be resolved from static appearance or geometry alone. We study temporal language-based 3D grounding, where the objective is to identify the referred object in the current frame by leveraging multi-frame observations. We propose TrackTeller, a temporal multimodal grounding framework that integrates LiDAR-image fusion, language-conditioned decoding, and temporal reasoning in a unified architecture. TrackTeller constructs a shared UniScene representation aligned with textual semantics, generates language-aware 3D proposals, and refines grounding decisions using motion history and short-term dynamics. Experiments on the NuPrompt benchmark demonstrate that TrackTeller consistently improves language-grounded tracking performance, outperforming strong baselines with a 70% relative improvement in Average Multi-Object Tracking Accuracy and a 3.15-3.4 times reduction in False Alarm Frequency.

📄 PDF Abstract BibTeX arXiv:2512.21641

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Object Tracking

Similar Papers 제목 키워드 기반

Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training

2024-05-09 · Sheng Yan, Xin Du, Zongying Li, Yi Wang 외

Temporal grounding is crucial in multimodal learning, but it poses challenges when applied to animal behavior data due to the sparsity and uniform distribution of moments. To address these challenges, we propose a novel …

Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

2024-02-25 · Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang 외

Despite progress in multimodal large language models (MLLMs), the challenge of interpreting long-form videos in response to linguistic queries persists, largely due to the inefficiency in temporal grounding and limited p…

Computational EfficiencyLanguage ModellingOptical Flow EstimationQuestion Answering+1

SimBase: A Simple Baseline for Temporal Video Grounding

2024-11-12 · Peijun Bao, Alex C. Kot

This paper presents SimBase, a simple yet effective baseline for temporal video grounding. While recent advances in temporal grounding have led to impressive performance, they have also driven network architectures towar…

Video Grounding

From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning

2026-05-19 · Ahmed Y. Gado, Omar Y. Goba, Alaa Hassanein, Catherine M. Elias 외 arxiv

Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and Large Multimodal Models (LMMs) continue to treat time as a secondar…

Autonomous Vehicles

MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI

2026-03-20 · Rozain Shakeel, Abdul Rahman Mohammad Ali, Muneeb Mushtaq, Tausifa Jan Saleem 외 arxiv

Despite the rapid progress of Multimodal Large Language Models (MLLMs), their ability to perform reliable visual grounding in high-stakes clinical software environments remains underexplored. Existing GUI benchmarks larg…

Visual Grounding