paper-with-me

Papers Object Localization

“Object Localization” 태그가 달린 논문 724편 · 필터 해제

Tail-Likelihood Reinforcement Learning

2026-09-02 · Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, Guanning Zeng 외 arxiv

Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of prod…

Reinforcement LearningObject Localization

WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes

2026-08-28 · Kishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Md. Sadman Haque 외 arxiv

Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask …

Object LocalizationObject Detection

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

2026-08-24 · Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li 외 arxiv

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires …

Object LocalizationSpatial Reasoning

Doomed to Re-Annotate, Forever: The ImageNet Story

2026-08-13 · Illia Volkov, Nikita Kisel, Tetiana Mishkina, Klara Janouskova 외 arxiv

Top-1 accuracy on ImageNet-1k remains the most commonly reported metric in visual recognition. Quality issues with the dataset have been repeatedly reported, yet the original 2012 noisy labels are still predominantly use…

Object Localization

DRPFNet: Dual-domain Residual Progressive Fusion Network for RGB-Thermal Object Detection

2026-08-04 · Zian Wang, Changchun Li arxiv

RGB-thermal (RGB-T) object detection aims to fuse complementary information from visible and thermal modalities to achieve robust detection under varying illumination and weather conditions. Current methods typically emp…

Object LocalizationObject Detection

RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI

2026-08-01 · Kai Geissler, Laurens Müller-Groh, Hans Meine arxiv

Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models carry a high computational cost, and only few aim to be broadly applica…

3D Object DetectionObject Localization

Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks

2026-07-26 · Qi Ming, Haitian Yang, Xudong Zhao, Mingjing Zhao 외 arxiv

Bounding boxes are fundamental for object localization in visual detection tasks. Among them, oriented bounding boxes are widely used in visual detection tasks, which provide a more precise directional representation. Ge…

Semantic SimilarityObject Localization

VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval

2026-07-09 · ZhiXin Sun arxiv

Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-language and reference-based approaches have significantly advanced this field, th…

Object LocalizationObject Detection

Exploring SAM Supervision for Fine-Grained UAV Target Segmentation under Data Scarcity

2026-07-04 · Le-Anh Tran arxiv

Unmanned aerial vehicle (UAV) target segmentation remains challenging due to the small size of objects, appearance variations, cluttered backgrounds, and the scarcity of densely annotated data. These factors hinder the p…

Computational EfficiencyObject Localization

Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics

2026-07-03 · Anvitha Ramachandran, Dhruv Parikh, Haoyang Fan, Rajgopal Kannan 외 arxiv

State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal sequence modeling. While effective for sequential data, directional sca…

Instance SegmentationSemantic SegmentationObject LocalizationObject Detection

Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models

2026-07-01 · Kensuke Nakamura, Byung-Woo Hong arxiv

Personalized object localization (POL) localizes an object instance in a query image based on a few reference images with bounding-box annotations and a target object label. The pioneering method, IPLoc, solves this task…

Few-Shot Object DetectionObject Localization

GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision

2026-07-01 · Dianyu Wang, Peirong Zhang, Xuyang Li, Xiaoxuan Liu 외 arxiv

Recent multimodal large language models (MLLMs) have shown strong cross-modal understanding and coordinate generation abilities in visual grounding. However, transferring these abilities to remote sensing visual groundin…

Object LocalizationVisual Grounding

Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization

2026-06-26 · Shiang-Feng Tsai, Jin-Cheng Jhang, Yen-Ling Tai, Jia-Hong Lai 외 arxiv

Vision-Language-Action (VLA) models leverage large-scale vision-language pretraining for flexible robot manipulation, yet at test time they remain brittle along two axes: spatial generalization, when object positions dif…

Object LocalizationRobot Manipulation

TACTFUL: Tactile-Driven Exploration For Object Localization and Identification in Confined Environments

2026-06-23 · Shivani Kamtikar, Chung Hee Kim, Camilla Tabasso, Tye Brady 외 arxiv

Humans effortlessly locate and identify objects by touch alone, even without vision. In contrast, robotic systems rely heavily on vision and struggle with autonomous tactile exploration and object identification. We pres…

Object Localization

UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving

2026-06-23 · Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu 외 arxiv

Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing methods still face a fundamental trade-off between temporal reasoning and spatial preci…

Zero-shot GeneralizationObject LocalizationScene UnderstandingAutonomous Driving

Few-class Fidelity: Evaluating Explanations of Real-conditions CNN classifiers with Optimized Perturbations

2026-06-23 · Wistan Marchadour, Pedro Soto Vega, Franck Vermet, Mathieu Hatt arxiv

The wide use of Convolutional Neural Networks (CNN) in numerous domains and real-world classification applications is justified by their high precision and automation speed, helping users concentrate on higher-expertise …

Object Localization

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

2026-06-22 · Anindya Mondal, Sauradip Nag, Anjan Dutta arxiv

ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is bu…

Object LocalizationImage GenerationObject CountingCrowd Counting

Motion-Aware Reinforcement Learning For Object Localization

2026-06-19 · Prithvi Raj Singh, Satyendra Singh arxiv

We present MARLNet (Motion-Aware Reinforcement Learning Network), a PPO-based bounding-box refinement agent that incorporates a constant-velocity motion prior into the observation state and an action smoothness penalty i…

Reinforcement LearningObject Localization

FEMOT: Multi-Object Tracking using Frame and Event Cameras

2026-06-12 · Shiao Wang, Xiao Wang, Chao Wang, Yitao Li 외 arxiv

Conventional RGB cameras have been widely used in multi-object tracking due to their ability to capture rich appearance and semantic information. However, their performance is often degraded under complex real-world chal…

Multi-Object TrackingObject Localization

Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

2026-06-10 · Lu Qiu, Yizhuo Li, Yi Chen, Yuying Ge 외 arxiv

World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a ph…

Object LocalizationRobot ManipulationVideo Generation
1–20 / 724 다음 →