paper-with-me

Papers Object Counting

“Object Counting” 태그가 달린 논문 214편 · 필터 해제

Counting Beyond Instances: A Benchmark for Group-Individual Object Counting

2026-09-04 · Rui Wang, Junyi Huang, Jiahui Li, Qiao Yu 외 arxiv

Visual counting is commonly formulated at the instance level, aiming to estimate how many objects of a queried category appear in an image. However, real-world counting often involves higher-level semantic units formed b…

Object Counting

Depth-Guided Video Object Counting in Crowded Scenes

2026-08-06 · Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou 외 arxiv

Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, li…

Object Counting

Spatially-Aware Class-Agnostic Object Counting

2026-07-18 · Robert Wijaya, Md. Tanvir Hossain, Amanda Kau, Ngai-Man Cheung arxiv

Generalised object counting aims to estimate the number of instances of an arbitrary object category from a single image, but many recent methods can struggle on structurally complex objects due to limited spatial modell…

Object Counting

The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs

2026-07-10 · Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh, Nikolai Rozanov 외 arxiv

Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting. We investigate whether this reflects missing internal knowledge or a gap between internal repr…

Object Counting

AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting

2026-07-02 · Muhammad Ibraheem Siddiqui, Muhammad Haris Khan arxiv

Zero-shot object counting (ZOC) aims to count instances of arbitrary object categories specified only through textual prompts. Recent training-free approaches leverage foundation models such as SAM to reformulate countin…

Object Counting

Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

2026-06-24 · Davie Chen arxiv

Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts. Yet existing …

Object Counting

TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

2026-06-24 · Yu-Yang Chen, Lan-Zhe Guo arxiv

Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexity remains poorly understood. We introdu…

Visual Question AnsweringVisual ReasoningObject Counting

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

2026-06-22 · Anindya Mondal, Sauradip Nag, Anjan Dutta arxiv

ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is bu…

Object LocalizationImage GenerationObject CountingCrowd Counting

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

2026-06-17 · Jinchao Ge, Lingqiao Liu, Shuwen Zhao, Lei Wang arxiv

Vision foundation tools such as open-vocabulary detectors, segmentation models, and post-processing operators are powerful building blocks for computer vision, but their effectiveness depends heavily on how they are orch…

Domain GeneralizationObject Counting

RT-Counter: Real-Time Text-Guided Open-Vocabulary Object Counting

2026-06-16 · Hao-Yuan Ma, Li Zhang, Zhiwei Zhu, Jie Gao arxiv

Text-guided open-vocabulary object counting (TOOC) aims to count objects belonging to the categories specified by natural language descriptions. Although vision-language pre-trained models have been successful applied to…

Computational EfficiencyObject Counting

Test-Time Training for Robust Text-Guided Open-Vocabulary Object Counting

2026-06-16 · Hao-Yuan Ma, Yuda Zou, Li Zhang, Yongchao Xu arxiv

Text-guided Open-vocabulary Object Counting (TOOC) enables counting arbitrary object categories specified by text prompts, offering substantially greater flexibility than conventional closed-set counting. However, existi…

Object Counting

MambaCount: Efficient Text-guided Open-vocabulary Object Counting with Spatial Sparse State Space Duality Block

2026-06-16 · Hao-Yuan Ma, Li Zhang, Minjie Qiang, Jie Gao arxiv

Text-guided Open-vocabulary Object Counting (TOOC) aims to estimate the number of objects described by text prompts, which is particularly challenging in dense scenes with large scale variations. Existing TOOC approaches…

Object Counting

Count Anything

2026-05-29 · Mengqi Lei, Shuokun Cheng, Wei Bao, Shaoyi Du 외 arxiv

Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehic…

Domain GeneralizationObject Counting

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models

2026-05-20 · Qirui Shen, Wenda Wang, Jiachen Lu, Zilong Huang 외 arxiv

Architectural spatial intelligence, the ability to recognize and infer architectural space, is fundamental to tasks such as robot navigation, embodied interaction, and 3D scene understanding and generation. Although exte…

Scene UnderstandingRobot NavigationObject Counting

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

2026-05-19 · Tianle Li, Xuyang Shen, Yan Ma, Rongxin Guo 외 arxiv

Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the level of individual visual claims. A good dense caption should be bot…

Reinforcement LearningScene RecognitionImage CaptioningObject Counting

The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting

2026-05-18 · Corentin Dumery, Niki Amini-Naieni, Shervin Naini, Pascal Fua arxiv

Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mixed-object setting that dominates real-world applications such as ind…

Object Counting

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

2026-05-12 · Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras 외 arxiv

Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) i…

Spatial ReasoningObject Counting

Count Anything at Any Granularity

2026-05-11 · Chang Liu, Haoning Wu, Weidi Xie arxiv

Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that a central reason is that counting granul…

Object CountingImage Editing

SketchVLM: Vision language models can annotate images to explain thoughts and guide users

2026-04-23 · Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri 외 arxiv

When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models (VLMs) such as Gemini-3-Pro and GPT-5 only respond with text, which ca…

Trajectory PredictionVisual ReasoningObject Counting

Learning to count small and clustered objects with application to bacterial colonies

2026-04-21 · Minghua Zheng, Na Helian, Peter C. R. Lane, Yi Sun 외 arxiv

Automated bacterial colony counting from images is an important technique to obtain data required for the development of vaccines and antibiotics. However, bacterial colonies present unique machine vision challenges that…

Feature EngineeringObject Counting
1–20 / 214 다음 →