paper-with-me

홈 › Papers

CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting

2025-10-27 · Md Tanvir Hossain, Akif Islam, Mohd Ruhul Ameen arxiv

Humans can often count unfamiliar objects by observing visual repetition and composition, rather than relying only on object categories. However, many exemplar-free counting models struggle in such situations and may overcount when objects contain symmetric components, repeated substructures, or partial occlusion. We introduce CountFormer, a controlled adaptation of a density-regression framework inspired by CounTR, where the image encoder is replaced with the self-supervised vision foundation model DINOv2. The resulting transformer features are combined with explicit two-dimensional positional embeddings and decoded by a lightweight convolutional network to produce a density map whose integral gives the final count. Our goal is not to propose a new counting architecture, but to study whether foundation-based representations improve structural consistency under a strictly exemplar-free setting. On FSC-147, CountFormer achieves competitive performance under the official benchmark (MAE 19.06, RMSE 118.45). Qualitative analysis suggests fewer part-level overcounting errors for some structurally complex objects, while overall error remains broadly consistent with prior approaches. Sensitivity analysis shows that evaluation metrics are strongly affected by a small number of extreme high-density scenes. Overall, the results highlight the role of representation quality in exemplar-free object counting.

📄 PDF Abstract BibTeX arXiv:2510.23785

Code (0)

등록된 구현이 없습니다.

Tasks

Exemplar-Free CountingObject Counting

Similar Papers 제목 키워드 기반

CountFormer: Multi-View Crowd Counting Transformer

2024-07-02 · Hong Mo, Xiong Zhang, Jianchao Tan, Cheng Yang 외

Multi-view counting (MVC) methods have shown their superiority over single-view counterparts, particularly in situations characterized by heavy occlusion and severe perspective distortions. However, hand-crafted heuristi…

Crowd Counting

Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication

2024-07-11 · Arif Ahmad, Mothika Gayathri Khyathi, Pushpak Bhattacharyya

Reduplication and repetition, though similar in form, serve distinct linguistic purposes. Reduplication is a deliberate morphological process used to express grammatical, semantic, or pragmatic nuances, while repetition …

token-classificationToken Classification

Prompt-oriented Output of Culture-Specific Items in Translated African Poetry by Large Language Model: An Initial Multi-layered Tabular Review

2025-01-29 · Adeyola Opaluwah

This paper examines the output of cultural items generated by Chat Generative PreTrained Transformer Pro in response to three structured prompts to translate three anthologies of African poetry. The first prompt was broa…

Language ModelingLanguage ModellingLarge Language ModelSpecificity+1

UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers

2025-12-04 · Min Zhao, Bokai Yan, Xue Yang, Hongzhou Zhu 외 arxiv

Recent image diffusion transformers achieve high-fidelity generation, but struggle to generate images beyond these scales, suffering from content repetition and quality degradation. In this work, we present UltraImage, a…

Self-Segregating and Coordinated-Segregating Transformer for Focused Deep Multi-Modular Network for Visual Question Answering

2020-06-25 · Chiranjib Sur

Attention mechanism has gained huge popularity due to its effectiveness in achieving high accuracy in different domains. But attention is opportunistic and is not justified by the content or usability of the content. Tra…

DiversityQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)+1