paper-with-me

홈 › Papers

T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting

2025-01-01 · CVPR 2025 1 · Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, Michael P. Pound

Zero-shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-language models like CLIP, but often exhibit limited sensitivity to text prompts. We present T2ICount, a diffusion-based framework that leverages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step denoising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the denoising U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/cha15yq/T2ICount

📄 PDF Abstract BibTeX

Code (1)

cha15yq/t2icount 공식 구현 pytorch

Tasks

DenoisingObject CountingSensitivityZero-Shot Counting

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Focus 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

OmniCount: Multi-label Object Counting with Semantic-Geometric Priors

2024-03-08 · Anindya Mondal, Sauradip Nag, Xiatian Zhu, Anjan Dutta

Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, …

ObjectObject CountingOpen-vocabulary object countingTraining-free Object Counting+1

MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos

2024-09-06 · Yin Tang, Wei Luo, Jinrui Zhang, Wei Huang 외

Multi-instance Repetitive Action Counting (MRAC) aims to estimate the number of repetitive actions performed by multiple instances in untrimmed videos, commonly found in human-centric domains like sports and exercise. In…

GPURepetitive Action Counting

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task

2025-04-20 · Ahmad Khalil, Mahmoud Khalil, Alioune Ngom

In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Large Language Model (LLM. ResNetVLLM addre…

Language ModelingLanguage ModellingLarge Language ModelVideo Understanding

MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding

2024-05-28 · Jiaze Wang, Yi Wang, Ziyu Guo, Renrui Zhang 외

We introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversi…

3D Classification3D Object Recognition3D Shape RetrievalContrastive Learning+6

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

2025-10-29 · Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi arxiv

Zero-shot scene understanding in real-world settings presents major challenges due to the complexity and variability of natural scenes, where models must recognize new objects, actions, and contexts without prior labeled…

Scene UnderstandingActivity DetectionObject Recognition