T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting
Zero-shot object counting aims to count instances of arbitrary object categories specified by text descriptions. Existing methods typically rely on vision-language models like CLIP, but often exhibit limited sensitivity to text prompts. We present T2ICount, a diffusion-based framework that leverages rich prior knowledge and fine-grained visual understanding from pretrained diffusion models. While one-step denoising ensures efficiency, it leads to weakened text sensitivity. To address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision signals by leveraging the cross-attention maps extracted from the denoising U-Net. Furthermore, we observe that current benchmarks mainly focus on majority objects in images, potentially masking models' text sensitivity. To address this, we contribute a challenging re-annotated subset of FSC147 for better evaluation of text-guided counting ability. Extensive experiments demonstrate that our method achieves superior performance across different benchmarks. Code is available at https://github.com/cha15yq/T2ICount
Code (1)
Tasks
DenoisingObject CountingSensitivityZero-Shot CountingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
OmniCount: Multi-label Object Counting with Semantic-Geometric Priors
Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, …
ObjectObject CountingOpen-vocabulary object countingTraining-free Object Counting+1MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
Multi-instance Repetitive Action Counting (MRAC) aims to estimate the number of repetitive actions performed by multiple instances in untrimmed videos, commonly found in human-centric domains like sports and exercise. In…
GPURepetitive Action CountingResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Large Language Model (LLM. ResNetVLLM addre…
Language ModelingLanguage ModellingLarge Language ModelVideo UnderstandingMM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
We introduce MM-Mixing, a multi-modal mixing alignment framework for 3D understanding. MM-Mixing applies mixing-based methods to multi-modal data, preserving and optimizing cross-modal connections while enhancing diversi…
3D Classification3D Object Recognition3D Shape RetrievalContrastive Learning+6Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
Zero-shot scene understanding in real-world settings presents major challenges due to the complexity and variability of natural scenes, where models must recognize new objects, actions, and contexts without prior labeled…
Scene UnderstandingActivity DetectionObject Recognition