paper-with-me

Papers

Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

2026-03-17 · Da Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao arxiv

Zero-shot object counting (ZSOC) aims to enumerate objects of arbitrary categories specified by text descriptions without requiring visual exemplars. However, existing methods often treat counting as a coarse retrieval task, suffering from a lack of fine-grained quantity awareness. Furthermore, they frequently exhibit spatial insensitivity and degraded generalization due to feature space distortion during model adaptation.To address these challenges, we present \textbf{QICA}, a novel framework that synergizes \underline{q}uantity percept\underline{i}on with robust spatial \underline{c}ast \underline{a}ggregation. Specifically, we introduce a Synergistic Prompting Strategy (\textbf{SPS}) that adapts vision and language encoders through numerically conditioned prompts, bridging the gap between semantic recognition and quantitative reasoning. To mitigate feature distortion, we propose a Cost Aggregation Decoder (\textbf{CAD}) that operates directly on vision-text similarity maps. By refining these maps through spatial aggregation, CAD prevents overfitting while preserving zero-shot transferability. Additionally, a multi-level quantity alignment loss ($\mathcal{L}_{MQA}$) is employed to enforce numerical consistency across the entire pipeline. Extensive experiments on FSC-147 demonstrate competitive performance, while zero-shot evaluation on CARPK and ShanghaiTech-A validates superior generalization to unseen domains.

📄 PDF Abstract BibTeX arXiv:2603.16129

Code (0)

등록된 구현이 없습니다.

Tasks

Object Counting

Similar Papers 제목 키워드 기반

Spatial-Aware Object Embeddings for Zero-Shot Localization and Classification of Actions

2017-07-28 · ICCV 2017 10 · Pascal Mettes, Cees G. M. Snoek

We aim for zero-shot localization and classification of human actions in video. Where traditional approaches rely on global attribute or object classification scores for their zero-shot knowledge transfer, our main contr…

Action LocalizationAttributeClassificationGeneral Classification+3

Locality-Aware Zero-Shot Human-Object Interaction Detection

2025-05-26 · CVPR 2025 1 · Sanghyun Kim, Deunsol Jung, Minsu Cho

Recent methods for zero-shot Human-Object Interaction (HOI) detection typically leverage the generalization ability of large Vision-Language Model (VLM), i.e., CLIP, on unseen categories, showing impressive results on va…

Human-Object Interaction DetectionObjectZero-Shot Human-Object Interaction Detection

MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness

2025-03-27 · Zihao Zheng, Xiuping Cui, Size Zheng, Maoliang Li 외

With the advances in artificial intelligence, Mix-of-Experts (MoE) has become the main form of Large Language Models (LLMs), and its demand for model compression is increasing. Quantization is an effective method that no…

Language ModelingLanguage ModellingModel CompressionQuantization

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

2026-06-23 · Tianyu Zhu, Yingping Liang, Hesong Li, Ying Fu arxiv

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive s…

Referring Video Object SegmentationZero-shot GeneralizationImage SegmentationVideo Segmentation

A Simple Framework for Open-Vocabulary Zero-Shot Segmentation

2024-06-23 · Thomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar 외

Zero-shot classification capabilities naturally arise in models trained within a vision-language contrastive framework. Despite their classification prowess, these models struggle in dense tasks like zero-shot open-vocab…

Representation Learningzero-shot-classificationZero-Shot LearningZero Shot Segmentation