Exploring Contextual Attribute Density in Referring Expression Counting
Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses challenges for prior arts, as they struggle to accurately align attribute information with correct visual patterns. Given the proven importance of "visual density", it is presumed that the limitations of current REC approaches stem from an under-exploration of "contextual attribute density" (CAD). In the scope of REC, we define the CAD as the measure of the information intensity of one certain fine-grained attribute in visual regions. To model the the CAD, we propose a U-shape CAD estimator in which referring expression and multi-scale visual features from GroundingDINO can interact with each other. With additional density supervisions, we can effectively encode CAD, which is subsequently decoded via a novel attention procedure with CAD-refined queries. Integrating all these contributions, our framework significantly outperforms state-of-the-art REC methods, achieves 30% error reduction in counting metics and a 10% improvement in localization accuracy. The surprising results shed lights on the significance of contextual attribute density for REC. Code will be at github.com/Xu3XiWang/CAD-GD.
Code (1)
Tasks
AttributeReferring ExpressionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exploring Contextual Attribute Density in Referring Expression Counting
Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses c…
AttributeReferring ExpressionA case study on context-bound referring expression generation
In recent years, Bayesian models of referring expression generation have gained prominence in order to produce situationally more adequate referring expressions. Basically, these models enable the integration of differen…
Referring ExpressionReferring expression generationReferring Expression Generation and Comprehension via Attributes
Referring expression is a kind of language expression that used for referring to particular objects.To make the expression without ambiguation, people often use attributes to describe the particular object. In this paper…
AttributeReferring ExpressionReferring expression generationJoint Visual Grounding with Language Scene Graphs
Visual grounding is a task to ground referring expressions in images, e.g., localize "the white truck in front of the yellow one". To resolve this task fundamentally, the model should first find out the contextual object…
Referring ExpressionVisual GroundingCorpus-based Referring Expressions Generation
In Natural Language Generation, the task of attribute selection (AS) consists of determining the appropriate attribute-value pairs (or semantic properties) that represent the contents of a referring expression. Existing …
AttributeReferring ExpressionText Generation