Context-Infused Visual Grounding for Art
Many artwork collections contain textual attributes that provide rich and contextualised descriptions of artworks. Visual grounding offers the potential for localising subjects within these descriptions on images, however, existing approaches are trained on natural images and generalise poorly to art. In this paper, we present CIGAr (Context-Infused GroundingDINO for Art), a visual grounding approach which utilises the artwork descriptions during training as context, thereby enabling visual grounding on art. In addition, we present a new dataset, Ukiyo-eVG, with manually annotated phrase-grounding annotations, and we set a new state-of-the-art for object detection on two artwork datasets.
Code (1)
Tasks
object-detectionObject DetectionPhrase GroundingVisual GroundingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
KSAT: Knowledge-infused Self Attention Transformer -- Integrating Multiple Domain-Specific Contexts
Domain-specific language understanding requires integrating multiple pieces of relevant contextual information. For example, we see both suicide and depression-related behavior (multiple contexts) in the text ``I have a …
SpecificityVision-Infused Deep Audio Inpainting
Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, \ie synthesizing missing audio segments that correspond to …
Audio inpaintingImage InpaintingKnowledge Infused Policy Gradients with Upper Confidence Bound for Relational Bandits
Contextual Bandits find important use cases in various real-life scenarios such as online advertising, recommendation systems, healthcare, etc. However, most of the algorithms use flat feature vectors to represent contex…
DescriptiveMulti-Armed BanditsMusic RecommendationRecommendation SystemsThe Context of Crash Occurrence: A Complexity-Infused Approach Integrating Semantic, Contextual, and Kinematic Features
Understanding the context of crash occurrence in complex driving environments is essential for improving traffic safety and advancing automated driving. Previous studies have used statistical models and deep learning to …
Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language ModelLearning Cross-modal Context Graph for Visual Grounding
Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the res…
Graph MatchingGraph Neural NetworkLanguage ModellingNatural Language Visual Grounding+2