paper-with-me

홈 › Papers

More Pictures Say More: Visual Intersection Network for Open Set Object Detection

2024-08-26 · Bingcheng Dong, Yuning Ding, Jinrong Zhang, Sifan Zhang, Shenglan Liu

Open Set Object Detection has seen rapid development recently, but it continues to pose significant challenges. Language-based methods, grappling with the substantial modal disparity between textual and visual modalities, require extensive computational resources to bridge this gap. Although integrating visual prompts into these frameworks shows promise for enhancing performance, it always comes with constraints related to textual semantics. In contrast, viusal-only methods suffer from the low-quality fusion of multiple visual prompts. In response, we introduce a strong DETR-based model, Visual Intersection Network for Open Set Object Detection (VINO), which constructs a multi-image visual bank to preserve the semantic intersections of each category across all time steps. Our innovative multi-image visual updating mechanism learns to identify the semantic intersections from various visual prompts, enabling the flexible incorporation of new information and continuous optimization of feature representations. Our approach guarantees a more precise alignment between target category semantics and region semantics, while significantly reducing pre-training time and resource demands compared to language-based methods. Furthermore, the integration of a segmentation head illustrates the broad applicability of visual intersection in various visual tasks. VINO, which requires only 7 RTX4090 GPU days to complete one epoch on the Objects365v1 dataset, achieves competitive performance on par with vision-language models on benchmarks such as LVIS and ODinW35.

📄 PDF Abstract BibTeX arXiv:2408.14032

Code (0)

등록된 구현이 없습니다.

Tasks

GPUobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Like Partying? Your Face Says It All. Predicting the Ambiance of Places with Profile Pictures

2015-05-28 · Miriam Redi, Daniele Quercia, Lindsay T. Graham, Samuel D. Gosling

To choose restaurants and coffee shops, people are increasingly relying on social-networking sites. In a popular site such as Foursquare or Yelp, a place comes with descriptions and reviews, and with profile pictures of …

All

Helping Visually Impaired People Take Better Quality Pictures

2023-05-14 · Maniratnam Mandal, Deepti Ghadiyaram, Danna Gurari, Alan C. Bovik

Perception-based image analysis technologies can be used to help visually impaired people take better quality pictures by providing automated guidance, thereby empowering them to interact more confidently on social media…

Multi-Task Learning

Coding of volumetric content with MIV using VVC subpictures

2022-06-06 · Maria Santamaria, Vinod Kumar Malamal Vadakital, Lukasz Kondrad, Antti Hallapuro 외

Storage and transport of six degrees of freedom (6DoF) dynamic volumetric visual content for immersive applications requires efficient compression. ISO/IEC MPEG has recently been working on a standard that aims to effici…

Decoder

What your Facebook Profile Picture Reveals about your Personality

2017-08-03 · Cristina Segalin, Fabio Celli, Luca Polonio, Michal Kosinski 외

People spend considerable effort managing the impressions they give others. Social psychologists have shown that people manage these impressions differently depending upon their personality. Facebook and other social med…

General Classification

Py-Feat: Python Facial Expression Analysis Toolbox

2021-04-08 · Jin Hyun Cheong, Eshin Jolly, Tiankang Xie, Sophie Byrne 외

Studying facial expressions is a notoriously difficult endeavor. Recent advances in the field of affective computing have yielded impressive progress in automatically detecting facial expressions from pictures and videos…

2D Cyclist Detection