paper-with-me

MAVL

Multiscale Attention ViT with Late fusion

2000년 도입 · 논문 4편에서 사용

Multiscale Attention ViT with Late fusion (MAVL) is a multi-modal network, trained with aligned image-text pairs, capable of performing targeted detection using human understandable natural language text queries. It utilizes multi-scale image features and uses deformable convolutions with late multi-modal fusion. The authors demonstrate excellent ability of MAVL as class-agnostic object detector when queried using general human understandable natural language command, such as "all objects", "all entities", etc.

출처: Class-agnostic Object Detection with Multi-modal Transformer

소개 논문: Class-agnostic Object Detection with Multi-modal Transformer

Multi-Modal Methods · Computer Vision