CoVR
Composed Video Retrieval
2000년 도입 · 논문 4편에서 사용
The composed video retrieval (CoVR) task is a new task, where the goal is to find a video that matches both a query image and a query text. The query image represents a visual concept that the user is interested in, and the query text specifies how the concept should be modified or refined. For example, given an image of a fountain and the text _during show at night_, the CoVR task is to retrieve a video that shows the fountain at night with a show.
출처: CoVR-2: Automatic Data Construction for Composed Video Retrieval
소개 논문: CoVR-2: Automatic Data Construction for Composed Video Retrieval
Video-Text Retrieval Models · Computer Vision