Combining Geometric, Textual and Visual Features for Predicting Prepositions in Image Descriptions
Code (0)
등록된 구현이 없습니다.
Tasks
Image RetrievalSimilar Papers 제목 키워드 기반
I Can Has Cheezburger? A Nonparanormal Approach to Combining Textual and Visual Information for Predicting and Generating Popular Meme Descriptions
2015-05-01 · HLT 2015 5
· William Yang Wang, Miaomiao Wen
Cultural Vocal Bursts Intensity PredictionImage RetrievalInformation RetrievalLanguage Modelling+2
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
2016-06-06 · EMNLP 2016 11
· Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach 외
Modeling textual or visual information with vector representations trained from large language or visual datasets has been successfully explored in recent years. However, tasks such as visual question answering require c…
Phrase GroundingVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)Lifting GIS Maps into Strong Geometric Context for Scene Understanding
2015-07-14
· Raúl Díaz, Minhaeng Lee, Jochen Schubert, Charless C. Fowlkes
Contextual information can have a substantial impact on the performance of visual tasks such as semantic segmentation, object detection, and geometric estimation. Data stored in Geographic Information Systems (GIS) offer…
Depth Estimationobject-detectionObject DetectionScene Understanding+2DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot Segmentation
2025-07-22
· Shuai Chen, Fanman Meng, Xiwei Zhang, Haoran Wei 외
arxiv
This paper presents DFR (Decompose, Fuse and Reconstruct), a novel framework that addresses the fundamental challenge of effectively utilizing multi-modal guidance in few-shot segmentation (FSS). While existing approache…
Contrastive LearningCombining Visual and Textual Features for Information Extraction from Online Flyers
2014-10-01 · EMNLP 2014 10
· Emilia Apostolova, Noriko Tomuro
Named Entity Recognition (NER)