paper-with-me

홈 › Papers

Multi-Resolution Language Grounding with Weak Supervision

2014-10-01 · EMNLP 2014 10 · R. Koncel-Kedziorski, Hannaneh Hajishirzi, Ali Farhadi
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Acquisition

Similar Papers 제목 키워드 기반

Finding "It": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos

2018-06-01 · CVPR 2018 6 · De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg 외

Grounding textual phrases in visual content with standalone image-sentence pairs is a challenging task. When we consider grounding in instructional videos, this problem becomes profoundly more complex: the latent tempora…

Multiple Instance LearningSentenceVisual Grounding

Generating Descriptions with Grounded and Co-Referenced People

2017-04-05 · CVPR 2017 7 · Anna Rohrbach, Marcus Rohrbach, Siyu Tang, Seong Joon Oh 외

Learning how to generate descriptions of images or videos received major interest both in the Computer Vision and Natural Language Processing communities. While a few works have proposed to learn a grounding during the g…

Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision

2024-10-10 · Shengcao Cao, Liang-Yan Gui, Yu-Xiong Wang

Current large multimodal models (LMMs) face challenges in grounding, which requires the model to relate language components to visual entities. Contrary to the common practice that fine-tunes LMMs with additional groundi…

Question AnsweringVisual Question Answering

Who are you referring to? Coreference resolution in image narrations

2022-11-26 · ICCV 2023 1 · Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen

Coreference resolution aims to identify words and phrases which refer to same entity in a text, a core task in natural language processing. In this paper, we extend this task to resolving coreferences in long-form narrat…

coreference-resolutionCoreference Resolution

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

2024-03-18 · CVPR 2024 1 · Chaolei Tan, JianHuang Lai, Wei-Shi Zheng, Jian-Fang Hu

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However, existing VPG a…

Multiple Instance Learning