paper-with-me

Papers

Focal Visual-Text Attention for Memex Question Answering

2018-12-14 · IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2018 12 · Junwei Liang, Lu Jiang, Liangliang Cao, Yannis Kalantidis, Li-Jia Li, and Alexander Hauptmann

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photo albums, we have to look at whole collections with sequences of photos. This paper proposes a new multimodal MemexQA task: given a sequence of photos from a user, the goal is to automatically answer questions that help users recover their memory about an event captured in these photos. In addition to a text answer, a few grounding photos are also given to justify the answer. The grounding photos are necessary as they help users quickly verifying the answer. Towards solving the task, we 1) present the MemexQA dataset, the first publicly available multimodal question answering dataset consisting of real personal photo albums; 2) propose an end-to-end trainable network that makes use of a hierarchical process to dynamically determine what media and what time to focus on in the sequential data to answer the question. Experimental results on the MemexQA dataset demonstrate that our model outperforms strong baselines and yields the most relevant grounding photos on this challenging task.

📄 PDF Abstract BibTeX

Code (1)

JunweiLiang/FVTA_memoryqa tf

Tasks

Memex Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Focal Visual-Text Attention for Visual Question Answering

2018-06-05 · CVPR 2018 6 · Junwei Liang, Lu Jiang, Liangliang Cao, Li-Jia Li 외

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collecti…

Memex Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

MemexQA: Visual Memex Question Answering

2017-08-04 · Lu Jiang, Junwei Liang, Liangliang Cao, Yannis Kalantidis 외

This paper proposes a new task, MemexQA: given a collection of photos or videos from a user, the goal is to automatically answer questions that help users recover their memory about events captured in the collection. Tow…

Memex Question AnsweringQuestion AnsweringVideo Question Answering

Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory

2026-03-04 · Zhenting Wang, Huancheng Chen, Jiayun Wang, Wei Wei arxiv

Large language model (LLM) agents are fundamentally bottlenecked by finite context windows on long-horizon tasks. As trajectories grow, retaining tool outputs and intermediate reasoning in-context quickly becomes infeasi…

Reinforcement Learning

Beyond Categories: The Visual Memex Model for Reasoning About Object Relationships

2009-12-01 · NeurIPS 2009 12 · Tomasz Malisiewicz, Alyosha Efros

The use of context is critical for scene understanding in computer vision, where the recognition of an object is driven by both local appearance and the objects relationship to other elements of the scene (context). Mos…

ObjectScene Understanding

MemX: An Attention-Aware Smart Eyewear System for Personalized Moment Auto-capture

2021-05-03 · Yuhu Chang, Yingying Zhao, Mingzhi Dong, Yujiang Wang 외

This work presents MemX: a biologically-inspired attention-aware eyewear system developed with the goal of pursuing the long-awaited vision of a personalized visual Memex. MemX captures human visual attention on the fly,…