paper-with-me

Papers

MemexQA: Visual Memex Question Answering

2017-08-04 · Lu Jiang, Junwei Liang, Liangliang Cao, Yannis Kalantidis, Sachin Farfade, Alexander Hauptmann

This paper proposes a new task, MemexQA: given a collection of photos or videos from a user, the goal is to automatically answer questions that help users recover their memory about events captured in the collection. Towards solving the task, we 1) present the MemexQA dataset, a large, realistic multimodal dataset consisting of real personal photos and crowd-sourced questions/answers, 2) propose MemexNet, a unified, end-to-end trainable network architecture for image, text and video question answering. Experimental results on the MemexQA dataset demonstrate that MemexNet outperforms strong baselines and yields the state-of-the-art on this novel and challenging task. The promising results on TextQA and VideoQA suggest MemexNet's efficacy and scalability across various QA tasks.

📄 PDF Abstract BibTeX arXiv:1708.01336

Code (1)

JunweiLiang/FVTA_MemexQA 공식 구현 tf

Tasks

Memex Question AnsweringQuestion AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

Focal Visual-Text Attention for Memex Question Answering

2018-12-14 · IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2018 12 · Junwei Liang, Lu Jiang, Liangliang Cao, Yannis Kalantidis 외

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collecti…

Memex Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Focal Visual-Text Attention for Visual Question Answering

2018-06-05 · CVPR 2018 6 · Junwei Liang, Lu Jiang, Liangliang Cao, Li-Jia Li 외

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collecti…

Memex Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory

2026-03-04 · Zhenting Wang, Huancheng Chen, Jiayun Wang, Wei Wei arxiv

Large language model (LLM) agents are fundamentally bottlenecked by finite context windows on long-horizon tasks. As trajectories grow, retaining tool outputs and intermediate reasoning in-context quickly becomes infeasi…

Reinforcement Learning

Beyond Categories: The Visual Memex Model for Reasoning About Object Relationships

2009-12-01 · NeurIPS 2009 12 · Tomasz Malisiewicz, Alyosha Efros

The use of context is critical for scene understanding in computer vision, where the recognition of an object is driven by both local appearance and the objects relationship to other elements of the scene (context). Mos…

ObjectScene Understanding

MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization

2023-05-25 · Shivam Sharma, Ramaneswaran S, Udit Arora, Md. Shad Akhtar 외

Memes are a powerful tool for communication over social media. Their affinity for evolving across politics, history, and sociocultural phenomena makes them an ideal communication vehicle. To comprehend the subtle message…

Common Sense Reasoning