paper-with-me

Papers

Beyond Categories: The Visual Memex Model for Reasoning About Object Relationships

2009-12-01 · NeurIPS 2009 12 · Tomasz Malisiewicz, Alyosha Efros

The use of context is critical for scene understanding in computer vision, where the recognition of an object is driven by both local appearance and the objects relationship to other elements of the scene (context). Most current approaches rely on modeling the relationships between object categories as a source of context. In this paper we seek to move beyond categories to provide a richer appearance-based model of context. We present an exemplar-based model of objects and their relationships, the Visual Memex, that encodes both local appearance and 2D spatial context between object instances. We evaluate our model on Torralbas proposed Context Challenge against a baseline category-based system. Our experiments suggest that moving beyond categories for context modeling appears to be quite beneficial, and may be the critical missing ingredient in scene understanding systems.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectScene Understanding

Similar Papers 제목 키워드 기반

MemexQA: Visual Memex Question Answering

2017-08-04 · Lu Jiang, Junwei Liang, Liangliang Cao, Yannis Kalantidis 외

This paper proposes a new task, MemexQA: given a collection of photos or videos from a user, the goal is to automatically answer questions that help users recover their memory about events captured in the collection. Tow…

Memex Question AnsweringQuestion AnsweringVideo Question Answering

Focal Visual-Text Attention for Memex Question Answering

2018-12-14 · IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2018 12 · Junwei Liang, Lu Jiang, Liangliang Cao, Yannis Kalantidis 외

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collecti…

Memex Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory

2026-03-04 · Zhenting Wang, Huancheng Chen, Jiayun Wang, Wei Wei arxiv

Large language model (LLM) agents are fundamentally bottlenecked by finite context windows on long-horizon tasks. As trajectories grow, retaining tool outputs and intermediate reasoning in-context quickly becomes infeasi…

Reinforcement Learning

Focal Visual-Text Attention for Visual Question Answering

2018-06-05 · CVPR 2018 6 · Junwei Liang, Lu Jiang, Liangliang Cao, Li-Jia Li 외

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collecti…

Memex Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning

2026-05-25 · Longteng Guo, Yifan Wang, Pengkang Huo, Tailai Chen 외 arxiv

Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning directly grounded in visual evidence. W…

Visual Reasoning