paper-with-me

홈 › Papers

Predicting Multiple Structured Visual Interpretations

2015-12-01 · ICCV 2015 12 · Debadeepta Dey, Varun Ramakrishna, Martial Hebert, J. Andrew Bagnell

We present a simple approach for producing a small number of structured visual outputs which have high recall, for a variety of tasks including monocular pose estimation and semantic scene segmentation. Current state-of-the-art approaches learn a single model and modify inference procedures to produce a small number of diverse predictions. We take the alternate route of modifying the learning procedure to directly optimize for good, high recall sequences of structured-output predictors. Our approach introduces no new parameters, naturally learns diverse predictions and is not tied to any specific structured learning or inference procedure. We leverage recent advances in the contextual submodular maximization literature to learn a sequence of predictors and empirically demonstrate the simplicity and performance of our approach on multiple challenging vision tasks including achieving state-of-the-art results on multiple predictions for monocular pose-estimation and image foreground/background segmentation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Pose EstimationScene SegmentationSegmentation

Similar Papers 제목 키워드 기반

How Useful Are the Machine-Generated Interpretations to General Users? A Human Evaluation on Guessing the Incorrectly Predicted Labels

2020-08-26 · Hua Shen, Ting-Hao Kenneth Huang

Explaining to users why automated systems make certain mistakes is important and challenging. Researchers have proposed ways to automatically produce interpretations for deep neural network models. However, it is unclear…

ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding

2025-05-09 · Shuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac 외

Understanding visual art requires reasoning across multiple perspectives -- cultural, historical, and stylistic -- beyond mere object recognition. While recent multimodal large language models (MLLMs) perform well on gen…

Image CaptioningObject RecognitionRAGRetrieval+1

Reasoning about Intent for Ambiguous Requests

2025-11-13 · Irina Saparina, Mirella Lapata arxiv

Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safety risks when that interpretation is wrong. We propose generating a single stru…

Conversational Question AnsweringReinforcement LearningSemantic Parsing

Towards Interpretable Counterfactual Generation via Multimodal Autoregression

2025-03-29 · Chenglong Ma, Yuanfeng Ji, Jin Ye, Lu Zhang 외

Counterfactual medical image generation enables clinicians to explore clinical hypotheses, such as predicting disease progression, facilitating their decision-making. While existing methods can generate visually plausibl…

counterfactualDecision MakingImage GenerationMedical Image Generation

MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks

2026-04-26 · Jui-Cheng Chiu, Yu-Chao Wang, Shengyang Luo, Tongyan Wang 외 arxiv

Appreciating multi-figure paintings requires understanding how characters relate through subtle cues like gaze alignment, gesture, and spatial arrangement. We present MIRAGE, an evidence-centric framework designed to sca…