Large Margin Learning of Upstream Scene Understanding Models
Upstream supervised topic models have been widely used for complicated scene understanding. However, existing maximum likelihood estimation (MLE) schemes can make the prediction model learning independent of latent topic discovery and result in an imbalanced prediction rule for scene classification. This paper presents a joint max-margin and max-likelihood learning method for upstream scene understanding models, in which latent topic discovery and prediction model estimation are closely coupled and well-balanced. The optimization problem is efficiently solved with a variational EM procedure, which iteratively solves an online loss-augmented SVM. We demonstrate the advantages of the large-margin approach on both an 8-category sports dataset and the 67-class MIT indoor scene dataset for scene categorization.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationPredictionScene ClassificationScene UnderstandingTopic ModelsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent complexity often requires multi-step reasoning that current large multimodal…
Video Question AnsweringInstruction Tuning Changes How Upstream State Conditions Late Readout: A Cross-Patching Diagnostic
Recent interpretability work has identified model-internal handles on post-trained behavior, including refusal directions, assistant/persona axes, and sparse chat-tuning features. These results localize where behaviors c…
EKTVQA: Generalized use of External Knowledge to empower Scene Text in Text-VQA
The open-ended question answering task of Text-VQA often requires reading and reasoning about rarely seen or completely unseen scene-text content of an image. We address this zero-shot nature of the problem by proposing …
Open-Ended Question AnsweringOptical Character Recognition (OCR)Question AnsweringVisual Question Answering (VQA)Multi-hop Upstream Anticipatory Traffic Signal Control with Deep Reinforcement Learning
Coordination in traffic signal control is crucial for managing congestion in urban networks. Existing pressure-based control methods focus only on immediate upstream links, leading to suboptimal green time allocation and…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningTraffic Signal ControlCAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multip…
Graph Neural NetworkScene Understanding