paper-with-me

Papers

Large Margin Learning of Upstream Scene Understanding Models

2010-12-01 · NeurIPS 2010 12 · Jun Zhu, Li-Jia Li, Li Fei-Fei, Eric P. Xing

Upstream supervised topic models have been widely used for complicated scene understanding. However, existing maximum likelihood estimation (MLE) schemes can make the prediction model learning independent of latent topic discovery and result in an imbalanced prediction rule for scene classification. This paper presents a joint max-margin and max-likelihood learning method for upstream scene understanding models, in which latent topic discovery and prediction model estimation are closely coupled and well-balanced. The optimization problem is efficiently solved with a variational EM procedure, which iteratively solves an online loss-augmented SVM. We demonstrate the advantages of the large-margin approach on both an 8-category sports dataset and the 67-class MIT indoor scene dataset for scene categorization.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationPredictionScene ClassificationScene UnderstandingTopic Models

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks

2026-04-25 · Jason Nguyen, Ameet Rao, Alexander Chang, Ishaan Kumar 외 arxiv

Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent complexity often requires multi-step reasoning that current large multimodal…

Video Question Answering

Instruction Tuning Changes How Upstream State Conditions Late Readout: A Cross-Patching Diagnostic

2026-05-08 · Yifan Zhou arxiv

Recent interpretability work has identified model-internal handles on post-trained behavior, including refusal directions, assistant/persona axes, and sparse chat-tuning features. These results localize where behaviors c…

EKTVQA: Generalized use of External Knowledge to empower Scene Text in Text-VQA

2021-08-22 · Arka Ujjal Dey, Ernest Valveny, Gaurav Harit

The open-ended question answering task of Text-VQA often requires reading and reasoning about rarely seen or completely unseen scene-text content of an image. We address this zero-shot nature of the problem by proposing …

Open-Ended Question AnsweringOptical Character Recognition (OCR)Question AnsweringVisual Question Answering (VQA)

Multi-hop Upstream Anticipatory Traffic Signal Control with Deep Reinforcement Learning

2024-11-10 · Xiaocan Li, Xiaoyu Wang, Ilia Smirnov, Scott Sanner 외

Coordination in traffic signal control is crucial for managing congestion in urban networks. Existing pressure-based control methods focus only on immediate upstream links, leading to suboptimal green time allocation and…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningTraffic Signal Control

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

2026-07-07 · He Liang, Chenyang Ma, Yiming Zhang, Sangyun Shin 외 arxiv

Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multip…

Graph Neural NetworkScene Understanding