paper-with-me

홈 › Papers

OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad

2025-03-24 · CVPR 2025 1 · Luyao Tang, Yuxuan Yuan, Chaoqi Chen, Zeyu Zhang, Yue Huang, Kun Zhang

Although foundation models (FMs) claim to be powerful, their generalization ability significantly decreases when faced with distribution shifts, weak supervision, or malicious attacks in the open world. On the other hand, most domain generalization or adversarial fine-tuning methods are task-related or model-specific, ignoring the universality in practical applications and the transferability between FMs. This paper delves into the problem of generalizing FMs to the out-of-domain data. We propose a novel framework, the Object-Concept-Relation Triad (OCRT), that enables FMs to extract sparse, high-level concepts and intricate relational structures from raw visual inputs. The key idea is to bind objects in visual scenes and a set of object-centric representations through unsupervised decoupling and iterative refinement. To be specific, we project the object-centric representations onto a semantic concept space that the model can readily interpret and estimate their importance to filter out irrelevant elements. Then, a concept-based graph, which has a flexible degree, is constructed to incorporate the set of concepts and their corresponding importance, enabling the extraction of high-order factors from informative concepts and facilitating relational reasoning among these concepts. Extensive experiments demonstrate that OCRT can substantially boost the generalizability and robustness of SAM and CLIP across multiple downstream tasks.

📄 PDF Abstract BibTeX arXiv:2503.18695

Code (1)

lytang63/OCRT 공식 구현 pytorch

Tasks

Domain GeneralizationRelationRelational Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
SAM 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

OCRTurk: A Comprehensive OCR Benchmark for Turkish

2026-02-03 · Deniz Yılmaz, Evren Ayberk Munis, Çağrı Toraman, Süha Kağan Köse 외 arxiv

Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is cruc…

A Real World Dataset for Multi-view 3D Reconstruction

2022-03-22 · Rakesh Shrestha, Siqi Hu, Minghao Gou, Ziyuan Liu 외

We present a dataset of 998 3D models of everyday tabletop objects along with their 847,000 real world RGB and depth images. Accurate annotations of camera poses and object poses for each image are performed in a semi-au…

3D ReconstructionMulti-View 3D ReconstructionObjectPose Estimation+1

DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction

2024-09-07 · YiHeng Wu, Roman Yangarber, Xian Mao

The remarkable capabilities of Large Language Models (LLMs) in text comprehension and generation have revolutionized Information Extraction (IE). One such advancement is in Document-level Relation Triplet Extraction (Doc…

Document-level Relation ExtractionReading ComprehensionRelationRelation Extraction+2

Boosting Segment Anything Model Towards Open-Vocabulary Learning

2023-12-06 · Xumeng Han, Longhui Wei, Xuehui Yu, Zhiyang Dou 외

The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM finding applications and adaptations in va…

modelObjectObject LocalizationRegion Proposal+1

Open World Object Detection in the Era of Foundation Models

2023-12-10 · Orr Zohar, Alejandro Lozano, Shelly Goel, Serena Yeung 외

Object detection is integral to a bevy of real-world applications, from robotics to medical image analysis. To be used reliably in such applications, models must be capable of handling unexpected - or novel - objects. Th…

Medical Image AnalysisObjectobject-detectionObject Detection+1