paper-with-me

홈 › Papers

HOLA: Holistic Multi-Modal Alignment for Open-Set 3D Recognition

2026-05-31 · Koby Aharonov, Oren Shrout, Ayellet Tal arxiv

Open-set 3D recognition requires models that generalize to rare or unseen categories. Recent approaches address this by distilling language-vision knowledge into 3D encoders, typically relying on heavy 2D ViTs and aligning each point cloud with a single image or caption, thus anchoring representations to partial views. We propose aligning each point cloud with multiple images and textual descriptions to capture a more holistic understanding of 3D objects. To realize this idea, it is essential to design a loss function capable of jointly aligning a 3D instance with multiple matched signals, multi-view images and multiple texts, while separating positive aggregation from negative competition. We introduce such a function, termed the decoupled multi-positive contrastive loss. Our formulation enhances the loss's hardness-aware focus on challenging negatives, avoiding the "spotlight crowding" that occurs when many positives share the same softmax with all the negatives. Complementing this, we present a lightweight text adapter applied only to web captions, reducing the domain gap to curated annotations and enabling effective use of large-scale unsupervised text. Our model demonstrates state-of-the-art open-vocabulary performance on long-tail benchmarks, yielding substantial zero-shot improvements while sustaining high frame rates.

📄 PDF Abstract BibTeX arXiv:2606.01334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis

2025-08-27 · Liana Patel, Negar Arabzadeh, Harshit Gupta, Ankita Sundar 외 arxiv

The ability to research and synthesize knowledge is central to human expertise and progress. A new class of AI systems--designed for generative research synthesis--aims to automate this process by retrieving information …

Open-set Cross Modal Generalization via Multimodal Unified Representation

2025-07-20 · Hai Huang, Yan Xia, Shulei Wang, Hanting Wang 외 arxiv

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in o…

Self-Supervised LearningContrastive Learning

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

2026-04-17 · Bo Gao, Chang Liu, Yuyang Miao, Siyuan Ma 외 arxiv

Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an open challenge, where prior works mainly decompose this task into se…

Story Visualization

Tuning Multi-mode Token-level Prompt Alignment across Modalities

2023-09-21 · NeurIPS 2023 11

Advancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily focus on single-mode (only one prompt for…

Open Research Knowledge Graph: Next Generation Infrastructure for Semantic Scholarly Knowledge

2019-01-30 · Mohamad Yaser Jaradeh, Allard Oelen, Kheir Eddine Farfar, Manuel Prinz 외

Despite improved digital access to scholarly knowledge in recent decades, scholarly communication remains exclusively document-based. In this form, scholarly knowledge is hard to process automatically. In this paper, we …

FormKnowledge Graphs