paper-with-me

홈 › Papers

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

2025-09-19 · Cristian Sbrolli, Matteo Matteucci arxiv

The whole is greater than the sum of its parts-even in 3D-text contrastive learning. We introduce SceneForge, a novel framework that enhances contrastive alignment between 3D point clouds and text through structured multi-object scene compositions. SceneForge leverages individual 3D shapes to construct multi-object scenes with explicit spatial relations, pairing them with coherent multi-object descriptions refined by a large language model. By augmenting contrastive training with these structured, compositional samples, SceneForge effectively addresses the scarcity of large-scale 3D-text datasets, significantly enriching data complexity and diversity. We systematically investigate critical design elements, such as the optimal number of objects per scene, the proportion of compositional samples in training batches, and scene construction strategies. Extensive experiments demonstrate that SceneForge delivers substantial performance gains across multiple tasks, including zero-shot classification on ModelNet, ScanObjNN, Objaverse-LVIS, and ScanNet, as well as few-shot part segmentation on ShapeNetPart. SceneForge's compositional augmentations are model-agnostic, consistently improving performance across multiple encoder architectures. Moreover, SceneForge improves 3D visual question answering on ScanQA, generalizes robustly to retrieval scenarios with increasing scene complexity, and showcases spatial reasoning capabilities by adapting spatial configurations to align precisely with textual instructions.

📄 PDF Abstract BibTeX arXiv:2509.15693

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringContrastive LearningSpatial ReasoningPoint Clouds

Similar Papers 제목 키워드 기반

SceneForge: Structured World Supervision from 3D Interventions

2026-05-14 · Jizhizi Li, Jiayang Ao, Danny Wicks, Petru-Daniel Tudosiu arxiv

Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such supervision is difficult to obtain from observation-level datasets, which …

Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios

2025-10-30 · Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi arxiv

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limi…

Zero-shot GeneralizationScene Understanding

Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition

2024-02-21 · Mingkun Yang, Biao Yang, Minghui Liao, Yingying Zhu 외

Scene text recognition is a rapidly developing field that faces numerous challenges due to the complexity and diversity of scene text, including complex backgrounds, diverse fonts, flexible arrangements, and accidental o…

DiversityScene Text Recognition

Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure

2026-01-31 · Trishna Chakraborty, Udita Ghosh, Aldair Ernesto Gongora, Ruben Glatt 외 arxiv

Laboratories are prone to severe injuries from minor unsafe actions, yet continuous safety monitoring -- beyond mandatory pre-lab safety training -- is limited by human availability. Vision language models (VLMs) offer p…

Image Generation

ALADIN:Attribute-Language Distillation Network for Person Re-Identification

2026-03-23 · Wang Zhou, Boran Duan, Haojun Ai, Ruiqi Lan 외 arxiv

Recent vision-language models such as CLIP provide strong cross-modal alignment, but current CLIP-guided ReID pipelines rely on global features and fixed prompts. This limits their ability to capture fine-grained attribu…

Person Re-IdentificationRepresentation Learning