paper-with-me

Papers

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

2026-07-16 · Saad Ejaz, Miguel Fernandez-Cortizas, Javier Civera, Holger Voos, Jose Luis Sanchez-Lopez arxiv

CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, with applications in robotics and augmented reality. Recent zero-shot methods use vision foundation models to match image regions to CAD models; yet their correspondences are typically appearance-driven or unreliable under occlusion or synthetic-to-real domain shift. To address these limitations, we introduce SUFLECA (Scaling Up Feature LEarning for CAD-to-image Alignment), a weakly supervised framework for zero-shot CAD alignment with two key contributions. First, SUFLECA scales up geometry-grounded feature learning from pretrained visual representations through Normalized Object Coordinates (NOCs) supervision on images spanning up to 12 real and synthetic datasets, learning compact geometry-aware features that generalize across domains. Second, we propose a geometrically consistent matching algorithm that establishes reliable CAD-to-image correspondences. Together, these contributions enable accurate, sub-second alignment per object instance without iterative pose refinement. On ScanNet25k, SUFLECA achieves 32.8%/42.6% category/instance accuracy, outperforming the strongest zero-shot baseline by 9.7/12.5 percentage points with a smaller computational footprint, and for the first time on this benchmark, even surpassing existing pose-supervised methods. Code is available at: https://github.com/snt-arg/SUFLECA

📄 PDF Abstract BibTeX arXiv:2607.15058

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Timestep-Aware Diffusion Model for Extreme Image Rescaling

2024-08-17 · Ce Wang, Zhenyu Hu, Wanjie Sun, Zhenzhong Chen

Image rescaling aims to learn the optimal low-resolution (LR) image that can be accurately reconstructed to its original high-resolution (HR) counterpart, providing an efficient image processing and storage method for ul…

DecoderImage Rescalingmodel

Depth Matching Method Based on ShapeDTW for Oil-Based Mud Imager

2025-12-01 · Fengfeng Li, Zhou Feng, Hongliang Wu, Hao Zhang 외 arxiv

In well logging operations using the oil-based mud (OBM) microresistivity imager, which employs an interleaved design with upper and lower pad sets, depth misalignment issues persist between the pad images even after vel…

On the Scalability of Diffusion-based Text-to-Image Generation

2024-04-03 · CVPR 2024 1 · Hao Li, Yang Zou, Ying Wang, Orchid Majumder 외

Scaling up model and data size has been quite successful for the evolution of LLMs. However, the scaling law for the diffusion based text-to-image (T2I) models is not fully explored. It is also unclear how to efficiently…

DenoisingDiversityImage GenerationText to Image Generation+1

CLIP-SR: Collaborative Linguistic and Image Processing for Super-Resolution

2024-12-16 · Bingwen Hu, Heng Liu, Zhedong Zheng, Ping Liu

Convolutional Neural Networks (CNNs) have advanced Image Super-Resolution (SR), but most CNN-based methods rely solely on pixel-based transformations, often leading to artifacts and blurring, particularly with severe dow…

Image Super-ResolutionSuper-Resolution

Scaling Laws for Task-Optimized Models of the Primate Visual Ventral Stream

2024-11-08 · Abdulkadir Gokce, Martin Schrimpf

When trained on large-scale object classification datasets, certain artificial neural network models begin to approximate core object recognition (COR) behaviors and neural response patterns in the primate visual ventral…

Brain DecodingInductive BiasObject Recognition