paper-with-me

Papers

VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation

2025-12-08 · Md Selim Sarowar, Sungho Kim arxiv

Vision Foundation Models (VFMs) and Vision Language Models (VLMs) have revolutionized computer vision by providing rich semantic and geometric representations. This paper presents a comprehensive visual comparison between CLIP based and DINOv2 based approaches for 3D pose estimation in hand object grasping scenarios. We evaluate both models on the task of 6D object pose estimation and demonstrate their complementary strengths: CLIP excels in semantic understanding through language grounding, while DINOv2 provides superior dense geometric features. Through extensive experiments on benchmark datasets, we show that CLIP based methods achieve better semantic consistency, while DINOv2 based approaches demonstrate competitive performance with enhanced geometric precision. Our analysis provides insights for selecting appropriate vision models for robotic manipulation and grasping, picking applications.

📄 PDF Abstract BibTeX arXiv:2512.07215

Code (0)

등록된 구현이 없습니다.

Tasks

3D Pose Estimation

Similar Papers 제목 키워드 기반

Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy

2024-02-11 · Simon Ging, María A. Bravo, Thomas Brox

The evaluation of text-generative vision-language models is a challenging yet crucial endeavor. By addressing the limitations of existing Visual Question Answering (VQA) benchmarks and proposing innovative evaluation met…

Language ModelingOpen Vocabulary Attribute DetectionVisual Question AnsweringVisual Question Answering (VQA)

CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models

2025-09-25 · Jie Cai, Kangning Yang, Lan Fu, Jiaming Ding 외 arxiv

We introduce CompareBench, a benchmark for evaluating visual comparison reasoning in vision-language models (VLMs), a fundamental yet understudied skill. CompareBench consists of 1000 QA pairs across four tasks: quantity…

Multimodal Reasoning

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

2026-05-04 · Hyobin Park, Minseok Seo, Dong-Geol Choi arxiv

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly …

Image Retrieval

VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework

2024-03-14 · Chris Kelly, Luhui Hu, Bang Yang, Yu Tian 외

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achieve open-world visual perception remains …

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering+1

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

2023-12-21 · CVPR 2024 1 · Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su 외

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical eleme…

Image RetrievalImage-to-Text RetrievalLanguage ModellingLarge Language Model+11