paper-with-me

홈 › Papers

Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects

2025-04-10 · CVPR 2025 1 · Shalini Maiti, Lourdes Agapito, Filippos Kokkinos

Rapid advancements in text-to-3D generation require robust and scalable evaluation metrics that align closely with human judgment, a need unmet by current metrics such as PSNR and CLIP, which require ground-truth data or focus only on prompt fidelity. To address this, we introduce Gen3DEval, a novel evaluation framework that leverages vision large language models (vLLMs) specifically fine-tuned for 3D object quality assessment. Gen3DEval evaluates text fidelity, appearance, and surface quality by analyzing 3D surface normals, without requiring ground-truth comparisons, bridging the gap between automated metrics and user preferences. Compared to state-of-the-art task-agnostic models, Gen3DEval demonstrates superior performance in user-aligned evaluations, placing it as a comprehensive and accessible benchmark for future research on text-to-3D generation. The project page can be found here: \href{https://shalini-maiti.github.io/gen3deval.github.io/}{https://shalini-maiti.github.io/gen3deval.github.io/}.

📄 PDF Abstract BibTeX arXiv:2504.08125

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationText to 3D

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Focus 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

QUDEVAL: The Evaluation of Questions Under Discussion Discourse Parsing

2023-10-23 · Yating Wu, Ritika Mangla, Greg Durrett, Junyi Jessy Li

Questions Under Discussion (QUD) is a versatile linguistic framework in which discourse progresses as continuously asking questions and answering them. Automatic parsing of a discourse to produce a QUD structure thus ent…

Discourse ParsingLanguage ModelingLanguage ModellingQuestion Generation+2

Low-Cost Generation and Evaluation of Dictionary Example Sentences

2024-04-09 · Bill Cai, Clarence Boon Liang Ng, Daniel Tan, Shelvia Hotama

Dictionary example sentences play an important role in illustrating word definitions and usage, but manually creating quality sentences is challenging. Prior works have demonstrated that language models can be trained to…

AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models

2025-01-23 · Yang Fan

As Large Language Models (LLMs) are pretrained on massive-scale corpora, the issue of data contamination has become increasingly severe, leading to potential overestimation of model performance during evaluation. To addr…

Fairness

Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents

2025-11-14 · Davide Napolitano, Luca Cagliero, Fabrizio Battiloro arxiv

The evolution of Visual Large Language Models (VLLMs) has revolutionized the automatic understanding of Visually Rich Documents (VRDs), which contain both textual and visual elements. Although VLLMs excel in Visual Quest…

Visual Question Answering

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

2025-05-17 · Xiechi Zhang, Zetian Ouyang, LinLin Wang, Gerard de Melo 외

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, wh…

Question Answering