paper-with-me

홈 › Papers

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities?

2025-02-12 · Jiahe Jin, Yanheng He, Mingyan Yang

In this work, we identify the "2D-Cheating" problem in 3D LLM evaluation, where these tasks might be easily solved by VLMs with rendered images of point clouds, exposing ineffective evaluation of 3D LLMs' unique 3D capabilities. We test VLM performance across multiple 3D LLM benchmarks and, using this as a reference, propose principles for better assessing genuine 3D understanding. We also advocate explicitly separating 3D abilities from 1D or 2D aspects when evaluating 3D LLMs.

📄 PDF Abstract BibTeX arXiv:2502.08503

Code (1)

llm-class-group/revisiting-3d-llm-benchmarks 공식 구현

Similar Papers 제목 키워드 기반

Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist

2024-07-11 · ZiHao Zhou, Shudong Liu, Maizhen Ning, Wei Liu 외

Exceptional mathematical reasoning ability is one of the key features that demonstrate the power of large language models (LLMs). How to comprehensively define and evaluate the mathematical abilities of LLMs, and even re…

GSM8KMathMathematical Reasoning

Revisiting Parameter-Efficient Tuning: Are We Really There Yet?

2022-02-16 · Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, Shangsong Liang

Parameter-Efficient Tuning (PETuning) methods have been deemed by many as the new paradigm for using pretrained language models (PLMs). By tuning just a fraction amount of parameters comparing to full model finetuning, P…

Do We Really Need to Collect Millions of Faces for Effective Face Recognition?

2016-03-23 · Iacopo Masi, Anh Tuan Tran, Jatuporn Toy Leksut, Tal Hassner 외

Face recognition capabilities have recently made extraordinary leaps. Though this progress is at least partially due to ballooning training set sizes -- huge numbers of face images downloaded and labeled for identity -- …

Face RecognitionFace Verification

Revisiting Long-tailed Image Classification: Survey and Benchmarks with New Evaluation Metrics

2023-02-03 · Chaowei Fang, Dingwen Zhang, Wen Zheng, Xue Li 외

Recently, long-tailed image classification harvests lots of research attention, since the data distribution is long-tailed in many real-world situations. Piles of algorithms are devised to address the data imbalance prob…

image-classificationImage ClassificationLong-tail Learning

What's the Meaning of Superhuman Performance in Today's NLU?

2023-05-15 · Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 외

In the last five years, there has been a significant focus in Natural Language Processing (NLP) on developing larger Pretrained Language Models (PLMs) and introducing benchmarks such as SuperGLUE and SQuAD to measure the…

PositionReading Comprehension