paper-with-me

홈 › Papers

Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models

2023-10-23 · Gonzalo Martínez, Javier Conde, Elena Merino-Gómez, Beatriz Bermúdez-Margaretto, José Alberto Hernández, Pedro Reviriego, Marc Brysbaert

Vocabulary tests, once a cornerstone of language modeling evaluation, have been largely overlooked in the current landscape of Large Language Models (LLMs) like Llama, Mistral, and GPT. While most LLM evaluation benchmarks focus on specific tasks or domain-specific knowledge, they often neglect the fundamental linguistic aspects of language understanding and production. In this paper, we advocate for the revival of vocabulary tests as a valuable tool for assessing LLM performance. We evaluate seven LLMs using two vocabulary test formats across two languages and uncover surprising gaps in their lexical knowledge. These findings shed light on the intricacies of LLM word representations, their learning mechanisms, and performance variations across models and languages. Moreover, the ability to automatically generate and perform vocabulary tests offers new opportunities to expand the approach and provide a more complete picture of LLMs' language skills.

📄 PDF Abstract BibTeX arXiv:2310.14703

Code (1)

wordsgpt/llm_vocabulary_evaluation 공식 구현

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…
Weight Decay 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

PETA: Evaluating the Impact of Protein Transfer Learning with Sub-word Tokenization on Downstream Applications

2023-10-26 · Yang Tan, Mingchen Li, Pan Tan, Ziyi Zhou 외

Large protein language models are adept at capturing the underlying evolutionary information in primary structures, offering significant practical value for protein engineering. Compared to natural language models, prote…

Language ModelingLanguage ModellingProtein Language ModelTransfer Learning

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

2025-10-09 · Haolin Yang, Yuxing Long, Zhuoyuan Yu, Zihan Yang 외 arxiv

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reason…

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

2026-09-09 · Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte arxiv

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, …

Audio-Visual Speech Recognition

The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding

2023-11-29 · CVPR 2024 1 · Lorenzo Bianchi, Fabio Carrara, Nicola Messina, Claudio Gennaro 외

Recent advancements in large vision-language models enabled visual object detection in open-vocabulary scenarios, where object classes are defined in free-text formats during inference. In this paper, we aim to probe the…

Objectobject-detectionObject DetectionOpen-vocabulary object detection+1

Assessing Language Models with Scaling Properties

2018-04-24 · Shuntaro Takahashi, Kumiko Tanaka-Ishii

Language models have primarily been evaluated with perplexity. While perplexity quantifies the most comprehensible prediction performance, it does not provide qualitative information on the success or failure of models. …