paper-with-me

홈 › Papers

A Critique of a Critique of Word Similarity Datasets: Sanity Check or Unnecessary Confusion?

2017-07-12 · Minh Le

Critical evaluation of word similarity datasets is very important for computational lexical semantics. This short report concerns the sanity check proposed in Batchkarov et al. (2016) to evaluate several popular datasets such as MC, RG and MEN -- the first two reportedly failed. I argue that this test is unstable, offers no added insight, and needs major revision in order to fulfill its purported goal.

📄 PDF Abstract BibTeX arXiv:1707.03819

Code (0)

등록된 구현이 없습니다.

Tasks

Word Similarity

Similar Papers 제목 키워드 기반

A critique of word similarity as a method for evaluating distributional semantic models

2016-08-01 · WS 2016 8 · Miroslav Batchkarov, Thomas Kober, Jeremy Reffin, Julie Weeds 외
Document ClassificationNatural Language InferenceWord Similarity

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

2026-06-29 · Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh 외 arxiv

Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): it has no single correct answer, and most aesthetic evaluation measures models against numeric scores rather than the written cri…

multimodal generation

The Critique of Critique

2024-01-09 · Shichao Sun, Junlong Li, Weizhe Yuan, Ruifeng Yuan 외

Critique, as a natural language description for assessing the quality of model-generated content, has played a vital role in the training, evaluation, and refinement of LLMs. However, a systematic method to evaluate the …

Question Answering

Training Language Models to Critique With Multi-agent Feedback

2024-10-20 · Tian Lan, Wenwei Zhang, Chengqi Lyu, Shuaibin Li 외

Critique ability, a meta-cognitive capability of humans, presents significant challenges for LLMs to improve. Recent works primarily rely on supervised fine-tuning (SFT) using critiques generated by a single LLM like GPT…

Reinforcement Learning (RL)

RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

2023-05-15 · Afra Feyza Akyürek, Ekin Akyürek, Aman Madaan, Ashwin Kalyan 외

Despite their unprecedented success, even the largest language models make mistakes. Similar to how humans learn and improve using feedback, previous work proposed providing language models with natural language feedback…

reinforcement-learningRetrievaltext similarity