paper-with-me

홈 › Papers

RISCORE: Enhancing In-Context Riddle Solving in Language Models through Context-Reconstructed Example Augmentation

2024-09-24 · Ioannis Panagiotopoulos, Giorgos Filandrianos, Maria Lymperaiou, Giorgos Stamou

Riddle-solving requires advanced reasoning skills, pushing LLMs to engage in abstract thinking and creative problem-solving, often revealing limitations in their cognitive abilities. In this paper, we examine the riddle-solving capabilities of LLMs using a multiple-choice format, exploring how different prompting techniques impact performance on riddles that demand diverse reasoning skills. To enhance results, we introduce RISCORE (RIddle Solving with COntext REcontruciton) a novel fully automated prompting method that generates and utilizes contextually reconstructed sentence-based puzzles in conjunction with the original examples to create few-shot exemplars. Our experiments demonstrate that RISCORE significantly improves the performance of language models in both vertical and lateral thinking tasks, surpassing traditional exemplar selection strategies across a variety of few-shot settings.

📄 PDF Abstract BibTeX arXiv:2409.16383

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceSentence

Similar Papers 제목 키워드 기반

CC-Riddle: A Question Answering Dataset of Chinese Character Riddles

2022-06-28 · Fan Xu, Yunxiang Zhang, Xiaojun Wan

The Chinese character riddle is a unique form of cultural entertainment specific to the Chinese language. It typically comprises two parts: the riddle description and the solution. The solution to the riddle is a single …

General KnowledgeLanguage ModellingMultiple-choiceQuestion Answering

BiRdQA: A Bilingual Dataset for Question Answering on Tricky Riddles

2021-09-23 · Yunxiang Zhang, Xiaojun Wan

A riddle is a question or statement with double or veiled meanings, followed by an unexpected answer. Solving riddle is a challenging task for both machine and human, testing the capability of understanding figurative, c…

Multiple-choiceQuestion Answering

VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation

2024-06-27 · Yixiao Song, Yekyung Kim, Mohit Iyyer

Existing metrics for evaluating the factuality of long-form text, such as FACTSCORE (Min et al., 2023) and SAFE (Wei et al., 2024), decompose an input text into "atomic claims" and verify each against a knowledge base li…

FormText Generation

Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

2024-07-28 · Nitzan Bitton-Guetta, Aviv Slobodkin, Aviya Maimon, Eliya Habba 외

Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person's discomfort, ther…

World Knowledge

The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles

2025-11-02 · Abhinav P M, Ojasva Saxena, Oswald C, Parameswari Krishnamurthy arxiv

The extent to which large language models (LLMs) can perform culturally grounded reasoning across non-English languages remains underexplored. This paper examines the reasoning and self-assessment abilities of LLMs acros…