paper-with-me

홈 › Papers

uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?

2024-04-03 · Pouya Sadeghi, Amirhossein Abaskohi, Yadollah Yaghoobzadeh

Inspired by human cognition, Jiang et al.(2023c) create a benchmark for assessing LLMs' lateral thinking-thinking outside the box. Building upon this benchmark, we investigate how different prompting methods enhance LLMs' performance on this task to reveal their inherent power for outside-the-box thinking ability. Through participating in SemEval-2024, task 9, Sentence Puzzle sub-task, we explore prompt engineering methods: chain of thoughts (CoT) and direct prompting, enhancing with informative descriptions, and employing contextualizing prompts using a retrieval augmented generation (RAG) pipeline. Our experiments involve three LLMs including GPT-3.5, GPT-4, and Zephyr-7B-beta. We generate a dataset of thinking paths between riddles and options using GPT-4, validated by humans for quality. Findings indicate that compressed informative prompts enhance performance. Dynamic in-context learning enhances model performance significantly. Furthermore, fine-tuning Zephyr on our dataset enhances performance across other commonsense datasets, underscoring the value of innovative thinking.

📄 PDF Abstract BibTeX arXiv:2404.02474

Code (1)

ipouyall/can-llms-be-lateral-thinkers 공식 구현 pytorch

Tasks

In-Context LearningPrompt EngineeringRAGRetrievalRetrieval-augmented GenerationSentence

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Multi-Head Attention 설명 없음
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense

2024-04-22 · Yifan Jiang, Filip Ilievski, Kaixin Ma

While vertical thinking relies on logical and commonsense reasoning, lateral thinking requires systems to defy commonsense associations and overwrite them through unconventional thinking. Lateral thinking has been shown …

Common Sense Reasoning

RACAI at SemEval-2022 Task 11: Complex named entity recognition using a lateral inhibition mechanism

2022-07-01 · SemEval (NAACL) 2022 7 · Vasile Pais

This paper presents RACAI’s system used for the shared task of “Multilingual Complex Named Entity Recognition (MultiCoNER)”, organized as part of the “The 16th International Workshop on Semantic Evaluation (SemEval 2022)…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

AmazUtah_NLP at SemEval-2024 Task 9: A MultiChoice Question Answering System for Commonsense Defying Reasoning

2024-05-16 · Mina Ghashami, Soumya Smruti Mishra

The SemEval 2024 BRAINTEASER task represents a pioneering venture in Natural Language Processing (NLP) by focusing on lateral thinking, a dimension of cognitive reasoning that is often overlooked in traditional linguisti…

Multiple-choiceQuestion AnsweringSentence

Learning to Think from Multiple Thinkers

2026-04-27 · Nirmit Joshi, Roey Magen, Nathan Srebro, Nikolaos Tsilivis 외 arxiv

We study learning with Chain-of-Thought (CoT) supervision from multiple thinkers, all of whom provide correct but possibly systematically different solutions, e.g., step-by-step solutions to math problems written by diff…

Active Learning

Mothman at SemEval-2024 Task 9: An Iterative System for Chain-of-Thought Prompt Optimization

2024-05-03 · Alvin Po-Chun Chen, Ray Groshan, Sean von Bayern

Extensive research exists on the performance of large language models on logic-based tasks, whereas relatively little has been done on their ability to generate creative solutions on lateral thinking tasks. The BrainTeas…

MemorizationPrompt Engineering