paper-with-me

Papers StrategyQA

“StrategyQA” 태그가 달린 논문 40편 · 필터 해제

Fusing Bidirectional Chains of Thought and Reward Mechanisms A Method for Enhancing Question-Answering Capabilities of Large Language Models for Chinese Intangible Cultural Heritage

2025-05-13 · Ruilin Liu, Zhixiao Zhao, Jieqiong Li, Chang Liu 외

The rapid development of large language models (LLMs) has provided significant support and opportunities for the advancement of domain-specific LLMs. However, fine-tuning these large models using Intangible Cultural Heri…

Knowledge DistillationLarge Language ModelQuestion AnsweringStrategyQA

Rule-Guided Feedback: Enhancing Reasoning by Enforcing Rule Adherence in Large Language Models

2025-03-14 · Aissatou Diallo, Antonis Bikakis, Luke Dickens, Anthony Hunter 외

In this paper, we introduce Rule-Guided Feedback (RGF), a framework designed to enhance Large Language Model (LLM) performance through structured rule adherence and strategic information seeking. RGF implements a teacher…

Checkmate In OneGSM8KLanguage ModelingLanguage Modelling+3

DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability

2025-03-04 · Yunzhen He, Yusuke Takase, Yoichi Ishibashi, Hidetoshi Shimodaira

Large Language Models (LLMs) are increasingly being used in real-world applications. However, concerns about the reliability of the content they generate persist, as it frequently deviates from factual correctness or exh…

GSM8KLogical ReasoningStrategyQATrajectory Prediction+1

Voting or Consensus? Decision-Making in Multi-Agent Debate

2025-02-26 · Lars Benedikt Kaesberg, Jonas Becker, Jan Philip Wahle, Terry Ruas 외

Much of the success of multi-agent debates depends on carefully choosing the right parameters. The decision-making protocol stands out as it can highly impact final model answers, depending on how decisions are reached. …

Decision MakingMMLUStrategyQA

Unraveling Indirect In-Context Learning Using Influence Functions

2025-01-01 · Hadi Askari, Shivanshu Gupta, Terry Tong, Fei Wang 외

This work introduces a novel paradigm for generalized In-Context Learning (ICL), termed Indirect In-Context Learning. In Indirect ICL, we explore demonstration selection strategies tailored for two distinct real-world sc…

In-Context LearningInformativenessMMLUStrategyQA

AutoReason: Automatic Few-Shot Reasoning Decomposition

2024-12-09 · Arda Sevinc, Abdurrahman Gumus

Chain of Thought (CoT) was introduced in recent research as a method for improving step-by-step reasoning in Large Language Models. However, CoT has limited applications such as its need for hand-crafted few-shot exempla…

StrategyQA

Dialectical Behavior Therapy Approach to LLM Prompting

2024-10-10 · Oxana Vitman, Nika Amaglobeli, Paul Plachinda

Large language models demonstrated state-of-the-art results on various reasoning tasks when applying the chain-of-thought (CoT) prompting technique. CoT prompting guides the model into breaking tasks into a few intermedi…

GSM8KStrategyQA

Rationale-Aware Answer Verification by Pairwise Self-Evaluation

2024-10-07 · Akira Kawabata, Saku Sugawara

Answer verification identifies correct solutions among candidates generated by large language models (LLMs). Current approaches typically train verifier models by labeling solutions as correct or incorrect based solely o…

ARCStrategyQAvalid

A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions

2024-09-30 · Laurène Vaugrante, Mathias Niepert, Thilo Hagendorff

In an era where large language models (LLMs) are increasingly integrated into a wide range of everyday applications, research into these models' behavior has surged. However, due to the novelty of the field, clear method…

Prompt EngineeringStrategyQA

Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning

2024-09-25 · Debargha Ganguly, Srinivasan Iyengar, Vipin Chaudhary, Shivkumar Kalyanaraman

Large Language Models (LLMs) have revolutionized natural language processing, yet they struggle with inconsistent reasoning, particularly in novel domains and complex logical sequences. This research introduces Proof of …

BenchmarkingFormal LogicMultimodal ReasoningProgram Synthesis+1

Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

2024-08-12 · Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang 외

This paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into…

GSM8KMathStrategyQA

Meta-prompting Optimized Retrieval-augmented Generation

2024-07-04 · João Rodrigues, António Branco

Retrieval-augmented generation resorts to content retrieved from external sources in order to leverage the performance of large language models in downstream tasks. The excessive volume of retrieved content, the possible…

Multi-hop Question AnsweringQuestion AnsweringRetrievalRetrieval-augmented Generation+1

Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks

2024-07-04 · Dharunish Yugeswardeenoo, Kevin Zhu, Sean O'Brien

Although LLMs have the potential to transform many fields, they still underperform humans in reasoning tasks. Existing methods induce the model to produce step-by-step calculations, but this research explores the questio…

GSM8KStrategyQA

Advancing Process Verification for Large Language Models via Tree-Based Preference Learning

2024-06-29 · Mingqian He, Yongliang Shen, Wenqi Zhang, Zeqi Tan 외

Large Language Models (LLMs) have demonstrated remarkable potential in handling complex reasoning tasks by generating step-by-step rationales.Some methods have proven effective in boosting accuracy by introducing extra v…

Binary ClassificationGSM8KMathStrategyQA

Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast

2024-05-23 · Chufan Shi, Cheng Yang, Xinyu Zhu, Jiahao Wang 외

Mixture-of-Experts (MoE) has emerged as a prominent architecture for scaling model size while maintaining computational efficiency. In MoE, each token in the input sequence activates a different subset of experts determi…

Computational EfficiencyGSM8KHumanEvalmbpp+2

Improving Attributed Text Generation of Large Language Models via Preference Learning

2024-03-27 · Dongfang Li, Zetian Sun, Baotian Hu, Zhenyu Liu 외

Large language models have been widely adopted in natural language processing, yet they face the challenge of generating unreliable content. Recent works aim to reduce misinformation and hallucinations by resorting to at…

MisinformationRetrievalStrategyQAText Generation

CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge

2024-03-03 · Willis Guo, Armin Toroghi, Scott Sanner

Knowledge graph question answering (KGQA) is a well-established field that seeks to provide factual answers to natural language (NL) questions by leveraging knowledge graphs (KGs). However, existing KGQA datasets suffer …

Claim VerificationGraph Question AnsweringHallucinationKnowledge Graphs+2

Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation

2024-02-21 · Phuc Phan, Hieu Tran, Long Phan

We propose a straightforward approach called Distillation Contrastive Decoding (DCD) to enhance the reasoning capabilities of Large Language Models (LLMs) during inference. In contrast to previous approaches that relied …

Arithmetic ReasoningGSM8KLanguage ModellingQuantization+1

Towards Uncertainty-Aware Language Agent

2024-01-25 · Jiuzhou Han, Wray Buntine, Ehsan Shareghi

While Language Agents have achieved promising success by placing Large Language Models at the core of a more versatile design that dynamically interacts with the external world, the existing approaches neglect the notion…

MMLUStrategyQAUncertainty Quantification

Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning

2024-01-19 · Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan 외

Self-consistency (SC) has been a widely used decoding strategy for chain-of-thought reasoning. Despite bringing significant performance improvements across a variety of multi-step reasoning tasks, it is a high-cost metho…

GSM8KMathStrategyQA
1–20 / 40 다음 →