paper-with-me

홈 › Papers

HopWeaver: Synthesizing Authentic Multi-Hop Questions Across Text Corpora

2025-05-21 · Zhiyu Shen, Jiyuan Liu, Yunhe Pang, Yanghui Rao

Multi-Hop Question Answering (MHQA) is crucial for evaluating the model's capability to integrate information from diverse sources. However, creating extensive and high-quality MHQA datasets is challenging: (i) manual annotation is expensive, and (ii) current synthesis methods often produce simplistic questions or require extensive manual guidance. This paper introduces HopWeaver, the first automatic framework synthesizing authentic multi-hop questions from unstructured text corpora without human intervention. HopWeaver synthesizes two types of multi-hop questions (bridge and comparison) using an innovative approach that identifies complementary documents across corpora. Its coherent pipeline constructs authentic reasoning paths that integrate information across multiple documents, ensuring synthesized questions necessitate authentic multi-hop reasoning. We further present a comprehensive system for evaluating synthesized multi-hop questions. Empirical evaluations demonstrate that the synthesized questions achieve comparable or superior quality to human-annotated datasets at a lower cost. Our approach is valuable for developing MHQA datasets in specialized domains with scarce annotated resources. The code for HopWeaver is publicly available.

📄 PDF Abstract BibTeX arXiv:2505.15087

Code (1)

zh1yushen/hopweaver 공식 구현 pytorch

Tasks

Multi-hop Question AnsweringQuestion Answering

Similar Papers 제목 키워드 기반

Side Auth: Synthesizing Virtual Sensors for Authentication

2023-01-27 · Yan Long, Kevin Fu

While the embedded security research community aims to protect systems by reducing analog sensor side channels, our work argues that sensor side channels can be beneficial to defenders. This work introduces the general p…

Enhancing Medical Imaging with GANs Synthesizing Realistic Images from Limited Data

2024-05-22 · Yinqiu Feng, Bo Zhang, Lingxi Xiao, Yutian Yang 외

In this research, we introduce an innovative method for synthesizing medical images using generative adversarial networks (GANs). Our proposed GANs method demonstrates the capability to produce realistic synthetic images…

Model Optimization

LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

2024-10-13 · Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang 외

With the rapid development of AI-generated content, the future internet may be inundated with synthetic data, making the discrimination of authentic and credible multimodal data increasingly challenging. Synthetic data d…

Multiple-choice

Persona Authentication through Generative Dialogue

2021-10-25 · Fengyi Tang, Lifan Zeng, Fei Wang, Jiayu Zhou

In this paper we define and investigate the problem of \emph{persona authentication}: learning a conversational policy to verify the consistency of persona models. We propose a learning objective and prove (under some mi…

Quantifying truth and authenticity in AI-assisted candidate evaluation: A multi-domain pilot analysis

2025-11-02 · Eldred Lee, Nicholas Worley, Koshu Takatsuji arxiv

This paper presents a retrospective analysis of anonymized candidate-evaluation data collected during pilot hiring campaigns conducted through AlteraSF, an AI-native resume-verification platform. The system evaluates res…