paper-with-me

Papers

PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models

2024-01-26 · Haochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song

Large Language Models (LLMs) have succeeded remarkably in understanding long-form contents. However, exploring their capability for generating long-form contents, such as reports and articles, has been relatively unexplored and inadequately assessed by existing benchmarks. The prevalent evaluation methods, which predominantly rely on crowdsourcing, are recognized for their labor-intensive nature and lack of efficiency, whereas automated metrics, such as the ROUGE score, demonstrate discordance with human judgment criteria. In this paper, we propose ProxyQA, an innovative framework dedicated to assessing long-text generation. ProxyQA comprises in-depth human-curated meta-questions spanning various domains, each accompanied by specific proxy-questions with pre-annotated answers. LLMs are tasked to generate extensive content in response to these meta-questions, by engaging an evaluator and incorporating the generated texts as contextual background, ProxyQA assesses the generated content's quality through the evaluator's accuracy in addressing the proxy-questions. We examine multiple LLMs, emphasizing ProxyQA's demanding nature as a high-quality assessment tool. Human evaluation demonstrates that the proxy-question method is notably self-consistent and aligns closely with human evaluative standards. The dataset and leaderboard is available at \url{https://proxy-qa.com}.

📄 PDF Abstract BibTeX arXiv:2401.15042

Code (1)

namco0816/proxyqa 공식 구현

Tasks

ArticlesFormText Generation

Similar Papers 제목 키워드 기반

SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification

2025-11-12 · Mohamed Elaraby, Jyoti Prakash Maheswari arxiv

Large Language Models (LLMs) with extended context windows promise direct reasoning over long documents, reducing the need for chunking or retrieval. Constructing annotated resources for training and evaluation, however,…

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

2026-07-23 · Mack Nixon, Liam Wright, Yevgeniya Kovalchuk, Alison Fang-Wei Wu 외 arxiv

Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governan…

Evaluating generative models in high energy physics

2022-11-18 · Raghav Kansal, Anni Li, Javier Duarte, Nadezda Chernyavskaya 외

There has been a recent explosion in research into machine-learning-based generative modeling to tackle computational challenges for simulations in high energy physics (HEP). In order to use such alternative simulators i…

Generative Adversarial NetworkVocal Bursts Intensity Prediction

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

2026-04-22 · Xiyang Wu, Zongxia Li, Guangyao Shi, Alexander Duffy 외 arxiv

Long horizon interactive environments are a testbed for evaluating agents skill usage abilities. These environments demand multi step reasoning, the chaining of multiple skills over many timesteps, and robust decision ma…

Decision Making

A long-term alternative formula for a stochastic stock price model

2019-04-09 · Takuya Okabe, Jin Yoshimura

This study presents a long-term alternative formula for stock price variation described by a geometric Brownian motion on the basis of median instead of mean or expected values. The proposed method is motivated by the ob…