paper-with-me

홈 › Papers

LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability

2025-10-28 · Zikai Xiao, Fei Huang, Jianhong Tu, Jianhui Wei, Wen Ma, Yuxuan Zhou, Jian Wu, Bowen Yu, Zuozhu Liu, Junyang Lin arxiv

Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce \textbf{LongWeave}, which balances real-world and verifiable assessment with Constraint-Verifier Evaluation (CoV-Eval). CoV-Eval constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and constraints based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs shows that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase.

📄 PDF Abstract BibTeX arXiv:2510.24345

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation

2023-10-18 · Shengqiang Zhang, Philipp Wicke, Lütfi Kerem Şenel, Luis Figueredo 외

The convergence of embodied agents and large language models (LLMs) has brought significant advancements to embodied instruction following. Particularly, the strong reasoning capabilities of LLMs make it possible for rob…

Caption GenerationInstruction Following

AudioStory: Generating Long-Form Narrative Audio with Large Language Models

2025-08-27 · Yuxin Guo, Teng Wang, Yuying Ge, Shijie Ma 외 arxiv

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To address this gap, w…

Audio Generation

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

2025-12-15 · Jianxiong Gao, Zhaoxi Chen, Xian Liu, Junhao Zhuang 외 arxiv

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: contr…

Video Generation

OsmT: Bridging OpenStreetMap Queries and Natural Language with Open-source Tag-aware Language Models

2025-12-04 · Zhuoyue Wan, Wentao Hu, Chen Jason Zhang, Yuanfeng Song 외 arxiv

Bridging natural language and structured query languages is a long-standing challenge in the database community. While recent advances in language models have shown promise in this direction, existing solutions often rel…

Infinite Motion: Extended Motion Generation via Long Text Instructions

2024-07-11 · Mengtian Li, Chengshuo Zhai, Shengxiang Yao, Zhifeng Xie 외

In the realm of motion generation, the creation of long-duration, high-quality motion sequences remains a significant challenge. This paper presents our groundbreaking work on "Infinite Motion", a novel approach that lev…

Motion GenerationMotion Synthesis