paper-with-me

홈 › Papers

OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation

2025-12-01 · Jinzheng Yu, Yang Xu, Haozhen Li, Junqi Li, Yifan Feng, Ligu Zhu, Hao Shen, Lei Shi arxiv

Online Public Opinion Reports consolidate news and social media for timely crisis management by governments and enterprises. While large language models have made automated report generation technically feasible, systematic research in this specific area remains notably absent, particularly lacking formal task definitions and corresponding benchmarks. To bridge this gap, we define the Automated Online Public Opinion Report Generation (OPOR-GEN) task and construct OPOR-BENCH, an event-centric dataset covering 463 crisis events with their corresponding news articles, social media posts, and a reference summary. To evaluate report quality, we propose OPOR-EVAL, a novel agent-based framework that simulates human expert evaluation by analyzing generated reports in context. Experiments with frontier models demonstrate that our framework achieves high correlation with human judgments. Our comprehensive task definition, benchmark dataset, and evaluation framework provide a solid foundation for future research in this critical domain.

📄 PDF Abstract BibTeX arXiv:2512.01896

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning from Label Proportion with Online Pseudo-Label Decision by Regret Minimization

2023-02-17 · Shinnosuke Matsuo, Ryoma Bise, Seiichi Uchida, Daiki Suehiro

This paper proposes a novel and efficient method for Learning from Label Proportions (LLP), whose goal is to train a classifier only by using the class label proportions of instance sets, called bags. We propose a novel …

Pseudo Label

TurkBench: A Benchmark for Evaluating Turkish Large Language Models

2026-01-11 · Çağrı Toraman, Ahmet Kaan Sever, Ayse Aysu Cengiz, Elif Ecem Arslan 외 arxiv

With the recent surge in the development of large language models, the need for comprehensive and language-specific evaluation benchmarks has become critical. While significant progress has been made in evaluating Englis…

Instruction Following

Efficient Online Data Mixing For Language Model Pre-Training

2023-12-05 · Alon Albalak, Liangming Pan, Colin Raffel, William Yang Wang

The data used to pretrain large language models has a decisive impact on a model's downstream performance, which has led to a large body of work on data selection methods that aim to automatically determine the most suit…

Language ModelingLanguage ModellingMMLU

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

2025-05-30 · Jiahao Ying, Wei Tang, Yiran Zhao, Yixin Cao 외

This paper introduces a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of LLMs. By decomposing the evaluation along the dimensions of linguistic medium and cultural context, this framew…

MOFSimBench: Evaluating Universal Machine Learning Interatomic Potentials In Metal--Organic Framework Molecular Modeling

2025-07-16 · Hendrik Kraß, Ju Huang, Seyed Mohamad Moosavi arxiv

Universal machine learning interatomic potentials (uMLIPs) have emerged as powerful tools for accelerating atomistic simulations, offering scalable and efficient modeling with accuracy close to quantum calculations. Howe…