paper-with-me

홈 › Papers

MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions

2026-05-10 · Yaoning Yu, Ye Yu, Haojing Luo, Haohan Wang arxiv

LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations remain fragmented, which makes it difficult to compare systems or measure progress. In this paper, we focus on Reddit discussions as a concrete first step toward evaluating real-world social simulation. Reddit threads provide public, topic-grounded, multi-party interactions where people share experiences, debate, seek advice, express emotion, and collectively respond to products, events, and social issues. These discussions offer an observable window into broader social behavior, making them a useful setting for testing whether LLM agents can reproduce not only fluent text, but also the distributional patterns and interaction dynamics of real online communities. We introduce MiroBench, a benchmark for Reddit discussion simulation built from 4,292 real Reddit threads. MiroBench uses statistical tests to compare generated and real discussions across four major aspects: repetition and semantic uniformity, narrative content, toxicity and aggression, and structural complexity. Experiments across five domains and five models show that current simulators remain distributionally mismatched with real Reddit threads, while a lightweight prompt-based improvement procedure provides only limited gains. MiroBench offers a concrete benchmark for measuring, diagnosing, and improving realism in LLM-based social simulation.

📄 PDF Abstract BibTeX arXiv:2606.14715

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Systematic Benchmarking of SUMO Against Data-Driven Traffic Simulators

2025-12-20 · Erdao Liang arxiv

This paper presents a systematic benchmarking of the model-based microscopic traffic simulator SUMO against state-of-the-art data-driven traffic simulators using large-scale real-world datasets. Using the Waymo Open Moti…

Autonomous Driving

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

2026-06-02 · Tong Nie, Yuewen Mei, Yihong Tang, Junlin He 외 arxiv

Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversariality to expose failures while preserving realism. Existing method…

Autonomous Driving

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

2026-06-09 · Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin, Swasthi P Rao 외 arxiv

Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain limited in task complexity, realism, and d…

3D Generation for Embodied AI and Robotic Simulation: A Survey

2026-04-29 · Tianwei Ye, Yifan Mao, Minwen Liao, Jian Liu 외 arxiv

Embodied AI and robotic systems increasingly depend on scalable, diverse, and physically grounded 3D content for simulation-based training and real-world deployment. While 3D generative modeling has advanced rapidly, emb…

Data AugmentationScene Generation3D Generation

SAGE: Scalable Agentic 3D Scene Generation for Embodied AI

2026-02-10 · Hongchi Xia, Xuan Li, Zhaoshuo Li, Qianli Ma 외 arxiv

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-…

Scene Generation