paper-with-me

홈 › Papers

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

2024-06-12 · Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, JianFeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causalities. To this end, we introduce MMWorld, a new benchmark for multi-discipline, multi-faceted multimodal video understanding. MMWorld distinguishes itself from previous video understanding benchmarks with two unique advantages: (1) multi-discipline, covering various disciplines that often require domain expertise for comprehensive understanding; (2) multi-faceted reasoning, including explanation, counterfactual thinking, future prediction, etc. MMWorld consists of a human-annotated dataset to evaluate MLLMs with questions about the whole videos and a synthetic dataset to analyze MLLMs within a single modality of perception. Together, MMWorld encompasses 1,910 videos across seven broad disciplines and 69 subdisciplines, complete with 6,627 question-answer pairs and associated captions. The evaluation includes 2 proprietary and 10 open-source MLLMs, which struggle on MMWorld (e.g., GPT-4V performs the best with only 52.3\% accuracy), showing large room for improvement. Further ablation studies reveal other interesting findings such as models' different skill sets from humans. We hope MMWorld can serve as an essential step towards world model evaluation in videos.

📄 PDF Abstract BibTeX arXiv:2406.08407

Code (1)

eric-ai-lab/mmworld 공식 구현

Tasks

counterfactualFuture predictionVideo Understanding

Similar Papers 제목 키워드 기반

An Annotation Scheme for Factuality and its Application to Parliamentary Proceedings

2025-09-30 · Gili Goldin, Shira Wigderson, Ella Rabinovich, Shuly Wintner arxiv

Factuality assesses the extent to which a language utterance relates to real-world information; it determines whether utterances correspond to facts, possibilities, or imaginary situations, and as such, it is instrumenta…

Fact Checking

Discovering patterns of online popularity from time series

2019-04-10 · Mert Ozer, Anna Sapienza, Andrés Abeliuk, Goran Muric 외

How is popularity gained online? Is being successful strictly related to rapidly becoming viral in an online platform or is it possible to acquire popularity in a steady and disciplined fashion? What are other temporal c…

ClusteringTime SeriesTime Series AnalysisTime Series Clustering

Multi-Faceted Question Complexity Estimation Targeting Topic Domain-Specificity

2024-08-23 · Sujay R, Suki Perumal, Yash Nagraj, Anushka Ghei 외

Question difficulty estimation remains a multifaceted challenge in educational and assessment settings. Traditional approaches often focus on surface-level linguistic features or learner comprehension levels, neglecting …

Information RetrievalQuestion GenerationQuestion-GenerationRetrieval+1

UniDWM: Towards a Unified Driving World Model via Multifaceted Representation Learning

2026-02-02 · Shuai Liu, Siheng Ren, Xiaoyao Zhu, Quanmin Liang 외 arxiv

Achieving reliable and efficient planning in complex driving environments requires a model that can reason over the scene's geometry, appearance, and dynamics. We present UniDWM, a unified driving world model that advanc…

Representation LearningTrajectory PlanningAutonomous Driving

mTrust: Discerning Multi-Faceted Trust in a Connected World

2012-02-08 · WSDM 2012 2 · Jiliang Tang, Huiji Gao, Huan Liu

Traditionally, research about trust assumes a single type of trust between users. However, trust, as a social concept, inherently has many facets indicating multiple and heterogeneous trust relationships between users. D…