paper-with-me

홈 › Papers

SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality

2026-03-28 · Qinghao Guan, Yuchen Pan, Donghao Li, Zishi Zhang, Yiyang Chen, Lu Li, Flaminia Canu, Emilia Volkart, Gerold Schneider arxiv

In religion and theology studies, spirituality has garnered significant research attention for the reason that it not only transcends culture but offers unique experience to each individual. However, social scientists often rely on limited datasets, which are basically unavailable online. In this study, we collaborated with social scientists to develop a high-quality multimedia multi-modal datasets, \textbf{SACRED}, in which the faithfulness of classification is guaranteed. Using \textbf{SACRED}, we evaluated the performance of 13 popular LLMs as well as traditional rule-based and fine-tuned approaches. The result suggests DeepSeek-V3 model performs well in classifying such abstract concepts (i.e., 79.19\% accuracy in the Quora test set), and the GPT-4o-mini model surpassed the other models in the vision tasks (63.99\% F1 score). Purportedly, this is the first annotated multi-modal dataset from online spirituality communication. Our study also found a new type of connectedness which is valuable for communication science studies.

📄 PDF Abstract BibTeX arXiv:2603.27331

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages

2024-05-22 · Corinne Aars, Lauren Adams, Xiaokan Tian, Zhaoyu Wang 외

This study presents the development and evaluation of a ByT5-based multilingual translation model tailored for translating the Bible into underrepresented languages. Utilizing the comprehensive Johns Hopkins University B…

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard

2025-11-13 · Yudong Yang, Xuezhen Zhang, Zhifeng Han, Siyin Wang 외 arxiv

Recent progress in LLMs has enabled understanding of audio signals, but has also exposed new safety risks arising from complex audio inputs that are inadequately handled by current safeguards. We introduce SACRED-Bench (…

Training Multimedia Event Extraction With Generated Images and Captions

2023-06-15 · Zilin Du, Yunxin Li, Xu Guo, Yidan Sun 외

Contemporary news reporting increasingly features multimedia content, motivating research on multimedia event extraction. However, the task lacks annotated multimodal training data and artificially generated training dat…

Event ExtractionStructured Prediction

Creating HAVIC: Heterogeneous Audio Visual Internet Collection

2012-05-01 · LREC 2012 5 · Stephanie Strassel, Am Morris, a, Jonathan Fiscus 외

Linguistic Data Consortium and the National Institute of Standards and Technology are collaborating to create a large, heterogeneous annotated multimodal corpus to support research in multimodal event detection and relat…

Event Detection

Joint Multimedia Event Extraction from Video and Article

2021-09-27 · Findings (EMNLP) 2021 11 · Brian Chen, Xudong Lin, Christopher Thomas, Manling Li 외

Visual and textual modalities contribute complementary information about events described in multimedia documents. Videos contain rich dynamics and detailed unfoldings of events, while text describes more high-level and …

Articlescoreference-resolutionCoreference ResolutionEvent Coreference Resolution+1