paper-with-me

홈 › Papers

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

2026-06-30 · Ziqi Li, Zijian Chen, Tingzhu Chen, Guangtao Zhai arxiv

Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contextual dependencies embedded in complete divination charges. Recently, the powerful visual perception capabilities of multimodal large language models (MLLMs) have opened new possibilities for OBI information processing. In this work, we introduce S-OBI, a novel benchmark for evaluating MLLMs in Sentence-level OBI understanding. Instead of using noisy and incomplete rubbings as the visual input, S-OBI synthesizes clear and standardized sentence-level OBI instances through glyph substitution and composition. According to 95 original rubbings with translations that have been identified, corrected, and verified by experts, we replace characters in the original rubbings with corresponding clean glyph samples sourced from existing OBI datasets while preserving the overall inscriptional structure and semantic organization. This mitigates the influence of low-level distortions and enables a more focused evaluation of sentence-level OBI understanding. Based on this, we design semantic matching, semantic slot extraction, and contextual reasoning tasks and obtain 695 question-answer pairs. Experiments reveal the inferiority of contemporary MLLMs on sentence-level OBI understanding. In particular, visual perception errors in unmasked regions propagate through the reasoning chain, leading to erroneous predictions for masked characters, which indicates that sentence-level OBI understanding in current models remains strongly dependent on character-level recognition. Overall, S-OBI provides a diagnostic benchmark for evaluating whether MLLMs can move beyond isolated character recognition toward structured inscription-level understanding.

📄 PDF Abstract BibTeX arXiv:2606.31169

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task

2025-05-28 · Yanbei Jiang, Yihao Ding, Chao Lei, Jiayang Ao 외

Current Multimodal Large Language Models (MLLMs) excel in general visual reasoning but remain underexplored in Abstract Visual Reasoning (AVR), which demands higher-order reasoning to identify abstract rules beyond simpl…

Visual Reasoning

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

2025-06-04 · Ziming Cheng, Binrui Xu, Lisheng Gong, Zuhe Song 외

With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously. However, existing MLLM benchmarks focu…

SentenceVisual Reasoning

Q-Bench+: A Benchmark for Multi-modal Foundation Models on Low-level Vision from Single Images to Pairs

2024-02-11 · ZiCheng Zhang, HaoNing Wu, Erli Zhang, Guangtao Zhai 외

The rapid development of Multi-modality Large Language Models (MLLMs) has navigated a paradigm shift in computer vision, moving towards versatile foundational models. However, evaluating MLLMs in low-level visual percept…

Image Quality AssessmentQuestion AnsweringVisual Question Answering

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

2025-05-27 · Yang Shi, Huanqian Wang, Wulin Xie, Huanyao Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such…

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

2026-08-12 · Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu 외 hf

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already achieved approximately 90% accuracy on the Video-MME le…

Video Question Answering