paper-with-me

홈 › Papers

OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion

2024-11-26 · Hanqi Jiang, Yi Pan, JunHao Chen, Zhengliang Liu, Yifan Zhou, Peng Shu, Yiwei Li, Huaqin Zhao, Stephen Mihm, Lewis C Howe, Tianming Liu

Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We introduce OracleSage, a novel cross-modal framework that integrates hierarchical visual understanding with graph-based semantic reasoning. Specifically, we propose (1) a Hierarchical Visual-Semantic Understanding module that enables multi-granularity feature extraction through progressive fine-tuning of LLaVA's visual backbone, (2) a Graph-based Semantic Reasoning Framework that captures relationships between visual components and semantic concepts through dynamic message passing, and (3) OracleSem, a semantically enriched OBS dataset with comprehensive pictographic and semantic annotations. Experimental results demonstrate that OracleSage significantly outperforms state-of-the-art vision-language models. This research establishes a new paradigm for ancient text interpretation while providing valuable technical support for archaeological studies.

📄 PDF Abstract BibTeX arXiv:2411.17837

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Better Visual Dialog Agents with Pretrained Visual-Linguistic Representation

2021-05-24 · CVPR 2021 1 · Tao Tu, Qing Ping, Govind Thattai, Gokhan Tur 외

GuessWhat?! is a two-player visual dialog guessing game where player A asks a sequence of yes/no questions (Questioner) and makes a final guess (Guesser) about a target object in an image, based on answers from player B …

Referring ExpressionReferring Expression ComprehensionVisual DialogVisual Grounding

ConsistCompose: Unified Multimodal Layout Control for Image Composition

2025-11-23 · Xuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 외 arxiv

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…

Visual GroundingImage Generation

AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control

2026-03-15 · Peng Xu, Zhengnan Deng, Jiayan Deng, Zonghua Gu 외 arxiv

Vision-Language Navigation (VLN) for Unmanned Aerial Vehicles (UAVs) demands complex visual interpretation and continuous control in dynamic 3D environments. Existing hierarchical approaches rely on dense oracle guidance…

Vision-Language NavigationContinuous Control

InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

2025-08-12 · Xiaolei Diao, Zhihan Zhou, Lida Shi, Ting Wang 외 arxiv

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effecti…

Data Augmentation

VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point Cloud

2023-03-25 · CVPR 2023 1 · Ziqin Wang, Bowen Cheng, Lichen Zhao, Dong Xu 외

The task of 3D semantic scene graph (3DSSG) prediction in the point cloud is challenging since (1) the 3D point cloud only captures geometric structures with limited semantics compared to 2D images, and (2) long-tailed r…

3D geometry3d scene graph generationPredictionRelation