paper-with-me

홈 › Papers

ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER

2025-09-13 · Jielong Tang, Shuang Wang, Zhenxing Wang, Jianxing Yu, Jian Yin arxiv

Grounded Multimodal Named Entity Recognition (GMNER) extends traditional NER by jointly detecting textual mentions and grounding them to visual regions. While existing supervised methods achieve strong performance, they rely on costly multimodal annotations and often underperform in low-resource domains. Multimodal Large Language Models (MLLMs) show strong generalization but suffer from Domain Knowledge Conflict, producing redundant or incorrect mentions for domain-specific entities. To address these challenges, we propose ReFineG, a three-stage collaborative framework that integrates small supervised models with frozen MLLMs for low-resource GMNER. In the Training Stage, a domain-aware NER data synthesis strategy transfers LLM knowledge to small models with supervised training while avoiding domain knowledge conflicts. In the Refinement Stage, an uncertainty-based mechanism retains confident predictions from supervised models and delegates uncertain ones to the MLLM. In the Grounding Stage, a multimodal context selection algorithm enhances visual grounding through analogical reasoning. In the CCKS2025 GMNER Shared Task, ReFineG ranked second with an F1 score of 0.6461 on the online leaderboard, demonstrating its effectiveness with limited annotations.

📄 PDF Abstract BibTeX arXiv:2509.10975

Code (0)

등록된 구현이 없습니다.

Tasks

Grounded Multimodal Named Entity RecognitionVisual Grounding

Similar Papers 제목 키워드 기반

RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and Intensity Responses

2021-11-01 · Shengyuan Xu, Wenxiao Zhao, Jing Guo

Most GAN(Generative Adversarial Network)-based approaches towards high-fidelity waveform generation heavily rely on discriminators to improve their performance. However, GAN methods introduce much uncertainty into the ge…

Audio GenerationGenerative Adversarial NetworkSinging Voice Synthesis

Compressed Sensing MRI Reconstruction using a Generative Adversarial Network with a Cyclic Loss

2017-09-03 · Tran Minh Quan, Thanh Nguyen-Duc, Won-Ki Jeong

Compressed Sensing MRI (CS-MRI) has provided theoretical foundations upon which the time-consuming MRI acquisition process can be accelerated. However, it primarily relies on iterative numerical solvers which still hinde…

compressed sensingGenerative Adversarial NetworkMRI Reconstruction

Synergizing Discriminative Exemplars and Self-Refined Experience for MLLM-based In-Context Learning in Medical Diagnosis

2026-03-29 · Wenkai Zhao, Zipei Wang, Mengjie Fang, Di Dong 외 arxiv

General Multimodal Large Language Models (MLLMs) often underperform in capturing domain-specific nuances in medical diagnosis, trailing behind fully supervised baselines. Although fine-tuning provides a remedy, the high …

Medical Diagnosis

An Expert is Worth One Token: Synergizing Multiple Expert LLMs as Generalist via Expert Token Routing

2024-03-25 · Ziwei Chai, Guoyin Wang, Jing Su, Tianjie Zhang 외

We present Expert-Token-Routing, a unified generalist framework that facilitates seamless integration of multiple expert LLMs. Our framework represents expert LLMs as special expert tokens within the vocabulary of a meta…

SOLID: a Framework of Synergizing Optimization and LLMs for Intelligent Decision-Making

2025-11-19 · Yinsheng Wang, Tario G You, Léonard Boussioux, Shan Liu arxiv

This paper introduces SOLID (Synergizing Optimization and Large Language Models for Intelligent Decision-Making), a novel framework that integrates mathematical optimization with the contextual capabilities of large lang…