paper-with-me

Papers

Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese

2024-04-11 · Yuichi Inoue, Kento Sasaki, Yuma Ochi, Kazuki Fujii, Kotaro Tanahashi, Yu Yamaguchi

Vision Language Models (VLMs) have undergone a rapid evolution, giving rise to significant advancements in the realm of multimodal understanding tasks. However, the majority of these models are trained and evaluated on English-centric datasets, leaving a gap in the development and evaluation of VLMs for other languages, such as Japanese. This gap can be attributed to the lack of methodologies for constructing VLMs and the absence of benchmarks to accurately measure their performance. To address this issue, we introduce a novel benchmark, Japanese Heron-Bench, for evaluating Japanese capabilities of VLMs. The Japanese Heron-Bench consists of a variety of imagequestion answer pairs tailored to the Japanese context. Additionally, we present a baseline Japanese VLM that has been trained with Japanese visual instruction tuning datasets. Our Heron-Bench reveals the strengths and limitations of the proposed VLM across various ability dimensions. Furthermore, we clarify the capability gap between strong closed models like GPT-4V and the baseline model, providing valuable insights for future research in this domain. We release the benchmark dataset and training code to facilitate further developments in Japanese VLM research.

📄 PDF Abstract BibTeX arXiv:2404.07824

Code (1)

turingmotors/heron 공식 구현 pytorch

Similar Papers 제목 키워드 기반

HeroNet: A Hybrid Retrieval-Generation Network for Conversational Bots

2023-01-29 · Bolin Zhang, Yunzhe Xu, Zhiying Tu, Dianhui Chu

Using natural language, Conversational Bot offers unprecedented ways to many challenges in areas such as information searching, item recommendation, and question answering. Existing bots are usually developed through ret…

Multi-Task LearningQuestion AnsweringRetrievalSentence

Lean Clients, Full Accuracy: Hybrid Zeroth- and First-Order Split Federated Learning

2026-01-14 · Zhoubin Kou, Zihan Chen, Jing Yang, Cong Shen arxiv

Split Federated Learning (SFL) enables collaborative training between resource-constrained edge devices and a compute-rich server. Communication overhead is a central issue in SFL and can be mitigated with auxiliary netw…

Federated Learning

AeTHERON: Autoregressive Topology-aware Heterogeneous Graph Operator Network for Fluid-Structure Interaction

2026-04-15 · Sushrut Kumar arxiv

Surrogate modeling of body-driven fluid flows where immersed moving boundaries couple structural dynamics to chaotic, unsteady fluid phenomena remains a fundamental challenge for both computational physics and machine le…

Heron Inference for Bayesian Graphical Models

2018-02-19 · Daniel Rugeles, Zhen Hai, Gao Cong, Manoranjan Dash

Bayesian graphical models have been shown to be a powerful tool for discovering uncertainty and causal structure from real-world data in many application fields. Current inference methods primarily follow different kinds…

Computational EfficiencyVariational Inference

Harnessing PDF Data for Improving Japanese Large Multimodal Models

2025-02-20 · Jeonghun Baek, Akiko Aizawa, Kiyoharu Aizawa

Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data. Current Japanese LMMs often rely on trans…

Optical Character Recognition (OCR)