paper-with-me

홈 › Papers

FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain

2025-05-23 · Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao

Retrieval-Augmented Generation (RAG) plays a vital role in the financial domain, powering applications such as real-time market analysis, trend forecasting, and interest rate computation. However, most existing RAG research in finance focuses predominantly on textual data, overlooking the rich visual content in financial documents, resulting in the loss of key analytical insights. To bridge this gap, we present FinRAGBench-V, a comprehensive visual RAG benchmark tailored for finance which effectively integrates multimodal data and provides visual citation to ensure traceability. It includes a bilingual retrieval corpus with 60,780 Chinese and 51,219 English pages, along with a high-quality, human-annotated question-answering (QA) dataset spanning heterogeneous data types and seven question categories. Moreover, we introduce RGenCite, an RAG baseline that seamlessly integrates visual citation with generation. Furthermore, we propose an automatic citation evaluation method to systematically assess the visual citation capabilities of Multimodal Large Language Models (MLLMs). Extensive experiments on RGenCite underscore the challenging nature of FinRAGBench-V, providing valuable insights for the development of multimodal RAG systems in finance.

📄 PDF Abstract BibTeX arXiv:2505.17471

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRAGRetrievalRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

2026-01-18 · Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou 외 arxiv

Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence…

FinSight: Towards Real-World Financial Deep Research

2025-10-19 · Jiajie Jin, Yuyao Zhang, Yimeng Xu, Hongjin Qian 외 arxiv

Generating professional financial reports is a labor-intensive and intellectually demanding process that current AI systems struggle to fully automate. To address this challenge, we introduce FinSight (Financial InSight)…

CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

2025-06-16 · Jiangtong Li, Yiyun Zhu, Dawei Cheng, Zhijun Ding 외

Multimodal Large Language Models (MLLMs) have rapidly evolved with the growth of Large Language Models (LLMs) and are now applied in various fields. In finance, the integration of diverse modalities such as text, charts,…

Decision MakingFinancial AnalysisLanguage ModelingLanguage Modelling+2

Ukrainian Visual Word Sense Disambiguation Benchmark

2026-03-24 · Yurii Laba, Yaryna Mohytych, Ivanna Rohulia, Halyna Kyryleyza 외 arxiv

This study presents a benchmark for evaluating the Visual Word Sense Disambiguation (Visual-WSD) task in Ukrainian. The main goal of the Visual-WSD task is to identify, with minimal contextual information, the most appro…

Word Sense Disambiguation

FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning

2025-10-09 · Shuangyan Deng, Haizhou Peng, Jiachen Xu, Rui Mao 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in recent years. However, their rigorous evaluation within specialized domains like finance is hindered by the absence of datasets characterized by …

Mathematical Reasoning