paper-with-me

홈 › Papers

BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment

2026-01-10 · Xin Guo, Rongjunchen Zhang, Guilong Lu, Xuntao Guo, Shuai Jia, Zhi Yang, Liwen Zhang arxiv

Large language models have undergone rapid evolution, emerging as a pivotal technology for intelligence in financial operations. However, existing benchmarks are often constrained by pitfalls such as reliance on simulated or general-purpose samples and a focus on singular, offline static scenarios. Consequently, they fail to align with the requirements for authenticity and real-time responsiveness in financial services, leading to a significant discrepancy between benchmark performance and actual operational efficacy. To address this, we introduce BizFinBench.v2, the first large-scale evaluation benchmark grounded in authentic business data from both Chinese and U.S. equity markets, integrating online assessment. We performed clustering analysis on authentic user queries from financial platforms, resulting in eight fundamental tasks and two online tasks across four core business scenarios, totaling 29,578 expert-level Q&A pairs. Experimental results demonstrate that ChatGPT-5 achieves a prominent 61.5% accuracy in main tasks, though a substantial gap relative to financial experts persists; in online tasks, DeepSeek-R1 outperforms all other commercial LLMs. Error analysis further identifies the specific capability deficiencies of existing models within practical financial business contexts. BizFinBench.v2 transcends the limitations of current benchmarks, achieving a business-level deconstruction of LLM financial capabilities and providing a precise basis for evaluating efficacy in the widespread deployment of LLMs within the financial domain. The data and code are available at https://github.com/HiThink-Research/BizFinBench.v2.

📄 PDF Abstract BibTeX arXiv:2601.06401

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs

2025-05-26 · Guilong Lu, Xuntao Guo, Rongjunchen Zhang, Wenqiao Zhu 외

Large language models excel in general tasks, yet assessing their reliability in logic-heavy, precision-critical domains like finance, law, and healthcare remains challenging. To address this, we introduce BizFinBench, t…

Question Answering

Towards a unified framework for bilingual terminology extraction of single-word and multi-word terms

2018-08-01 · COLING 2018 8 · Jingshu Liu, Emmanuel Morin, Pe{\~n}a Saldarriaga

Extracting a bilingual terminology for multi-word terms from comparable corpora has not been widely researched. In this work we propose a unified framework for aligning bilingual terms independently of the term lengths. …

Word Embeddings

About Evaluating Bilingual Lexicon Induction

2022-06-01 · LREC (BUCC) 2022 6 · Martin Laville, Emmanuel Morin, Phillippe Langlais

With numerous new methods proposed recently, the evaluation of Bilingual Lexicon Induction have been quite hazardous and inconsistent across works. Some studies proposed some guidance to sanitize this; yet, they are not …

Bilingual Lexicon Induction

Duality Regularization for Unsupervised Bilingual Lexicon Induction

2019-09-03 · Xuefeng Bai, Yue Zhang, Hailong Cao, Tiejun Zhao

Unsupervised bilingual lexicon induction naturally exhibits duality, which results from symmetry in back-translation. For example, EN-IT and IT-EN induction can be mutually primal and dual problems. Current state-of-the-…

Bilingual Lexicon InductionTranslation

ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

2025-05-08 · Wanjiang Weng, Xiaofeng Tan, Hongsong Wang, Pan Zhou

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critic…

Motion Generation