paper-with-me

Papers

MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

2026-06-23 · Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang arxiv

Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text. A benchmark over these four signals is hard to build because finance violates four assumptions of time-series evaluation: text must be gated by its publication date to prevent look-ahead, quarterly fundamentals are reported with a one- to ninety-day lag, filing text is partly redundant with the numerical statement fields it accompanies, and macroeconomic regimes leak across calendar splits. No public benchmark addresses all four signals jointly. MacroLens covers 4,416 U.S. small- and micro-cap equities over 2021-2026. Seven tasks share one point-in-time panel of prices, 46.8M XBRL accounting facts, 53 macroeconomic series, 295,860 SEC filings, and 215,882 news articles, plus a scenario layer of 1,130 macroeconomic events across 49 types automatically detected and rendered as natural language. Tasks span contextual forecasting, public and private valuation, statement generation from fundamentals and descriptions, scenario-conditioned returns, and real-estate valuation. We evaluate 19 methods across six families spanning naive heuristics through time-series foundation models, fine-tuned LLM-based time-series models, and zero-shot large language models (LLMs), plus a five-step feature-context ablation on two frontier LLMs and a gradient-boosted baseline. MacroLens is released at https://huggingface.co/datasets/DeepAuto-AI/MacroLens.

📄 PDF Abstract BibTeX arXiv:2606.24950

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning

2025-10-09 · Shuangyan Deng, Haizhou Peng, Jiachen Xu, Rui Mao 외 arxiv

Multimodal Large Language Models (MLLMs) have made substantial progress in recent years. However, their rigorous evaluation within specialized domains like finance is hindered by the absence of datasets characterized by …

Mathematical Reasoning

FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs

2025-05-18 · Junzhe Jiang, Chang Yang, Aixin Cui, Sihan Jin 외

Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, and tool limitations. Although large langu…

SuperCLUE-Fin: Graded Fine-Grained Analysis of Chinese LLMs on Diverse Financial Tasks and Applications

2024-04-29 · Liang Xu, Lei Zhu, Yaotong Wu, Hang Xue

The SuperCLUE-Fin (SC-Fin) benchmark is a pioneering evaluation framework tailored for Chinese-native financial large language models (FLMs). It assesses FLMs across six financial application domains and twenty-five spec…

Computational EfficiencyLogical ReasoningManagement

Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering

2025-10-28 · Michail Dadopoulos, Anestis Ladas, Stratos Moschidis, Ioannis Negkakis arxiv

Retrieval-Augmented Generation (RAG) struggles on long, structured financial filings where relevant evidence is sparse and cross-referenced. This paper presents a systematic investigation of advanced metadata-driven Retr…

Question Answering

All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection

2026-01-07 · Yuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He 외 arxiv

We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures the contextual complexity of financial new…