Search arXivSearch

arXiv · 2510.02906

FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence

Abstract

Multi-hop reasoning over financial disclosures is often a retrieval problem before it becomes a reasoning or generation problem: relevant facts are dispersed across sections, filings, companies, and years, and LLMs often expend excessive tokens navigating noisy context. Without precise Knowledge Graph (KG)-guided selection of relevant context, even strong reasoning models either fail to answer or consume excessive tokens, whereas KG-linked evidence enables models to focus their reasoning on composing already retrieved facts. We present FinReflectKG - MultiHop, a benchmark built on FinReflectKG, a temporally indexed financial KG that links audited triples to source chunks from S&P 100 filings (2022-2024). Mining frequent 2-3 hop subgraph patterns across sectors (via GICS taxonomy), we generate financial analyst style questions with exact supporting evidence from the KG. A two-phase pipeline first creates QA pairs via pattern-specific prompts, followed by a multi-criteria quality control evaluation to ensure QA validity. We then evaluate three controlled retrieval scenarios: (S1) precise KG-linked paths; (S2) text-only page windows centered on relevant text spans; and (S3) relevant page windows with randomizations and distractors. Across both reasoning and non-reasoning models, KG-guided precise retrieval yields substantial gains on the FinReflectKG - MultiHop QA benchmark dataset, boosting correctness scores by approximately 24 percent while reducing token utilization by approximately 84.5 percent compared to the page window setting, which reflects the traditional vector retrieval paradigm. Spanning intra-document, inter-year, and cross-company scopes, our work underscores the pivotal role of knowledge graphs in efficiently connecting evidence for multi-hop financial QA. We also release a curated subset of the benchmark (555 QA Pairs) to catalyze further research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abhinav Arun, Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali. 2025-10-03. FinReflectKG -- MultiHop: Financial QA Benchmark for Reasoning with Knowledge Graph Evidence. https://arxiv.org/abs/2510.02906

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Loss Choice or Model Choice? The Role of Forecast Level in Cryptocurrency Volatility Forecasting

Volatility forecasts play a central role in financial risk management because their overall level and day-to-day movements affect downstream decisions. Most studies compare forecasting models while keeping the training loss fixed. Yet losses emphasise different errors and can target different properties of future volatility, so raw comparisons may combine persistent forecast-level differences with differences in daily forecast movements. This leaves unresolved whether the importance of loss choice comes mainly from the forecast level it targets or from differences that remain after level adjustment. We address this gap through a comparison of seven losses and five models across major cryptocurrencies. Validation-based alignment adjusts the forecast level before the raw and aligned forecasts are evaluated using statistical scores and one-day Value-at-Risk. Before alignment, marginal score variation is greater across losses. After alignment, model choice becomes the larger source of variation in the full five-model comparison, while cross-loss differences in VaR breach rates narrow substantially. Our contribution is a comprehensive evaluation of loss and model choice that shows why losses can appear so influential in raw comparisons and how this interpretation changes when forecast level and downstream risk are considered explicitly.

q-fin.CP

CFOs Meet LLMs

Business sentiment is a closely watched economic signal, but measuring it is slow and costly: surveys typically reach only a few hundred firms, arrive periodically, and take time to compile. We show that large language models hold the potential to address these shortcomings. We prompt an LLM to role-play as the CFO of a specific company on a specific date, for every public-company CFO who responded to the Duke--Federal Reserve CFO Survey between 2002 and 2025, and answer a question about economy-wide optimism. The LLM-generated optimism score predicts the individual CFO's actual answer, even in specifications that include firm and year-quarter fixed effects as well as a control variable measuring the human CFO's lagged response. Accuracy increases with the information provided to the LLM, and the relation persists under quarterly aggregation. We find the same patterns hold for two other questions measuring CFO expectations: the respondent's optimism about their own firm and their expectation of own-firm revenues. With appropriate conditioning, LLMs may in the future be able to serve as digital twins of executives, offering scalable, high-frequency expectations data for financial research and policy.

q-fin.CP

The Physical Crash Frontier: What Finite Option Quotes Can and Cannot Reveal

Physical crash probabilities recovered from option prices depend on a pricing kernel and on a risk-neutral distribution that finitely many bid and ask quotes do not identify. For a power utility investor, we characterize the pairs of physical crash probability and expected loss below the crash threshold that the quotes admit; the boundary of this set is the physical crash frontier. Both coordinates are ratios of moments, yet when the index is bounded above the set is convex, and second-order cone programs compute it exactly at the calibrated risk aversion of two. In a decade of weekly S&P 500 cross sections, the quotes beyond the two puts nearest a 10 percent decline shrink the range of admissible crash probabilities by about 80 percent, yet its upper end remains two to three times its lower end. That lower end exists only because the index is bounded. Otherwise, for any investor more risk averse than the log investor, a vanishing probability far in the right tail inflates the denominator and drives the crash probability to zero while every quote stays inside its spread. A positive floor is therefore a joint statement about prices and a tail restriction; anything tighter than the frontier is an assumption.

q-fin.CP