Source-linked AI summary
FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review
Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
TL;DR
Existing financial benchmarks measure broad competence but are often organized around datasets or task formats rather than deployed professional decisions and evidence states. FinRiskAtlas evaluates operation execution under fixed evidence and evidence-state control under evolving conditions, finding that operation-level rankings diverge and knowledge-based selection can incur operation-specific regret. The framework therefore supports evaluation units aligned with professional decisions and the evidence states from which they are made.
Problem
Existing financial benchmarks provide general competence measurements but often organize evaluation around datasets or task formulations rather than professional operations and evidence boundaries.
Method
FinRiskAtlas evaluates operation execution under fixed evidence, while FinRisk-Ask replays professional trajectories to evaluate proceeding or requesting evidence without exposing future information during inference.
Results
Across 33 configurations, downstream-operation rankings had mean pairwise Spearman correlation 0.42, and Domain Knowledge shortlisting incurred up to 18.01 points of regret on an individual operation.
Takeaways & Limitations
Broad financial capability scores do not fully capture where models are reliable in professional workflows, supporting evaluation aligned with decisions and evidence states.
Takeaways & Limitations
FinRisk-Ask requires later evidence to be expert-verified as unavailable at the decision point and relevant to an unresolved review need.
Abstract
from arXiv · showhide
Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision. Existing financial benchmarks cover knowledge, reasoning, compliance, and professional tasks, but their evaluation units are often organized around datasets or task formulations rather than the decisions that deployed systems support. We introduce FinRiskAtlas, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions. The static benchmark contains 9,742 instances across 53 task families, including 42 Domain Knowledge families and eleven downstream review operations defined by explicit evaluation contracts. FinRisk-Ask extends this framework through offline replay of 680 pre-action states from 104 de-identified professional trajectories, withholding future evidence during inference and using it only to construct expert-verified evidence targets. Across 33 model configurations, operation-level evaluation yields non-redundant rankings (mean pairwise Spearman correlation 0.42 across downstream operations), and knowledge-based shortlisting can incur up to 18.01 points of regret on individual operations. FinRisk-Ask further shows that entering the Ask branch more frequently does not necessarily improve request targeting or end-to-end evidence acquisition. These results show that broad financial capability scores do not fully capture where models are reliable in professional workflows, motivating evaluation units aligned with the decisions and evidence states that deployed systems must support.
1. Introduction
Professional financial review consists of multiple evidence-dependent decisions, not a single prediction task. Deploying LLMs therefore requires evaluating both the specific review operation and whether the available evidence supports proceeding.
- Financial review may require identifying parties, extracting evidence, verifying quantities, applying provisions, assigning risks, or producing supported recommendations.
- Operation execution asks which model configuration can reliably perform a required review operation using the currently visible record.
- Evidence-state control asks whether a workflow should proceed with available information or acquire additional evidence before deciding.
KNOWLEDGE 2,202
FinRiskAtlas combines a broad knowledge benchmark with operation-specific and trajectory-based evaluation. It measures both execution under fixed evidence and evidence acquisition under evolving review states.
- 9,742 static instances span 53 task families, including 42 Domain Knowledge families and eleven downstream operation families.
- FinRiskAtlas defines each downstream family through an explicit contract covering visible evidence, the decision object, the reviewer artifact, and scoring protocol.
- FinRisk-Ask replays 680 pre-action states from 104 completed, de-identified professional trajectories while withholding later evidence during inference.
- Later-observed evidence in retained Ask states is used only to construct expert-verified unresolved evidence targets.
- Across 33 configurations, downstream-operation rankings have mean pairwise Spearman correlation 0.42, while Domain Knowledge shortlisting can incur up to 18.01 points of regret.
- Action-selection behavior can still produce substantially different end-to-end evidence-acquisition performance.
2. Related Work
Prior financial benchmarks broaden coverage of knowledge, reasoning, applications, compliance, and safety, while diagnostic and information-acquisition research studies specialized behaviors. FinRiskAtlas addresses the remaining need for workflow-aligned evaluation of professional operations and evidence boundaries.
- Financial benchmarks progressed from focused language and numerical reasoning tasks to broad suites covering heterogeneous capabilities and Chinese financial scenarios.
- Recent application benchmarks examine document errors, tool-supported research, compliance, safety, and finance-specific red-team behavior.
- Structured diagnostic frameworks show that aggregate scores can hide meaningful differences across capabilities and scenarios.
- Existing diagnostic frameworks primarily organize evaluation around capabilities or scenarios rather than deployed professional operations and evidence boundaries.
- Clarification, selective answering, abstention, deferral, and help-seeking methods study how models respond to ambiguity or insufficient information.
- Recent information-acquisition benchmarks evaluate missing-variable identification and what or when information should be requested under incomplete conditions.
- FinRisk-Ask evaluates recorded professional boundaries offline, measuring both whether models should acquire evidence and whether requests address unresolved review needs.
3. The FinRiskAtlas Benchmark
FinRiskAtlas treats the decision context—not an isolated task format—as the fundamental unit for evaluating professional financial workflows.
- A useful benchmark unit specifies available information, the decision to resolve, the expected reviewer artifact, and the evaluation procedure.
2 Capability Taxonomy
FinRiskAtlas organizes evaluation around professional review functions, evidence regimes, decision objects, and reviewer artifacts rather than surface task formats. Its static benchmark separates supporting knowledge from evidence processing and applied review, while FinRisk-Ask evaluates evidence-state control as a complementary setting.
- Benchmark structure: 53 task families and 9,742 instances comprise the static benchmark, including 42 Domain Knowledge families and eleven downstream operation families.FinRisk-Ask is a separate state-level setting and is not counted as an additional static family.
- Evaluation contracts: Each task family is governed by an evaluation contract specifying capability, visible information, decision object, reviewer-artifact schema, and scoring protocol.The information regime includes evidence sources, fields, and the temporal boundary available to the model.
- Evaluation contracts: Family boundaries follow professional function: different decisions or artifacts define distinct families, even when tasks share source cases or surface formats.The contract defines the measurement boundary rather than a mandatory execution order for every review case.
- Capability taxonomy: The taxonomy covers Domain Knowledge, Evidence-Grounded Processing, and Applied Review as nested capability layers for financial risk and compliance review.Domain Knowledge supports review, Evidence-Grounded Processing structures heterogeneous records, and Applied Review integrates evidence and rules into professional outputs.
- Benchmark construction: Construction is capability-first and model-independent, with expert-defined contracts, candidate review, and quality checks for evidence support, ambiguity, leakage, parsing, and provenance.Items lacking stable references or reproducible scoring procedures are removed.
- Evidence-state evaluation: FinRisk-Ask complements fixed-evidence operation evaluation by replaying evolving evidence states and assessing whether models should proceed or request additional evidence.Later evidence is withheld during inference and used only to construct expert-verified request targets.
4. Experiments
Experiments show that downstream review operations produce heterogeneous configuration rankings, making broad knowledge scores insufficient for reliable operation selection. Evidence-state performance likewise depends on both choosing the correct action branch and requesting evidence that addresses the unresolved need.
- Experimental setup: 33 model configurations were evaluated using zero-shot direct-answer inference under archived configurations, with malformed outputs retained in the denominator.Judge-scored comparisons exclude the evaluator configuration and therefore use 32 configurations.
- RQ1: Operation-level evaluation: ρ = 0.42 mean pairwise Spearman correlation across downstream operations, with 37 of 55 operation pairs below 0.5.Institution matching and risk classification are nearly independent at ρ = −0.03, while the first five correlation components explain 85.0% of eigenvalue mass.
- RQ1: Operation-level evaluation: Domain Knowledge correlation ranges from 0.33 for institution matching to 0.89 for case classification.Gemini-3-Flash-preview and Kimi-K3 differ by 0.01 Domain Knowledge points but lead different downstream operations by 14.10 and 8.47 points.
- RQ2: Knowledge-based screening: 18.01 points of information-extraction score are forgone at k = 1 under Domain-Knowledge-based shortlisting.Regret also reaches 11.21 points for quantitative reasoning and 8.47 points for legal-outcome prediction; decision-view generation retains 6.70 points at k = 15.
- RQ3: Evidence-state control: 28.31 points is the largest ERA difference among configurations whose BAcc values differ by no more than 0.1 points.High end-to-end alignment requires both reliable entry into the Ask branch and request targeting aligned with the unresolved evidence need.
- RQ3: Evidence-state control: 46.26 points is the median Ask–Proceed recall gap across the 32 non-evaluator configurations.AskR exceeds ProceedR, and 24 configurations have gaps larger than 20 points, motivating separate branch-specific diagnostics alongside BAcc and ERA.
5. Conclusion
The paper frames financial LLM evaluation around professional review operations and evolving evidence states rather than source datasets or task formats. It constructs FinRiskAtlas through explicit contracts and complements it with trajectory-grounded FinRisk-Ask, while keeping static instances and active-review states separate.
- Conclusion: FinRiskAtlas evaluates whether configurations perform required review operations and control the evidence state underlying decisions.The static benchmark uses fixed evidence, while FinRisk-Ask evaluates evolving states requiring decisions about additional evidence.
- Conclusion: Static instances and active-review states are different evaluation units and are reported separately.The benchmark composition therefore does not combine fixed-evidence instances with trajectory-based review states.
- Conclusion: FinRiskAtlas construction defines capabilities and evaluation contracts before selecting sources, constructing instances, and conducting expert validation.Direct normalization, structured transformation, and source-conditioned construction obtain items without determining their evaluation family.
- Conclusion: Family sizes are intentionally unequal because construction prioritizes capability coverage rather than balanced sampling.Aggregation is performed only over explicitly compatible family sets.
A.3. Data Sources and Expert Involvement
FinRiskAtlas combines expert-defined financial review coverage with contract-based, reproducible evaluation procedures. Its construction emphasizes expert validation, consensus references, evidence-grounded scoring, and controlled response processing.
- Data sources and expert involvement: Six domain experts defined capabilities, validated family boundaries, reviewed candidate items, analyzed ambiguity, and performed quality control.
- Data sources and expert involvement: Every retained instance received independent review from at least two relevant financial or legal experts, with adjudication and removal of items lacking stable consensus.The released benchmark contains consensus references rather than unresolved annotations.
- Quality control: Family-level quality control verifies shared capabilities, visible information regimes, decision objects, reviewer artifacts, and scoring protocols.Families are split when decision objects differ and merged when their contracts are equivalent.
- Quality control: Instance-level checks enforce evidence support, ambiguity control, duplicate control, leakage control, parser validation, and provenance tracking.Knowledge evaluation is distinguished from record-grounded review, where correctness depends on evidence available in context.
- Evaluation protocol: FinRiskAtlas uses fixed prompts, output schemas, parsers, and scoring implementations, retaining raw responses, parsed predictions, and final metrics for verification.Experiments use zero-shot direct-answer inference, archived run manifests, and a denominator that includes empty, failed, or unparseable responses.
- Evaluation protocol: The eleven downstream operations instantiate explicit contracts covering visible information, decision objects, reviewer-artifact schemas, and scoring protocols.
B.4. Semantic Evaluation Protocol
The semantic protocol evaluates open-ended review outputs and FinRisk-Ask requests with fixed rubric-based instrumentation. Its interpretation is bounded by evaluator-specific effects and by offline replay of recorded workflow states rather than live interaction.
- Semantic evaluation: Three Applied Review operations use a fixed semantic evaluator for legal-judgment, decision-view, and disputed-issue generation.The evaluator receives the task input, reference information, candidate response, and task-specific rubric without candidate-model identity.
- Semantic evaluation: Evaluator configuration, prompts, and parsing procedures are fixed independently of candidate generation and applied uniformly across responses.The evaluator's own outputs are retained for transparency but excluded from best-value annotations and judge-scored comparisons.
- FinRisk-Ask request evaluation: FinRisk-Ask scores request alignment against expert-verified future evidence targets using three levels: 1 for direct, 0.5 for relevant but incomplete, and 0 for unrelated.The Ask-or-Proceed reference comes from the recorded workflow transition, not the semantic evaluator.
- Interpretation boundary: The protocol measures recorded transitions and trajectory-supported evidence needs offline, without executing generated requests or estimating their causal value after acquisition.It evaluates evidence-state control under a fixed historical decision boundary rather than optimizing a dialogue policy or long-horizon interaction.
- FinRisk-Ask request evaluation: FinRisk-Ask replays 680 pre-action states from 104 de-identified trajectories, withholding subsequent evidence, actions, and outcomes during inference.The released states include 583 recorded Ask states and 97 recorded Proceed states; ten Ask-labeled states were removed without admissible future targets.
- FinRisk-Ask request evaluation: Generated requests are evaluated against underlying evidence needs rather than exact historical wording, so semantically equivalent requests can receive credit.Targets are retained only when experts verify temporal unavailability, unresolved-need relevance, and meaningful acquisition objectives.
C.5. Metrics
FinRisk-Ask separates action agreement from evidence-request targeting through branch-specific and end-to-end metrics. These measures characterize performance under the benchmark’s recorded workflow and released evidence-target protocol, not complete professional review quality.
- Action metrics: FinRisk-Ask evaluates parsed model actions against recorded reviewer actions at each of 680 released states.
- Action metrics: The reference-state subsets contain 583 Ask states and 97 Proceed states, enabling branch-specific recall and balanced recorded-action agreement.Balanced agreement gives equal weight to the two reference branches.
- Request metrics: Evidence-Request Alignment measures end-to-end request alignment over all recorded Ask states, while Conditional Request Alignment evaluates requests conditional on entering the Ask branch.Conditional Request Alignment is undefined and reported as “–” when no states enter Ask.
- Request metrics: ERA = AskR × CRA expresses end-to-end evidence acquisition as the product of entering the Ask branch and conditionally aligning the request.
- Interpretation boundary: BAcc measures agreement with recorded workflow behavior, whereas ERA and CRA measure alignment with the released evidence-target set.Together, these metrics characterize evidence-state control under the benchmark protocol rather than complete professional review quality.
D.1. Robustness of Operation-Specific Capability Profiles
FinRiskAtlas shows that downstream operations preserve distinct capability profiles even when configurations have similar broad knowledge scores. Evidence-state control adds further non-identical dimensions, separating action reproduction from request alignment and end-to-end evidence acquisition.
- Operation-specific profiles: 25.5 percentile points is the median downstream profile gap among configuration pairs with Domain Knowledge difference ≤0.5.Across all 496 pairs, the median profile gap is 34.2 percentile points.
- Operation-specific profiles: 49.9% of total eigenvalue mass is captured by the leading component, while the first five capture 85.0%.Operation rankings share common structure but are not reducible to a single latent ordering.
- Knowledge-based selection: A Domain Knowledge shortlist of size one recovers 2 of 11 operation maxima, whereas size ten recovers 8 of 11.Full recovery requires a shortlist of size 23, so knowledge ranking is useful but incomplete for operation-specific selection.
- Evidence-state control: Legal-judgment generation correlates more strongly with CRA (ρ = 0.78) than BAcc (ρ = 0.14), and decision-view generation shows the same pattern (ρ = 0.79 versus ρ = 0.07).These are descriptive associations within the evaluated configuration pool and do not imply causality.
- Evidence-state control: Ask recall is generally higher than Proceed recall across 32 non-evaluator configurations.Action agreement can therefore differ from balanced agreement, while rankings can also shift between BAcc and ERA.
- Evidence-state control: FinRisk-Ask separates reproducing a recorded transition, recognizing when evidence is needed, and requesting evidence aligned with the unresolved decision need.The three request-alignment examples distinguish direct alignment, partial alignment, and missed evidence acquisition.
G. Limitations and Responsible Use
FinRiskAtlas is scoped to Chinese-language financial risk-control and compliance review under released operation contracts and offline replay protocols. It is intended for evaluation and comparison, not autonomous financial or legal decision-making.
- Scope: FinRiskAtlas measures performance on released Chinese-language review operations, not general capability across all institutions, jurisdictions, or workflows.Its design studies how evaluation units influence model selection in professional review settings.
- Evidence-state boundary: FinRisk-Ask evaluates alignment with verified evidence needs from completed trajectories rather than executing requests in live workflows.Future evidence is used to construct evaluation targets but withheld during inference.
- Evidence-state boundary: Ask-or-Proceed labels reflect one enterprise risk-control workflow and do not define a universal review policy.Organizations, policies, and operating conditions may differ from the released workflow behavior.
- Responsible use: The benchmark should not replace qualified reviewers or support autonomous financial or legal decisions.Real deployment also requires institution-specific policies, regulatory requirements, human oversight, and operational constraints beyond the benchmark contracts.