Source-linked AI summary
FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models
Dmitry Stanishevskii, Nini Kamkia, Alexey Khoroshilov, Dmitry Zmitrovich, Denis Kokosinskii, Zhirayr Hayrapetyan, Andrei Kalmykov
TL;DR
Financial LLM evaluation lacks a professionally grounded hierarchy spanning foundational to expert reasoning and applied financial domains. FINESSE-Bench addresses this with eight specialized benchmarks totaling 3,993 questions, and shows that strong public-benchmark performance does not always transfer while performance declines with increasing difficulty.
Problem
Existing financial benchmarks lack the simultaneous combination of explicit difficulty gradation, professionally recognizable expertise levels, and broad applied-domain coverage needed for reliable evaluation.
Method
FINESSE-Bench combines eight specialized datasets totaling 3,993 questions with certification-inspired difficulty levels, technical analysis, applied derivatives trading, and Russian-language olympiad problems.
Results
Strong public-benchmark performance does not always transfer to FINESSE-Bench, while models often show performance drops from CFA-like Level 1 to more difficult levels.
Takeaways & Limitations
More complete financial LLM evaluation requires testing transferability, difficulty hierarchy, domain specialization, and robustness across distinct subject groups.
Takeaways & Limitations
The substantial multiple-choice component may simplify tasks through option structure and elimination heuristics, so it does not guarantee equivalence to real professional reasoning.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly being applied to financial analysis, reporting, investment decision support, risk management, compliance, and professional training. However, robust evaluation of their domain competence in finance remains incomplete. Widely used open benchmarks such as FinQA, ConvFinQA, and TAT-QA have played an important role in advancing financial question answering and numerical reasoning, but they focus primarily on question answering over financial reports and do not provide an explicit hierarchy of professional difficulty. Broader resources, including FinanceBench, PIXIU, FinBen, and FLaME, expand the coverage of financial tasks, yet the problem of evaluating the transition from foundational knowledge to expert-level financial reasoning remains open. In this work, we present FINESSE-Bench, a suite of eight specialized benchmarks comprising 3,993 questions for hierarchical evaluation of financial competencies in LLMs. FINESSE-Bench combines exam-oriented datasets inspired by professional certifications (CFA-like Levels 1-3, CMT-like Level 2, and CFTe-like Level 1), applied trading task collections, and a Russian-language olympiad benchmark. This design enables evaluation of domain breadth, performance degradation as difficulty increases, the ability to solve computational tasks, and model behavior in specialized financial domains. We also describe a unified evaluation protocol covering multiple-choice questions, numerical answers, and short open-ended responses, together with an automated scoring scheme for freeform answers based on the LLM-as-judge paradigm. FINESSE-Bench is intended both as a complement to existing open financial benchmarks and as a tool for more substantive evaluation of professionally relevant financial competencies in large language models.
1 Introduction
FINESSE-Bench addresses gaps in financial-LLM evaluation by combining explicit professional difficulty levels with broader coverage of specialized financial domains. It provides eight datasets totaling 3,993 questions and a unified protocol for multiple-choice, numerical, and free-form tasks.
- Motivation: Existing benchmarks established financial question answering and numerical reasoning, then broadened coverage to public-company documents and financial NLP tasks.FinQA, ConvFinQA, and TAT-QA focus on financial documents and hybrid table-text data, while FinanceBench, PIXIU, FinBen, and FLaME expand task coverage.
- Motivation: Existing resources underrepresent technical analysis, derivatives trading, and scenario-based portfolio management, while generally lacking an explicit difficulty hierarchy.These limitations restrict evaluation of practically important domains and progression from foundational to expert-level competence.
- FINESSE-Bench: 3,993 questions span eight specialized datasets in FINESSE-Bench, which combines certification-inspired difficulty levels with domain specialization.The suite measures transitions from foundational to advanced and expert competence while incorporating specialized financial tasks.
- Evaluation protocol: A unified evaluation protocol covers multiple-choice, numerical, short free-form, and case-linked questions, using judge-model scoring when exact matching is insufficient.The protocol also combines fixed prompting templates and deterministic inference settings where applicable.
- FINESSE-Bench: Technical analysis, derivatives trading, and a Russian-language olympiad block broaden FINESSE-Bench beyond standard financial-report question answering.This design targets specialized financial competencies that are underrepresented in existing benchmarks.
2 Related Work
Prior financial benchmarks advanced financial question answering, numerical reasoning, and broad task coverage, but did not fully provide hierarchical evaluation of professional competence. FINESSE-Bench addresses this gap through complementary certification-oriented and practice-oriented benchmarks, alongside scalable but qualified model-as-judge evaluation.
- Existing financial QA benchmarks: FinQA, ConvFinQA, and TAT-QA established financial question answering and numerical reasoning over reports, conversational contexts, and combined tabular-textual sources.
- Broader financial resources: FinanceBench, PIXIU, FinBen, and FLaME broadened financial benchmark coverage across public-company documents, task types, datasets, and financial domains.
- Unresolved evaluation gap: Existing resources often lack the simultaneous combination of explicit difficulty gradation, professionally recognizable expertise levels, and broad applied-finance coverage.
- Automated evaluation: LLM-based judging offers scalable evaluation for open-ended responses, but remains a practical compromise rather than a complete substitute for expert annotation because of bias and prompt sensitivity.
- FINESSE-Bench design: FINESSE-Bench constructs complementary benchmarks spanning progression from foundational preparation to expert-level tasks and practice-oriented domains underrepresented in existing open resources.
3 FINESSE-Bench: Design Principles
FINESSE-Bench treats financial competence as multidimensional and evaluates it across task types, professional difficulty levels, and specialized domains. Its design combines realistic, diverse, multilingual, and verifiable tasks to support broader and more reproducible assessment.
- Design principle: Financial competence is multidimensional, so the benchmark evaluates average accuracy alongside error patterns across task types and difficulty levels.A model may perform well on financial reporting yet worse on portfolio construction, technical analysis, or derivatives trading.
- Realism: Realism grounds questions in financial statement interpretation, company valuation, risk management, investment decisions, technical indicators, and option-strategy calculations.These skills are intended to reflect real financial practice and professional training.
- Difficulty hierarchy: Explicit difficulty gradation spans foundational, intermediate, and expert levels, enabling evaluation of transfer from basic knowledge to complex scenario-based and multi-step tasks.The hierarchy is inspired by multi-level professional certifications.
- Domain breadth: Domain breadth extends beyond financial reporting and financial NLP to technical analysis, derivatives trading, and Russian-language olympiad problems.The suite complements existing strengths in question answering over financial reporting and finance-related NLP.
- Format diversity and verifiability: Format diversity and multilinguality combine multiple-choice, numerical, free-form, linked case-based, and Russian-language tasks, while verifiable answers support automated scoring and reproducible error analysis.The varied formats make narrow optimization for one evaluation format more difficult and better approximate educational and professional scenarios.
4 Dataset Description … 4.3 CFA-like Level 3
FINESSE-Bench comprises eight specialized datasets totaling 3,993 questions and organizes financial competence hierarchically from foundational literacy to expert-level synthesis. The CFA-like levels progress from basic disciplines through complex application scenarios to strategic portfolio, wealth, risk, and ethical analysis.
- 4 Dataset Description: 3,993 questions comprise FINESSE-Bench’s eight specialized datasets, which collectively support hierarchical evaluation of financial competencies.The datasets are described according to their purpose and role within the benchmark suite’s overall hierarchy.
- 4.1 CFA-like Level 1: 1,069 questions assess foundational finance disciplines and basic financial literacy through predominantly multiple-choice items.Covered areas include ethics, quantitative methods, economics, financial reporting, corporate finance, and investment fundamentals.
- 4.1 CFA-like Level 1: CFA-like Level 1 measures applied competence across ethics, quantitative methods, economics, reporting, corporate finance, and investment fundamentals.Its stated purpose is to measure basic financial literacy and applied competence.
- 4.2 CFA-like Level 2: 293 questions in linked item sets evaluate complex application scenarios based on shared cases and interrelated questions.The level emphasizes multi-step calculations, advanced financial statement analysis, valuation, fixed income, and derivatives.
- 4.2 CFA-like Level 2: CFA-like Level 2 emphasizes multi-step calculations, advanced financial statement analysis, valuation, fixed income, and derivatives.Several interrelated questions rely on a common case within each linked item set.
- 4.3 CFA-like Level 3: 318 questions target expert competence in portfolio management, private wealth planning, risk management, and complex ethical case analysis.The level requires strategic thinking and synthesis across multiple areas of finance.
4.4 CMT-like Level 2 · 4.5 CFTe-like Level 1 · 4.6 VLigaBench-ru
The section introduces three specialized datasets spanning applied technical analysis, foundational technical-analysis concepts, and Russian-language olympiad-style economic and financial reasoning. Together, they assess market-signal skills, technical-analysis foundations, and calculation-intensive multilingual problem solving.
- 4.4 CMT-like Level 2: CMT-like Level 2 contains 251 questions focused on applied technical analysis and market statistics.Topics include technical-analysis theory, chart patterns, indicators, volume, open interest, trading-system testing, and risk management.
- 4.4 CMT-like Level 2: The CMT-like Level 2 dataset diagnoses skills related to working with market signals.Its coverage connects technical-analysis knowledge with trading-system testing and risk management.
- 4.6 VLigaBench-ru: VLigaBench-ru is a Russian-language olympiad-style dataset containing 324 problems in microeconomics, macroeconomics, financial mathematics, and game theory.It broadens the benchmark beyond conventional financial question answering.
- 4.6 VLigaBench-ru: Unlike typical financial QA tasks, VLigaBench-ru emphasizes reasoning, calculation, and careful handling of Russian-language problem statements.Its olympiad-style design tests problem solving across economics, financial mathematics, and game theory.
4.7 Trading_TA · 4.8 Trading_derivatives · 4.9 Dataset Statistics
The benchmark includes applied technical-analysis and specialized derivatives trading tasks, alongside dataset statistics and format definitions. Together, these components emphasize practice-oriented, calculation-intensive financial competence across varied task types.
- 4.7 Trading_TA: Trading_TA contains 413 applied technical-analysis tasks covering pattern recognition, trading strategies, trade management, backtesting, and multi-timeframe analysis.The block assesses practice-oriented competence in trading contexts.
- 4.7 Trading_TA: Its tasks address momentum and mean-reversion strategies, entry and exit rules, and stop management.
- 4.8 Trading_derivatives: Trading_derivatives consists of 544 tasks on options, synthetic positions, put-call parity, arbitrage, Greeks, hedging, pricing, and futures strategies.It is one of FINESSE-Bench’s most specialized and calculation-intensive components.
- 4.8 Trading_derivatives: The derivatives block combines conceptual and computational topics, including arbitrage, Greeks, hedging, pricing, and futures strategies.
- 4.9 Dataset Statistics: Table 1 reports the core statistics of the FINESSE-Bench datasets.
- 4.9 Dataset Statistics: Dataset formats distinguish MCQ multiple-choice questions, NAQ numerical-answer questions, and SAQ short-answer questions.
4.10 Data Collection and Curation
FINESSE-Bench questions were collected from public educational, exam-style, training, internet, and olympiad sources, then normalized, structurally aligned, and manually checked. Incomplete provenance, possible training-data overlap, and uneven topic/source coverage motivate cautious non-commercial release with a removal mechanism.
- Collection and curation: Questions came from publicly available internet sources, educational materials, training problems, exam-style explanations, and olympiad problems, followed by format normalization and manual correctness checks.Answer structures were also aligned during curation.
- Provenance and release: Incomplete provenance limits full traceability, prompting a cautious distribution policy, non-commercial licensing, and repository-based removal of disputed materials.The removal mechanism addresses disputes over individual materials.
- Known limitations: Potential limitations include overlap between some questions and models’ training data and biases from uneven topic and source coverage.These limitations are discussed further in Section 8.
5 Evaluation Protocol
FINESSE-Bench evaluates diverse model types under a unified, mostly deterministic prompting protocol with fixed configurations and medium reasoning effort where supported. It uses GPT-5.2 as an LLM judge for binary correctness, reports accuracy, and supports per-dataset and grouped evaluation across exam-like, public-benchmark, and trading/TA directions.
- Inference protocol: A unified fixed prompt template without few-shot demonstrations is used for each task type, with temperature 0.0 wherever applicable.Model configurations are fixed before evaluation and remain unchanged during the benchmark run.
- Inference protocol: Models supporting controllable reasoning are evaluated with medium reasoning effort during scoring.
- Scoring: GPT-5.2 serves as the judge model, receiving the question, reference answer, and tested response before assigning a binary correctness score.The pipeline follows the LLM-as-judge paradigm and extends arena-hard-auto for FINESSE-Bench.
- Metrics and aggregation: Accuracy is the primary metric, with results reported both for individual datasets and aggregated across exam-like, public-benchmark, and trading/TA directions.Exam-like aggregation covers CFA-like Levels 1–3, CMT-like Level 2, and VLigaBench-ru; trading/TA covers Trading_derivatives, Trading_TA, and CFTe-like Level 1.
- Uncertainty reporting: 95% confidence intervals are computed by bootstrap for each model and benchmark, using stratified bootstrap with dataset-size-proportional weights for aggregated groups.The main text presents point estimates, while confidence intervals and standard errors are provided in the accompanying repository.
6 Main Results
FINESSE-Bench reveals distinctions among models that classical public benchmarks often compress: exam-oriented tasks produce clearer capability stratification and difficulty-related degradation, while applied trading and technical-analysis tasks expose specialized weaknesses and non-transferability.
- Classical public benchmarks: Classical open financial benchmarks show modest gaps among leading models, limiting separation despite their continued value as a common reference.Several strong models achieve similar accuracy values, so these benchmarks provide limited differentiation among contemporary systems.
- Exam-oriented benchmarks: Exam-oriented FINESSE-Bench results produce more pronounced model differences and clearer stratification than classical open benchmarks.This pattern suggests stronger discrimination of professionally oriented financial knowledge and reasoning.
- Exam-oriented benchmarks: Performance often drops from CFA-like Level 1 to more difficult exam-like levels, indicating that the suite measures quality retention as task complexity increases.The observed heterogeneity across difficulty levels supports FINESSE-Bench’s hierarchical evaluation objective.
- Applied trading and technical analysis: Applied technical-analysis and derivatives-trading datasets can induce rankings that differ from public benchmarks, showing that standard performance does not automatically transfer to specialized settings.These domains add diagnostic information about financial competence that traditional benchmark formats may weakly reflect.
- Aggregated benchmark groups: Aggregated results show that strong classical-benchmark performance does not always transfer to exam-oriented and applied professional tasks.The reported gaps between public-benchmark performance and exam-like or trading/TA performance support FINESSE-Bench’s motivation as a complementary evaluation suite.
7 Analysis of Results
FINESSE-Bench reveals performance gaps that classical public financial benchmarks often miss, especially on exam-like and trading/technical-analysis tasks. Its hierarchical and model-scaling analyses show that advanced difficulty and professional task structure provide stronger differentiation among models.
- Transfer gaps: Most models perform worse on FINESSE-Bench than on classical open financial benchmarks, showing that public-benchmark strength does not guarantee professional-task competence.The degradation generally appears simultaneously for exam-like and trading/TA groups.
- Transfer gaps: Qwen3.5-Plus-02-15 has gaps of 0.0080 on exam-like tasks and 0.0381 on trading/TA, while Fino1-8B, Fin-R1-7B, and Fin-o1-8B show gaps of 0.27–0.36.GLM-5 has a slightly negative exam-like gap of −0.0003, whereas Claude Sonnet 4.6, GPT-5.2, and Kimi K2.5 also maintain comparatively small gaps.
- Difficulty hierarchy: CFA-like Level 3 is harder than Level 1 for every model, although Level 1-to-Level 2 degradation is not strictly monotonic.Claude Sonnet 4.6, Kimi K2.5, GPT-5.2, MiniMax M2.5, and GLM-5 can have negative ∆L1→L2 values, but still degrade on Level 3 relative to Level 1.
- Difficulty hierarchy: DeepSeek-V3.2 declines sequentially from Level 1 through Level 3, with a total extreme-level gap of 0.2175.GPT-5.4 and Llama 4 Maverick also show sequential declines across the three CFA-like levels.
- Model discrimination and scaling: Qwen3 public-benchmark scores cluster at 0.8506–0.8699, while exam-like scores rise from 0.6891 to 0.8381 and trading/TA scores from 0.6744 to 0.8027 across model sizes.Within the Qwen3 family, FINESSE-Bench improvements from 8B through 235B are almost monotonic, exposing scaling differences that public benchmarks compress.
8 Practical Implications and Limitations
FINESSE-Bench supports practical model development through reproducible MCQ evaluation, stage-wise diagnosis, and comparison of balanced performance across financial subdomains. Its interpretation is limited by MCQ artifacts, incomplete item provenance and possible contamination, and gaps in coverage of several financial industry areas.
- Practical advantages: MCQ tasks support rapid evaluation loops, reproducible checkpoint comparison, and, when needed, logit-based evaluation without complex post-processing.The format is widely used in knowledge and reasoning evaluation for language models and is particularly useful for fine-tuning practice and intermediate validation.
- Practical advantages: Stage-wise evaluation can reveal substantial performance drops on exam-like or trading/TA groups despite competitive results on classical open financial benchmarks.The authors argue that relying on one popular benchmark may overestimate true domain competence, whereas grouped difficulty levels support intermediate model selection and troubleshooting.
- Practical advantages: Balanced models can be distinguished from local leaders because consistently strong performance across financial subdomains may matter more than a peak score on one benchmark.Some systems perform very strongly on one task group but behave less stably on others, making profile consistency relevant to practical model selection.
- Evaluation limitations: MCQ structure may simplify model performance through option structure and elimination heuristics, so it is an engineering advantage rather than an ideal substitute for professional reasoning.FINESSE-Bench therefore combines MCQ with numerical-answer, short-answer, applied-scenario, and linked-case tasks to balance reproducibility, development usability, and realistic reasoning coverage.
- Data limitations: Incomplete provenance limits full traceability of individual items and necessitates a cautious release policy, leading to release for non-commercial research use.Questions came from public internet sources, educational materials, training tasks, and publicly available preparation formats, but individual provenance was not documented completely.
- Data limitations: Partial contamination cannot be completely ruled out, and FINESSE-Bench is presented as a diagnostic step rather than a benchmark fully protected against contamination.Some questions may have appeared in certain models’ training data; the authors emphasize benchmark fidelity and design quality alongside coverage breadth.
- Domain coverage: FINESSE-Bench does not cover the entire financial industry, lacking full benchmark blocks for regulatory compliance, banking risk modeling, insurance analytics, corporate liquidity management, and financial law.Its existing scope spans financial reporting, corporate finance, technical analysis, derivatives trading, and Russian-language olympiad problems.
9 Conclusion
FINESSE-Bench is a hierarchical suite of eight specialized financial benchmarks designed to evaluate domain breadth, professional applications, difficulty progression, transferability, and robustness. The authors present it as a useful first-step instrument while identifying expansion, multilingual, open-ended, and validation directions for future work.
- Contributions: FINESSE-Bench combines difficulty hierarchy, broad domain coverage, and professionally oriented applied scenarios across eight specialized financial benchmarks.It is intended to complement benchmarks focused on financial-report question answering and broader but less structured financial task collections.
- Findings: Strong performance on public financial benchmarks does not consistently transfer to FINESSE-Bench task groups, particularly exam-like and trading/technical-analysis tasks.The analysis identifies this transfer gap as new diagnostic information beyond classical open financial benchmarks.
- Findings: The CFA-like hierarchy captures performance degradation as difficulty increases, measuring foundational financial literacy alongside robustness on advanced tasks.This structure supports evaluation of how model behavior changes across professional difficulty levels.
- Findings: Many FINESSE-Bench subsets concentrate questions in an informative difficulty zone where contemporary models begin to diverge in performance.The saturation and discriminative-power analysis supports the suite’s ability to distinguish model behavior meaningfully.
- Implications: A complete assessment of financial competence requires testing transferability, difficulty hierarchy, domain specialization, and robustness across distinct subject groups, not only popular open benchmarks.The conclusion retains that existing resources remain important and useful while positioning FINESSE-Bench as a broader assessment suite.
- Future work: Future work should expand financial subdomains, strengthen multilingual coverage, increase open-ended tasks, and further validate judge-model-based scoring.The current benchmark is presented as a first step that already supports model comparison and intermediate fine-tuning-pipeline validation.