Source-linked AI summary

FinForge: Semi-Synthetic Financial Benchmark Generation

Glenn Matlin, Akhil Theerthala, Anant Gupta, Anirudh JM, Rayan Castilla, Yi Mei Ng, Sudheer Chava

arXiv:2601.06747v2cs.AI

TL;DR

Financial LM evaluation lacks open, domain-specific benchmarks that capture both conceptual and quantitative reasoning. FinForge combines expert-guided curation with controlled LM synthesis to generate dynamic finance evaluations; FinForge-5k reveals weaknesses in conceptual reasoning and limitations in automated validation.

  • Problem

    Specialized financial LM evaluation lacks open benchmarks with sufficient domain depth for conceptual and quantitative reasoning.

  • Method

    FinForge combines authoritative-source curation with controlled LM-based question synthesis and validation to create semi-synthetic finance benchmarks.

  • Results

    FinForge-5k evaluations identify high-impact weaknesses in financial reasoning, particularly conceptual failures rather than simple arithmetic errors.

  • Takeaways & Limitations

    FinForge provides a framework for systematic, reproducible, and evolving evaluation of language models in finance and other expert-driven disciplines.

  • Takeaways & Limitations

    Reliance on Gemini 2.5 Flash for generation and evaluation can produce ambiguous questions and distort assessment of complex reasoning.

Abstract

from arXiv · show

Evaluating Language Models (LMs) in specialized, high-stakes domains such as finance remains a significant challenge due to the scarcity of open, high-quality, and domain-specific datasets. Existing general-purpose benchmarks provide broad coverage but lack the depth and domain fidelity needed to assess LMs' capabilities for real-world financial reasoning, which requires both conceptual understanding and quantitative rigor. To address this gap, we introduce FinForge, a scalable, semi-synthetic pipeline for constructing finance-specific evaluation benchmarks through a hybrid of expert-guided data curation and controlled LM-based synthesis. FinForge combines manual and programmatic corpus construction from authoritative financial sources with structured question generation and validation using Gemini 2.5 Flash. To demonstrate the pipeline's efficacy, we produce FinForge-5k, a snapshot benchmark comprising over 5,000 human-validated question-answer pairs across 11 finance subdomains, derived from a curated corpus of 100,000 verified documents totaling 143M tokens. Evaluation of state-of-the-art open-source and closed-source models on FinForge-5k reveals significant differences in financial reasoning, with leading models achieving accuracy levels near 80%. These findings underscore the framework's utility for diagnosing current model limitations and guiding future improvements in financial domain competence. All code and data are available at https://github.com/gtfintechlab/FinForge.

Introduction

FinForge addresses the difficulty of evaluating financial reasoning by combining expert-guided curation with controlled LM synthesis to create dynamic, domain-specific benchmarks. Its pipeline grounds challenging, self-contained questions in verified financial content and supports continual updating.

  • The Reasoning Engine applies domain-specific schemas to infer implicit financial mechanisms before synthesizing complex multiple-choice questions and answers.
  • Finance evaluation requires both broad domain knowledge and quantitative problem-solving in a rapidly evolving, highly regulated field.
  • FinForge targets contamination and limited domain fidelity by generating finance questions from real-world content on demand.
  • The framework combines human-guided curation of authoritative sources with a multi-stage LM workflow for question synthesis and validation.
  • FinForge-5k contains 5,000 expert-level finance question–answer pairs derived from more than 100,000 verified articles totaling 143M tokens across 11 subdomains.

Related Works

Prior financial benchmarks are narrow in coverage, recency, scale, or reasoning scope, while broader static benchmarks risk memorization. FinForge extends dynamic benchmark generation by grounding larger, more difficult question sets in verified real financial texts.

  • Existing financial QA datasets provide limited scope, focus, and recency despite contributing to specialized subtasks.
  • Static test suites risk memorization because language models train on extensive public internet data, motivating refreshable or semi-synthetic benchmarks.
  • General-purpose suites such as MMLU, Big-Bench, and ARC assess broad knowledge or reasoning but are not specialized financial evaluations.
  • Earlier work includes LM-generated variants, real-time reading-comprehension sets, and multi-agent benchmark creation with human involvement.
  • FinQA and TAT-QA emphasize structured numerical reasoning over financial reports, whereas FinanceBench contains only 150 questions.
  • FinForge complements prior benchmarks by increasing scale and difficulty while grounding each question in current, verified financial text with source evidence.

Methodology

FinForge builds a verified finance corpus through hybrid expert and automated curation, then transforms documents into structured, self-contained questions using a controlled five-stage workflow. Rubric-based filtering and independent expert review refine the generated benchmark.

  • Data Curation: The corpus is designed to include both quantitative financial calculations and qualitative economic, market, and institutional knowledge.
  • Data Curation: A taxonomy decomposes finance into 11 subdomains, balancing conceptual completeness with scalable coverage.
  • Data Curation: Human source selection and automated domain filtering restrict the corpus to authoritative, verifiable content and exclude informal sources.
  • Question Generation: The five-stage generation workflow extracts salient information, creates answer plans, produces grounded questions, labels them, and validates them.
  • Question Generation: Generated questions embed necessary context and include plausible distractors, concise explanations, and domain-specific requirements.
  • Validation and Filtering: Validation checks relevance, self-sufficiency, logical consistency, clarity, and complexity, rejecting any question that fails one criterion.
  • Validation and Filtering: Expert feedback independently validates iterations and refines generation and filtering logic toward higher-quality, more challenging questions.

Results

FinForge-5k was constructed from a large verified corpus and evaluated models across a consistent 5,000-question multiple-choice benchmark. Results show differences by model availability and scale, while expert review exposes weaknesses in automated validation.

  • Benchmark Construction: More than 100,000 verified documents produced a 143M-token corpus across 11 financial subdomains during a seven-day workflow.
  • Benchmark Construction: 10,000 generated question–answer pairs were filtered to create the 5,000-pair FinForge-5k benchmark.
  • Evaluation: Accuracy was measured in a consistent multiple-choice setup across the 5,000 benchmark questions.
  • Evaluation: GPT-4o and Claude Sonnet 4 achieved 73.4% and 72.6% accuracy, while same-generation open-source models demonstrated superior performance.
  • Evaluation: Qwen3-Next-80B was within 1% of DeepSeek, GPT-4o, and Sonnet models, indicating that sheer model scale was not the sole determinant of performance.
  • Expert Evaluation: 70% of 500 expert-reviewed samples were clear, accurate, and relevant, while the remaining 30% mainly involved ambiguity or missing assumptions.
  • Expert Evaluation: The automated LM judge approved 100% of the identical sample, producing a 30-point discrepancy with expert review.
  • Evaluation: Gemini 2.5 Flash and Gemma 27B scored 79.3% and 74.0%, but both were omitted from the primary benchmark because of circular contamination risk.

Discussion

FinForge-5k exposes systematic weaknesses in financial reasoning, especially on complex personal and corporate finance tasks and quantitative questions. The findings identify conceptual misinterpretation as a central failure mode and motivate targeted improvements in financial-sector models.

  • FinForge-5k reveals systematic weaknesses in current model capabilities through proportional error analysis.The analysis examines reasoning deficiencies beyond aggregate accuracy scores.
  • Personal Finance & Wealth Management and Corporate Finance & Valuation are notably challenging relative to other benchmark topics.These tasks require simultaneously satisfying constraints involving tax liabilities, liquidity needs, and risk.
  • Models perform comparatively better on Markets and Derivatives and Portfolio Management than on complex multi-constraint finance tasks.
  • Quantitative questions are the most difficult across subjects, followed by counterfactual inquiries.For this snapshot, multi-hop questions were frequently derived from distinct sections of single documents, representing a specific reasoning-difficulty tier.
  • Incorrect quantitative answers reflect either conceptual failures or arithmetic errors.Conceptual failures involve incorrect methodologies, assumptions, or logic, whereas arithmetic failures preserve the correct steps but miscalculate.
  • These weaknesses indicate that average users cannot currently rely on the models for important financial decisions.The gap is especially evident in Personal Finance & Wealth Management and has been overlooked by conventional retrieval-focused benchmarks.
  • FinForge identifies high-impact weaknesses in contemporary LMs and provides a framework for generating nuanced, domain-specific benchmarks.The framework is presented as a basis for guiding future advancements and focused improvements to financial-sector models.

Conclusion

FinForge combines expert oversight, verified financial sources, controlled LM synthesis, and multi-stage validation to generate scalable, domain-grounded benchmarks. FinForge-5k reveals that state-of-the-art models particularly struggle with conceptual financial reasoning, establishing a foundation for evolving evaluation in finance and other expert disciplines.

  • Scalable, domain-grounded benchmark generation is achievable through expert oversight and controlled LM synthesis.
  • FinForge integrates verified financial sources with structured question generation and multi-stage validation.This combination is used to produce datasets reflecting the reasoning depth and quantitative rigor demanded in economic analysis.
  • FinForge-5k evaluations reveal specific, high-impact weaknesses in state-of-the-art models, particularly failures of conceptual reasoning rather than simple arithmetic.
  • FinForge offers a generalizable methodology for transparent, extendable evaluation pipelines in specialized domains.Future work described in the paper includes broader subfields, temporal and dynamic data, and continual-learning benchmarks.
  • FinForge lays groundwork for systematic, reproducible, and evolving evaluation of LMs in finance and other expert-driven disciplines.

Limitations

The study’s main limitations concern its reliance on Gemini 2.5 Flash for generation and evaluation, evaluator transparency, and contamination risks involving models from the generator’s family.

  • The study relies on Gemini 2.5 Flash for both question generation and evaluation.Manual verification found that generated questions often lacked contextual assumptions needed for definitive answers, creating ambiguity that can compromise assessment reliability.
  • Gemini 2.5 Flash’s speed-oriented design may lack the sophistication needed to evaluate complex reasoning models accurately.The resulting capabilities mismatch may distort performance metrics.
  • The boolean validator is a black box whose lack of transparency obstructs analysis of failed questions and improvement of the generation pipeline.
  • Data contamination and circular dependency remain concerns when evaluating models from the same family as the generator.The study omits Gemini and Gemma from primary results to mitigate this issue, while future work is expected to develop more robust isolation strategies.
Loading 2601.06747v2…