Source-linked AI summary

FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning

Liang Hu, Jianpeng Jiao, Jiashuo Liu, Yanle Ren, Zhoufutu Wen, Kaiyuan Zhang, Xuanliang Zhang, Xiang Gao, Tianci He, Fei Hu, Yali Liao, Zaiyuan Wang, Chenghao Yang, Qianyu Yang, Mingren Yin, Zhiyuan Zeng, Ge Zhang, Xinyi Zhang, Xiying Zhao, Zhenwei Zhu, Hongseok Namkoong, Wenhao Huang, Yuwen Tang

arXiv:2509.13160v1cs.LGcs.AI

TL;DR

FinSearchComp addresses the absence of open, end-to-end benchmarks for realistic financial search and reasoning, where agents must handle time-sensitive, domain-specific analyst tasks. It introduces a 635-question benchmark across three task families, curated with professional expertise and open evaluation tooling. Grok 4 (web) leads the Global subset, DouBao (web) leads Greater China, and web search plus financial plugins improve performance, although systems remain fragile and below human performance in important settings.

  • Problem

    Existing open financial datasets do not evaluate end-to-end agent search on realistic, complicated, time-sensitive analyst tasks.

  • Method

    The paper builds FinSearchComp with 635 expert-curated questions across three analyst-style task families, two market subsets, deterministic answers, and an open evaluation harness.

  • Results

    Grok 4 (web) leads the Global subset, DouBao (web) leads Greater China, and equipping agents with web search and financial plugins improves FinSearchComp performance.

  • Takeaways & Limitations

    FinSearchComp offers a high-difficulty testbed for measuring realistic financial search and reasoning while exposing weaknesses in search depth, freshness awareness, and evidence integration.

Abstract

from arXiv · show

Search has emerged as core infrastructure for LLM-based agents and is widely viewed as critical on the path toward more general intelligence. Finance is a particularly demanding proving ground: analysts routinely conduct complex, multi-step searches over time-sensitive, domain-specific data, making it ideal for assessing both search proficiency and knowledge-grounded reasoning. Yet no existing open financial datasets evaluate data searching capability of end-to-end agents, largely because constructing realistic, complicated tasks requires deep financial expertise and time-sensitive data is hard to evaluate. We present FinSearchComp, the first fully open-source agent benchmark for realistic, open-domain financial search and reasoning. FinSearchComp comprises three tasks -- Time-Sensitive Data Fetching, Simple Historical Lookup, and Complex Historical Investigation -- closely reproduce real-world financial analyst workflows. To ensure difficulty and reliability, we engage 70 professional financial experts for annotation and implement a rigorous multi-stage quality-assurance pipeline. The benchmark includes 635 questions spanning global and Greater China markets, and we evaluate 21 models (products) on it. Grok 4 (web) tops the global subset, approaching expert-level accuracy. DouBao (web) leads on the Greater China subset. Experimental analyses show that equipping agents with web search and financial plugins substantially improves results on FinSearchComp, and the country origin of models and tools impact performance significantly.By aligning with realistic analyst tasks and providing end-to-end evaluation, FinSearchComp offers a professional, high-difficulty testbed for complex financial search and reasoning.

1 Introduction

FinSearchComp addresses the lack of realistic end-to-end financial search benchmarks by modeling time-sensitive retrieval, historical lookup, and multi-source analyst workflows. Its evaluation shows strong but uneven progress: web search and financial plugins help, while leading systems still trail experts and remain vulnerable to freshness and reconciliation failures.

  • Results: On the Global subset, Grok 4 (web) scores 68.9%, beats GPT-5-Thinking (web) by 5.0 pp, and trails human experts by 6.1 pp; DouBao (web) leads Greater China, where all models remain over 34 pp below humans.These results come from the overall comparison in Figure 1.
  • Benchmark design: FinSearchComp evaluates realistic financial search through 635 expert-curated questions spanning three analyst-style task families and global and Greater China markets.The tasks cover Time-Sensitive Data Fetching, Simple Historical Lookup, and Complex Historical Investigation, with multi-stage verification and rubric-based scoring.
  • Contribution: FinSearchComp provides an open-source dataset with deterministic gold answers and an open evaluation harness for end-to-end financial search assessment.The benchmark is intended to measure proximity to expert-level competence in realistic financial search.
  • Analysis: Across 21 evaluated models, web search and financial plugins improve performance, but recurring errors include shallow search, stale evidence, incorrect extraction, and cross-unit or calendar misalignment.The analysis identifies specialized-tool use and evidence freshness as concrete improvement targets.
  • Motivation: The benchmark targets capabilities that general browsing datasets omit, including temporal validity, unit alignment, reporting-calendar alignment, and provenance reconciliation across sources.These requirements reflect finance workflows combining real-time signals, historical disclosures, and unstructured context.

2 FinSearchComp

FinSearchComp is designed around realistic financial analyst workflows, spanning increasing levels of freshness, historical fidelity, and multi-period synthesis. Its evaluation combines rubric-guided judgment with tolerance-aware scoring and human validation.

  • Construction and Quality Control: FinSearchComp applies source selection, ambiguity mitigation, multi-expert verification, and rubric-based quality control to improve question reliability.The construction process uses separate pipelines for the three tasks while applying uniform quality control across them.
  • Design Principles: The benchmark is intended to reflect professional financial work across markets, languages, reporting conventions, and regulatory settings.Its Global and Greater China subsets use English and Chinese questions, mirrored task templates, and balanced entity coverage by sector and size.
  • Task Design: T1 retrieves rapidly changing data, T2 retrieves fixed historical facts, and T3 aggregates or synthesizes information across long periods.The tasks respectively stress freshness and calendar handling, reporting conventions and unit fidelity, and long-horizon retrieval with multi-step reasoning.
  • Task Design: FinSearchComp organizes analyst-style search into three tasks that progress from time-sensitive fetching to simple lookup and complex historical investigation.The task families cover freshness management, point-in-time fidelity, and multi-period synthesis, with difficulty scaling from T1 to T3.

3 Experiments

FinSearchComp evaluates 22 mainstream products and a human baseline across Global and Greater China subsets, revealing task-, region-, and model-dependent performance differences.

  • 3.1 Overall Results: Grok-4 (web) and GPT-5-Thinking form the leading global tier, while Chinese products lead on the Greater China subset but remain below human experts.The reported rankings differ across subsets, with Grok-4 (web) securing the global top score and DouBao (web) leading Greater China.
  • 3.2 Results Across Different Tasks: Performance declines monotonically from T1 to T2 to T3, reflecting increasingly demanding analyst workflows involving multi-hop retrieval, temporal reasoning, entity resolution, and evidence reconciliation.The hardest tasks additionally require finance-specific interpretation of filings, disclosures, accounting terminology, and corporate actions.
  • 3.2 Results Across Different Tasks: US models lead on the Global set and Chinese models on Greater China, a pattern attributed to regional corpus coverage, linguistic conventions, and alignment or recency effects.These factors are described as improving home-field performance without implying data leakage.
  • 3.2 Results Across Different Tasks: Grok-4 (web) and GPT-5-Thinking outperform other systems increasingly as task difficulty rises, with the largest margin on T3.The paper links these gains to multi-step reasoning, timeline alignment, and entity disambiguation; Grok-4 (web) also tops T3 in Greater China.

4 Case Study

The case studies show that search and financial-plugin access materially affect performance, while task structure and model reasoning capability shape the remaining differences.

  • 4.1 Search Capability: Search-enabled models gain 40.8, 29.0, and 8.1 points on T1, T2, and T3 respectively, whereas models without search score 0 on T1.Search remains beneficial on historical tasks, but the gains are smaller as tasks require deeper reasoning and synthesis.
  • 4.2 Financial Plugins: Financial plugins improve DeepSeek R1 performance on YuanBao, including a 31.9 pp T1 improvement over the comparison setting.Plugins provide direct access to current and historical financial data, but YuanBao-R1 remains suboptimal, showing that intrinsic model capability also matters.
  • 4.3 Model Origin: US models generally perform better on global assets and Chinese models on Chinese assets, although most models exceed a 100% Global-to-Chinese score ratio on T3.Doubao and Kimi k2 have the highest ratios among Chinese models, suggesting comparatively balanced regional performance.
  • 4.1 Search Capability: Grok 4 (web) ranks highest on T2 by using diverse reliable sources, whereas parametric-memory answers and news-only retrieval often miss granular official-filing details.The cited examples include historical financial disclosures such as income-statement information.
  • 4.1 Search Capability: No product exceeds 30 on T3 except Grok 4 (web) and GPT-5-Thinking (web), because complex investigation requires structured retrieval through APIs or SQL.Successful attempts are largely limited to queries requiring fewer than five data points.
  • 4.4 Reasoning Capability: Reasoning capacity reduces T1 performance by an average of 7.0 points, while its effect on T2 and T3 is negligible.The paper attributes the T1 decline to the task’s low complexity and possible overthinking by reasoning models.

5 Related Work

Prior financial benchmarks measure domain knowledge and reasoning but commonly provide the relevant data, while agentic benchmarks broaden tool interaction without fully covering realistic open-domain financial search.

  • Financial Benchmarks: FinQA and ConvFinQA evaluate numerical reasoning over annual reports by composing multi-step programs from textual and tabular evidence.Other suites expand task and language coverage across classification, extraction, generation, and related finance capabilities.
  • Financial Benchmarks: Existing financial datasets substantially reduce the search challenge by supplying relevant financial data, and Finance Agent Benchmark is limited to static historical search.This leaves room for memorization and does not fully test time-sensitive open-domain retrieval.
  • Agentic Benchmarks: Agentic benchmarks evaluate goal-directed interaction with external tools, while BrowseComp variants test persistent web navigation and creative search strategies.The cited finance and general benchmarks emphasize planning, API use, long-horizon reasoning, or web navigation.

6 Conclusion

FinSearchComp fills a gap in end-to-end evaluation of realistic financial data search with an expert-curated, open benchmark spanning demanding tool-orchestrated tasks.

  • 6 Conclusion: FinSearchComp provides 635 expert-curated questions across three demanding tasks requiring SQL, APIs, and web search to obtain verifiable financial answers.The benchmark is designed for realistic, context-free financial data search by LLM-based agents.
  • 6 Conclusion: State-of-the-art agents significantly underperform humans, often because of insufficient search depth and outdated information.The benchmark is released to support development of more robust and reliable financial agents.

7 Contributions

The paper lists its core contributors and identifies corresponding authors and ByteDance Seed affiliations.

  • Liang Hu and Zhoufutu Wen are marked as core contributors.
  • Liang Hu and Zhoufutu Wen are identified with dagger markers as corresponding authors.
  • Contributors without explicit affiliations are from ByteDance Seed, while Xuanliang Zhang and Yanle Ren were interns there.

8 Xpert Platform

Xpert is an expert-level platform for specialized training data and evaluation, focused on complex real-world tasks and distributed benchmark artifacts.

  • 8.1 What is Xpert Platform: Xpert is an expert-level data service platform for specialized training data and evaluation solutions.
  • The platform emphasizes AI evaluation on expert-level complex real-world tasks rather than mainstream exam-oriented assessments.
  • FinSearchComp is distributed as JSONL files with questions, answers, tools, and traces, alongside a sandboxed trace-replay harness and documentation.

A.2 Illustration of Inconsistent Calculation Methods for the Same Metric

The appendix illustrates that financial metrics can vary across institutions and data providers because calculation conventions are not uniform.

  • Forward- and backward-adjusted stock prices can differ substantially across databases, so the benchmark queries non-adjusted prices.
  • PE (TTM) is ambiguous because institutions may use different definitions of earnings.
  • Dual-listed market capitalization depends on the specified formula, such as summing listing values or multiplying total shares.
  • Futures continuous contracts vary with main-contract switching and construction algorithms across institutions.
  • Cryptocurrency prices differ across exchanges, creating source-dependent values.

A.3 Guide for Mitigating Ambiguity

FinSearchComp mitigates ambiguity by specifying standards, units, currencies, precision, and calculation rules in questions and by accommodating legitimate answer variation.

  • Table 4 consolidates the annotation guidance for mitigating ambiguity in FinSearchComp.
  • A.3 Guide for Mitigating Ambiguity: Questions should specify accounting standards, such as GAAP or Non-GAAP, so answers use comparable ground-truth conventions.
  • Questions should state currencies, units, and required numerical precision to make answers directly comparable.
  • Industry and dual-listing questions must specify classification standards and market-capitalization formulas rather than relying on ambiguous wording.
  • Futures answers may accept both decimal and hexadecimal-style quote formats when they represent the same value.

B Detailed Scores on FinSearchComp

This section reports detailed FinSearchComp scores for various models in Table 5.

  • Table 5 presents detailed scores for various models evaluated on FinSearchComp.
  • The table is intended to support model-level performance inspection beyond the headline comparisons.
  • These detailed results complement the benchmark’s broader evaluation of model performance.

C Prompt

The judge scores financial answers against trusted reference information, checking required content and applying task-specific numerical accuracy rules. Examples show that complete content can still receive zero for an inconsistent value, while values within an allowed range receive one.

  • Scoring rules: The judge treats the Real-time Authentic Information as the correct ground truth and scores required content against its accuracy criteria.Missing required information yields 0; extra information is not penalized.
  • Accuracy criteria: Rounding-only criteria accept values differing through decimal-place rounding, while high-low criteria accept any value within the reference range.The examples also define absolute-error and expanded-range criteria.
  • Accuracy criteria: Dates are not scored directly, so mismatches involving latest-versus-specific dates or time zones do not reduce credit when numerical values are correct.The judge focuses on numerical values such as prices and indices.
  • Examples: In the NVDA example, correct opening, high, and low prices cannot offset an incorrect change percentage, producing a final score of 0.The USD/CNY and COMEX gold examples score 1 because their answers fall within the specified acceptable ranges.
  • Scoring rules: A response receives 1 only when it includes all required key points and every supplied value satisfies the applicable accuracy requirement.Partial answers or missing content receive 0.

D Case

The cases illustrate how search-tool use and source verification affect financial question answering. Examples include a failed historical lookup without search, successful plugin-assisted retrieval, and cross-verified stock-split research.

  • Case studies: The Walmart case combines current quote data with forecasts, options commentary, financial-database information, and other web sources.The listed sources include Traders Union, Benzinga, Tencent Financial Industry Database, and market data pages.
  • Case studies: A current Walmart lookup reports a $96.08 close, 0.03% daily gain, $96.32 high, $95.60 low, and $1.241 billion turnover on August 27, 2025.The response also reports a $96.05 previous close and a 6.33% five-day decline.
  • Case studies: A web-based product fails a simple historical retrieval question because it does not use a search tool.The case asks for Apple’s 2021 investing cash flow in millions of dollars.
  • Case studies: A web-based product successfully answers a complex historical investigation by using Nasdaq as a reliable source and reverse-searching to verify retrieved data.The Apple stock-split case searches historical closing and opening prices across Nasdaq and additional sources before reporting the result.
  • Case studies: The Apple stock-split response reports a 4-for-1 split effective August 31, 2020, with opening price $127.58 versus prior close $499.23, a -$371.65 change.These figures are presented as the final response to the investigation question.
Loading 2509.13160v1…