Source-linked AI summary
FinBen: A Holistic Financial Benchmark for Large Language Models
Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, Yijing Xu, Haoqiang Kang, Ziyan Kuang, Chenhan Yuan, Kailai Yang, Zheheng Luo, Tianlin Zhang, Zhiwei Liu, Guojun Xiong, Zhiyang Deng, Yuechen Jiang, Zhiyuan Yao, Haohang Li, Yangyang Yu, Gang Hu, Jiajia Huang, Xiao-Yang Liu, Alejandro Lopez-Lira, Benyou Wang, Yanzhao Lai, Hao Wang, Min Peng, Sophia Ananiadou, Jimin Huang
TL;DR
Financial LLM evaluation lacks broad benchmarks suited to the complexity of finance. FinBen addresses this gap with 36 datasets across 24 tasks and seven aspects, and its evaluation finds strong performance on information extraction and textual analysis but weaker performance on complex tasks such as generation and forecasting. The benchmark also supports stock-trading and agent-based evaluation, with shared-task solutions outperforming GPT-4.
Problem
Financial LLM capabilities and limitations remain insufficiently understood because existing benchmarks are limited in breadth and financial tasks are complex.
Method
FinBen is an open-source benchmark with 36 datasets spanning 24 financial tasks across seven aspects, including stock trading, agent/RAG evaluation, and three novel datasets.
Results
LLMs perform strongly in information extraction and textual analysis but underperform on complex reasoning tasks such as text generation and forecasting; GPT-4 and Gemini lead in different task areas.
Takeaways & Limitations
FinBen provided the basis for a 12-team financial LLM shared task whose proposed methods outperformed GPT-4, supporting its use for advancing financial LLM research.
Takeaways & Limitations
FinBen is limited by restricted dataset size, evaluation of only LLaMA 70B among larger models, and reliance on American-market data and English texts.
Abstract
from arXiv · showhide
LLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive evaluation benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluation benchmark, including 36 datasets spanning 24 financial tasks, covering seven critical aspects: information extraction (IE), textual analysis, question answering (QA), text generation, risk management, forecasting, and decision-making. FinBen offers several key innovations: a broader range of tasks and datasets, the first evaluation of stock trading, novel agent and Retrieval-Augmented Generation (RAG) evaluation, and three novel open-source evaluation datasets for text summarization, question answering, and stock trading. Our evaluation of 15 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals several key findings: While LLMs excel in IE and textual analysis, they struggle with advanced reasoning and complex tasks like text generation and forecasting. GPT-4 excels in IE and stock trading, while Gemini is better at text generation and forecasting. Instruction-tuned LLMs improve textual analysis but offer limited benefits for complex tasks such as QA. FinBen has been used to host the first financial LLMs shared task at the FinNLP-AgentScen workshop during IJCAI-2024, attracting 12 teams. Their novel solutions outperformed GPT-4, showcasing FinBen's potential to drive innovation in financial LLMs. All datasets, results, and codes are released for the research community: https://github.com/The-FinAI/PIXIU.
1 Introduction
FinBen addresses the limited breadth of existing financial LLM benchmarks with a comprehensive evaluation framework spanning diverse tasks and datasets. Evaluations reveal strong performance on information extraction and textual analysis but persistent weaknesses on complex reasoning, forecasting, and generation.
- Motivation: Existing benchmarks cover few tasks and focus mainly on financial NLP, leaving forecasting, risk management, and decision-making insufficiently evaluated.This narrow coverage limits comprehensive assessment across complex financial applications.
- Benchmark scope: FinBen comprises 36 datasets spanning 24 financial tasks across seven financial aspects.The aspects are information extraction, textual analysis, question answering, text generation, risk management, forecasting, and decision-making.
- Innovations: FinBen adds broader task coverage, stock-trading evaluation, agent- and RAG-based evaluation, and three novel open-source datasets.The new datasets target text summarization, question answering, and stock trading.
- Community impact: 12 teams participated in FinNLP-AgentScen at IJCAI-2024, and their proposed methods outperformed GPT-4.The shared task illustrates FinBen’s use in developing financial LLM solutions.
- Findings: Across 15 evaluated LLMs, models excelled in information extraction and textual analysis but underperformed on complex reasoning tasks including text generation and forecasting.GPT-4 was strongest in several analysis and trading tasks, while Gemini was stronger in text generation and forecasting.
2 FinBen
FinBen organizes financial evaluation through a seven-domain taxonomy and combines reformulated existing datasets, benchmark data, and newly created datasets. Its tasks span extraction, analysis, question answering, generation, forecasting, risk management, and trading-oriented decision-making.
- Taxonomy: FinBen uses a seven-domain taxonomy covering information extraction, textual analysis, question answering, text generation, risk management, forecasting, and decision-making.The taxonomy provides a structured framework for evaluating financial LLM capabilities.
- Data sources: FinBen draws data from open-source studies, existing evaluation benchmarks, and novel datasets introduced in this paper.Experts reformulate some existing datasets into instruction-response pairs for zero-shot evaluation.
- Novel datasets: The novel datasets include EDTSum for financial-news summarization, FinTrade for stock trading, and Regulations for long-form financial-regulation question answering.FinTrade combines historical prices, news, and sentiment data for 10 stocks over one year.
- Additional tasks: The benchmark also covers extraction, textual classification, summarization, stock-movement forecasting, fraud detection, financial distress, and strategic trading decisions.Strategic decision-making is evaluated with the FinMem financial LLM agent on FinTrade.
- Question answering: Question answering includes multi-step numerical reasoning, multi-turn questions, and long-form responses involving financial reports, tables, and regulations.FinQA, TATQA, ConvFinQA, and Regulations are evaluated with F1, EMAcc, ROUGE, and BERTScore as applicable.
3 Evaluation
The evaluation measures zero-shot and few-shot performance across 15 representative general-purpose and financial LLMs. Experiments used standardized generation and batching settings on 16 NVIDIA A100 GPUs, requiring substantial computational resources and cost.
- Models: 15 representative general and financial LLMs were evaluated on FinBen.The models include ChatGPT, GPT-4, Gemini Pro, and LLaMA2 chat variants.
- Experimental settings: All experiments used a maximum of 1024 generation tokens and batch size 20,000 on 16 NVIDIA A100 80G GPUs.The experiments took approximately 600 hours and cost about $51,000 including GPT-4 API costs.
4 Results
FinBen results show strong performance on selected information extraction, textual analysis, and question answering tasks, but persistent weaknesses in complex reasoning, forecasting, risk management, and trading instruction adherence. GPT-4 leads stock trading, while model performance varies substantially by task, language, and model scale.
- Information Extraction: GPT-4 leads named entity recognition, while Gemini only slightly improves on complex extraction and remains below expectations.GPT-4 performs best on NER, FINER-ORD, and FinRED; Gemini performs only slightly better on causal detection and numerical understanding.
- Textual Analysis: Instruction-tuned FinMA 7B leads financial sentiment analysis but generalizes poorly across diverse textual analysis tasks.FinMA 7B performs best on FPB, FiQA-SA, and Headlines, yet can underperform general-domain models elsewhere.
- Question Answering: GPT-4 and Gemini lead across question answering datasets, while FinMA 7B remains constrained by model size and numeric reasoning bottlenecks.GPT-4 also handles the regulations dataset requiring both financial and legal knowledge.
- Text Generation: Gemini leads abstractive summarization, whereas all models struggle with extractive summarization and CFGPT sft-7B-Full declines relative to InternLM 7B.Among open-source models, LLaMA2 70B stands out in text summarization.
- Forecasting: All LLMs lag traditional methods in forecasting, with GPT-4 and Gemini performing only slightly better than random guessing.The results identify forecasting as a major weakness requiring complex reasoning.
- Risk Management: MCC score 0 occurs when low instruction-following models collapse imbalanced risk-management cases into a single class.The issue appears in credit scoring, fraud detection, and financial-distress identification.
- Decision Making: GPT-4 achieves the highest Sharpe Ratio over 1 and the lowest Maximum Drawdown in stock trading, outperforming reinforcement-learning baselines and Buy & Hold.Gemini ranks second with lower risk and volatility, while LLaMA-70B has the lowest profit among open-source models.
- Decision Making: Models below 70 billion parameters struggle to follow trading instructions consistently across transactions.The paper attributes this limitation to constrained comprehension, extraction capabilities, and context windows.
5 Conclusion
FinBen is a comprehensive financial benchmark spanning diverse tasks and datasets, with evaluations of 15 LLMs revealing both capabilities and limitations. The benchmark is intended to expand its language and task coverage over time.
- FinBen evaluates LLMs across 36 datasets and 24 financial tasks organized into seven critical aspects.The aspects include information extraction, textual analysis, question answering, text generation, risk management, forecasting, and decision-making.
- Evaluation of 15 LLMs, including GPT-4, ChatGPT, and Gemini, reveals key advantages and limitations across financial tasks.
- FinBen aims to incorporate additional languages and expand its financial task range to enhance applicability and impact.
- The benchmark’s applicability may be constrained by dataset size, model-size coverage, and reliance on American-market data and English texts.The paper also notes responsible-use concerns, including potential financial misinformation or unethical market influence.
Checklist
The checklist records that the paper reports contributions, limitations, ethical considerations, reproducibility materials, and evaluation details. It also references supplementary benchmark results and prompt documentation.
- The paper reports its contributions and discusses limitations, including through references to the limitations section.
- The authors state that they addressed potential negative societal impacts and ethical considerations in the Ethical Statement.
- The paper provides code, data, and reproduction instructions through a URL, while training details are not applicable because the benchmark only evaluates models.
- The checklist states that error bars and compute-resource information are reported in the experimental sections.
- The supplementary material includes additional LLM performance results and detailed dataset instructions and prompts.
D.2 Financial Evaluation Benchmarks
Financial evaluation benchmarks have expanded beyond early NLP tasks, but the supplied passages emphasize the need to cover realistic financial problems and provide prompt overviews for quantification datasets.
- FLUE covers five financial NLP tasks, including sentiment analysis, headline classification, named entity recognition, structure boundary detection, and question answering.
- Table 7 provides an overview of prompts for quantification-task datasets.
- Later benchmark work added event detection and realistic financial tasks, including stock trading-related evaluation efforts.
- Table 8 documents example prompts for remaining tasks, including financial sentiment analysis and stock movement prediction.
E Trading Accumulative Returns
The trading section reports detailed performance comparisons across stocks and LLM strategies, while the FinLLM Challenge evaluates classification, summarization, and single-stock trading. The supplied figure captions identify accumulated-return plots for eight stocks.
- Trading performance: Table 9 compares overall trading performance for large LLMs across various stocks, restricting results to models with at least 70B parameters.Smaller-context models are excluded because they reportedly struggle with instructions and producing a static holding strategy.
- FinLLM Challenge: The FinLLM Challenge targets financial classification, financial text summarization, and single-stock trading through three corresponding datasets.
- FinLLM Challenge: The challenge’s structured assessment is described as covering diverse financial scenarios holistically and effectively.
F.1 Tasks and Datasets
The FinLLM Challenge evaluates financial classification, text summarization, and single-stock trading, with task-specific datasets, metrics, and leakage testing.
- Task 1: Financial Classification: Financial classification categorizes texts as premises or claims using 7.75k training examples and 969 test examples.
- Task 1: Financial Classification: F1 and Accuracy evaluate classification, with F1 serving as the final ranking metric.
- Task 2: Financial Text Summarization: Financial text summarization abstracts financial news into concise summaries using 8k training examples and 2k test examples.
- Task 2: Financial Text Summarization: ROUGE-1, ROUGE-2, ROUGE-L, and BERTScore evaluate summary relevance, with ROUGE-1 used for final ranking.
- Task 3: Single Stock Trading: Single-stock trading evaluates sophisticated trading decisions on 291 examples, assessing profitability, risk management, and decision-making.
- The Data Leakage Test compares training- and test-set perplexity differences; larger differences indicate lower likelihood of model cheating.
F.3 Participants and Automatic Evaluation
The FinLLM Challenge involved 35 registered teams and evaluated submitted systems across classification, summarization, and stock-trading tasks. Reported results show strong team performance, including outcomes surpassing comparison models on summarization but not GPT-4 on trading.
- Participants: 35 teams registered for the FinLLM Challenge, with 11 teams submitting system description papers.
- Task 1: Financial Classification: Top-three classification teams achieved F1 scores comparable to LLaMA3-8B, below GPT-4 and LLaMA2-70B, while outperforming FinMA and other models.
- Task 2: Financial Text Summarization: In financial summarization, the three teams surpassed all other models on ROUGE-1.
- Task 3: Single Stock Trading: Top-1 Wealth Guide achieved the strongest Sharpe Ratio among teams, outperforming other large models but not GPT-4.
- Non-LLM Comparisons: The benchmark section also presents task-oriented non-LLM baselines for stock movement prediction and financial NLP tasks.
G.2 Financial NLP tasks
This section reports financial NLP performance for BERT-based models and stock-movement prediction performance for non-LLM models using tabulated comparison metrics.
- Financial NLP task performance for BERT-based models is presented in Table 15, with the best performance shown in bold.
- Stock-movement prediction performance for non-LLM models is presented in Table 14 using Accuracy and Matthews correlation coefficient.
Limitations
FinBen’s effectiveness and applicability are constrained by dataset size, model-size coverage, market and language scope, and responsible-use concerns. The accompanying material also carries academic-use restrictions and liability disclaimers.
- Dataset Size Limitations: Restricted open-source financial datasets may limit models’ financial understanding and generalization across contexts.
- Model Size Limitations: Computational constraints limited evaluation to LLaMA 70B, potentially excluding larger or differently architected models.
- Generalizability: Trading and forecasting tasks use American market data and English texts, potentially limiting applicability to global financial markets.
- Potential Negative Impacts: The paper warns that financial LLM advances could be misused to propagate misinformation or exert unethical market influence.
- Use Restrictions: The material is designated for academic and educational use and does not provide financial, legal, or investment counsel.
- Disclaimers: The authors provide no warranty regarding completeness or suitability and disclaim liability for consequences arising from reliance on the material.