Source-linked AI summary
A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges
Yuqi Nie, Yaxuan Kong, Xiaowen Dong, John M. Mulvey, H. Vincent Poor, Qingsong Wen, Stefan Zohren
TL;DR
Financial applications need methods that can handle complex, rapidly evolving information while addressing finance-specific practical and ethical concerns. This survey synthesizes financial LLM models, applications, resources, and challenges across major task areas. It concludes that LLMs show broad potential for improving the efficiency and accuracy of financial processes, but deployment remains constrained by interpretability, cost, computational complexity, and ethical requirements.
Problem
Existing surveys do not consistently provide a deep treatment of finance-specific practical challenges or the broader implications for financial decision-making and industry practice.
Method
The survey organizes financial LLM research across models, applications, datasets, code, benchmarks, and future challenges, covering linguistic, sentiment, time-series, reasoning, and agent-based tasks.
Results
LLMs demonstrate potential to improve the efficiency and accuracy of financial processes through contextual understanding and real-time analysis across diverse financial tasks.
Takeaways & Limitations
The survey provides researchers and practitioners with a holistic view of financial LLM applications, resources, and practical insights for further development.
Takeaways & Limitations
Financial LLM deployment remains challenged by limited interpretability, high inference costs, computational complexity, and ethical and regulatory concerns.
Abstract
from arXiv · showhide
Recent advances in large language models (LLMs) have unlocked novel opportunities for machine learning applications in the financial domain. These models have demonstrated remarkable capabilities in understanding context, processing vast amounts of data, and generating human-preferred contents. In this survey, we explore the application of LLMs on various financial tasks, focusing on their potential to transform traditional practices and drive innovation. We provide a discussion of the progress and advantages of LLMs in financial contexts, analyzing their advanced technologies as well as prospective capabilities in contextual understanding, transfer learning flexibility, complex emotion detection, etc. We then highlight this survey for categorizing the existing literature into key application areas, including linguistic tasks, sentiment analysis, financial time series, financial reasoning, agent-based modeling, and other applications. For each application area, we delve into specific methodologies, such as textual analysis, knowledge-based analysis, forecasting, data augmentation, planning, decision support, and simulations. Furthermore, a comprehensive collection of datasets, model assets, and useful codes associated with mainstream applications are presented as resources for the researchers and practitioners. Finally, we outline the challenges and opportunities for future research, particularly emphasizing a number of distinctive aspects in this field. We hope our work can help facilitate the adoption and further development of LLMs in the financial sector.
1 INTRODUCTION
LLMs offer broad capabilities for financial applications, while this survey addresses gaps in existing reviews through a holistic, practice-oriented examination of models, applications, resources, and challenges.
- Motivation: Financial LLMs support contextual understanding, scalable analysis, and complex emotion detection across financial tasks.These capabilities are presented as valuable for understanding market sentiment and supporting informed decisions.
- Challenges: The survey highlights lookahead bias, legal concerns, data pollution, signal decay, inference cost, uncertainty, interpretability, safety, and privacy as deployment challenges.Addressing these issues is framed as essential for ethical and effective financial use.
- Research gap: Existing surveys often underdevelop finance-specific practical challenges and broader implications for financial decision-making and industry practice.The survey positions itself as addressing these limitations through a broader review.
- Contributions: Its contributions include a holistic perspective connecting academic research with practical implementation and relevance for researchers and practitioners.The survey emphasizes real-world financial applications and practical insights.
- Scope: The survey examines LLM applications across linguistic tasks, sentiment analysis, financial time series, reasoning, and agent-based modeling.It also organizes the field around models, applications, data, code, benchmarks, and future challenges.
2 MODELS
Financial LLMs span general-purpose foundations and specialized variants, with adaptation strategies ranging from zero-shot use to efficient fine-tuning for domain-specific tasks.
- Collections of Models: Financially specialized LLMs are built from foundational families including GPT, BERT, T5, ELECTRA, BLOOM, and Llama.The survey catalogs models by their foundational architectures and financial adaptations.
- Collections of Models: Llama-based financial variants target investment, sentiment, and other finance tasks, while FinGPT emphasizes accessible and transparent open-source development.InvestLM is described as offering investment recommendations comparable to cutting-edge commercial models.
- Collections of Models: Financial models also include variants based on Mistral, Qwen, Baichuan, InternLM, and LLaVA for specialized and multimodal applications.FinVIS-GPT is identified as a multimodal model for financial chart analysis.
- Zero-shot vs Fine-tuning: Zero-shot learning relies on pre-existing knowledge for unseen tasks, whereas fine-tuning adjusts a pretrained model on task- or domain-specific data.Fine-tuning is favored when domain accuracy, customization, privacy, or adaptation to real-time changes matters.
- Zero-shot vs Fine-tuning: Instruction tuning, retrieval augmentation, LoRA, quantization, and smaller models are presented as approaches for more efficient financial adaptation.Instruction tuning can improve target-task performance and zero-shot or few-shot capabilities.
- Zero-shot vs Fine-tuning: Zero-shot or few-shot learning is preferred when labeled data is limited, rapid deployment is important, or modularity and interpretability are prioritized.The choice between adaptation strategies depends on the application’s data and operational requirements.
3.1 Linguistic Tasks
LLMs advance financial linguistic tasks by handling long-range context, diverse document structures, entity extraction, classification, and knowledge-graph querying. The surveyed methods address domain-specific data, multimodal layouts, computational cost, and limited labeled examples.
- Textual Work: Transformer self-attention helps LLMs manage long-term dependencies and contextual information that challenged earlier recurrent models in lengthy financial documents.Earlier RNN/LSTM models struggled with long sequences, complex expressions, large datasets, and unstructured financial data.
- Textual Work: Financial-document summarization and extraction research addresses long inputs through document segmentation, specialized models, structural chunking, multilingual adaptation, and domain-specific fine-tuning.Retrieval-augmented generation chunking can use structural elements rather than paragraphs, while other work targets multilingual and cryptocurrency settings.
- Textual Work: Layout-aware models address PDF understanding challenges by preserving spatial information from images, charts, and tables instead of relying only on plain-text conversion.DocLLM uses bounding-box information to model spatial arrangement, whereas text conversion can lose layout-dependent information.
- Textual Work: LLM-based NER extracts companies, financial terms, stock symbols, indicators, and monetary values for downstream financial applications.UniversalNER uses targeted distillation and mission-focused instruction tuning to reduce computational burden while maintaining NER accuracy.
- Knowledge-based Analysis: Knowledge-graph methods support financial information retrieval and classification, including NL-to-GQL generation and knowledge-enriched text representations.A self-instruction pipeline generates NL-GQL pairs without labeled data and fine-tunes LLMs with LoRA; KGEB achieves 91.98% accuracy and a 90.89% F1 score on a NEEQ dataset.
3.2 Sentiment Analysis
Financial sentiment analysis has evolved from interpretable lexicons and conventional machine learning toward LLM-based analysis across diverse market data and economic documents. The surveyed applications use contextual models to capture complex sentiment and support forecasting, portfolio decisions, and policy analysis, while retaining domain and data challenges.
- Overview: Sentiment analysis quantitatively examines opinions, subjectivity, emotions, and market-related sentiment in financial text.Its financial importance follows from the use of market sentiment in forecasting and actions.
- Pre-LLM Sentiment Analysis: Lexicon-based methods are simple and interpretable but can miss context-dependent sentiment, sarcasm, irony, and complex linguistic constructs.They include General Inquirer, LIWC, SO-CAL, and Loughran–McDonald word lists and remain used for financial news and social media.
- Pre-LLM Sentiment Analysis: Machine-learning sentiment methods capture patterns beyond lexicons and support market-movement prediction from financial news and social media, but require large datasets and remain domain-limited.Supervised approaches include SVM, Naive Bayes, KNN, Random Forests, and MLPs; unsupervised methods do not require labeled data.
- Data-driven Applications: LLM sentiment applications span social media, news, earnings calls, corporate communications, regulatory filings, and monetary-policy documents.These sources provide real-time public sentiment, objective events, management perspectives, corporate signals, legal or accounting information, and policy tone.
- Data-driven Applications: Financial LLMs improve contextual sentiment analysis, with reported gains in stock prediction, portfolio optimization, adversarial robustness, FOMC analysis, and complex financial sentences.FinBERT outperforms traditional techniques for negative FOMC sentiment, while Sentiment Focus improves analysis of sentences containing contradictory sentiments.
- Challenges: Applying LLMs to financial sentiment remains challenging because economic indicators, policy documents, and other domain-specific materials require further methodological enhancement and comprehensive interpretation.The survey also identifies broader deployment concerns including lookahead bias, signal decay, inference cost, uncertainty, interpretability, safety, and privacy.
3.3 Financial Time Series Analysis
LLMs are being applied to financial time series as both supportive tools and direct analytical models, leveraging sequential-data processing, Transformer architectures, and multimodal capabilities. Research spans forecasting, anomaly detection, classification, augmentation, and imputation, while financial forecasting evidence remains mixed and several areas require further refinement.
- LLMs for Time Series: LLMs can augment time series models by generating textual features and descriptive statistics that incorporate information beyond the original data.
- LLMs for Time Series: LLMs are also used directly for time series analysis because they process sequential data, rely on effective Transformer architectures, and support multimodal inference.
- LLMs for Time Series: LLM-based methods have demonstrated applications in forecasting, anomaly detection, classification, and imputation, while time series foundation models seek to capture complex temporal dependencies.
- Forecasting: Financial forecasting studies report benefits from diverse data, instruction-based fine-tuning, chain-of-thought reasoning, GNN integration, and multimodal inputs including text, images, audio, news, and market series.
- Forecasting: Evidence on financial forecasting is mixed: ChatGPT underperforms in zero-shot multimodal prediction, whereas GPT-4 outperforms traditional models and earlier LLMs when using news headlines.
- Challenges and Opportunities: The survey identifies explainability, comprehensive news understanding, and multimodal integration as areas requiring continued research to realize LLMs’ potential in financial time series.
- Anomaly Detection: In anomaly detection, an LLM-based multi-agent framework combining statistical methods with AI analytics improved efficiency, precision, and automation for S&P 500 data.
- Other Applications: Other financial time series applications include trend or regime classification, synthetic-data augmentation, and missing-value imputation.
3.4 Financial Reasoning
LLMs support financial reasoning across planning, investment recommendations, market analysis, auditing, compliance, and reconciliation. The surveyed applications show personalization and efficiency gains, while also exposing concerns about bias, fairness, reliability, and regulation.
- Planning: LLMs can analyze financial situations, goals, and risk tolerance to provide personalized advice and streamline financial planning.They can also support client communication, education, and responses to common financial concerns.
- Investment Recommendations: ChatGPT achieved a monthly three-factor alpha of up to 3% when generating portfolios, particularly from policy-related news.The study also found that parameter choices such as temperature affect recommendation creativity and accuracy.
- Financial Statement Analysis: GPT-4 Turbo outperformed human analysts in predicting earnings changes and matched specialized machine-learning models using anonymized financial statements.The study designed inputs to reduce identity inference and look-ahead bias.
- Challenges: Financial reasoning applications remain constrained by implicit model biases, inconsistent advice across users and languages, and the need for regulation and human oversight.These concerns affect fairness, reliability, consumer protection, and informed decision-making.
- Auditing and Compliance: LLM-based systems can improve financial auditing, regulatory compliance, fraud detection, and risk management by interpreting records and identifying inconsistencies.Applications include regulatory text matching, compliance verification, and automated inspection.
3.5 Agent-based Modeling
LLM-enhanced agent-based models support adaptive trading, economic simulation, and multi-agent financial analysis by combining agent interactions with large-scale data processing. The section also highlights computational, behavioral-design, and validation challenges.
- Overview: LLM-integrated agent-based models let agents interpret financial news, reports, and social media, producing more realistic and adaptive simulations.This integration supports trading, investment, and economic modeling contexts.
- Trading and Investments: StockAgent uses LLM-powered agents to simulate investor behaviors and assess macroeconomic events, policy changes, and financial reports.GPT-3.5 Turbo agents showed more diverse and independent trading styles, whereas Gemini agents were more homogeneous and trend-following.
- Trading and Investments: FinAgent combines textual, numerical, and visual data with diversified memory retrieval and tool augmentation for quantitative and high-frequency trading.The system supports stocks and cryptocurrency in dynamic trading environments.
- Trading and Investments: FINMEM and QuantAgent use continuous or iterative learning to refine strategies, extract financial signals, and identify trading opportunities.QuantAgent’s inner loop refines responses through a knowledge base, while its outer loop uses real-world testing and knowledge enhancement.
- Simulating Markets and Economic Activities: LLM-based economic simulations model human-like decisions, behavioral factors, and competitive adaptation, but may require substantial computation, expertise, calibration, and sensitivity analysis.EconAgent models macroeconomic activities, Homo Silicus agents combine rational and emotional factors, and CompeteAI studies strategy adaptation under competition.
- Multi-agent Systems: Multi-agent frameworks support financial sentiment analysis and information verification by assigning specialized agents to errors or document-matching tasks.HAD targets sarcasm, aspect mismatches, and temporal expressions in financial sentiment analysis.
3.6 Other Applications
Cloud computing is presented as a way to scale LLM-based financial automation, customer interaction, and decision support. Serverless architectures may improve efficiency and cost-effectiveness across financial sectors.
- Cloud Computing: Cloud computing can improve the scalability, efficiency, and cost-effectiveness of LLM applications across financial sectors.The survey discusses automation, customer interactions, and decision support in banking.
- Cloud Computing: Serverless cloud architectures may provide scalable platforms for deploying LLM-based financial services and operations.The survey links this deployment model with potential innovation, operational efficiency, and customer-centricity.
4 DATASETS, CODE AND BENCHMARK
The survey catalogs datasets, benchmarks, and open resources for financial NLP, reasoning, and prediction. It emphasizes standardized evaluation, multilingual coverage, and broader assessment beyond NLP-only tasks.
- Datasets: Financial datasets support training and evaluation for sentiment analysis, question answering, relation extraction, and numerical reasoning.Examples include Financial PhraseBank, FiQA, and FinQA.
- Benchmarks and Code: Benchmarks provide standardized comparisons intended to improve reliability, accuracy, transparency, reproducibility, and continuous model improvement.Sharing code and methodologies is presented as a way to promote collaboration and practical implementation.
- Benchmarks and Code: FLUE evaluates financial sentiment analysis, news classification, named entity recognition, structure boundary detection, and question answering.Its associated financial models use pretraining objectives incorporating financial keywords, phrases, and span boundaries.
- Benchmarks and Code: PIXIU combines the FinMA financial LLM, a large-scale multitask instruction dataset, and the FLARE evaluation benchmark as publicly available resources.FLARE assesses diverse financial capabilities more broadly than benchmarks focused solely on NLP.
- Language Impact: Language-specific and bilingual benchmarks examine financial model performance across Chinese, Japanese, Spanish, and English contexts.These resources cover tasks including sentiment analysis, entity recognition, relation extraction, summarization, auditing, and financial examinations.
5 CHALLENGES AND OPPORTUNITIES
The survey identifies deployment challenges spanning data quality, temporal validity, computational cost, evaluation, interpretability, ethics, accountability, and privacy. It also describes opportunities for adaptive models, improved benchmarks, optimization, and stronger safeguards.
- Data Issues: Financial LLMs face data pollution from inaccurate information and from models learning from LLM-generated content.These conditions may degrade performance, reduce data relevance, and produce rigid learning.
- Modeling and Benchmarking Issues: Signal decay and changing market conditions can reduce the effectiveness of LLM-generated trading strategies and complicate their evaluation.The survey calls for adaptive models and benchmarks aligned with markets shaped by widespread LLM use.
- Modeling Issues: High computational demands create a trade-off between inference speed, cost, and performance when processing large financial datasets.Model optimization and hardware advances are identified as opportunities to reduce costs and improve speed.
- Modeling Issues: Lookahead bias can make LLM financial backtests overly optimistic by incorporating future information during training.Avoiding future information is necessary for more reliable time-series forecasting and dynamic financial modeling.
- Ethical and Legal Issues: Hallucinated or factually incorrect financial outputs raise legal and reliability concerns because financial reporting is subject to strict standards.The survey connects generated-content errors with potentially severe organizational consequences.
- Modeling Issues: Uncertainty estimation is important because LLM responses are sampled and repeated queries can produce different outputs, including errors.Single-sample outputs can therefore be misleading for financial decision-making or forecasting.
- Ethical Issues: Interpretability, benign alignment, legal responsibility, safety, and privacy remain central requirements for trustworthy financial deployment.The survey discusses faithfulness and informativeness for rationales, ethical outputs, accountability frameworks, and protection of sensitive financial data.
6 CONCLUSION
The survey synthesizes how LLMs can enhance diverse financial tasks while identifying barriers to responsible deployment. It concludes that addressing these limitations and continued research are needed to advance financial-sector integration.
- The survey covers LLM applications in linguistic tasks, sentiment analysis, financial time series analysis, financial reasoning, and agent-based modeling.
- LLMs show potential to improve the efficiency and accuracy of financial processes through contextual understanding and real-time analysis.
- Data privacy, interpretability, and computational costs remain challenges for responsible and effective deployment in finance.
- The survey aims to stimulate further research and innovation by summarizing the current state, advantages, and limitations of LLMs in financial applications.
- Continued exploration may support LLM integration into finance for more strategic investment and efficient decision-making.