Source-linked AI summary
Large Language Models in Law: A Survey
Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, Philip S. Yu
TL;DR
Legal LLM research remains limited while judicial systems face growing demands and unresolved questions about AI’s role in legal decision-making. This paper systematically surveys legal LLM technologies, research, applications, challenges, and recommendations. It concludes that legal LLMs can support judicial work, but their development requires attention to data, algorithms, judicial practice, and human oversight.
Problem
Legal LLM applications are still developing, with unresolved questions about their characteristics, shortcomings, future role, and relationship to judicial decision-making.
Method
The paper conducts a systematic review of legal LLM research and synthesizes their applications, challenges, optimization strategies, and future directions.
Results
The survey covers legal LLM applications in legal consultation and judge-assisted trials, while identifying challenges involving data, algorithms, judicial practice, and ethics.
Takeaways & Limitations
The paper recommends optimizing legal LLMs and clarifying their role within judicial systems to support fairer judgments while preserving judicial decision-making.
Takeaways & Limitations
Legal LLMs are not sufficient to replace judges, whose independence includes legal interpretation, fact-finding, discretion, and experiential judgment.
Abstract
from arXiv · showhide
The advent of artificial intelligence (AI) has significantly impacted the traditional judicial industry. Moreover, recently, with the development of AI-generated content (AIGC), AI and law have found applications in various domains, including image recognition, automatic text generation, and interactive chat. With the rapid emergence and growing popularity of large models, it is evident that AI will drive transformation in the traditional judicial industry. However, the application of legal large language models (LLMs) is still in its nascent stage. Several challenges need to be addressed. In this paper, we aim to provide a comprehensive survey of legal LLMs. We not only conduct an extensive survey of LLMs, but also expose their applications in the judicial system. We first provide an overview of AI technologies in the legal field and showcase the recent research in LLMs. Then, we discuss the practical implementation presented by legal LLMs, such as providing legal advice to users and assisting judges during trials. In addition, we explore the limitations of legal LLMs, including data, algorithms, and judicial practice. Finally, we summarize practical recommendations and propose future development directions to address these challenges.
1. Introduction
The introduction frames legal LLMs as an emerging response to growing judicial demands and surveys their technologies, applications, limitations, and future directions.
- Growing case volumes and limited human resources have prolonged judicial processes, motivating the application of AI in law.
- The paper identifies open questions about legal LLMs’ characteristics, shortcomings, future role, and ability to replace judges.
- The survey systematically reviews legal LLM research, including company and university studies, fine-tuning techniques, and evaluation strategies.
- It examines legal LLM applications in judicial decision-making and describes how they may help judges make fairer decisions.
- The paper summarizes key challenges and proposes future directions and improvement suggestions for legal LLMs.
2. Key Technologies of LLMs
The section introduces foundational AI, machine-learning, neural-network, deep-learning, NLP, and LLM concepts, then traces language models’ evolution toward more capable legal applications.
- Related Concepts: AI, machine learning, neural networks, deep learning, and NLP provide the conceptual and technical foundations for understanding LLMs.NLP connects machines with human language, while neural networks and deep learning support data processing, feature learning, classification, and prediction.
- Core Technologies: Self-attention assigns weights across sequence elements and uses multi-head processing to capture relationships and long-range dependencies.The mechanism outputs a weighted sequence representation and is widely used in NLP models.
- Core Technologies: Multi-task learning shares representations across tasks, supporting generalization to unknown datasets while improving efficiency and reducing resource consumption.It replaces multiple independent models with one model handling multiple tasks.
- Related Concepts: Foundation models are large-scale pretrained models that can be fine-tuned for multiple downstream tasks.They serve as a basis for constructing specific AI applications.
- Related Concepts: LLMs are large-parameter language models pretrained on extensive unlabeled data and refined through fine-tuning and reward modeling.They support tasks including automatic text generation and translation.
- Evolution of LLMs: Language models progressed from statistical N-gram and bag-of-words approaches to neural models, Word2Vec, GPT-3, and later GPT-3.5 and GPT-4.The cited evolution reflects increasing expressive, processing, scale, and generation capabilities.
3. Evolution of Judicial Technology
The section contrasts traditional, resource-intensive judicial processes with AI-supported approaches built around legal big data and LLM capabilities. It highlights efficiency gains alongside data, privacy, and implementation constraints.
- Traditional Judiciary: Traditional judiciary relies on human decision-making, precedents, contextual legal judgment, and substantial time and resources.High case volumes and limited personnel can prolong hearings, evidence collection, and trials.
- LLM Judiciary: Legal LLMs can assist professionals by identifying similar cases, summarizing case details, supporting decisions, and drafting repetitive documents.These functions are presented as ways to address the imbalance between case volume and available personnel.
- Legal Big Data: Legal big data is unstructured and requires NLP and text-analysis techniques to convert documents into AI-processable structured data.Sources include legal concepts, legislation, judgments, and commentaries with inconsistent textual formats.
- Legal Big Data: Legal data spans languages, cultures, domains, and document types, creating translation, terminology, integration, and specialized-analysis challenges.Examples include comparing multilingual European Union regulations and combining federal and state court records.
- Legal Big Data: Legal data must be regularly updated because laws and regulations frequently change.Tax regulations are cited as an example of provisions requiring current research and advice.
- Legal Big Data: Legal data may contain sensitive personal and case information, requiring anonymization and strict privacy and security protection.Criminal court documents may include identifying information and defendants’ criminal histories.
- LLM Judiciary: Judicial AI applications are described as simplifying repetitive procedures and improving efficiency by reducing administrative work and increasing trial efficiency.These systems incorporate more comprehensive legal big data into judicial workflows.
4. Recent applications
The section surveys fine-tuned legal LLMs, evaluation approaches, and judicial AI deployments. These systems support consultation, case analysis, reasoning, document generation, and decision assistance across legal settings.
- Fine-tuned Models: The survey examines ten fine-tuned legal LLMs, model-evaluation metrics and methods, and AI legal case studies.The analysis focuses on distinct characteristics in handling legal matters and feasible evaluation practices.
- Fine-tuned Models: Chinese legal LLMs use domain-specific pretraining or fine-tuning to support legal consultation, question answering, case analysis, reasoning, search, and document generation.Examples include LawGPT_zh, JurisLMs, Fuzi.mingcha, LaWGPT, LexiLaw, Lawyer LLaMA, HanFei, ChatLaw, Lychee, and WisdomInterrogatory.
- Fine-tuned Models: ChatLaw includes ChatLaw-13B and ChatLaw-33B, while ChatLaw-Text2Vec uses 930,000 court cases for similarity matching.The models are trained on legal news, forums, judicial interpretations, and court-case data.
- Evaluation: A proposed legal-LLM evaluation system combines prompt-based assessment with subjective legal-expert judgments and weighted objective indicators.The proposal describes two evaluation levels and combines weights with ratings.
- Law+AI Cases: Judicial AI deployments include Shanghai’s “206” criminal case assistance system, Zhejiang’s “Mobile Micro-Court,” Hangzhou’s smart judgment system, and ChatLaw.These examples illustrate early integration of AI into judicial practice.
- Law+AI Cases: AI is used internationally to assist sentencing, criminal-conviction decisions, recidivism assessment, and the handling of small cases.Examples include Colombia’s ChatGPT-assisted sentencing, the UK’s HART system, and the Netherlands’ ProKid 12-SI system.
5. Challenges
Legal LLMs face interconnected challenges involving data quality and availability, algorithmic transparency and fairness, and their role in judicial practice. These constraints affect reliability, trust, fairness, and judicial independence.
- Data challenges: Legal LLM development depends on representative, comprehensive, high-quality annotated datasets that address bias, privacy, timeliness, and scalability.Datasets may be incomplete, imbalanced, outdated, or difficult to extend beyond specific periods, courts, or legal provisions.
- Data challenges: Insufficient acquisition and sharing of judicial data limits coverage because courts restrict access and use inconsistent permissions and formats.Non-standard legal documents and difficulties integrating data across court levels further reduce completeness and accuracy.
- Data challenges: Legal concepts are inherently uncertain, and current AI systems may misinterpret their boundaries or generate inappropriate assumptions.Conceptual meanings can also evolve over time and vary by location, making stale training data unsuitable for later legal environments.
- Algorithmic challenges: Complex LLM structures make decisions difficult to predict or explain, while biased data and outsourced algorithms can undermine fairness, transparency, and public trust.The paper highlights risks from racial or gender bias, opaque algorithmic operations, and reduced transparency in legal interpretations.
- Judicial-practice challenges: Legal LLMs are not sufficient to replace judges, and excessive intervention may encourage reliance on AI while weakening judicial independence and party participation.Administrative use of AI-assisted decision indicators can also encourage abuse of AI and undermine judges’ status in trials.
- Judicial-practice challenges: AI-assisted judicial systems may create imbalances between public and private power when authorities control data and defendants cannot access equivalent information.Unequal data access can disadvantage the defense during evidence collection and restrict trial fairness.
6. Future Directions
The paper proposes improving legal LLMs through broader and better-governed data, stronger infrastructure and long-text algorithms, clearer legal boundaries, and constrained transparency. It also emphasizes preserving judicial independence and expanding legal consultation.
- 6.1. Data and Infrastructure: Courts should broaden legal-data acquisition by improving privacy protection, sharing rules, secure platforms, document standards, and preprocessing.The paper also recommends adapting pretraining models such as Lawformer to long legal documents.
- 6.2. Algorithm Level: The paper recommends defining ambiguous legal-concept boundaries and limiting legal LLMs’ decision-making proportion in cases involving those concepts.These limits are intended to preserve judicial authority where concepts remain difficult to formalize.
- 6.1. Data and Infrastructure: Legal LLM infrastructure can be strengthened with high-performance computing, distributed training, scalable storage, and deployment techniques such as compression and containerization.These measures target faster training, efficient data management, and more practical deployment.
- 6.2. Algorithm Level: Long-text performance can be improved through external recall, model optimization, and attention-mechanism optimization using legal knowledge bases and specialized legal models.External recall supplies legal documents, precedents, regulations, and explanations to support long-text processing.
- 6.2. Algorithm Level: Algorithmic bias and opacity should be addressed through public evaluation, protective algorithms, explainable AI, and legal indicators for assessment.The proposals include differential privacy, improved bias identification, and greater transparency in decision-making algorithms.
- 6.2. Algorithm Level: Algorithmic transparency should be limited and context-sensitive, with confidentiality agreements or closed hearings when trade secrets or national security are involved.The paper presents limited transparency as a way to support supervision while allowing protected information to remain confidential.
- 6.3. Dealing with Traditional Judiciary: Legal models should have clearly defined roles that uphold judges’ independence, while their consultation functions can expand access to legal advice for people unfamiliar with the law.The paper separates consulting assistance from unrestricted judicial decision-making.
7. Conclusions
The paper concludes that legal LLMs offer opportunities for legal consultation and judge-assisted trials while remaining constrained by technical, privacy, ethical, and case-specific challenges. It recommends improving data, model capabilities, governance, and cooperation as these systems develop.
- 7. Conclusions: Legal LLMs can provide legal advice, assist judges, accelerate case processing, reduce workload, and improve decision-making accuracy and consistency.These benefits are presented as opportunities rather than established universal outcomes.
- 7. Conclusions: Current challenges include insufficient long-text processing, limited adaptation to individual cases, and privacy and ethical issues.The paper identifies these constraints as central concerns for legal LLM deployment.
- 7. Conclusions: Future directions include stronger data quality and privacy protection, improved model adaptability, ethical and regulatory frameworks, international cooperation, and multi-party collaboration.The proposed cooperation aims at sustainable development and social benefits within the judicial field.