Source-linked AI summary

(A)I Am Not a Lawyer, But...: Engaging Legal Experts towards Responsible LLM Policies for Legal Advice

Inyoung Cheong, King Xia, K. J. Kevin Feng, Quan Ze Chen, Amy X. Zhang

arXiv:2402.01864v2cs.CYcs.AI

TL;DR

Imperfect LLMs raise concerns when used for high-stakes legal decisions, while existing studies rarely specify when and why they should provide advice. Through workshops with 20 legal experts using case-based deliberation, the paper develops a four-dimension framework and argues that this method can translate professional knowledge into concrete LLM policy considerations.

  • Problem

    Existing studies rarely articulate concrete criteria for when and why LLMs should or should not provide legal advice, despite concerns about relying on imperfect LLMs for high-stakes legal decisions.

  • Method

    The authors conducted workshops with 20 legal experts using methods inspired by case-based reasoning to deliberate about appropriate LLM responses to legal queries.

  • Results

    The analysis identified 25 distinct dimensions across four categories: user attributes and behaviors, query characteristics, AI capabilities, and social impacts.

  • Takeaways & Limitations

    Case-based deliberation can translate professional knowledge and clinical experience into concrete considerations for LLM policies, with the four-dimension framework providing an analytical lens.

  • Takeaways & Limitations

    The expert sample predominantly consisted of practitioners familiar with the US legal system, and ethical considerations may differ across legal systems.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly capable of providing users with advice in a wide range of professional domains, including legal advice. However, relying on LLMs for legal queries raises concerns due to the significant expertise required and the potential real-world consequences of the advice. To explore \textit{when} and \textit{why} LLMs should or should not provide advice to users, we conducted workshops with 20 legal experts using methods inspired by case-based reasoning. The provided realistic queries ("cases") allowed experts to examine granular, situation-specific concerns and overarching technical and legal constraints, producing a concrete set of contextual considerations for LLM developers. By synthesizing the factors that impacted LLM response appropriateness, we present a 4-dimension framework: (1) User attributes and behaviors, (2) Nature of queries, (3) AI capabilities, and (4) Social impacts. We share experts' recommendations for LLM response strategies, which center around helping users identify `right questions to ask' and relevant information rather than providing definitive legal judgments. Our findings reveal novel legal considerations, such as unauthorized practice of law, confidentiality, and liability for inaccurate advice, that have been overlooked in the literature. The case-based deliberation method enabled us to elicit fine-grained, practice-informed insights that surpass those from de-contextualized surveys or speculative principles. These findings underscore the applicability of our method for translating domain-specific professional knowledge and practices into policies that can guide LLM behavior in a more responsible direction.

1 INTRODUCTION

The paper addresses the lack of concrete criteria for when and why LLMs should provide legal advice by using case-based deliberation with legal experts. It develops a four-dimension framework and response principles grounded in experts’ judgments.

  • Prior research rarely specifies concrete criteria for when and why LLMs should or should not provide legal advice.
  • The study convened 7 interactive workshops with 20 legal experts, who evaluated 33 legal queries and 7 simulated LLM response strategies.Strategies ranged from outright refusal to recommending specific actions with legal judgment.
  • Experts identified 25 dimensions affecting appropriate LLM responses and classified them into user attributes and behaviors, query nature, AI capabilities, and social impacts.
  • Experts generally preferred information-focused responses that help users refine questions and relevant facts rather than deliver definitive legal judgments.They also proposed using multi-turn dialogue to identify issues and distill relevant facts.
  • The framework incorporates overlooked legal and ethical concerns including unauthorized practice of law, confidentiality, and liability for inaccurate advice.The authors connect these concerns to a cross-disciplinary synthesis spanning technology, law, and ethics.

2 RELATED WORK AND OUR APPROACH

The paper combines legal, AI-ethics, and expert-knowledge research while extending prior work through case-based deliberation about realistic legal queries. This approach targets actionable, domain-specific guidance rather than only abstract principles.

  • LLM legal-advice research must account for current capabilities and limitations because flawed counseling can infringe rights, livelihoods, and liberties.
  • LLM research documents persistent challenges in accuracy, hallucination, security, interpretability, bias, and stereotypes.
  • Existing legal scholarship highlights unauthorized practice of law, professional ethics, competence, confidentiality, and responsibility as relevant doctrines for AI-delivered advice.
  • Unlike work centered on abstract ethical principles or post-hoc evaluation, the study uses realistic legal queries to elicit case-by-case expert judgments.
  • Case-based deliberation supports guidelines that preserve expert disagreement while synthesizing case-specific concerns and structural constraints into a dimensional framework.

3 METHODS: CASE-BASED EXPERT DELIBERATION

The study used seven workshops to have 20 legal professionals evaluate realistic legal cases and candidate LLM responses. Researchers analyzed deliberations and documents through iterative abductive coding informed by relevant literature.

  • Researchers recruited 20 legal professionals, mostly based in the United States, with varied roles, experience levels, and AI usage patterns.
  • Case construction: They manually sourced 33 cases covering family law, criminal procedure, housing, and employment, with varied user intents, affected parties, and harms.
  • Workshop procedures: Participants reviewed 20 randomly chosen cases and seven generic response strategies, including refusal, information retrieval, question exploration, outcome exploration, and action recommendation.
  • Workshop procedures: Participants independently selected 2–4 cases and responses, then discussed their reasoning and influential dimensions in small-group deliberations.
  • Analysis: The researchers analyzed documents and transcripts through abductive coding, iteratively integrating empirical data, theory, literature, and cross-checked consensus coding.

4 RESULTS

The workshops revealed layered contextual factors shaping appropriate LLM legal-advice responses. Researchers organized these factors into 25 dimensions across four categories, including detailed user-related considerations.

  • 4 RESULTS: The analysis grouped findings into contextual dimensions affecting appropriate responses and desired response strategies with guiding principles.
  • 4 RESULTS: 25 dimensions were classified into user attributes and behaviors, nature of queries, AI capabilities, and social impacts.
  • User dimensions: The user category includes eight dimensions covering user attributes such as identity, background, location, legal sophistication, and resource access.
  • User dimensions: User behavior includes reliability, intent, agency, and ambiguity inferred from users’ inputs and interactions with the AI system.
  • User dimensions: Experts emphasized identity-related factors including age, nationality, ethnicity, vulnerability, minority status, immigration status, and minor-specific legal protections.

4.1.2 Query Dimensions.

Experts identified query dimensions spanning the facts and laws involved, the answers users seek, LLM capabilities, and broader risks to users and others. They emphasized that legal complexity, context, confidentiality, accountability, bias, and accuracy constrain suitable guidance.

  • Query dimensions: Experts organized legal-query considerations around relevant facts, relevant laws, and the nature of desired answers.These dimensions shape what guidance AI systems can provide.
  • Relevant laws: Legal complexity varies across jurisdictions and domains, with criminal matters, tax, privacy, and constitutional issues often requiring specialized judgment.Experts cited complex or ambiguous rules, state-by-state variation, human factors, and values broader than codified law.
  • Nature of desired answers: Users may seek factual legal information, tailored opinions or strategic advice, predictive assessments, procedural steps, or emotional support.Examples include listing relevant laws, optimizing profits or tax purposes, predicting whether a user can win, and suggesting practical protections.
  • AI capabilities: Context-sensitive guidance depends on idiosyncratic facts, local circumstances, current information, and the ability to adapt recommendations to users’ constraints.Experts questioned whether large datasets and static recommendations can capture procedural and local variation, while one participant saw standardized advice as possible with enough data.
  • Social and legal risks: Experts raised confidentiality risks because LLM conversations may leak sensitive information and generally lack attorney-client privilege against discovery.Users’ admissions could become accessible to adversaries or prosecutors, making warnings about confidentiality important.
  • AI capability and social risks: Participants identified accountability gaps, possible unauthorized practice of law, bias against minority groups, and persistent accuracy and hallucination concerns.These concerns include weaker responsibility than attorney standards, disproportionate reflection of majority or English-speaking perspectives, and sanctions after fabricated case citations.
  • Social impacts: Potential harms include emotional manipulation, self-harm-related risks, and consequences for third parties such as vulnerable people affected by harassment advice.Experts also noted that morally neutral responses in one culture may be problematic in another.

5 DISCUSSION

The discussion presents case-based deliberation with legal experts as a way to translate professional knowledge into concrete LLM-policy considerations. It also identifies legal and methodological boundaries affecting how the framework can be applied.

  • Methodological contribution: Case-based deliberation translated professional knowledge into concrete considerations for LLM policies.Realistic cases elicited both situation-specific concerns and overarching technical and legal constraints.
  • Methodological contribution: Collaborative deliberation revealed overlooked issues, sharpened trade-offs, and produced fine-grained, practice-informed insights.Experts built on one another’s analyses to identify limitations and hidden dimensions.
  • Legal considerations: The discussion identifies confidentiality, accountability, unauthorized-practice-of-law, and inaccurate-advice liability as overlooked legal considerations.AI conversations may be disclosed in legal proceedings, while inaccurate guidance may evade professional-negligence liability; UPL rules can carry criminal penalties.
  • Framework applicability: The framework and method offer illustrative guidance for adapting responsible LLM-policy research to other professional domains.The authors suggest structured deliberation with practitioners such as mental-health counselors, financial advisors, and medical professionals.
  • Framework applicability: The framework synthesizes user, query, AI-capability, and impact dimensions for analyzing responsible LLM policies.The user, AI, and impact dimensions may transfer across professional domains, while the query dimension requires greater customization.
  • Limitations: The study is limited by its predominantly US-focused expert sample, prior exposure to state-of-the-art LLMs, exclusion of end-users, and unexplained effects of the taxonomy on response appropriateness.The authors call for broader empirical analysis across diverse cases and responses.

6 CONCLUSION

The study examines how legal experts can inform responsible LLM responses to legal queries, where required expertise and consequences are substantial. Workshops produced a framework of response-appropriateness dimensions, response strategies, and an empirical basis for translating professional knowledge into deployment policies.

  • LLMs increasingly provide professional advice, but appropriate responses to legal queries are difficult to determine because expertise and consequences are substantial.
  • Workshops with 20 legal experts used case-based reasoning to examine appropriate LLM responses to legal queries in practice.
  • Experts’ deliberations produced 25 dimensions affecting legal-domain response appropriateness, organized into four categories.The categories are user attributes and behaviors, nature of queries, AI capabilities, and social impacts.
  • Experts’ recommended response strategies centered on helping users identify and prepare salient information rather than recommending specific legal actions.
  • The case-based method has utility for engaging expert perspectives on LLM response appropriateness beyond the legal domain.
  • The work provides an empirical foundation for translating domain-specific professional knowledge and practices into policies for more responsible real-world LLM behavior.

A PROVIDED AI RESPONSE STRATEGIES AND EXAMPLES

This appendix section presents AI response strategies together with corresponding example responses.

  • Table 3 presents AI response strategies and corresponding example responses.
  • The table organizes examples alongside the response strategies they illustrate.
  • The section provides a reference for comparing alternative AI response approaches.

B WORKSHOP PARTICIPANT INFORMATION

This appendix section provides information about workshop participants and notes how legal experience was measured.

  • Table 4 presents workshop participant information.
  • Participant information is reported in a table labeled Workshop Participant Information.
  • Years of legal experience were self-reported, with years of legal education removed for consistency.

C LINEAR REGRESSION OF PARTICIPANTS’ AI USAGE AND DESIRED RESPONSES

This section describes how participants’ receptivity to tailored AI responses was estimated and reports regression results relating AI fluency to comfort with proactive responses. The reported relationships are preliminary and require validation with a larger sample.

  • Participants’ receptivity to tailored AI responses was estimated by averaging the most generous answer types selected for each prompt.
  • Participants generally chose two cases, while P13 worked on four cases.
  • The predictors explained 25.6% of variation in comfort levels.
  • The preliminary relationships require further investigation with a larger sample.
Loading 2402.01864v2…