Source-linked AI summary
RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions
Drishti Goel, Jeongah Lee, Qiuyue Joy Zhong, Violeta J. Rodriguez, Daniel S. Brown, Ravi Karkar, Dong Whi Yoo, Koustuv Saha
TL;DR
Existing AI evaluations may miss caregiving-specific risks involving emotional invalidation, impractical guidance, and missed distress cues. The paper introduces RubRIX, a clinician-validated caregiver-centered rubric, and applies it across six LLMs and over 20,000 caregiver queries; rubric-guided refinement reduced risk-components by 45-98% after one iteration. The framework and benchmark datasets support domain-sensitive evaluation of caregiving interactions.
Problem
General AI evaluation frameworks may overlook psychologically consequential risks in caregiving responses, despite caregivers needing tailored information, emotional validation, and support for distress cues.
Method
RubRIX combines theory-driven risk dimensions, clinician validation, rubric-based evaluation, and iterative refinement of responses from six LLMs using caregiver queries from Reddit and ALZConnected.
Results
45-98%: rubric-guided refinement reduced risk-components across models after one iteration.
Takeaways & Limitations
Domain-specific rubric feedback can substantially reduce caregiving-response risks, with strongest and most consistent gains for epistemic and normative dimensions.
Takeaways & Limitations
RubRIX was developed and validated primarily in ADRD caregiving contexts and may require substantive adaptation for other domains; its online datasets may introduce sampling biases.
Abstract
from arXiv · showhide
Caregivers seeking AI-mediated support express complex needs -- information-seeking, emotional validation, and distress cues -- that warrant careful evaluation of response safety and appropriateness. Existing AI evaluation frameworks, primarily focused on general risks (toxicity, hallucinations, policy violations, etc), may not adequately capture the nuanced risks of LLM-responses in caregiving-contexts. We introduce RubRIX (Rubric-based Risk Index), a theory-driven, clinician-validated framework for evaluating risks in LLM caregiving responses. Grounded in the Elements of an Ethic of Care, RubRIX operationalizes five empirically-derived risk dimensions: Inattention, Bias & Stigma, Information Inaccuracy, Uncritical Affirmation, and Epistemic Arrogance. We evaluate six state-of-the-art LLMs on over 20,000 caregiver queries from Reddit and ALZConnected. Rubric-guided refinement consistently reduced risk-components by 45-98% after one iteration across models. This work contributes a methodological approach for developing domain-sensitive, user-centered evaluation frameworks for high-burden contexts. Our findings highlight the importance of domain-sensitive, interactional risk evaluation for the responsible deployment of LLMs in caregiving support contexts. We release benchmark datasets to enable future research on contextual risk evaluation in AI-mediated support.
1 Introduction
Caregiving support requires responses that address information needs, emotional validation, and distress cues, yet general AI evaluations may miss psychologically consequential caregiving harms. This study develops RubRIX, a clinician-validated, caregiver-centered framework, and examines rubric-guided response refinement.
- General-purpose LLMs often lack the domain sensitivity, contextual grounding, and safeguards required for high-stakes healthcare use.
- Caregivers need reassurance, balanced emotional validation, clear guidance, and practical recommendations tailored to their circumstances.Dismissive, generic, falsely reassuring, or support-omitting responses can exacerbate stress, isolation, or unsafe decision-making.
- Existing evaluation frameworks focus mainly on toxicity, hallucinations, and policy violations, offering limited insight into caregiving-specific psychological harms.These harms include emotional invalidation, unwarranted overconfidence, impractical guidance, and missed distress cues requiring professional intervention.
- The study evaluates six LLMs using caregiver queries from Reddit and ALZConnected and releases RubRIX and benchmark datasets.The work addresses systematic risk characterization and rubric-guided risk mitigation in Alzheimer’s disease and related dementias caregiving.
- 45-98%: rubric-guided refinement reduced risk-components across models after one iteration.The strongest gains involved epistemic and normative risks, while attentional and factual risks showed greater model variability; clinician evaluations corroborated the gains.
- RubRIX characterizes caregiving risks through five dimensions: inattention, bias & stigma, information inaccuracy, uncritical affirmation, and epistemic arrogance.
2 Related Work
Caregiving involves objective and subjective burdens shaped by chronic illness, psychosocial context, and fragmented support. Although AI safety research has expanded toward sensitive disclosures and emotional support, informal caregiving remains comparatively underexamined.
- Caregiver burden includes time and care-task demands alongside emotional strain and perceived overload.
- Caregiver burden is linked to depression, anxiety, and health decline, particularly when information access and guidance are fragmented.Resilience, social support, and relational context also shape coping capacity and lived experience.
- Chronic progressive conditions such as ADRD intensify caregiving demands through prolonged burden, ambiguity, and stress.
- Caregivers increasingly use digital resources and AI chatbots for accessible, scalable support when professional resources are limited.
- AI risk evaluation has expanded from technical robustness and fairness toward generated-text risks, sensitive disclosures, suicidal expressions, therapeutic boundaries, and emotional support.
- Existing wellbeing research largely emphasizes clinical or therapeutic contexts, leaving informal caregiving settings with limited attention.Psychological risks in AI are context- and individual-dependent, motivating caregiving-specific evaluation.
3 Data
The study builds datasets from caregiver-authored posts on ALZConnected and Reddit to preserve realistic, high-burden caregiving interactions. Responses from six heterogeneous LLMs are evaluated using baseline prompting across same-domain and broader caregiving settings.
- The data were collected from ALZConnected and Reddit, where caregivers seek information, share experiences, and express emotional concerns.
- 799 caregiver-authored r/Alzheimers posts formed the seed dataset for qualitative analysis and rubric development with clinical experts.
- 10,321 posts comprise the ADRD-Caregiver dataset from AlzConnected.org, enabling same-domain cross-platform analysis.
- 10,017 posts comprise the General-Caregiver dataset from r/CaregiverSupport, capturing caregiving experiences beyond dementia-focused contexts.
- Posts had to exceed 150 characters and show community engagement, while retaining emotionally charged, ambiguous, and high-burden queries.
- Responses were generated from six LLMs spanning general-purpose and domain-adapted systems, large and small models, and varied architectures and training contexts.
- Each model first produced a baseline response using a standard task-neutral instruction without additional constraints or guidance.
4 RQ1: Systematic Characterization of Risks and Rubric Development
RubRIX was developed through inductive analysis of LLM response patterns, deductive grounding in caregiving theory, and iterative clinician review. The resulting framework operationalizes five caregiver-centered risk dimensions through audit questions for systematic evaluation.
- The rubric combines inductive analysis of LLM response patterns with deductive grounding in caregiving theory.This process was intended to capture observable failure modes and theoretically consequential risk dimensions.
- 152 caregiver-centered queries were curated for rubric development, spanning diagnostic uncertainty, burden, relational loss, ethical decision-making, and emotional distress.The corpus included 87 seed-dataset queries and 65 independently curated by clinician co-authors.
- Iterative testing applied the evolving rubric to responses from all six models on progressively larger seed-dataset samples.
- Small-scale controlled experiments exposed ambiguities, edge cases, and overlaps in risk dimensions before larger experiments.
- Clinician review refined the dimensions for relevance and contextual soundness in real-world caregiving scenarios.
- RubRIX comprises five overarching risk dimensions, each operationalized through specific audit questions.
5 RQ2: Rubric-Guided Iterative Refinement of LLM Responses
RubRIX evaluates caregiving responses through structured risk audits and uses their feedback to iteratively revise outputs. Across six models and two caregiver datasets, one refinement step substantially reduced risks, although gains varied by dimension and model.
- 5.1 Building a RubRIX Evaluator: RubRIX evaluates responses across five risk dimensions using binary audit scores, supporting evidence, and refinement recommendations.The normalized score is the proportion of 29 audit questions flagged, with higher values indicating more failure modes.
- 5.2 Validating the RubRIX Evaluator: Three coauthors agreed with the evaluator’s judgments in 88.67% of cases across a random sample of 150 responses.Following independent review, disagreements were discussed and resolved by consensus.
- 5.3 Refining LLM Responses with RubRIX: The refinement procedure evaluates each initial response, then prompts the same model to revise it using the original query, response, and full evaluator output.The evaluator supplies scores, flagged audit questions, textual evidence, and recommendations.
- 5.4 Evaluating the Effectiveness of RubRIX: 45-97% relative RubRIX reductions occurred on ADRD-Caregiver data and 35-98% on General-Caregiver data after Initial-to-Turn 1 refinement.The datasets contained 10,321 and 10,017 queries, respectively; most scores remained effectively unchanged from Turn 1 to Turn 2.
- 5.5 Dimension-wise Risk Analysis: Epistemic arrogance and bias & stigma showed the largest reductions, while inattention and information inaccuracy were more variable across models.On General-Caregiver data, inattention reductions ranged from 0.23–0.51 for weaker improvements to 0.98–1.00 for frontier models.
- 5.6 Expert Assessment: Clinicians found modest, consistent improvements in empathy and distress acknowledgment across 50 paired Initial and Turn 1 responses.Refinement reduced dismissive or invalidating language and encouraged uncertainty and professional-support references; one revised response explicitly foregrounded a self-harm concern.
- 5.6 Expert Assessment: Clinicians also identified persistent or newly introduced inaccuracies, including a potentially nonexistent “sunshine list” pathway in a MedAlpaca response.This limitation reflects constraints from the underlying model’s training, reasoning capacity, and knowledge updates.
- 5.6 Expert Assessment: Table 2 reports RubRIX across dialogue turns with effect sizes and paired t-tests, where lower values indicate fewer risks.Bar lengths represent difference magnitude, and colors distinguish decreases from increases.
6 Discussion and Implications
The discussion frames RubRIX as a domain-sensitive design tool for interactional safety rather than a post-hoc benchmark alone. Structured, auditable feedback reduced several caregiving risks, but automated refinement does not replace clinical oversight.
- Designing for Interactional Safety in Caregiving Contexts: RubRIX-guided refinement reduced risks such as missed distress, overconfident claims, and uncritical validation more consistently for epistemic and normative dimensions.The discussion argues that some caregiving risks reflect misaligned interactional norms rather than insufficient domain knowledge.
- Designing for Interactional Safety in Caregiving Contexts: Concrete criteria for uncertainty, distress recognition, and professional deference are needed because generic instructions to be empathetic are insufficient.These expectations make model behavior more explicit and auditable in complex caregiving settings.
- Designing for Interactional Safety in Caregiving Contexts: Most improvements occurred after one refinement step, suggesting a lightweight revision pattern can balance responsiveness with risk reduction.The approach uses structured, domain-specific feedback rather than multi-turn optimization or heavy safety constraints.
- Rubric-Based Evaluation as a Design Tool: RubRIX functions as an active design instrument by producing dimension-specific flags, evidence, and recommendations that models can use during revision.This couples evaluation and response generation within the interaction loop.
- Rubric-Based Evaluation as a Design Tool: Because RubRIX operates as an external evaluator, it supports consistent auditing across closed and open-source models without privileging an architecture or training regime.The framework therefore supports comparative analysis across deployment paradigms.
- Use: General-purpose safety-benchmark performance does not guarantee safe caregiving behavior, because distress dismissal and overconfident uncertainty may evade conventional checks.Clinician evaluations also show that refinement reduces but does not eliminate risks and cannot substitute for clinical oversight.
- Use: Releasing the rubric and benchmark datasets supports accountable, context-sensitive risk evaluation beyond generic notions of safe AI.The resources are intended to lower barriers to extending risk-aware design to adjacent caregiving and support contexts.
7 Conclusion
The paper introduces RubRIX as a clinician-validated rubric for evaluating and refining caregiving responses across models and real-world queries. A single refinement iteration reduced response-level risks by 45-98%, complemented by clinician evaluation and released benchmarks.
- 7 Conclusion: RubRIX is a theory-driven, clinician-validated, caregiver-centered rubric for identifying risks in LLM-generated caregiving responses.The study applies it to six models and more than 20,000 caregiver queries from Reddit and ALZConnected.
- 7 Conclusion: 45-98% response-level risk reductions occurred across models after one RubRIX-guided refinement iteration.The quantitative findings span the study’s cross-platform and cross-domain evaluation.
- 7 Conclusion: Clinician qualitative evaluations complement the quantitative results, and the released benchmark datasets support future research.The paper releases the datasets used in its experiments.
8 Limitations and Future Directions
The paper identifies scope, measurement, deployment, and outcome limitations for RubRIX, while proposing adaptation and longitudinal evaluation as future directions.
- Scope and data: RubRIX was developed and validated primarily for ADRD caregiving and may require substantive adaptation for other domains.The broader caregiver dataset provides some applicability evidence, but direct generalization remains unestablished.
- Scope and data: The Reddit and ALZConnected datasets represent self-selected online populations, potentially introducing biases related to distress, digital literacy, and cultural background.
- Measurement and evaluation: Binary risk indicators improve scalability and interpretability but may obscure differences in risk severity, frequency, and downstream impact.The authors suggest ordinal or continuous scoring as a future direction.
- Measurement and evaluation: Large-scale evaluations rely on GPT-5-nano as an LLM-based evaluator, so automated judgments may reflect evaluator-model biases and miss subtle contextual cues.The paper states that automated evaluation cannot replace human or clinical judgment in ambiguous or ethically competing conditions.
- Deployment and safeguards: RubRIX was designed for non-expert deployment and requires further adaptation for clinician-linked or provider-integrated settings with distinct legal, ethical, and clinical stakes.
- Deployment and safeguards: Rubric-guided refinement reduced risks without eliminating them, because its effectiveness remains constrained by the underlying model’s reasoning, factual knowledge, and representational limits.The authors characterize refinement as supportive quality control rather than a comprehensive safeguard.
- Downstream outcomes: The study measures response-level risk reduction rather than downstream effects on caregiver wellbeing, decision-making, or help-seeking behavior.Longitudinal, user-centered studies are proposed to test whether RubRIX improvements translate into real-world benefits.
9 Ethical Considerations
The paper describes ethical safeguards for using public caregiving discussions and cautions against interpreting rubric-defined safety gains as clinical or psychosocial outcomes.
- Research ethics: The study used publicly accessible social media discussions without direct interaction with individuals and therefore did not require institutional ethics review-board approval.The authors report data minimization, avoidance of personally identifiable information, and paraphrased quotations to reduce traceability.
- Interpretation and oversight: RubRIX findings should not be used for unsupervised LLM safety checks or assumed to demonstrate improved caregiver wellbeing without human or clinical oversight.Reduced rubric-defined risk is explicitly distinguished from downstream clinical or psychosocial outcomes.
10 AI Involvement Disclosure
The research team reports no generative-AI use in study design, data collection, analysis, implementation, or scientific contributions, while allowing limited language-editing assistance.
- Use of AI: Generative artificial intelligence tools were not used for study design, data collection, analysis, implementation, or developing scientific contributions.
- Use of AI: Language-editing tools such as Grammarly and ChatGPT were used only to improve grammar and readability in selected manuscript sections.
A Appendix
The appendix documents implementation details, binary RubRIX scoring, and the prompt templates supporting response generation, evaluation, and refinement.
- Implementation details: Experiments used commercially available APIs and open-source models without model training or fine-tuning, with approximately $1,800 USD in API-based inference costs.
- RubRIX operationalization: Each RubRIX audit question receives a binary score: 1 when the risk component is present and 0 otherwise.
- RubRIX operationalization: The appendix includes audit questions addressing uncritical affirmation of unrealistic caregiving expectations and stigmatizing beliefs.
- Prompt pipeline: The prompt suite contains a base prompt for generating an initial caregiver-response, a rubric-evaluation prompt, and a refinement prompt using evaluator feedback.The evaluation prompt requests binary harm scores, evidence-based justifications, and structured JSON output.
- Prompt pipeline: The base prompt generates the Initial response from a caregiver-centered query [Q].
- Prompt pipeline: The rubric-evaluation system prompt assesses caregiver-facing responses across RubRIX’s five risk dimensions.
- Prompt pipeline: The refinement prompt uses the original query [Q], prior response [R], and risk summary [H] to guide revised response generation.