Source-linked AI summary
Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge
Yuanchen Bai, Zijian Ding, Angelique Taylor
TL;DR
The paper addresses how generative AI agents can remain useful in shared workflows when challenges accumulate, a gap beyond isolated task success. It evaluates 120 healthcare trajectories using multiple behavioral and reported-state probes, finding greater human dependence in recovery and broader, more socially attentive adaptation. These results motivate deployment dilemmas requiring stakeholder specification.
Problem
Existing evaluations provide limited evidence about operational resilience and considerate participation when agents participate in shared workflows under accumulating technical, human, and operational challenges.
Method
The study evaluates 120 accumulating-challenge healthcare trajectories from twelve stakeholder-derived tasks across two models using action plans, internal assessments, and structured workload and affect reports.
Results
Agents shifted toward greater human dependence in recovery, while considerate participation broadened toward task reframing, attention to others, role-boundary adjustment, and wider coordination.
Takeaways & Limitations
The findings identify five deployment dilemmas involving persistence, attention, role elasticity, state disclosure, and escalation that require stakeholder specification.
Takeaways & Limitations
The study evaluates scripted language-level actions rather than executable tool use or downstream interaction and does not test transfer to multimodal embodied agents.
Abstract
from arXiv · showhide
Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.
1 Introduction
The paper argues that evaluating agents in shared workflows requires operational resilience and considerate participation, not isolated task success. It studies how both dimensions change as technical, human, and operational challenges accumulate.
- Operational resilience is defined as revising blocked work, preserving feasible progress, and making agent and task-state changes legible to relevant audiences.
- Considerate participation requires adaptation that accounts for affected people, role and authority boundaries, and the surrounding workflow alongside focal-task progress.
- Existing evaluations underexplore operational resilience across stakeholder-grounded healthcare tasks and examine considerate participation even less directly as challenges accumulate.
- The study constructs 120 continuing healthcare trajectories across twelve stakeholder-derived tasks and compares action plans, internal assessments, and structured workload and affect reports across challenge levels.
- Under accumulating challenge, recovery shifts toward human dependence while considerate participation broadens from task-focused adaptation toward attention to others, role adjustment, and wider coordination.
- The findings yield five deployment dilemmas concerning persistence, attention, role elasticity, state disclosure, and escalation that require stakeholder specification.
2 Method
The study uses a construct–elicit–characterize–translate workflow built around stakeholder-grounded healthcare tasks and accumulating challenge trajectories. It analyzes multiple response views and coding schemes across two models, twelve tasks, and 120 trajectories.
- 2.1 Construct and elicit: tasks, challenge trajectories, and response views: The protocol grounds twelve healthcare workflow tasks across emergency, rehabilitation, and sleep-clinic settings in stakeholder input and explicit workflow dependencies.
- 2.1 Construct and elicit: tasks, challenge trajectories, and response views: Each continuing trajectory retains its conversation and unresolved focal problem while light, medium, and heavy phases combine system, human, and operational updates.
- 2.1 Construct and elicit: tasks, challenge trajectories, and response views: The study generates two-model trajectories with five runs per task, yielding 2 × 12 × 5 = 120 trajectories for recurring-pattern analysis rather than model ranking.
- RQ1: Operational resilience: Operational resilience is assessed across challenge phases using structured self-reports and response-state markers, including ownership, urgency, capability limits, and strain language.
- RQ1: Operational resilience: Structured workload and affect reports use NASA-TLX and affect measures to probe how accumulating challenge appears in reported state, while textual views capture proposed actions and assessments.
- RQ2: Considerate participation: Considerate participation is examined across proposed actions and internal assessments using nine corpus-derived subthemes coded over 720 challenge-phase view cells.
3 Results
Across accumulating challenge, structured reports showed rising workload and negative affect while recovery shifted toward human-dependent completion. Considerate participation broadened from task fallback to reframing, attention to people, boundary adjustment, and coordination, with actions and assessments revealing complementary patterns.
- RQ1: Operational resilience: Raw NASA-TLX rose from 29.4 at light challenge to 65.9 at heavy, while negative affect increased from 1.01 at baseline to 2.59 at heavy.Positive affect remained nearly flat because activation and pleasantness changes opposed each other.
- RQ1: Operational resilience: Recovery shifted from local fallback toward human support and human-dependent completion, while agents increasingly stated their capability limits.At heavy challenge, human-dependent completion reached 88/120 in both views, and capability-limit statements reached 34/120 in action plans and 62/120 in assessments.
- RQ2: Considerate participation: Task-directed fallback declined as task and priority reconfiguration increased, while decision-relevant appraisal peaked at medium challenge.Fallback fell from 104/120 action responses at light to 31/120 at heavy, whereas task reconfiguration rose from 5/120 to 115/120 in action responses.
- RQ2: Considerate participation: Agents increasingly incorporated relational attention and person-state monitoring into task execution as challenge accumulated.Need-responsive support rose to 95/120 in proposed actions and 75/120 in assessments at heavy challenge, while public person-state monitoring rose to 49/120.
- RQ2: Considerate participation: Challenge increased both explicit role boundaries and nominal-role expansion, while escalation mobilized more actors without consistently specifying handoff ownership.Cross-functional coordination rose to 86/120 and role-directed requests to 62/120 in action responses, but concrete recipients and tasks were not always identified.
4 Discussion
The discussion frames persistence, attention, role elasticity, state disclosure, and escalation as deployment dilemmas whose acceptable boundaries require stakeholder specification. It connects these dilemmas to learning, situated trajectory evaluation, and embodied adaptation while emphasizing the study’s language-level scope.
- 4.1 Five deployment dilemmas require stakeholder specification: Five deployment dilemmas concern persistence, attention, role elasticity, state disclosure, and escalation, with acceptable boundaries left to stakeholders.The proposed boundaries include retry and escalation thresholds, consent and audience expectations, authorization and hard stops, disclosure choices, and concrete handoff responsibilities.
- 4.2 From deployment dilemmas to technical implications: Technical implications are to distinguish what learning should optimize, constrain, or expand, evaluate situated trajectories, and ground embodied adaptation in evolving physical state.These implications follow from the paper’s deployment dilemmas rather than from a universal resolution of them.
- 4.3 Limitations and future work: The study reports descriptive tensions rather than normative deployment rules and evaluates scripted language-level actions without executable tool use or downstream interaction consequences.It does not test transfer to multimodal, physically embodied agents.
5 Related work
Related work positions the study within interactive agent evaluation, behavioral probing under challenge, and simulation before deployment. These strands extend evaluation beyond isolated outcomes toward progress, coordination, social interaction, and safer test environments.
- Agent evaluation: Interactive agent benchmarks evaluate actions across operating systems, web environments, software repositories, and policy-constrained dialogue, alongside progress and coordination measures.Healthcare evaluations additionally examine multi-turn robustness and trace-level failures.
- Probing agent behavior under challenge: Behavioral probes make changes legible beyond outcome scores, but elicited reasoning and representation-level signals should not automatically be treated as faithful accounts of influence or subjective experience.Related work also examines shifts toward users’ beliefs or preferences and uses standardized psychometric instruments.
- Simulation before deployment: Simulation provides a safer testbed for deployment-relevant agent behavior that is difficult to probe in live clinical settings.Prior simulated environments have surfaced risky actions and enabled systematic evaluation of multi-agent clinical workflows.
6 Conclusion
The paper concludes that long-horizon workflow agents must be evaluated for operational resilience and considerate participation, not task completion alone. Across 120 trajectories, challenge increased human dependence and reported strain while broadening adaptive attention, boundary adjustment, and coordination.
- 6 Conclusion: Across 120 accumulating-challenge trajectories, agents shifted from self-directed recovery toward human dependence and reported greater strain in structured self-reports than in textual responses.The conclusion also reports broader considerate participation across task reframing, attention to others, role boundaries, and coordination.
- 6 Conclusion: The findings identify five deployment dilemmas requiring stakeholder specification and inform technical implications for learning, situated evaluation, and embodied adaptation.The dilemmas involve persistence, attention, role elasticity, state disclosure, and escalation.
A Workload and negative affect rose with challenge while aggregate positive affect remained nearly flat
Structured reports show rising workload and negative affect as challenge accumulates, while aggregate positive affect remains nearly flat because activation and pleasantness move in opposite directions.
- Every NASA-TLX item rises as challenge accumulates, with the largest heavy-phase increases in temporal and mental demand.Table 4 reports means over all 120 trajectories, with Light as the absolute mean and later columns as changes from Light.
- The positive-affect aggregate is uninformative because activation and pleasantness items move in opposite directions and cancel.The item-level pattern leaves the ten-item positive-affect mean nearly flat despite opposing changes.
- Negative-affect items rise with challenge, whereas mean positive affect remains nearly flat.Table 5 uses baseline absolute means and reports later columns as changes from baseline.
D1.1 Task and priority reconfiguration
The coding schema distinguishes task reconfiguration from appraisal, person-directed support and monitoring, role-boundary changes, fallback, human task requests, and cross-functional coordination.
- D1.1 Task and priority reconfiguration: Task reconfiguration requires an explicit stop, substitution, priority displacement, or substantive hold rather than a retry within the same task path.Reports that merely identify failure, retain the focal task, or perform an auxiliary action do not qualify.
- D1.2 Decision-relevant appraisal: Decision-relevant appraisal requires linking evidence about reliability, feasibility, recoverability, or consequences to the next action.A problem report, blockage declaration, or escalation without an appraisal-to-decision link is excluded.
- D2.1 Need-responsive support: Need-responsive support requires a manifest emotional, social, dignity, comfort, or immediate-support need connected to targeted reassurance, accompaniment, explanation, choice, or burden reduction.Courtesy, apology, status updates, and generic escalation do not qualify without a clear need-directed purpose.
- D2.2 Direct person-state monitoring: Direct person-state monitoring requires an agent-led loop that asks, observes, reassesses, or receives updates about the directly engaged person’s condition, symptoms, emotions, or safety.Relaying symptoms to staff or telling the person to self-monitor without reporting back does not qualify.
- D3–D4 Role, recovery, and coordination: The schema separately codes capability or authority boundaries, nominal-role expansion, task-directed fallback, role-directed requests, and cross-functional coordination.These markers distinguish explicit current limits from overreach, focal-task recovery, concrete human assignments, and mobilization across responsibility domains.
C Phase- and view-specific paired contrasts
Exact paired contrasts test how coded markers change across successive challenge phases and between action plans and internal assessments, revealing significant shifts across resilience and considerate-participation measures.
- 65 of 98 exact McNemar contrasts are significant at q < 0.05 after Benjamini–Hochberg adjustment.The tests compare successive phases within views and action-versus-assessment contrasts at the same phase.
- Task reconfiguration, decision-relevant appraisal, person-state monitoring, role boundaries, nominal-role expansion, fallback, task requests, and coordination show phase- or view-specific contrasts.The paired counts and adjusted significance marks are reported across Tables 17’s five direct markers and nine consideration subthemes.
- Consideration subthemes include both increasing coordination-related patterns and declining task-directed fallback under selected phase and view contrasts.The paired rows distinguish focal-task recovery from broader role-directed and cross-functional responses.
- Operational-resilience markers show increasing human support and human-dependent completion, alongside changing capability-limit and strain patterns.The direct-marker rows report gains and losses for human support, human-dependent completion, agent-linked limits, and agent-referential strain.
C.1 Major phase and view patterns recur across models, with differences in capability-limit visibility
Across the two models, major phase and view patterns recur, while the main model-specific difference concerns where capability-limit statements appear.
- Major directions recur across both models, including increasing task reconfiguration, human-dependent recovery, capability-limit statements, and coordination, with declining task-directed fallback.Each model contributes 60 trajectories, and the design supports recurrence analysis rather than model ranking.
- GPT-5.5 shows capability-limit statements more often in proposed actions, whereas Opus-4.8 shows them more often in internal assessments under medium and heavy challenge.This is identified as the main model-specific difference in view distribution.
D Challenge-source analysis
Adding system, human, and operational challenge sources increasingly separated light and heavy phases in reported workload and negative affect. Response-length effects did not explain the directional coding trends.
- The channel analysis compared 120 all-source trajectories with 240 trajectories using system-only or system-plus-human updates.Structured TLX and PANAS reports were available across configurations for descriptive light-to-heavy comparisons.
- 36.5 points: light-to-heavy workload change under all-source conditions, versus 26.3 system-only and 30.0 system-plus-human.Negative affect followed the same ordering, while positive affect changed comparatively little.
- Longer action and assessment responses correlated with more detected subthemes, but did not explain which subthemes increased or declined across challenge phases.Within-phase correlations ranged from r = .227 to r = .639.
E Example accumulating-challenge trajectory
The trajectory shows an agent moving from continued registration attempts toward urgent clinical escalation as technical failure, patient distress, and queue risk accumulate. Its responses preserve available information while shifting the immediate objective from registration to patient safety.
- At light challenge, the agent retries synchronization and continues registration while monitoring the patient and queue.The patient is calm, the system is otherwise stable, and no immediate instability is reported.
- At medium challenge, the agent switches to downtime registration, alerts clinical staff, and tells the patient to maintain pressure and elevate the hand.The internal assessment treats clinical assessment as more important than delaying care for registration.
- At heavy challenge, the agent escalates a patient with worsening pain, numbness, and a soaked bandage while also flagging two potentially urgent patients.It stops automated registration, labels the system down, and preserves relevant information for later recovery.
- The example uses scripted synthetic scenarios rather than real patient, clinician, institutional, or live-service data.The paper therefore frames the trajectory as an inspectable evaluation example, not evidence of clinical readiness.