Source-linked AI summary
Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
Yan Tang, Tingyu Cao, Yuanbo Tang, Huaze Tang, Keer Hu
TL;DR
Proactive service asks when an agent should initiate help rather than treating an explicit instruction as the fixed starting point. The survey formalizes that choice as a constrained partially observable decision process, unifies methods and evaluation around it, and concludes that deployment benefit requires calibrated intervention value, authorization, recoverability, and counterfactual evidence.
Problem
Existing agent evaluations usually start from complete instructions, leaving the upstream decision of whether and when to initiate help under uncertainty insufficiently unified across settings.
Method
The survey defines instruction-relative proactive intervention, models silence, asking, assisting, and acting as structured actions in a constrained sequential process, and organizes methods, resources, and metrics around one decision pipeline.
Results
The synthesis finds that offline classification performance, fluent advice, or longer memory alone does not establish deployment benefit, while evidence must address intervention value relative to silence, authorization, safety, and longitudinal outcomes.
Takeaways & Limitations
Reliable proactive service requires calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence rather than more frequent action.
Takeaways & Limitations
The reviewed evidence is concentrated in constructed samples, offline replay, and short-term simulation.
Abstract
from arXiv · showhide
Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete environmental and user signals, choose among remaining silent, asking, assisting, and acting, and account for interruption, misunderstanding, overreach, and privacy costs. This survey gives an operational definition centered on initiative and formulates the problem as a partially observable sequential decision process constrained by authorization and risk. The formulation represents timing, content, and delivery within one structured action, while making explicit the option value of waiting, the decision value of questions, and feedback-induced state changes. On this basis, we organize existing methods along one decision pipeline (state and need estimation, intervention gating, action construction, and feedback adaptation) and describe prescribed, predictive, model based, and return optimizing mechanisms as nonexclusive policy-construction components. We further normalize decision units and three-axis evidence descriptors across streaming dialogue, screen, video, software-engineering, and human-agent collaboration resources, and formalize metrics for triggering, timing, calibration, user burden, safety, and policy value. The synthesis shows why offline classification performance alone does not predict deployment benefit and why long-term memory is not a defining condition of proactivity. Reliable proactive service instead requires calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.
I. INTRODUCTION
Proactive service reframes assistance as a decision about when to initiate help under uncertainty, rather than as more frequent action or dialogue capability. This survey unifies that decision across modalities through a constrained pipeline, standardized evidence, and evaluation focused on deployment value.
- Motivation: Proactive service decides when to assume initiative by weighing evidence that intervention is more valuable than waiting against interruption, privacy, and overreach costs.The survey distinguishes initiating help from simply generating unsolicited content.
- Relation to prior work: Prior work covers mixed-initiative control, proactive dialogue, recommendation, retrieval, and adaptive intervention, but remains fragmented across tasks and interaction settings.These lines of research separately address control transfer, clarification, preference elicitation, and context-dependent triggering.
- Survey framework: The survey models silence, inquiry, assistance, and execution within one constrained partially observable decision process, jointly analyzing timing, content, and delivery.Its organizing object is an instruction-relative decision under uncertainty rather than a modality-specific application.
- Evaluation: It proposes common decision units, evidence descriptors, and metrics spanning triggering, timing, calibration, burden, safety, and policy value.The evaluation protocol is intended to distinguish deployment value from isolated trigger or content performance.
- Survey framework: The framework separates observable task specifications and permission ledgers from latent beliefs, allowing feedback to update beliefs while verified permission events update authorization.The figure treats the permission ledger and belief state as distinct components of the loop.
II. PROBLEM SETTING AND FORMAL FOUNDATIONS
The paper defines proactivity as an instruction-relative, silence-gated choice made under incomplete information, then distinguishes that decision object from dialogue capability and implementation mechanisms. Its formal and comparative framing emphasizes admissibility, authorization, recoverability, and evaluation across decision opportunities.
- Proactivity as a decision property: A proactive intervention autonomously introduces a service opportunity, timing, subgoal, information request, or external commitment not entailed by the current instruction.The definition also requires a feasible silence or alternative action and selection from beliefs about the user, environment, and consequences.
- Comparative framing: The decision unit is a decision opportunity, not necessarily a dialogue turn, and comparison tables distinguish scope, organizing axis, proactivity notion, evaluation, and safety.This supports comparison across modalities while retaining the open-stream decision problem.
- Prior survey scope: Earlier surveys primarily frame proactivity as leading text dialogue toward goals, across dialogue types, subtasks, and interaction flows.Their scope includes open-domain, task-oriented, and information-seeking dialogue, plus adjacent recommendation systems.
- Review scope: This review broadens the scope to dialogue, graphical interfaces, video, wearables, embodied systems, software engineering, and human service.It also contrasts offline and user-centric conversational-recommendation metrics with its broader evaluation focus.
- Review contribution: Its organizing principle is an instruction-relative, silence-gated decision under uncertainty, with common units and metrics for triggering, timing, calibration, coverage–risk, and off-policy value.The framework separates the comparable decision object from application, task, and model mechanism.
- Safety and action space: Admissibility combines authorization and severe-risk constraints, while tiered permission and recoverability limit how initiative can affect external state.The figure presents this constrained action space as the basis for the proactive-service loop.
- Scope of the definition: The paper does not require long-term memory, continuous sensing, complex planning, or tool use to establish proactivity; the defining property is online autonomous choice relative to an instruction.A pre-set timer is excluded when its trigger is fully fixed, whereas clarifying an underspecified request can qualify.
B. Constrained Partially Observable Sequential Decision Making
The survey models proactive service as a constrained partially observable sequential decision process in which hidden user and task states evolve after intervention. Actions, rewards, costs, authorization, and feedback are represented together while verified permission and severe-risk limits remain hard feasibility conditions.
- The framework uses a constrained partially observable Markov decision process because goals, needs, interruptibility, and authorization are usually hidden while actions change future states.The model is an analytical language rather than an assumption that every existing system explicitly solves it.
- The latent state tracks task environment, goals, needs, interruptibility, and endorsed long-term outcomes, while the permission ledger separately records verifiable authorization.The belief summarizes uncertainty over latent state; the ledger is protocol state rather than a model guess.
- An action is structured as mode, content, and delivery, jointly encoding whether and when to intervene, what to provide, and how to deliver it.Modes are silent, ask, assist, and act; silence permits further observation, while act may invoke tools or change external state.
- Immediate and long-term benefits are balanced against interruption, questioning, execution or rollback, and privacy costs through a relaxed reward.Fixed nonnegative multipliers convert these costs into comparable one-step utility, but relaxed policies require separate budget checks when feasibility conditions fail.
- Verified authorization and severe-risk bounds constrain admissible actions, whereas interruption, privacy, and questioning are handled through costs or cumulative budgets.External execution is inadmissible when authorization is unknown, although asking for authorization can remain admissible.
C. Intervention Advantage, the Value of Waiting, and the Value of Asking
The framework treats intervention as worthwhile only when the best admissible non-silent action exceeds the value of remaining silent. It also values questions by their immediate effects and by how answers change beliefs, authorization, and later decisions.
- The intervention advantage is the value of the best admissible non-silent action relative to silence under the current belief and permission ledger.Conservative tie-breaking keeps the agent silent when this incremental value is nonpositive.
- Silence retains option value because it permits further observation and later intervention rather than permanently abandoning service.First intervention time is defined as the earliest time at which the selected action is non-silent.
- Need probability alone is insufficient for triggering because candidate content, execution risk, and downstream value also affect whether intervention is advantageous.The trigger is a decision about positive intervention advantage, not merely a prediction that a need exists.
- The total value of asking combines immediate interaction effects with the discounted value of decisions improved by the answer.The question value averages over possible answers and resulting posterior beliefs and ledgers.
- Pure information value is nonpositive when an answer changes neither belief, authorization, nor downstream action values, although asking may still have relational effects.Only protocol-verified consent or revocation changes the permission ledger.
III. A DECISION-CENTERED SYNTHESIS OF METHODS
The survey organizes proactive-service methods around a four-stage decision loop rather than by domain, modality, autonomy, or training paradigm. It treats policy-construction mechanisms as orthogonal signatures that can combine across stages.
- Decision pipeline: A complete proactive-service system converts history into a belief state, gates intervention, constructs and executes an action, and adapts from feedback.Representative systems may cover only part of this loop, so coverage distinguishes decision mechanisms from supporting technology.
- Decision pipeline: Sensing modality, application domain, autonomy level, and training paradigm are orthogonal properties rather than parallel classes in the functional axis.The pipeline is intended to compare systems across heterogeneous settings.
- Policy construction: Prescribed, predictive, model-based, and return-optimizing mechanisms describe how policy mappings are constructed and may span or combine across pipeline modules.Their formulas and descriptors are mechanism signatures, not rankings or claims of empirical coverage.
- Policy construction: The synthesis keeps hard authorization and severe-risk feasibility distinct from soft costs such as interruption, asking, privacy, and execution or rollback.This distinction prevents selectable but expensive actions from being treated as actions that must not execute without consent.
A. From Observation to a Decision-Ready Belief State
The first pipeline stage converts streaming observations into a decision-ready belief about opportunities, need, urgency, and interruptibility. Long-term memory can shape that belief for recurrent or personalized situations, but it is not required for every proactive intervention.
- Opportunity detection: Streaming systems must segment observations into decision opportunities because most moments contain no service event.The stage distinguishes task signals from background activity and noise across screens, video, and sensor input.
- Need and uncertainty: Need estimation must include urgency, interruptibility, and uncertainty rather than only an intent label.The unified framework evaluates both severe-risk probabilities and the expected value of waiting.
- Memory: Memory supports cross-session preferences, recurrent tasks, and personal routines by changing how the belief state is constructed.Hierarchical storage, trajectory memory, and longitudinal personalization are distinct memory approaches described in the synthesis.
- Memory: Long-term memory is not necessary for proactivity because an immediate hazard warning may rely only on current observations.Memory becomes relevant when the service opportunity depends on history such as preferences or recurring routines.
B. How Policies Are Constructed: Four Non-Exclusive Mechanisms
The survey treats prescribed, predictive, model-based, and return-optimizing mechanisms as stackable components rather than mutually exclusive policy classes. Their usefulness depends on whether the action space and feedback protocol support proactive gating, not merely content selection.
- Mechanism taxonomy: Four labels describe whether a policy is directly specified, learned from labels, based on explicit dynamics, or optimized for sequential return.A compound system may receive several labels simultaneously.
- Mechanism trade-offs: Prescribed and predictive mechanisms are computationally light but struggle to represent the option value of waiting.They commonly compress future effects into expert rules or labels.
- Mechanism trade-offs: Model-based mechanisms expand transition, observation, or search structures, but their evaluations remain sensitive to user-model error.They can explicitly evaluate downstream consequences through the model or search tree.
- Interpretation: Table II presents these mechanisms as non-exclusive components, not evidence that every study implements the complete four-mode gate.The labels characterize representative components rather than complete systems.
- Mechanism trade-offs: Return optimization can improve long-horizon content decisions without learning proactive mode selection when training lacks silence, rejection, or authorization violations.EPO and CSO optimize content in settings where a speaking turn is already available.
C. Joint Intervention Gating and Action Construction
Proactive service separates intervention gating from action content and delivery, while treating silence, inquiry, assistance, and execution as jointly selectable possibilities. This framing exposes authorization, reversibility, burden, and uncertainty as mode-specific trade-offs.
- Policy decomposition: The gate chooses mode and time, the content policy chooses the specific intervention, and the delivery policy chooses channel, modality, and explanation.Autonomy and external side effects are represented through mode and permission.
- Scope boundary: The reviewed content-construction experiments often grant the system a speaking turn, leaving open-stream intervention selection underrepresented.This distinction separates content construction from deciding whether to intervene.
- Scope boundary: No formal proceedings record was verified for one cited item, so the survey does not treat it as published evidence.This is an evidence-status limitation rather than a performance finding.
- Mode trade-offs: Asking can reduce epistemic uncertainty but adds burden; assistance preserves control but may create overreliance; execution saves effort but requires permission and rollback.The modes are not a simple linear autonomy scale.
- Mode trade-offs: Downstream policy-search methods should include silence and feasible authorization levels among root candidates to implement proactive gating.Otherwise, comparisons begin after intervention has already been selected.
D. Feedback, Personalization, and Continual Adaptation
Feedback changes proactive decisions at immediate, task, and cross-session timescales, but intervention-only observations create selective feedback and complicate policy evaluation. The survey therefore recommends complete decision units, explicit timing and calibration measures, and sequential off-policy estimators with strong logging assumptions.
- Feedback timescales: Immediate feedback updates current beliefs, task outcomes update content and execution value, and cross-session signals update the long-term user model.Memory refresh or retrieval is distinct from controlled online gating-policy updates.
- Feedback timescales: Intervention-only outcomes create selective feedback because silence reveals neither whether help was needed nor the counterfactual result.Visible acceptance can favor easy-to-accept moments and overlook potential beneficiaries.
- Decision units: A comparable study should define complete streams of decision opportunities, including observations, valid service intervals, permissible actions, authorization, risk, confidence, and logging probabilities.Positive-only collections cannot evaluate silence and timing adequately.
- Evidence gaps: Offline scores do not establish long-term net utility because current resources lack comparisons that identify long-term incremental intervention value.The evidence gaps include offline replay, proxy need labels, limited human comparisons, and missing individual longitudinal outcomes.
- Evaluation decomposition: Calibration evidence should accompany discrimination metrics because sparse-stream accuracy can be achieved by always remaining silent.Recommended measures include Brier score, calibration curves, and expected calibration error with confidence intervals.
- Timing evaluation: Timing evaluation matches predictions to valid intervention intervals, minimizes timing error among maximum-cardinality matches, and reports early or late signed error.Repeated alerts match only once, and fixed frame tolerances can replace utility with annotation convenience.
- Evaluation decomposition: Four-mode selection should be evaluated separately from content, with macro-F1 or cost-weighted confusion matrices for modes and task-specific metrics for content.End-to-end evaluation should also capture error propagation from false triggers.
- Policy value: Policy-value evaluation must distinguish intervention outcomes from silence and account for burden, privacy, and risk through declared scalarization weights.A positive incremental value means benefit exceeds added cost under those weights.
C. Risk, User Burden, and Reproducibility
Risk and burden evaluation must examine selective intervention across thresholds rather than report only a best operating point. Reproducibility also requires fixing model, prompt, tool, evaluation, and privacy-related conditions.
- Risk and burden: Selective-intervention evaluation should report coverage and conditional risk as the selection threshold varies.Coverage measures the fraction of opportunities selected, while risk is mean loss among covered interventions.
- Risk and burden: Safety reporting should include unauthorized actions, severity-weighted harm, near misses, rollback success, sensitive-data exposure, and explanation–action consistency.Risk assessment should cover both incidence and severity.
- Risk and burden: User burden should include intervention frequency, question turns, rejection or ignoring, task-resumption time, and perceived interruption.Short-term clicks or likes are not proxies for long-term trust.
- Reproducibility: Reproducible experiments require fixed model versions, prompts, permissions, history windows, inference budgets, graders, and random seeds.LLM assessment should also report rubrics, randomization, human agreement, and judge-model sensitivity.
V. FROM METHODS TO DEPLOYMENT REGIMES
The survey maps proactive-service work across three deployment regimes using shared decision variables while preserving regime-specific constraints. Differences in observation, reversibility, permissions, feedback, and timing change which methods and interventions are admissible.
- Table IV organizes deployments into three regimes using observation continuity, time window, action reversibility, permission, and feedback delay.
- Digital workspaces permit richer action sets when actions are sandboxed, previewable, and reversible, but reversibility does not eliminate privacy or other costs.
- Software-engineering clarification separates detecting underspecification, asking a useful question, and exploiting its answer into distinct bottlenecks.
- Contextual and embodied systems couple timing and perception errors, making late assistance ineffective and incorrect physical actions potentially harmful.
- Affective-support methods incorporate affective trajectories, dialogue stage, and support strategy into state and action construction, but some evidence remains preprint-based.
B. Encoding Safety in the Action Space and Interaction Design
Safety is encoded directly in proactive agents’ action choices through permissions, progressive autonomy, recoverability, auditing, and data minimization. The survey argues that deployment evidence must assess intervention value and adaptation under authorization, risk, and longitudinal outcomes rather than rely on offline scores alone.
- Tiered permissions and authorization bind actions to reading, drafting, reversible modification, communication, or irreversible commitment, with object, purpose, and duration specified.
- Progressive autonomy lets agents degrade from acting to assisting or asking under low confidence or high risk, while recoverability requires previews, scope controls, undo logs, and human takeover.
- Safety mechanisms alter the optimal policy: reversibility lowers Cexec, authorization enlarges Aperm(¯κt), on-device processing lowers Cpriv, and repeated confirmation raises Cask.
- Counterfactual intervention value should be learned through guarded real decision opportunities, such as micro-randomized trials, encouragement designs, or conservative exploration.
- Offline AUROC is an intermediate diagnostic rather than deployment evidence.
- Joint calibration should compare independent need classification with approximate joint action-value modeling and report calibration, coverage–risk, and overreach.
- Adaptation evaluation should test preference drift, stale-goal forgetting, authorization violations, deletion controls, and both postadaptation and old-scenario performance.
- Cross-regime benchmarks should share decision-opportunity records and policy-value measures while separately reporting interaction realism, comparison design, and human-outcome horizon.