Source-linked AI summary

Toward Personal Intelligence Through Cooperative Observation

Yashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, Osmar R. Zaiane

arXiv:2608.17128v1cs.AIcs.CYcs.HC

TL;DR

Personal AI is limited by incomplete, task-dependent observation of users’ goals and commitments, while the governance of richer observation remains unresolved. The paper proposes cooperative observation as a user-governed framework and reports a preliminary six-month Organizm deployment whose single-user record cannot identify causal effects or support generalization.

  • Problem

    Personal AI lacks complete, task-relevant context, and richer observation leaves unresolved how information is selected, compressed, and governed under user control.

  • Method

    The paper proposes cooperative observation, formalizes task-conditioned context allocation, and describes Organizm’s user-owned memory and feedback-driven planning.

  • Results

    Organizm’s six-month single-user deployment records changes in reporting and planning inputs but cannot determine why those changes occurred.

  • Takeaways & Limitations

    Evaluation should measure whether task-relevant context improves assistance and how assistance quality affects later sharing, correction, and revocation.

  • Takeaways & Limitations

    The six-month naturalistic deployment covers one user, so its observations do not identify causal effects or support generalization.

Abstract

from arXiv · show

A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a bounded system must select and compress information for the task at hand. We argue that this observation bottleneck has a cooperative structure: the system builds a partial model of the user's changing life, the user evaluates its actions, and the user's consent and control shape what it can observe next. Useful and inspectable behavior can give users a reason to maintain or expand the observation channel, while failures can lead them to correct, narrow, revoke, or abandon it. We use the term cooperative observation for this feedback loop among usefulness, trust, and future access, and propose it as a framework for personal intelligence. We report a preliminary single-subject account from Organizm, a prototype used over six months, and outline evaluation directions for measuring how observation quality shapes personal AI.

1 Introduction

Personal intelligence is bounded by the lossy observations through which an assistant learns a user’s goals, constraints, habits, and commitments. The paper proposes cooperative observation as a user-governed framework linking assistance quality with changes in future access to information.

  • Observation bottleneck: Personal intelligence depends on user context, but the assistant’s model is bounded by a lossy stream of reports, sensor readings, and interactions.Relevant context supports planning, inconsistency detection, and decisions across time.
  • Observation bottleneck: Agents need richer context to act personally, while emerging devices capture context that becomes useful when an agent can act on it.This creates an observation bottleneck approached from both software and sensing hardware.
  • Observation bottleneck: Existing health data can aid personal assistance, but platform-specific exports, APIs, and permissions often restrict an assistant’s access to users’ bodily and daily-life signals.Richer sensing devices expand what can enter the observation channel, while practical transfer may remain unavailable.
  • User governance: Richer observation can deepen surveillance, manipulation, and dependency when systems optimize platform objectives such as engagement, retention, and revenue rather than user-set goals.The paper contrasts these risks with modeling capacity directed toward goals the user sets.
  • Contribution: Cooperative observation frames personal intelligence as user-governed: assistance quality and changes to the observation channel are evaluated together across a spectrum from deliberate reports to continuous and neural observation.The framework also treats context allocation as selecting and representing user information for decisions under processing and disclosure constraints.

2 The Rise of Personal AI Systems

Personal AI systems are emerging from advances in tool-using language models, persistent memory, planning, feedback, efficient deployment, and richer sensing. Their effectiveness also depends on continual learning and a user-controlled observation channel that can expand with useful behavior or narrow after failures.

  • Capabilities: Tool-using language models now support assistants that plan, maintain memory, incorporate feedback, and act across sustained interactions.These capabilities build on general-purpose interfaces over tools, files, and communication channels, alongside behavior shaped by human preferences and instructions.
  • Capabilities: Recent agent systems combine language models with memory, planning, and reflection, while personal-agent projects add persistent memory, inspectable execution, and user-chosen models.OpenClaw-style ecosystems also share reusable skills that extend agent capabilities.
  • Deployment: More capable open-source models and architecture-aware scaling can improve personal agents’ interpretation and action over fixed histories while keeping more inference on user-controlled devices.An analysis of 51 open-source base models found that maximum capability per parameter increased sharply over time, while matched-budget experiments improved accuracy and inference throughput.
  • Observation: Smart glasses, wearable audio, health sensors, and neural interfaces widen the observation channel available to personal AI beyond typed text.Neural interfaces have begun decoding attempted speech directly from neural activity.
  • Continual learning: Personal intelligence requires continual learning because users’ goals, preferences, and constraints change, making persistent memory and feedback important for separating stable preferences from temporary noise.Continual-learning research addresses absorbing new information without overwriting existing knowledge.
  • Cooperative observation: Users may expand observation from explicit tasks to priorities, constraints, project history, calendars, wearables, or continuous capture after useful plans, but narrow or close it after failures.The framework treats user control over these observation changes as a requirement.

3 The Cooperative Observation Framework

The cooperative observation framework models personal AI as acting under partial observability while learning which parts of the user’s changing state matter through feedback. It links observation limits to a user-controlled feedback loop in which useful, reflective-beneficial assistance can sustain or expand access, while failures can narrow or revoke it.

  • State: Personal AI must infer task-relevant goals, constraints, commitments, and changing internal or external conditions from partial observations of the user’s evolving state.Which state components matter depends on the task, creating an information-bottleneck problem.
  • Framework: The framework combines a decision model, an information-theoretic observation bound, and a feedback loop through which user experience changes future observability.The user evaluates actions through immediate feedback and can change the observation channel for subsequent steps.
  • Information bound: State-dependent actions are bounded by the information history available to the system, so expanding the observation channel can raise the ceiling on personalized behavior.Changing the model or policy can better use information already present but cannot add user-specific information to the existing history.
  • Cooperative observation: Cooperative observation is joint adjustment of access to task-relevant context because doing so benefits the user, with evaluation and access decisions remaining under the user’s control.The user may provide a task, correction, or sensor stream, then share more, correct the model, or withhold data after evaluating the system’s action.
  • Feedback dynamics: Useful assistance may motivate users to maintain or expand observation, whereas effort, privacy costs, inaccurate actions, or limited trust may lead them to narrow or revoke access.The loop can be cooperative only when access remains user-controlled and improves reflective benefit; continued use or disclosure alone is insufficient evidence.

4 Observation Channels and the Complementary Self-Model

Section 4 frames personal-AI observation as a spectrum from deliberate self-report to automated, continuous multimodal, and neural channels, with broader access increasing both available information and privacy, consent, and autonomy demands. These observations can support a complementary self-model, but the system must select and compress history for each task while the user retains control over what it observes, remembers, and does.

  • Observation-channel spectrum: Broader or more continuous observation does not guarantee usefulness because added bandwidth may not contain information relevant to the task.Observation fidelity can increase through frequency, additional modalities, or direct measurement, yet task-specific usefulness still requires selecting relevant information.
  • Observation-channel spectrum: Observation channels range from manual reports through automated behavioral context and continuous multimodal capture to neural interfaces.Later phases can broaden the observation stream and reduce reliance on self-report, but they also increase privacy, consent, and governance demands.
  • Complementary self-model: Persistent records can complement self-knowledge by revealing changes over time, recurring problems, neglected projects, and gaps between stated priorities and observed behavior.Calendars, activity traces, and multimodal capture strengthen comparisons across weeks or months by adding evidence and context that written reports may omit.
  • Complementary self-model: A complementary self-model may support remembering, reflecting, and deciding, but the user must control what the system observes, remembers, and does.The Good Regulator theorem motivates modeling the goals, commitments, and constraints relevant to the user’s decisions.
  • Complementary self-model: A personal AI must select and compress its observation history because broader observation can exceed what it can use in a single decision.The information bottleneck frames this as a trade-off between retaining a compact representation and handling a world larger than the system can fully perceive or represent.

5 Organizm: A Prototype for Cooperative Observation

Organizm is a user-controlled prototype that couples hierarchical agents with hierarchical personal records to allocate task-appropriate context. Its six-month single-subject deployment illustrates expanding, voluntary observation and user correction, but cannot establish causation or generalization.

  • User control: The system stores logs, notes, and summaries as ordinary files in a user-owned hierarchy that users can inspect, edit, or delete.Folder- and project-level indexes provide compressed representatives before agents read more detailed files.
  • Architecture: Organizm couples hierarchical multi-agent personas with matching information tiers, assigning task-appropriate context scopes from immediate processing to long-term reflection.Index-first traversal supplies finer detail when the initially matched tier is insufficient.
  • Cooperative observation: Daily plans, weekly reviews, and user corrections turn self-report into a recurring observation channel that calibrates planning and deadline escalation.Rejecting overambitious plans teaches planning limits, while feedback on anxiety-inducing reminders teaches deadline escalation.
  • Six-month deployment: During January–June 2026, manual logging became near daily by May–June as weekly reviews and monthly summaries were added, but the single-subject record cannot identify causes or support generalization.The report describes this as an illustrative sketch of one author’s first six months using Organizm.
  • Six-month deployment: The observation channel expanded voluntarily within Phases 1–2 through calendar events, a deadline registry, review layers, and a finance channel, without wearables or continuous capture.Each addition had a specific purpose and remained in user-owned files that could be inspected or deleted.

6 Evaluation Directions

The paper proposes three complementary evaluation studies for cooperative observation: observation-channel ablation, longitudinal usefulness and channel-change tracking, and self-model evaluation. Together, these studies examine how personal context affects assistance, how assistance affects future observation, and how system and user knowledge change.

  • Evaluation Directions: Three studies evaluate cooperative observation through assistance quality, longitudinal changes in sharing and consent, and changes in system and user self-knowledge.The proposed studies are observation-channel ablation, longitudinal deployment, and self-model evaluation.
  • Observation-channel ablations: An observation-channel ablation would hold the model, tools, prompt template, and sampling settings fixed while testing participant-specific tasks across predefined categories.Tasks would have checkable constraints or outcomes and include categories such as scheduling, prioritization, planning, and forecasting.
  • Usefulness and channel change: A longitudinal deployment would relate perceived usefulness and task performance to whether users later add, retain, reduce, or revoke information sources.The study would also track changes in the types of tasks users bring to the system, because the feedback loop may alter the observed task distribution.
  • Complementary self-model evaluation: A self-model study would compare task-only and task-plus-record predictions of verifiable outcomes while assessing users’ forecasts, planned work, and stated goals.User self-knowledge measures include forecast accuracy for time allocation and project delays, plus consistency between planned work and stated goals.

7 Safety and Governance

Safety and governance for personal AI must address the expanded consequences of richer observation, including privacy, inference, manipulation, delegated authority, and user control. The paper calls for purpose-scoped permissions, inspectable execution, safeguards against dependence and multi-agent risks, and durable revocation and portability.

  • Safety and Governance: Richer observation expands both personal AI capabilities and the consequences of misuse or failure, making retained inferences, channel integrity, feedback, authority sharing, and access withdrawal central safety concerns.Because users supply feedback and control access, personal AI can serve as a small-scale testbed for alignment under uncertainty about user values.
  • Observation privacy and integrity: Raw records and derived inferences require equal privacy protection, while continuous audio and visual observation also raises consent concerns for bystanders.Users may share sleep data without anticipating health-condition inference, and breaches can expose both source data and conclusions.
  • Observation privacy and integrity: Permissions should be scoped to purposes and inference types, disclose retained conclusions, and support pausing, deletion, and revocation for raw signals and derived representations.Raw neural access does not authorize decoding inner speech, attention, affect, or other latent attributes.
  • Observation privacy and integrity: Every added channel broadens the attack surface, so personal AI needs code review, signing, sandboxing, revocable permissions, and containment against malicious observations and shared skills.Malicious content can redirect behavior, while shared skill instructions may request access to tools or data.
  • Execution transparency and auditability: Users need inspectable execution records and safeguards against optimizing immediate approval, including auditing, delayed-outcome evaluation, correction, and revocation.Immediate approval can reward flattery, avoidance of hard truths, or increased dependence; execution provenance connects evidence, tools, memory, and agent actions.
  • Multi-agent composition: Governance must account for multi-agent capabilities and risks, affected people’s consent, and durable control through revocation propagation, memory correction, and portable user data.Interacting agents can create system-level properties and infer information about communities or people who never contributed data.

8 Limitations

The framework’s limitations concern burdensome and ambiguous feedback, underspecified evaluation and context allocation, changing task distributions, unevaluated architectural coupling, and limited evidence from one user. Future work requires provenance-aware ablations and controlled multi-user studies to isolate effects and assess generalization.

  • Feedback and channel attribution: Cooperative observation depends on consistent user feedback, but burden and privacy concerns can restrict sharing and obscure which channel requires correction.Multiple active channels and sparse or inconsistent feedback make channel changes harder to interpret.
  • Feedback and channel attribution: Channel-level provenance combined with controlled ablations could help isolate the effects of individual observation channels.This is proposed as a direction for future studies.
  • Evaluation and allocation: Reflective benefit, trust, and task relevance require operational definitions specific to the user, task, and evaluation horizon.Task-conditioned context allocation also remains a design problem requiring comparisons under processing and disclosure constraints.
  • Evaluation and allocation: The delegated task distribution may change as users learn which tasks the system handles well, complicating performance comparisons over time.This feedback between system behavior and later inputs is related to performative prediction.
  • Architectural evaluation: Organizm’s coupled agent and memory hierarchies remain unevaluated against flat memory or a single agent operating over the same files.Future ablations should test whether coupling helps and by how much.
  • Evidence scope: The six-month deployment sketch covers one user under naturalistic conditions, so its reporting frequency, channel additions, and corrections identify neither causal effects nor generalization.Controlled multi-user studies are required for both.

9 Conclusion

The conclusion frames personal intelligence as cooperative observation: task-conditioned context allocation supports action, while user evaluation shapes future access. It calls for evaluation of assistance, sharing, correction, revocation, and changes in system and user knowledge, while noting Organizm’s single-user deployment cannot establish causal explanations.

  • Cooperative observation: Cooperative observation links task-conditioned context allocation, system action, and user evaluation in a feedback loop.The system selects sources and representations under processing and disclosure constraints, then user evaluation shapes what it can observe next.
  • Evaluation: Evaluation should test whether task-relevant context improves assistance and how assistance quality affects later sharing, correction, and revocation.It should also examine how the system’s knowledge and the user’s self-knowledge change over time.
  • Organizm and future directions: Organizm’s six-month single-user deployment records changes in reporting and planning inputs but cannot determine why those changes occurred.The conclusion calls for controlled multi-user studies to test these causes and identifies open software, local-first memory, portable data, and auditable execution as supports for user control.
Loading 2608.17128v1…