Source-linked AI summary

Why Organizational Rules Fail AI: O-I-B-A-R and the Externalization of Decision Boundaries

Chao Li, Chunyi Zhao

arXiv:2608.29055v1cs.CY

TL;DR

Organizational AI failures can arise when systems receive explicit procedures without the negative boundaries, runtime judgments, responsibilities, and learning history used in situated practice. The paper introduces O-I-B-A-R to externalize these elements through concrete failures, decision dimensions, suspensions, actions, and results. It argues that the framework's usefulness depends on preserving candid failure histories despite their organizational costs and incentives.

  • Problem

    Organizational AI can fail when formal rules omit contextual boundaries and the runtime values, responsibilities, and learning history required for real-world decisions.

  • Method

    O-I-B-A-R externalizes decision boundaries by comparing applicable judgments with concrete failures, identifying changing dimensions, and representing unresolved values as actionable queries.

  • Results

    The framework distinguishes negative-boundary cases from unresolved-value cases, which one-sided rule representations collapse by construction.

  • Takeaways & Limitations

    O-I-B-A-R treats human–AI handoffs as judgments over specific unresolved decision values rather than as undefined escalation to a human.

  • Takeaways & Limitations

    The framework's decomposition target is operational rather than proof that failures have one cause, and durable failure records may suppress candid reporting.

Abstract

from arXiv · show

AI systems increasingly enter organizations through policies, procedures, playbooks, prompts, and other explicit representations of work. Yet formal descriptions often differ from situated practice, and captured know-what can omit the contextual know-how experts use when judgments are uncertain. We argue that a recurring class of organizational AI failures arises partly from a knowledge representation problem at the sociotechnical interface: the AI receives the procedure, while the organization operates on the procedure plus negative boundaries, runtime judgments, responsibility assignments, and learning history. We introduce O-I-B-A-R (OPEN, IS, BUT, ACTION, RESULT), a scaffold for externalizing these missing decision boundaries. IS records when a judgment holds. BUT records a concrete failure containing information beyond the logical negation of IS. Comparable success and failure cases are decomposed toward a minimally sufficient changing variable, which becomes a value-bearing decision dimension. A suspension represents the state in which the dimension is known but its current value is unresolved, specifying what must be measured, asked, retrieved, or escalated to a human. RESULT confirms a boundary, shifts a threshold, or exposes a new dimension. Incidents can generate new dimensions, unresolved values can define human-AI handoffs, and feedback can expand the decision space. We also identify a sociotechnical tension: durable and attributable failure histories can suppress the candor on which useful boundary knowledge depends. Externalization must therefore be designed as an organizational intervention with real costs and incentives.

1 Introduction

Organizational AI can fail because explicit procedures omit the boundaries, runtime values, responsibilities, and learning history that guide situated practice. O-I-B-A-R addresses this representation problem by externalizing concrete failures, decision dimensions, unresolved values, and accountable human judgments.

  • Externalization can improve organizational learning while increasing attribution exposure and incentives to sanitize incidents, making candor part of the system's data-generating process.
  • Organizational AI receives canonical procedures, while experienced practice also depends on exceptions, contextual judgments, and repair history.
  • O-I-B-A-R records when a judgment holds, where it concretely fails, what changed, what value must be resolved, and how results revise the representation.
  • The paper identifies four representational gaps: boundary, runtime-value, responsibility, and learning.
  • Suspensions turn known-but-unresolved decision dimensions into operational queries and can specify what a human must substantively judge.

3 The O-I-B-A-R Framework

O-I-B-A-R externalizes organizational decision boundaries by starting with a testable judgment and domain, contrasting applicability with concrete failure, and converting the contrast into actionable dimensions and queries. The loop then commits a decision and uses evidence to refine or expand the represented decision space.

  • OPEN: OPEN specifies a decision-relevant judgment and the domain in which it can be tested.Topic labels are insufficient because the domain constrains which distinctions are useful and where abstraction should stop.
  • IS: IS records the conditions under which the current judgment is presently justified.It expresses an applicability boundary without claiming logical necessity.
  • BUT: BUT is a concrete failure whose explanatory information can exceed the logical negation of IS.A useful BUT identifies a condition and mechanism that help explain why the same judgment crosses into an undesirable region.
  • Boundary pairing: Comparable IS and BUT cases are decomposed toward one minimally sufficient changing variable, which becomes a value-bearing decision dimension.Pairs with no changing variable usually restate IS, while pairs with multiple changing variables should be decomposed.
  • Suspension: A suspension turns a known dimension with an unresolved current value into a query answerable through measurement, retrieval, questioning, observation, or human judgment.The query specifies the missing information and can define the substantive content of a human role.
  • ACTION and RESULT: ACTION commits the current boundary model to a decision, while RESULT can confirm a boundary, refine a threshold, or introduce a new dimension.Action uses measured or assumed dimension values; results provide structured feedback beyond a simple success/failure label.

4 The O-I-B-A-R Loop and Local Stability

The O-I-B-A-R loop feeds results back into new OPEN states, allowing boundaries and dimensions to evolve through observed cases. Local stability concerns a bounded case stream, but no-dimension runs do not establish completeness when exposure or reporting conditions fail.

  • The loop: The loop proceeds from OPEN through IS/BUT, dimension discovery, query, ACTION, and RESULT, with later results generating new OPEN states.A failed result may open a more specific question or make a newly discovered variable the next analysis object.
  • Local stability: Structural stability means observed results no longer expand the active dimension set or materially move existing boundaries for a specified case stream.Stability can coexist with unresolved runtime variables and does not claim that every world-relevant dimension has been discovered.
  • Local stability: No new dimensions across cases do not by themselves demonstrate convergence because rare dimensions can produce long waiting times.A stopping statement requires explicit assumptions about the exposure process.
  • Local stability: A declining discovery rate can reflect silence-induced false stability when failure incidents are not reported, rather than a maturing representation.The bound becomes invalid under distribution shift, dependent observations, selective exposure, or missing failure reports.

5 Diagnostics: When the Scaffold Fails Informatively

O-I-B-A-R treats incomplete fields as diagnostic signals about the quality of operational knowledge, rather than merely formatting defects. The resulting diagnoses direct narrower judgments, concrete incidents, dimension separation, typed questions, or explicit human responsibility.

  • Diagnostic principle: An incomplete O-I-B-A-R field provides diagnostic information instead of constituting a formatting failure.The observed pattern indicates what kind of operational knowledge or modeling problem requires attention.
  • Observed patterns: An empty IS and BUT indicate either malformed OPEN or absent operational knowledge, requiring a concrete judgment and experienced cases.A topic must be made testable before boundary analysis can proceed.
  • Observed patterns: A fluent IS with an empty BUT may indicate an overbroad domain, imported slogan, or definition rather than a falsifiable decision model.Narrowing the domain can make concrete failures recallable.
  • Observed patterns: A rich BUT with an empty IS indicates experience without a positive model, so analysis should work backward from failures to success conditions.When multiple variables change, the boundary pair should be split into separate pairs.
  • Observed patterns: A suspension phrased as “it depends” leaves the unknown untyped, whereas an unresolved dimension identifies the fact that would decide the case.When the suspension cannot be pre-resolved, runtime human judgment may be required.
  • Observed patterns: Dimension names such as “factor” or “attribute” signal that abstraction has lost a value-bearing type.The method instead defines a role by the variable being judged.

6 Worked Examples

The worked examples show how concrete failures reveal decision dimensions that formal rules omit. Customer-service cases turn negative boundaries into human-AI handoffs, while heap-selection cases demonstrate sequential dimension growth without a predefined taxonomy; research examples move such variables into construct formation during inquiry.

  • Customer service: In customer service, repeated empathy can fail when the customer needs a concrete refund remedy, revealing need type as the changing dimension.The action is held constant while the suspension asks whether the customer needs to be heard or needs a concrete remedy or number.
  • Customer service: When the current need is a concrete remedy, ACTION moves directly to amount, conditions, and timing.A successful response supports the boundary; a rejection based on comparisons may require a new dimension such as fairness perception.
  • Top-K selection: Heap-selection analysis successively exposes scale, input form, ownership, copying, and resource budget as separate decision dimensions.Each dimension arises from a distinct failure boundary, such as streaming input, mutation of caller-owned data, copying, or memory limits.
  • Top-K selection: The heap example generates dimensions from cases rather than beginning with a predefined taxonomy.The resulting dimensions emerge from separate failure boundaries.
  • Research claims: For research claims, concrete non-improvement cases can expose teaching practice or teacher quality as dimensions absent from a generic scope disclaimer.O-I-B-A-R uses such negative evidence for construct formation during the research process rather than only as post-hoc qualification.

7 Intellectual Lineage and Relationship to Prior Work

O-I-B-A-R synthesizes established contrastive elicitation, near-miss learning, boundary refinement, and incident-based knowledge elicitation around organizational decision boundaries. Its distinctive target is an open set of decision dimensions whose unresolved values support human–AI coordination and whose feedback can expand the representation.

  • Contrastive elicitation: Unlike Kellyian elicitation, O-I-B-A-R contrasts an applicability case with a consequential failure and couples the contrast to a query for an unresolved value.Kellyian methods reveal constructs through similarity and difference, whereas O-I-B-A-R uses outcome-asymmetric contrasts tied to future decision-making.
  • Near-miss learning: O-I-B-A-R adapts the near-miss intuition to messy organizational incidents by decomposing mixed descriptions until one decision-relevant change can be examined.Its target is not a new learning principle, but an elicitation procedure for identifying a minimally sufficient changing variable.
  • Feedback and coordination: O-I-B-A-R differs from Winston’s learner by representing unresolved values as suspensions for human–AI handoff and allowing RESULT to add dimensions absent from the current set.The framework therefore couples boundary discovery to an intermediate coordination state and feedback-driven expansion.
  • Representation construction: Unlike version-space learning, O-I-B-A-R helps construct the decision-variable representation when the organization does not yet know the relevant feature or vocabulary.Conventional version-space methods presuppose an instance representation and hypothesis language; O-I-B-A-R treats the boundary pair as a way to build them.
  • Incident-based elicitation: Relative to critical decision methods, O-I-B-A-R imposes a specific transformation: pair success with failure, identify a value-bearing variable, suspend unresolved values, and update the dimension set through results.This places incident-based expert elicitation inside an explicit representation and feedback cycle.
  • Contribution relative to prior work: O-I-B-A-R’s contribution is the synthesis of established contrastive mechanisms around organizational decision dimensions, unresolved values, and human–AI coordination.The framework narrows its novelty claim rather than presenting contrastive elicitation or boundary refinement as new.

8 Sociotechnical Implications for Human–AI Organizations

O-I-B-A-R reframes organizational rules as representations of applicability, failure boundaries, and unresolved facts rather than abstract instructions alone. Its suspension state makes human involvement a specific information-acquisition handoff that can move dynamically between people and AI.

  • Rule authoring: A machine-usable organizational rule should represent when an action is justified, a verified or testable failure boundary, and the fact needed before choosing a side.The framework does not require enumerating every exception; it requires enough negative evidence to identify the dimension on which a future case turns.
  • Suspension and oversight: A suspension turns an unresolved decision value into an actionable query for retrieval, measurement, interpretation, or escalation.This distinguishes a case requiring more information from one that simply falls outside the current applicability region.
  • Human–AI coordination: The resulting human–AI boundary is a dynamically located information-acquisition boundary rather than a fixed division of tasks.Human involvement is identified by the kind of information the current system cannot reliably obtain or interpret.
  • Epistemic division of labor: The proposed epistemic division of labor assigns humans ownership of claims, incident validity, and result interpretation while AI proposes counterexamples and questions.AI activity is favored where outputs can be checked against records, while substantive organizational claims remain human-owned.
  • Tacit knowledge transfer: O-I-B-A-R provides an auditable target for extracting tacit expertise from behavioral traces while preserving explicit boundaries, queries, and feedback rules.This connects recent LLM-based tacit-knowledge extraction to a structured organizational representation.

9 The Politics of Externalizing Failure Knowledge

Externalizing failure knowledge can make contextual boundaries reusable, but durable and attributable records may reduce the candor needed to produce those failures. The governance of identity, access, accountability, contestation, and learning therefore becomes part of the organizational decision problem.

  • From war story to record: Turning situated diagnostic stories into durable O-I-B-A-R records can improve reuse while altering incentives to disclose failures.The same contextual material that remains useful in local peer practice may become a dated, attributable record of decisions and outcomes.
  • Candor and attribution: O-I-B-A-R needs truthful failure histories, but durable and personally attributable records may reduce contributors’ willingness to produce them.The tension is between making organizational knowledge legible to AI and preserving the candor required to generate that knowledge.
  • Silence-induced distortion: Reduced disclosure can distort the observed BUT distribution and therefore which decision dimensions appear to matter.A decline in recorded failures may reflect silence rather than improved stability or convergence.
  • Governance: The framework does not resolve the conflict among accountability, privacy, candor, due process, and learning; the record architecture itself becomes part of the decision problem.Governance must specify how raw narratives, reusable abstractions, identity, and access are handled.
  • Implementation boundaries: A serious implementation must preserve corrections, contestation, and alternative interpretations while separating learning records from performance or disciplinary records where appropriate.Organizations must also avoid interpreting an empty incident stream as evidence that the model has converged.

10 Representational Sanity Check: Demonstrating State Aliasing

The sanity check shows that an IS-only representation cannot distinguish a case requiring an alternative action from one requiring information gathering or escalation. Collapsing both into “not IS” therefore aliases organizationally distinct states without making a real-world performance claim.

  • State construction: An IS-only representation maps a BUT case and a suspension case to the same “not IS” state.The construction compares a case requiring an alternative action with one requiring querying or escalation.
  • Representational consequence: A deterministic policy conditioned only on IS must choose the same behavior for both cases even though their required behaviors differ.The representation therefore aliases two organizationally distinct states by construction.
  • Scope of the check: The proposition establishes a representational limitation, not an empirical performance result for real AI systems.Whether incidents yield stable BUT states, decision dimensions, and suspensions remains an empirical question.
  • Executable check: The executable check only verifies that the software representation matches the state-aliasing construction and reports no synthetic performance percentages.Its purpose is consistency with the stated proposition rather than evaluation against a benchmark.

11 Testable Hypotheses and Evaluation Agenda

The paper proposes testable hypotheses about whether OIBAR improves boundary decisions, dimension discovery, human handoffs, transfer, and organizational learning. It outlines small empirical studies while emphasizing that reporting incentives and comparison with existing elicitation methods must be evaluated.

  • The framework’s value should be evaluated empirically rather than assumed from its formal structure.The paper states that its value depends on whether the proposed distinctions improve real organizational work.
  • OIBAR hypotheses target fewer inappropriate boundary actions, better discovery of unanticipated dimensions, more specific handoffs, and stronger transfer to novel cases.The hypotheses compare explicit boundaries and suspension queries with one-sided rules, fixed checklists, vague escalation criteria, and action-level admonitions.
  • Incident reviews distinguishing threshold revision from new-dimension discovery are hypothesized to yield more reusable updates and fewer action-specific guardrails.
  • Stability bounds require approximately stationary exposure assumptions, while reporting incentives may alter the observed incident stream.
  • A small study can compare the original rule, OIBAR, and optionally a taxonomy baseline on ordinary, boundary, and withheld-value cases.Proposed outcomes include inappropriate action, missed questions, unnecessary escalation, and transfer to novel cases.
  • Elicitation studies should compare OIBAR with Kelly-style, CTA/CDM, and fixed-taxonomy approaches on time cost, dimensions, analyst stability, and downstream usefulness.The comparison is needed because taxonomy-free dimension elicitation alone does not establish OIBAR’s novelty.

12 Limitations

The paper qualifies OIBAR’s decomposition, naming, learning, and externalization claims. Its scope depends on expert judgment, stable and candid incident reporting, and the ability to express decision-relevant expertise without treating the framework as a complete account of organizational failures.

  • The one-changing-variable criterion is an operational decomposition target, not a claim that real failures have a single cause.Whether a decomposition is causally or decision-relevantly adequate remains a human judgment.
  • OIBAR’s novelty rests on integrating open-world dimension generation, typed suspensions and handoffs, domain-relative naming, and RESULT-driven expansion.Existing methods provide precedents for contrastive elicitation, positive/negative learning, and counterexample-guided refinement.
  • New vocabulary in a BUT is only heuristic evidence of a new dimension; explanatory and decision relevance to the boundary transition is the substantive criterion.
  • Domain-relative naming can be unstable in underspecified domains or overlapping professional vocabularies, and fixed-point notation does not prove unique semantic convergence.
  • A RESULT does not automatically determine whether to shift a threshold or introduce a new dimension, because that distinction may remain contested expert judgment.
  • Apparent representational stability can reflect rare dimensions, distribution shift, selective observation, or suppressed reporting rather than completeness.The exposure-rate bound is conditional and cannot be interpreted as a completeness guarantee.
  • Durable failure records may improve learning while increasing attribution exposure, interpersonal risk, or incentives to sanitize incidents.Organizational silence and psychological safety affect the incident data from which BUT cases are elicited.
  • Not all tacit expertise can or should be fully externalized because some expertise is embodied, socially distributed, contested, private, or costly to articulate.The narrower claim concerns expertise expressible through decision-relevant success and failure boundaries, subject to governance preserving accountability and candid learning.

13 Conclusion

The paper concludes that organizational AI failures can reflect incomplete representations of work rather than insufficient system intelligence alone. OIBAR externalizes boundaries, unresolved values, responsibilities, and learning through incidents, while its practical value remains an empirical question.

  • Formal organizational knowledge often preserves positive rules while omitting negative boundaries, runtime values, responsibility, and accumulated failure history.
  • OIBAR represents work as evolving value-bearing decision dimensions and boundaries rather than as an expanded rule alone.Its questions cover the rule, applicability, concrete failure, changed variable, required value and responsible party, and revision forced by results.
  • OIBAR synthesizes contrastive elicitation, near-miss learning, positive/negative boundary refinement, critical-incident analysis, and situated-action research for a new deployment context.
  • A suspension identifies what the system does not know and can specify whether a machine or human should produce the unresolved value.This makes human involvement a decision-specific responsibility rather than a generic human-in-the-loop slogan.
  • Failure analysis should discover the decision axis a rule omitted and record what must be known before that axis guides action again.
  • Whether OIBAR improves human–AI systems depends on empirical transfer of expert judgment, more specific handoffs, and learning from failures without adding brittle rules.Formal elegance alone cannot establish that value.
  • The procedure begins by stating a judgment and domain, then recording IS, a concrete BUT, the changed variable, the judgment error, and the required suspension.
  • The accompanying sanity-check script distinguishes IS, BUT, and suspension but is not empirical evidence or a benchmark.
Loading 2608.29055v1…