Source-linked AI summary

Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events

Heyang Gong

arXiv:2608.25118v1cs.LGcs.AIstat.ML

TL;DR

Machine learning often leaves the persistent individuals linking multiple events implicit, so the paper makes the task-declared unit an explicit primitive and studies supervised unit-conditioned response laws. It learns a tokenizer and shared response-law form, distinguishing world-side heterogeneity from learner-side representation and identifying limits of single-row evidence and factorization. The resulting framework separates unit information’s predictive value from its accessibility and approximation, while restricting its formal guarantees to the supervised specialization.

  • Problem

    Machine-learning notation often identifies records without explicitly representing the persistent individuals they concern or whether response laws differ across individuals.

  • Method

    The paper declares units and unit-conditioned response laws, then learns a pair (Tϕ, Rθ) consisting of a tokenizer and one shared response-law form reading its contextual token.

  • Results

    The framework separates oracle predictive value, evidence accessibility, learner approximation, and what marginal prediction can certify about the internal token-and-form decomposition.

  • Takeaways & Limitations

    Unit information is sufficient to omit when preserving it changes none of the declared target, estimand, admissible answer, or evaluation protocol.

  • Takeaways & Limitations

    The factorization can be misspecified, joint training need not identify either module, and the paper does not prove approximation error vanishes with more training information.

Abstract

from arXiv · show

Machine learning is usually formalized through samples, while the persistent individual to which multiple observed or possible events refer often remains implicit. We propose the \emph{unit} as an explicit primitive at the level of task semantics. A learning task first declares a population of persistent referents and a sameness criterion; the realized value $u$ denotes the selected referent. Supervised learning is the main formal specialization. Its semantic object is a family of unit-conditioned response laws. Homogeneity is the special case in which those laws coincide; a sample-only conditional is silent as to whether the world is homogeneous or the observed law is only the marginal of a heterogeneous family. What is learned from data is a pair $(T_φ,R_θ)$: a tokenizer that produces a contextual unit token and one shared response-law form that reads it. The structured class takes that form to be a simple relation in the token; a linear predictor is the running instance. The token is the learner-side representation through which the task-side unit affects prediction, while a learner specification that omits unit information is unit-insensitive; homogeneity remains a property of the world-side response family. When identity is unresolved, the world-side law mixes unit-conditioned targets, while the learner composes its shared form with a token. A trusted resolver may fix the unit and supply a lookup token; otherwise \emph{unit abduction} forms a token of the same type from factual evidence. Unlinked single-row observations can fail to distinguish a heterogeneous unit world from a homogeneous pooled world; trusted same-unit pairs separate a restricted witness. The formal results concern this supervised specialization.

1 Introduction

The paper elevates the persistent referent linking events to a task-declared unit, then formalizes supervised learning over unit-conditioned response laws. It proposes shared token-and-response-form learning and separates unit information, evidence access, and evaluation boundaries.

  • 1 Introduction: Ordinary sample notation identifies records but leaves event-to-individual relations and cross-unit response-law heterogeneity implicit.This omission matters when events share an individual, queries vary while the individual is fixed, or attribution is uncertain.
  • 1 Introduction: Supervised learning treats a fixed unit as selecting an entire unit-conditioned response law rather than merely adding an observation identifier.Holding the unit fixed preserves the referent while context and event outcomes may vary.
  • 1 Introduction: The unit is a task-declared population primitive linking event records to the persistent referent they concern.The declaration specifies the population, sameness criterion, and retention span before choosing the learning object or model semantics.
  • 1 Introduction: Joint learning uses a tokenizer that produces a contextual unit token and one shared response-law form that reads it.The learned object is the pair (Tϕ, Rθ), with a simple relation in the token and a linear predictor as the running instance.
  • 1 Introduction: The paper distinguishes oracle unit-information value, evidence access, learner approximation, and evaluation effects involving row weighting and unit-disjoint splits.It also identifies a single-row impossibility boundary for distinguishing heterogeneous from pooled homogeneous worlds.

2 The Unit as a Machine-Learning Primitive

This section defines the unit as a referential object distinct from event content, prediction, and learner-side representation. It explains direct attribution and unresolved identity while preserving stochastic, context-dependent response behavior.

  • 2.1 Samples, Events, and Task-Declared Units: The event pair (x_i, y_i) carries event content, while u_i identifies the persistent individual to which that event is attributed.The sample index is bookkeeping rather than an identity marker.
  • 2.1 Samples, Events, and Task-Declared Units: Equality u_i = u_j records that two events concern the same individual, whereas predictive equivalence can hold for distinct units.Thus, referential sameness is stronger than having the same response law.
  • 2.1 Samples, Events, and Task-Declared Units: With direct access, a trusted key resolves the referent and lookup supplies a learner-side token for the shared response form.The token may be an embedding, preference factor, random effect, or other model-internal parameter block.
  • 2.1 Samples, Events, and Task-Declared Units: When attribution is unavailable to the learner, the complete-data representation differs from observations with learner-visible persistent-unit labels or trusted resolution.The unit declaration remains a task-side semantic commitment even when learned representations are not observed data.
  • 2.1 Samples, Events, and Task-Declared Units: Holding u fixed selects the same referent while allowing repeated measurements, choices, outcomes, and other event variation to remain stochastic.Independence, causal semantics, and deterministic responses are separate modeling commitments.

3 Learning a Family of Unit-Conditioned Response Laws

The supervised specialization treats the target as a family of unit-conditioned response laws and learns shared structure through contextual unit tokens. Homogeneity is a world-side property distinct from a unit-insensitive learner, while structured heterogeneity constrains one shared response form to read the token simply.

  • 3.1 Response Heterogeneity Across Units: The supervised semantic object is the family of unit-conditioned response laws indexed by task-declared units.
  • 3.2 Shared response-law form and contextual unit tokens: A unit-insensitive learner uses a common token or ignores the token, so its specification does not encode unit-specific predictive differences.
  • 3.1 Response Heterogeneity Across Units: Homogeneity means unit-conditioned response laws coincide almost everywhere, whereas a heterogeneous family may yield the same marginalized row-level conditional.
  • 3.2 Shared response-law form and contextual unit tokens: The structured class requires Rθ to read tokens through a declared simple relation, with a finite-dimensional linear predictor as the running instance.
  • 3.2 Shared response-law form and contextual unit tokens: Unresolved identity produces a token from factual evidence, while the learner applies the shared response form to that token rather than fitting unrelated unit-specific mechanisms.
  • 3.2 Shared response-law form and contextual unit tokens: The learned object is a tokenizer–response-form pair (Tϕ, Rθ), with every path from unit to prediction mediated by a contextual token and one shared response-law form.

4 Learner Access to the Unit

Learner access separates how a persistent unit is identified from how its information enters prediction. Trusted resolution supplies a lookup token, whereas unit abduction forms a same-typed contextual token from evidence; both use the shared response model.

  • 4.1 Access modes: Attribution determines the realized unit, while learner access determines how that unit enters downstream learning as a contextual token.
  • 4.1 Direct unit access: Direct unit access uses a trusted key to fix the referent and stable address, then reads or updates its learned token.
  • 4.1 Unit abduction: Unit abduction applies without a resolver by forming a contextual token from factual evidence; alternative queries reuse that token through Rθ.
  • 4.1 Access modes: Tokens are learner-side representations rather than units or necessarily identity posteriors, and direct access and abduction return tokens of the same type.
  • 4.2 Evidence and response axes: If evidence is uninformative about U, evidence-based individualization is unavailable and a calibrated learner remains at population-level uncertainty or abstains.
  • 4.2 Evidence and response axes: Learner access and unit-response dependence form independent axes, while record-wise evaluation can mix known-unit and unseen-unit prediction.

5 Unit Information: Predictive Value, Access, and Deployment

This section separates the predictive value of knowing a unit from the evidence and learned pair available at deployment, then characterizes what marginal observations and evaluation protocols can certify.

  • The learned object is the pair (Tϕ, Rθ), whose tokenizer organizes unit information and whose shared response form predicts from the resulting token.The theory distinguishes this learned pair from the unit-conditioned world-side response family.
  • Under log loss, oracle unit value decomposes into accessible value from O and residual value unavailable after the declared evidence cutoff.The residual term is I(U; Y Q | W), while the accessible term uses O and W.
  • The three information gaps vanish under the stated conditional independences, and a trusted key that uniquely resolves U zeros the residual term.Estimating the lookup token and shared response form remains an approximation problem.
  • In the binary example, oracle value is log 2, accessible value is log 2 − h(δ), and residual value is h(δ) nats.The unit remains predictive for every δ, while evidence quality ranges from perfectly informative to useless.
  • The learned pair realizes accessible response value minus predictive excess, whose bound separates unit-belief error from response-law approximation error.The KL result measures deployed log-loss penalty, while the total-variation result makes belief sensitivity depend on response heterogeneity.
  • Single-row marginals cannot distinguish homogeneous from heterogeneous unit worlds, whereas trusted same-unit pairs reveal within-unit covariance in the restricted witness.The witness has fixed-unit success probabilities 1/2 ± 1/4, with population variance and linked-pair covariance both 1/16.
  • Row and unit empirical risks agree exactly only when all observed units contribute equally many records, and record-wise splits mix known-unit and unseen-unit prediction.An unseen-unit claim therefore requires an explicitly unit-disjoint protocol, while the deployment estimand determines the appropriate weighting.

6 Recommendation as a Worked Setting

Recommendation illustrates the unit primitive when many interaction records share one user while candidate items change, with direct lookup and abducted tokens providing alternative access modes.

  • Recommendation uses persistent users as units across interaction records while candidate items or slates vary.The worked setting instantiates the declared unit boundary with users.
  • Direct unit access uses a trusted key to resolve the active user and read a stable token address.The lookup token is typically a learned row zθ(k).
  • Unit abduction forms a token of the same type from factual evidence, differing from direct access only in how the token is obtained.In the linear instance, prediction combines a user token with an item map through an inner product.
  • Recommendation models such as factorization and neural collaborative filtering instantiate the ID-indexed formulation with user-specific tokens.A deeper query encoder may be nonlinear while the running readout remains an inner product in the user token.

7 Relations to Established Learning Traditions

The unit formulation connects to established approaches that model persistent referents, supplied linkage, identity uncertainty, or shared response forms, while distinguishing units from environments.

  • Classical statistical learning is retained; the unit formulation adds the question of organizing persistent-unit regularities for one shared response-law form.This reframes the issue at the task-semantic level rather than replacing finite-sample generalization theory.
  • Longitudinal random-effects models closely combine supplied same-subject linkage, a shared response form, and a subject-specific token.Other traditions infer attribution or organize prediction around episodes, environments, or supplied group boundaries.
  • Persistent-unit variation is distinct from environment variation, though responses may vary along both axes.Covariate shift, domain adaptation, and invariant prediction primarily organize environmental or selection-regime variation.
  • In the worked task, users instantiate the declared unit boundary, and direct lookup or abduction supplies the learner-side user token.The table distinguishes the user referent from its indexed token and from evidence used by the tokenizer.

8 Discussion and Limitations

The discussion distinguishes the unit-sensitive formal framework from ordinary row-level learning and identifies boundaries on interpretation, learnability, attribution, and empirical testing.

  • Scope: Row-level deployments without same-unit questions can appropriately use ordinary sample formulations, but this does not establish homogeneous units.Without a declared unit axis, the row-level law remains silent about unit-response homogeneity.
  • Learnability limits: Assumption 1 is a computational interface rather than a learnability theorem, and unrestricted or misspecified classes weaken its interpretation.The paper does not prove vanishing approximation error or identify the tokenizer and response modules through joint training.
  • Representation and attribution: Unit boundaries and token representations are task-relative: trusted keys resolve referents, but token coordinates and similarity relations require separate choices.Token reparameterization can preserve predictions, while borrowing across units needs an explicit metric, kernel, graph, hierarchy, or other coupling.
  • Empirical boundary: Unlinked single-row observations cannot distinguish a restricted heterogeneous witness from a pooled homogeneous world, whereas trusted same-unit pairs can.The paper also notes that record-wise splits and unequal multiplicities can misalign evaluation with unseen-unit generalization and unit weighting.
  • Causal scope: Causal estimands require intervention, counterfactual, and identification assumptions beyond the unit primitive.The primitive preserves a referent across a query family but does not itself supply causal semantics.

9 Conclusion

The paper formalizes the unit as a task-level primitive and supervised learning as prediction from unit-conditioned response laws. It then separates world-side mixtures, learner-side token compositions, and the restrictions needed for shared structure and transfer.

  • Core formulation: The learned object is the pair (Tϕ, Rθ): a contextual tokenizer and one shared response-law form using a simple relation in the token.A finite-dimensional linear predictor is the running instance of the structured class.
  • Semantic and computational layers: A unit-omitting learner is unit-insensitive, while world-side homogeneity concerns whether unit-conditioned response laws coincide.The same row-level conditional can also arise by marginalizing a heterogeneous family.
  • Unresolved identity: When identity is unresolved, the exact world law mixes unit-conditioned targets, while deployed learners may instead compose a shared form with a contextual token.The theorem specialization uses an identity-mixture learner, which is distinct from the general token-composition formulation.
  • Task semantics: Unit declarations define populations, sameness criteria, and persistence spans before task-specific learning objects, without imposing a response law or model class.Dataset indices label records rather than creating new unit domains.
  • Unit access: A trusted resolver fixes the referent and lookup selects its token; factual evidence can otherwise support unit abduction returning the same type of token.Neither access route makes the token itself the unit or determines its parameterization.
  • Shared structure and transfer: Shared restrictions are necessary for transfer beyond observed units; a saturated unit-to-law class permits heterogeneity but does not by itself support unseen-unit transfer.A full learnability analysis additionally specifies loss, sampling, complexity, and learning rule.

C.4 Oracle Value Under Proper Scoring Rules

The paper quantifies the value of observing unit identity under proper scoring rules and distinguishes this oracle comparison from causal specializations that add structural and cross-world assumptions.

  • Oracle value under proper scoring rules: A proper-score proposition defines the Bayes-risk gap between predictors observing (U, W) and those observing only W.The gap is nonnegative under the stated admissibility and finiteness conditions.
  • Oracle value under proper scoring rules: For logarithmic loss, the oracle gap equals IΛ(U; Y | W) whenever the conditional mutual information is well defined.Strict propriety makes equality hold exactly when the pooled and unit-conditioned laws agree almost surely.
  • Oracle value under proper scoring rules: The result is an oracle, deployment-relative magnitude and does not establish that unit identity is observed, identifiable, causally useful, or learnable from finite selective observations.Its interpretation is therefore restricted to the declared deployment design and query.
  • Causal specialization: DiscoSCM’s causal decomposition abducts an identity posterior from factual evidence, evaluates fixed-unit counterfactual laws, and integrates them into a population-level answer.The identity posterior is one admissible token, not the definition of unit abduction in the general formulation.
  • Causal specialization: DiscoSCM specializes the generic unit-conditioned response family into structural causal mechanisms with intervention and counterfactual semantics.It separates the selected individual from event-level exogenous variation and adds assumptions beyond the generic unit primitive.
  • Scope and provenance: The causal hierarchy adds a specialized learned object, distribution-consistency and cross-world noise assumptions, then derives DiscoSCM-specific causal consequences.The unit primitive supplies the persistent referent and population-to-individual decomposition; DiscoSCM supplies the additional causal structure and dependent theorems.

E Additional Formal Results

The appendix formalizes value–access identities, single-row nonidentifiability, attribution and homogeneity distinctions, deployment risks, and repeated-unit evaluation protocols.

  • E Additional Formal Results: The appendix’s proofs retain unit identity and use the identity-belief specialization Qϕ(du | O) for value–access, mixture stability, and single-row collapse results.This specialization is not claimed to describe every tokenizer.
  • E.3 Single-Row Marginal Collapse and Linked-Pair Separation: Marginal response fit alone does not select a formed token, a which-unit belief, or a fixed-unit response decomposition.An unrestricted response-kernel class can reproduce any sample-only conditional after integration against every learner belief.
  • E.3 Single-Row Marginal Collapse and Linked-Pair Separation: Single-row observations can make heterogeneous and homogeneous unit worlds observationally identical, forcing every test’s maximum error to be at least 1/2.The equality of observable laws yields αn + βn = 1.
  • E.3 Single-Row Marginal Collapse and Linked-Pair Separation: For linked pairs, the heterogeneous and homogeneous worlds have variance difference 1/16 and pair probabilities shown in the appendix.The listed pair probabilities are H: 5/16, 3/16, 3/16, 5/16 and P: 1/4 for each outcome pair.
  • E.3 Single-Row Marginal Collapse and Linked-Pair Separation: The linked-pair separation requires trusted same-unit linkage and conditional product structure; linkage error, dependence, temporal state, or informative observation can invalidate it.It identifies neither the full mixing law nor membership of a realized unit in either response class.
  • E.4 Attribution information and response dependence are independent: Attribution informativeness and response homogeneity are independent: either can hold without the other, and mixture cancellation does not establish homogeneity.The appendix distinguishes these collapse conditions from single-row marginal collapse.
  • E.6 Oracle, deployed, row-weighted, and unit-weighted risks: Deployment risk depends on the declared experiment and weighting target; random-row evaluation can induce a size-biased unit law and need not estimate the intended population target.Subpopulation-law, oracle fixed-unit, and deployed marginalized quality are distinct quantities.
  • E.7 Repeated-unit weighting and splitting: Record-wise splitting can mix new observations from known units with new-unit cases, so unseen-unit generalization requires an explicitly unit-disjoint test protocol.Whole-unit splits, external new-unit cohorts, or another declared protocol provide such separation.

E.1 Proof of the Value–Access Decomposition

The proof decomposes value and access through conditional mutual information, yielding a risk ladder and bounds governed by response sufficiency and identity information.

  • E.1 Proof of the Value–Access Decomposition: The chain rule expands I(U, O; Y Q | W) into unit and evidence contributions in two equivalent orders.These expansions connect identity information and evidence information conditionally on the query context.
  • E.1 Proof of the Value–Access Decomposition: Response sufficiency makes O → U → Y Q a conditional Markov chain and yields the risk ladder through conditional data processing.It also bounds I(O; Y Q | W) by I(U; O | W).
  • E.1 Proof of the Value–Access Decomposition: The mutual-information decomposition requires response sufficiency, while representing the exact response conditional with P(U | O) additionally requires the external-query condition.The appendix states standard-Borel, regular-conditional, and finite-risk qualifications for the formal result.
  • E.1 Proof of the Value–Access Decomposition: The learner’s excess log loss is a KL divergence between marginalized response laws, and averaging proves the stated end-to-end risk identities.The component bound follows by applying relative-entropy chain rule and data processing to latent-joint laws.
  • E.1 Proof of the Value–Access Decomposition: The resulting upper bound is a marginal contraction of a latent-joint discrepancy and need not be tight because distinct latent decompositions can share response marginals.This is the appendix’s explicit limitation on interpreting the bound.

E.5 Proof and qualifications for mixture stability

The mixture-stability result separates attribution error from response-kernel approximation while stating support, observability, and stochastic fixed-unit boundaries.

  • E.5 Proof and qualifications for mixture stability: The stability proof assumes standard-Borel spaces, measurable Markov kernels, compatible regular conditionals, and a common full-mass support restriction.These conditions ensure the mixtures, integrals, and total-variation expressions are well defined.
  • E.5 Proof and qualifications for mixture stability: The proof bounds mixture discrepancy by combining response-kernel error with attribution-belief discrepancy through total variation and the triangle inequality.Jordan decomposition controls the attribution term, and the diameter is taken over the relevant unit region.
  • E.5 Proof and qualifications for mixture stability: The diameter term is query-specific and vanishes exactly when the selected true response laws agree across included units.Support outside the scientifically specified kernel region is a support failure, and total-variation stability alone does not bound unrestricted log-loss excess risk.
  • E.5 Proof and qualifications for mixture stability: The world target and learner pair remain distinct across token supervision, oracle-attribution response training, and marginalized end-to-end training regimes.Any history-derived state used at answer time must enter the tokenizer context, response context, or shared form for response sufficiency to hold.
  • E.5 Proof and qualifications for mixture stability: Population support does not ensure concentrated identity belief, and distinct units can remain observationally indistinguishable when they induce identical response laws under the study.Fixed-unit kernels may retain stochastic exogenous variation after identity is fixed.
  • E.5 Proof and qualifications for mixture stability: Within a declared same-unit query family, the unit and evidence cutoff remain fixed while queries vary, but shared attribution does not imply conditional independence across rows.A new factual event may update the formed token, whereas alternative response queries use the token at the fixed cutoff.

G Evaluation Checklist

A reproducible evaluation should pre-specify the unit population, evidence and access regime, target and split, baselines, and the unit-conditioned specification. It should also state which relation properties vary or remain invariant across units, how observations constrain tokens, and how evaluation separates unit-dependent learning from identifier memorization.

  • Evaluation checklist: Specify the unit population, identity-persistence span, factual-evidence cutoff, response query and context, and current target.
  • Evaluation checklist: Declare the access regime, available attribution truth, evaluated target, and whether the split tests known-unit/new-event or new-unit generalization.
  • Evaluation checklist: Include a matched unit-omitting baseline and any negative control required by the claim.
  • Evaluation checklist: Pre-specify which query–response properties vary with the unit, which remain invariant, and how same-unit observations constrain a unit’s token and response law.
  • Unit-conditioned specification: Assess whether the shared form and token space support prediction for an unseen unit and distinguish unit-dependent learning from identifier memorization.
  • Assumptions: Treat the response law as observational by default; interventional readings require assignment and identification assumptions, with pre-answer evidence excluding current targets and unavailable post-query measurements.
Loading 2608.25118v1…