Source-linked AI summary
Event-Time Confounding Under Bursty Human Dynamics
Michael Iannelli, Alan Ai
TL;DR
User-timed event windows can mistake continuation of an already active episode for an event effect. The paper formalizes this endogenous time-zero problem and tests it with known-null timestamps, showing that zero-effect moments reproduce a substantial share of apparent post-event lift.
Problem
Event-window studies lack a defensible time zero when user-timed events occur inside ongoing episodes already generating the outcome.
Method
The paper formalizes episode-selection bias and tests it using same-user, cross-surface logs and engineered pseudo-events guaranteed to cause nothing.
Results
Known-null timestamps reproduce a substantial share of real events’ apparent post-event lift, while activity generally peaks before user-timed events.
Takeaways & Limitations
Studies should compare similar episodes and use pre-event diagnostics and negative controls before giving event-window contrasts a causal interpretation.
Takeaways & Limitations
The evidence is conditional on the narrow active population studied and does not decompose real post-event associations into confounding and causal effects.
Abstract
from arXiv · showhide
Studies of digital behavior often align users at moments they choose, such as opening an AI assistant, clicking a recommendation, or visiting a product page, and interpret higher activity afterward as an event effect. We show how this creates an endogenous time zero: the event occurs during an ongoing task episode, so the aligned curve can trace episode continuation rather than a response to the event. In same-user, cross-surface web logs, AI, shopping, news, coding, and reference events are all preceded by broad activity increases that peak before time zero. Our strongest test uses known-null timestamps that cause nothing. Among the 5.8% of AI responses meeting strict pre-event activity and washout criteria, these timestamps show 3.42 times the post-event search activity of a within-user placebo, compared with 4.32 times for real events. The fraction of excess reproduced by the known null falls from 0.56 at detectably active moments to -0.04 at quiet moments, where the design detects none. We formalize this episode-selection bias, prove that a single-surface event window cannot separate it from a genuine effect without additional assumptions, and show in zero-effect simulations why user fixed effects and coarse activity matching can fail: the confound is within-user and time-varying. We provide a diagnostic protocol, public-data benchmarks, and burstcheck, a lightweight audit tool. User-timed events may have real effects, but post-event volume does not identify them by default; studies should compare similar episodes with and without the event.
1 Introduction
User-timed events can occur within already-active task episodes, so post-event activity may reflect episode continuation rather than an event effect. The paper formalizes this within-user, time-varying confounding and shows why single-surface event windows require additional assumptions.
- Motivation: Browsing and search peak roughly eight minutes before conversational-AI responses, reaching 3.2× and 3.3× placebo levels pre-event, then declining through the event.The event sits at the tail of a rising trajectory rather than establishing a defensible time zero.
- Mechanism: Episode-selection bias occurs when focal events are selected into latent task episodes whose continuation is misread as the event’s consequence.Bursty human activity makes this risk plausible, but the confound is the latent state—such as intent or task engagement—not burstiness itself.
- Novelty: The confound is within-user and time-varying, so fixed effects, within-person windows, and matching on recent history can compare the same user’s busy moments with quiet ones.This distinguishes episode-selection bias from between-person activity bias and explains why standard within-user defenses may fail.
- Contributions: The paper formalizes endogenous time zero and proves that single-surface event-aligned data cannot identify the event effect without further assumptions.The event-window contrast mixes the treatment effect with a backdoor path through latent episode intensity.
- Scope of the claim: The claim is conditional, not universal: user-timed event estimates may be dominated by episode shape when events fall inside task episodes, unless designs separate shape from treatment response.The burden of proof lies with the study, while adjustment is not ruled out categorically.
2 Related Work
The paper connects event-time confounding to established work on endogenous timing, negative controls, proxy-based identification, and bursty human activity. Its contribution is to link episodic temporal clustering to causal bias when latent episode state jointly determines event timing and outcomes.
- Endogenous treatment timing and time-varying confounding: Endogenous treatment timing can align treatment with an evolving outcome process, motivating designs such as timing-of-events models, g-methods, and target-trial emulation.These approaches obtain identification through explicit assumptions; the paper situates its event-time problem within this established tradition.
- Negative controls and proxies for unmeasured confounding: Negative-control outcomes can reveal shared confounding when treatment cannot plausibly affect them, although a still negative control does not validate a design.The paper distinguishes falsification by a moving negative control from validation by a still one.
- Negative controls and proxies for unmeasured confounding: Cross-surface activity streams serve as proxies for latent episode state, but conditioning on them reduces confounding without alone satisfying proximal identification conditions.Proximal causal inference requires treatment-side and outcome-side proxies with completeness conditions; the paper notes that panel event-study work goes further.
- Bursty human dynamics: Human activity is heavy-tailed and episodically clustered rather than Poisson [5] [20] [25], allowing latent state to jointly determine event timing and outcomes.The paper identifies the missing bridge as connecting descriptive burstiness to causal bias.
3 Known-Null Timestamps on Real Behavioral Trajectories
Known-null timestamps—moments that cause nothing by construction—still reproduce most of real AI events’ apparent post-event activity. This demonstrates that event-anchored analyses can mistake continuation of dense behavioral episodes for causal effects, especially at detectably active moments.
- Known-null design: Known-null timestamps reproduce most of real AI events’ apparent post-event excess despite having zero causal effect by construction.The primary comparison uses real AI responses and equally active, same-user moments with no AI response, with outcomes compared against within-user random-moment placebos.
- Known-null design: The primary null construction uses only strictly pre-event information, requiring active landmarks from prior non-search, non-AI page views and excluding recent AI responses.Pseudo moments are sampled from the same users under the same per-user cap, preserving the known-zero causal status without consulting future activity.
- Control ladder: 5.57× association arises from naive event anchoring alone, while pre-event state matching holds or slightly raises it to 6.11–6.44×.The pattern is consistent with an inspection-paradox mechanism in which event-anchored moments oversample locally dense spells; matching active pre-windows can select even denser moments.
- Where the risk lives: Episode-selection risk varies with pre-event context: naive lifts are 3.63× for landmark-active responses, 3.07× for mildly active responses, and 2.60× for quiet responses.All classes use the same impoverished clock-only matching, so the reported 0.60 span reflects population differences rather than matching richness.
- Where the risk lives: The reproduced fractions are floors rather than decompositions, falling from most of the association at active anchors to nothing detectable at quiet anchors.Quiet means quiet on the observed surfaces, so episodes contained in search or assistant activity may remain undetected; the section cautions that the headline and class results use different populations and matching schemes.
- Sensitivity analysis: Pseudo-events inserted into completed AI-free episodes reproduce 73% of real in-episode excess, rising to an upper-bound 108% with whole-episode intensity matching.Because completed episodes are defined using post-event information, this construction is a sensitivity analysis rather than the primary design.
4 Endogenous Time Zero
User-timed events can be selected within latent high-activity episodes, so event-window contrasts combine treatment effects with episode-selection bias that same-user designs and observed-history matching may not remove. The section formalizes when counts and shares differ, proves single-surface non-identification, and distinguishes temporally valid proxies from causally sufficient adjustment.
- Episode-selection bias: Episode-selection bias is non-zero when events and non-events differ under no effect, requiring both outcome-moving episode states and event selection into those states.The bias can be small even in highly bursty processes if events occur randomly within bursts.
- Same-user designs: User fixed effects remove stable between-user differences but not within-user busy-versus-quiet gaps, so event selection on latent activity can leave same-user contrasts biased.Matching on recent observed activity can likewise fail when the relevant state is latent and time-varying.
- Count/share divergence: 0.70→2.52 raw-count activity can coexist with a 0.389→0.304 activity-share decline when a small local substitution is masked by a burst.Counts and shares therefore represent different estimands; a share is valid as an alternative mix-effect estimand only under proportional bursting, not as a debiased count.
- Non-identification: A single-surface event-aligned process cannot distinguish a zero-effect latent-episode model from an effect model with exchangeable treatment and no latent confounding.The two models can induce the same distribution, including the pre-event path, making falsification rather than direct estimation the honest empirical strategy.
- Proxy adjustment: Past-only state estimates avoid post-treatment contamination but identify causal effects only when they achieve conditional mean exchangeability; noisy proxies generally do not.Thus, past-only filtering is necessary hygiene, while adjusted residuals remain associations of potentially unverified sign.
5 Events Sit Inside Cross-Surface Episodes
User-timed events are selected into intense, bursty within-user episodes that span multiple surfaces, so elevated post-event activity can reflect episode continuation rather than an event effect. Cross-surface activity provides observable proxies for this shared state, while negative controls qualify the episode interpretation.
- Cross-surface episode selection: The logs are bursty, with B=0.72, a median inter-event gap under one minute, and a long tail beyond the 90th percentile.Activity arrives in tight episodes separated by long quiet gaps.
- Cross-surface episode selection: 3.6× for shopping, 2.3× for news, 5.3× for coding/docs, 3.8× for reference, and 3.2× for AI: pre-event browsing exceeds each user’s placebo in every event domain.These elevations show that events concentrate in unusually active within-user moments across event types, not only around AI use.
- Cross-surface episode selection: The episode interpretation is deliberately qualified: focal events are selected into high-intensity within-user states that predict outcomes, even when co-active surfaces do not form one coherent task.Leisure is an imperfect negative control; the engineered pseudo-event exposure remains the strongest negative control because its irrelevance holds by construction.
- Cross-surface episode selection: 3 to 6× across all five focal domains: cross-surface activity is elevated before and after events, revealing a synchronized episode rather than an isolated surface pattern.Figure 4 masks each anchor surface and treats sparse cells as unestimable; a news→coding/docs cell remains near zero, supporting a nonmechanical pattern.
- Cross-surface episode selection: Cross-surface burst indices attenuate most of the naive association, because activity breadth provides a non-circular proxy for the shared state.The index excludes both the search outcome and the focal AI surface, while the same breadth that creates apparent effects also makes the latent activity state observable.
6 A Zero-Effect Simulation
In a zero-effect simulation, common event-window adjustments manufacture large effects because past activity imperfectly proxies the latent burst state. More state-aware estimators reduce but do not always eliminate this episode-selection bias, while diagnostic performance depends on the control’s burst loading.
- Common adjustments: +2.99 to +3.41 estimated effects persist against a true 0 across user fixed effects, activity matching, pre-activity stratification, and recent-activity intensity, near the naive +3.30.Each adjustment conditions on past observed activity, an errored proxy for the short latent burst governing selection.
- Alternative estimands: +0.07 [0.06, 0.08] share bias is smaller than count bias but not null, because outcome and other activity scale 30× and 20× across quiet-to-burst states.The share bias crosses zero at equal scaling (+0.004), so share is an alternative mix estimand rather than an automatic repair for counts.
- Negative control: −0.004 at zero control loading and +3.31 at full loading show that the negative control is silent when unrelated to bursts and tracks contamination when fully loaded.The control is therefore sensitive to the episode-intensity mechanism rather than a universal alarm.
- State-aware estimators: +0.03 from the true-state oracle and +0.25 from a two-sided smoothed latent-state estimate recover the null far better than the past-only forecast’s +1.93.The corresponding recoveries are approximately 92% and 42%; the key distinction is treatment-relevant state capture, not simply forward versus backward information.
- Pre/post comparison: −0.04 [−0.11, 0.04] from pre/post comparison is accidental under symmetric bursts, becoming −1.07 [−1.15, −0.99] when events concentrate in burst tails.This shows that estimator behavior depends on burst shape and event timing, not merely whether the estimator uses future information.
7 What Changes the Answer
The paper shifts the default unit of analysis from user-timed events to episodes, recommending comparisons of similar episodes with and without the event plus falsification diagnostics. These tools can expose and reduce episode-selection bias, but causal identification still requires exogenous timing or design-based variation.
- A diagnostic protocol: Falsification tests include pre-event trajectories, negative-control outcomes, active-window placebos, and pseudo-event experiments; Figure 6 packages them into a diagnostic protocol.The protocol lets readers judge whether a user-timed event is a treatment or a marker of an unfolding episode.
- What changes the answer: Alternative estimands and bias-reduction methods can lessen exposure to episode volume but do not by themselves identify causal effects.Options include activity share, episode-level outcomes, first-observed events, same-user active-window matching, whole-episode matching, lead adjustment, inverse-intensity weighting, latent-state filtering, and cross-surface proxies.
- What the diagnostics can and cannot conclude: Identification requires exogenous or randomized timing, an instrument, a rollout, or verified proximal proxy conditions.A single-surface observational log does not supply this variation on its own, and adjusted residuals can mix true effects with unmeasured state.
- The episode-level comparison, audited: Episodes with AI present have 1.2–1.9 percentage points higher within-episode search share across matched and unmatched constructions and both share estimands.This audited episode-level comparison is observational rather than causal.
- The practical shift: The recommended default shifts from event time versus ordinary time to similar episodes with and without the event, interpreted alongside a presence placebo.Cross-surface panels make the confounding episode visible, while pseudo-events with known zero effect reproduce most apparent lift in the real-data experiment.
- Scope: The headline result applies only to washout-aligned, landmark-active responses with positive baseline and matched support: eligible anchors are 5.8% of all in-panel AI responses.Landmark-active responses comprise 31% of the per-user-capped response set before the other restrictions, so the result does not describe all assistant use.
8 Conclusion
User-timed events may mark ongoing episodes rather than exogenous time zero, allowing zero-effect timestamps to reproduce substantial apparent post-event lift. Credible causal interpretation therefore requires episode-matched comparisons and design-based or otherwise verified identification.
- Conclusion: Zero-effect timestamps reproduce a substantial share of the apparent post-event lift in the studied active population, demonstrating what episode timing alone can produce without decomposing real associations into confounding and causal effects.A user-timed event may occur while an episode is already raising the outcome, so the known-null result does not identify the causal component of the real association.
- Conclusion: Studies should compare similar episodes, report count and share estimands, inspect pre-event activity, and apply negative controls and active-window placebos before interpreting event-window contrasts causally.These diagnostics can reveal unsafe designs, but only design-based variation or verified identification conditions can establish the remaining effect.
9 Ethics and Human-Subjects Statement
The study uses previously collected, opt-in, de-identified behavioral data without participant intervention or new data collection, but it was not submitted for IRB review. Protections include limiting analysis to timestamps and domains, reporting aggregate results, and disclosing the authors’ commercial conflict of interest.
- Human subjects: The analysis uses previously collected, opt-in, de-identified web, search, and conversational-AI events; authors neither intervened with users nor collected new data.Consent and withdrawal operate through the panel provider.
- Human subjects: No IRB approval or exemption was sought or obtained, and the study explicitly reports that it was not submitted for institutional review.The authors are not affiliated with a university.
- Privacy protections: Because pseudonymous clickstreams remain quasi-identifying despite de-identification, analyses use only event timestamps and domains, excluding page content and conversation text, with aggregate reporting.The stated protection is that sensitive material never leaves the analysis environment.
- Competing interests: The authors disclose employment by Scrunch AI, a commercial AI-traffic measurement company, and apply the paper’s critique to their own product category.Appendix C audits a stylized finding of this kind and reports that its naive form does not survive.
A Proofs of Propositions … E Reading the Diagnostics
The proofs show that latent, time-varying episode selection can create apparent effects despite zero causal impact, while the empirical audits distinguish confounded volume lifts from more defensible compositional, discrete, and routing estimands. Validation on Wikipedia spikes and burstcheck diagnostics provide concrete tests for pre-event elevation, placebo movement, burst scaling, and window dependence.
- A Proofs of Propositions: A Proofs of Propositions: User fixed effects remove stable heterogeneity but not within-user episode-selection bias, which remains positive under state monotonicity and positive event selection even when the true effect is zero.The within-user contrast factors into a positive state contrast and a positive selection gap.
- A Proofs of Propositions: A Proofs of Propositions: A count can rise while outcome share falls when a burst scales all activity and the event locally substitutes for the outcome.Thus raw counts and shares can diverge even when they arise from one underlying process.
- A Proofs of Propositions: A Proofs of Propositions: Pre-treatment filtered proxies avoid conditioning on treatment descendants, but temporal admissibility alone does not ensure adjustment validity without recovery of the treatment-relevant state or verified proxy conditions.Two-sided smoothing can use post-treatment variables and bias the contrast even under the null.
- B The Completed-Episode Sensitivity: B The Completed-Episode Sensitivity: The pseudo-event experiment inserted events into AI-free completed episodes at the positions occupied by real events, using runs with under-30-minute gaps, at least five page views, and at least ten minutes’ duration.The segmentation convention used the 30-minute web-log sessionization threshold attributed to Catledge and Pitkow.
- C A Worked Case: an AI “Demand Lift” That Does Not Survive the Audit: C A Worked Case: an AI “Demand Lift” That Does Not Survive the Audit: The apparent downstream lift was preceded by rising trajectories, had flat activity share, and was reproduced by a same-user null-burst placebo.A same-day pre-conversation window showed the same elevation, so the volume increase is consistent with episode continuation.
- C A Worked Case: an AI “Demand Lift” That Does Not Survive the Audit: C A Worked Case: an AI “Demand Lift” That Does Not Survive the Audit: Compositional, discrete, and routing estimands remain more defensible, whereas unaudited post-event volume lifts should be treated as upper bounds consistent with pure episode continuation.The worked case constrains the class of estimands rather than retracting any specific prior finding.
- D Discriminant Validation on Wikipedia Spikes: D Discriminant Validation on Wikipedia Spikes: Anticipated events had 4.8× versus 1.0× pre-event elevation for surprise events, with AUC 0.97 and balanced accuracy 0.93 at a 2.2× threshold.The event list was mechanically collected and the classifier was blinded to pageview series, though the threshold and accuracy were evaluated in-sample.
- E Reading the Diagnostics: E Reading the Diagnostics: In a zero-effect synthetic log, burstcheck flagged five of six checks, including 1.50× pre-event elevation, 4.05× negative-control movement, and 8.6× window-length sensitivity.It also reported count up while share stayed flat at 4.1/1.0×, indicating burst volume rather than an outcome-specific effect.
F Simulation Design and Latent-State Detail
The simulations use a two-state burst process where events are more likely during bursts but have no causal effect on outcomes. Oracle adjustment recovers the null, while feasible latent-state estimators reduce but do not fully eliminate bias.
- Simulation design: 4,000 trajectories of 480 five-minute bins follow a two-state Markov process with rare mean-five-step bursts, and event probability rises from 0.002 in quiet periods to 0.15 in bursts.Outcomes and activity are Poisson, with rates increasing substantially during bursts; the event affects neither the outcome nor the negative-control outcome.
- Latent-state results: +0.03 is the oracle-adjusted estimate, showing that conditioning on the true burst state recovers the null when positivity holds.The remaining failure is therefore an estimation problem when the state is latent rather than a failure of adjustment with a well-defined observed confounder.
- Latent-state results: +3.00 remains for an exponentially weighted intensity using the same past, whereas a Poisson hidden Markov model removes 42–47% of bias across four seeds.The limited recovery reflects difficulty forecasting a short, fast-mixing burst into the realized post-event state; the naive bias changes by 0.04 while the oracle remains null.
- Latent-state results: +0.25, approximately 92% of the null, is recovered by a two-sided smoothed estimate, but it is inadmissible because it uses post-event activity.This illustrates the tradeoff between better recovery and leakage of post-event information into the estimator.
G What the Observed Law Does Pin Down · H Cross-surface matrix: confidence intervals · I Weighting Sensitivity
The observed law leaves the ATT unpoint-identified, but yields assumption-dependent bounds and a breakdown value for episode selection. Cross-surface lifts remain robust to weighting, while confidence intervals require conservative multiplicity and cautious interpretation of sparse cells.
- G What the Observed Law Does Pin Down: The ATT is not point-identified: with Y≥0, the identified set is (−∞, E[Y | A=1]], unbounded below.Model B’s positive effect lies in the interior of this set; the result follows from persistence and positive selection on a latent state, not specifically burstiness or cross-surface data.
- G What the Observed Law Does Pin Down: The observed rate gap is 0.89, implying breakdown values m̄*=0.30, 0.22, and 0.15 when Δ̄=3, 4, and 6.The breakdown value is the residual state-posterior gap required for episode selection to explain the entire apparent effect, so it is assumption-dependent rather than a point estimate.
- G What the Observed Law Does Pin Down: In a true-zero-effect simulation, matching an observable proxy left m=0.29 at Δ=8.1, corresponding to a breakdown value of 0.11.This demonstrates that the frontier describes what would need to be true, not that the required discrepancy is empirically false; the bound applies to the matched landmark-active population under a binary-state model.
- H Cross-surface matrix: confidence intervals: Cross-surface post-event activity is reported relative to stabilized within-user placebos with 95% user-clustered bootstrap intervals, while the search column is precisely estimated.The AI→search interval differs slightly from §5 because that analysis caps events per user and weights users equally; surface assignment uses narrow exact-host allowlists.
- H Cross-surface matrix: confidence intervals: Multiplicity correction covers all 26 matrix cells, although only 24 were tested, yielding an honest rate of 14 of 24 tested.The argument focuses on the search column, where every domain clears 1, rather than sparse individual cells; News→coding/docs is treated as unestimable when its lower endpoint is pinned at zero.
- I Weighting Sensitivity: The naive AI→search lift remains large and above 1 under all four weighting schemes using stabilized within-user placebos.The one-random-event interval is wider and asymmetric because its effective event count is much smaller, not because of a transcription error.
- I Weighting Sensitivity: The mean of ratios is 1.07× versus 1.17× for the ratio of means because empty windows are dropped, disproportionately affecting placebo windows.Empty windows comprise 74% of placebo windows and 41% of post-event windows, so conditioning on active windows pulls the mean-of-ratios estimate toward 1.
J Known-Null Construction, Sensitivity, and Inference
This section defines the matched estimand and secondary pooled construction, then tests inference validity through variance checks, synthetic known-truth coverage, and sensitivity to influential users. The estimate varies with episode age and influential-user trimming, while the broader conclusion remains separated from point-estimate robustness.
- Construction: The headline estimand is a ratio of user-equal-weighted means, and matching covers 97.9% of observations while the pooled construction only partly matches panel-level marginals.Users with larger real excess receive more weight in the ratio; pooled weights equalize cell mix within users.
- Sensitivity: 5.09× apparent lift in the first ten minutes declines to 3.18× after half an hour, whereas matched null moments show no comparable gradient.Matched null values across the three age bins are 2.96×, 2.52×, and 2.91×.
- Inference: 0.086, 0.087, and 0.092 standard errors from the bootstrap, delta method, and jackknife differ by at most a 1.07 maximum-to-minimum ratio.These are within-design standard errors, not the headline total after design combination.
- Inference: 0.9 coverage of the computable probability limit across 100 synthetic replications supports the user-clustered bootstrap under the stated zero-effect data-generating process.The exercise used 200 synthetic users and 150 bootstrap draws per replication.
- Inference: 0.78 is the finite-sample shortfall from one in the zero-effect replication mean, because matching on a noisy observable proxy does not deliver conditional mean exchangeability.The probability limit itself is 0.21, distinguishing the estimand’s selection bias from finite-sample bias.
K A Public Benchmark for Episode-Selection Estimators · L Extended Related Work · M Ethics Detail: Review Status, Consent, Withdrawal, and Privacy
The public benchmark shows that episode-selection estimators can substantially reduce, but not always eliminate, excess activity on real bursty series, while related work situates the problem within endogenous timing and time-varying confounding. The study reports contractual rather than ethics review and limits privacy exposure through aggregate analysis and omission of raw content and individual-level outputs.
- K A Public Benchmark for Episode-Selection Estimators: On MovieLens, burstiness is nearly the panel’s (B=0.71 versus 0.72), while a focal rating is followed by 2.1× placebo activity and preceded by 1.8×.The comparison is qualitative because the two burstiness estimates use different tail truncations.
- K A Public Benchmark for Episode-Selection Estimators: At two, four, and eight latent states, the residual excess is 4.6×, 3.7×, and 2.9×, with eight states removing 58% of the excess.The remaining bias reflects approximation of a continuous latent state by a discrete model and is presented as an expected envelope for real aggregates.
- K A Public Benchmark for Episode-Selection Estimators: Table 4 defines a public, reproducible benchmark using plasmode tests that inject known effects, including zero-effect nulls, into real bursty processes.The feasible estimator is genuinely past-only, except on MovieLens where the feasible form uses a directly countable observable both-sides local intensity.
- L Extended Related Work: Related work connects the problem to endogenous treatment timing, self-controlled designs requiring exchangeability, and g-methods that model time-varying treatment processes.These frameworks address related identification problems by imposing explicit assumptions such as no anticipation, exchangeability, or a modeled treatment process.
- L Extended Related Work: Prior work explains bursty timing through task-order queues, session cascades, and self-exciting point processes, while modern difference-in-differences addresses heterogeneous and staggered timing.For behavioral logs, a pre-event rise conflicts with the simplest interpretation that treatment begins the change at time zero; endogenous timing is identified as the leading explanation here.
- M Ethics Detail: Review Status, Consent, Withdrawal, and Privacy: The study received contractual data-use review rather than ethics review, and consent was obtained by the panel provider for behavioral measurement and analysis.Readers requiring IRB oversight should discount the study accordingly.
- M Ethics Detail: Review Status, Consent, Withdrawal, and Privacy: Pseudonymous clickstreams remain quasi-identifying, so privacy protection relies on keeping page content, conversation text, and individual-level outputs out of the environment.Linking AI use with browsing and search is more sensitive than either stream alone, and the linkage is confined to aggregate measurement of a methodological failure mode.
N An Attempted Prevalence Audit, and Why It Does Not Settle the Question
The audit did not establish prevalence: among papers surfaced by the screening frame, externally assigned timing was common, but the screen missed known examples and was too limited to support broader conclusions. An industry sample likewise used varied designs rather than uniformly adopting event alignment.
- Prevalence audit: The screen could not estimate prevalence because it missed two hand-coded papers, concentrated yield in queries, left 19% unclear, and overrepresented computer science.These failures affected both retrieval and abstract-level coding, while arXiv underrepresents journal fields where natural-experiment designs are standard.
- Prevalence audit: Among the surfaced papers, externally assigned timing was the common choice, but the audit provides no evidence about settings where researchers lack that option.The authors removed prevalence language rather than reversing it because the audit neither supported widespread use nor its opposite.
- Industry audit: The 29-publication industry sample showed heterogeneous practices: 9 compared users by traffic source without event alignment, while 8 made quantitative claims without an identifiable comparison.The sample comprised vendor blogs, agency case studies, and analytics reports making quantitative claims about how AI use changes behavior.