Source-linked AI summary

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

Masahiro Kato, Daiki Honma, Taka Kato

arXiv:2609.11915v1stat.MLcs.AIcs.LGecon.EMstat.ME

TL;DR

Standard marketing data do not measure how often users encounter and notice firms in generated answers. GMMM constructs GEO and GEM inputs from generated-answer, placement, query, system-share, and notice data, then compares complete treatment sequences and establishes identification conditions. In simulations, estimating carryover and saturation improves coefficient recovery, while allowing response data to revise measurement parameters has no consistent benefit for estimating the GEO effect.

  • Problem

    Standard marketing data do not record nonsponsored firm-name occurrences in generated answers or whether users noticed sponsored placements.

  • Method

    GMMM constructs expected noticed GEO and GEM inputs, applies MMM transformations, compares complete treatment sequences, and derives identification conditions.

  • Results

    Estimating carryover and saturation explains most improvement in coefficient recovery over plug-in estimation, while response-data updating has no consistent benefit for estimating ∆G.

  • Takeaways & Limitations

    GMMM can determine treatment effects in some designs even when component coefficients cannot be recovered separately.

  • Takeaways & Limitations

    Identification requires occurrence probabilities that apply to the target population, a correctly specified conditional response mean, and control of confounding.

Abstract

from arXiv · show

Generative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm's name in generated answers. We develop Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization (GEO) and Generative Engine Marketing (GEM). For GEO, GMMM combines repeated generated answers with question counts, shares of use across generative systems, and notice probabilities. For GEM, it combines records of sponsored placements with notice probabilities. GMMM compares expected business responses under alternative treatment sequences and establishes sufficient conditions for identifying the resulting effects. We investigate the empirical performance of the proposed method using simulated answers to product recommendation in English and Japanese.

1 Introduction

GMMM extends marketing mix modeling to generative-media inputs that conventional records do not directly measure. It constructs expected noticed occurrences or placements, estimates treatment effects under complete sequences, and provides identification conditions.

  • 1 Introduction: Conventional logs omit nonsponsored name occurrences in generated answers and whether users noticed sponsored placements.Referral sessions also measure a different event from reading and noticing generated-answer content.
  • 1 Introduction: GMMM defines treatment effects by comparing expected business responses under complete GEO or GEM treatment sequences.Carryover means disabling an intervention changes inputs across affected periods, so a regression coefficient alone is insufficient.
  • 1.1 Contributions: GMMM connects measurements of generated answers and sponsored placements to marketing mix modeling.For GEO it models occurrence and notice; for GEM it relates spending to shown placements and notice before applying standard media transformations.
  • 1.1 Contributions: No universal ranking emerges between cut and joint posterior procedures for coefficient or treatment-effect estimation.The joint posterior improves coefficient RMSE in one condition but does not improve effect RMSE there and performs worse than cut in another.
  • 1.1 Contributions: Identification analysis allows treatment effects to remain identifiable when GEO-related component coefficients cannot be recovered separately.It also characterizes when occurrence probabilities differ from market counts only by scale and when market composition prevents that reduction.
  • Empirical illustration: The empirical collection contains 2,240 answers from two GPT models across 56 English and Japanese questions.Glasp occurs in 33.8% of GPT-5.6 Luna answers and 27.8% of GPT-4o answers, with the ordering reversed for English questions.

2 Setup

The setup defines responses, treatments, media inputs, and treatment-sequence effects for extending MMM with GEO and GEM. The constructed generative-media inputs must share the observational unit and evaluation framework with the response model.

  • 2 Setup: A treatment sequence specifies GEO or GEM settings across periods, including earlier values needed to initialize carryover.Treatment effects are evaluated over complete sequences because carryover links current responses to earlier media inputs.
  • 2 Setup: GMMM adds GEO and GEM to established media, using expected noticed generated-answer occurrences and sponsored placements as inputs.These inputs are constructed from generated answers, platform records, market counts, and attention observations rather than observed directly.
  • 2 Setup: GMMM uses repeated answers, question counts, system shares, notice indicators, spending, placement delivery, and response data.A randomized estimate of the same treatment effect can enter as an additional data source when the population and period are comparable.
  • 2.3 Data Sources: The public referral series cannot be combined with the generated-answer collection because it lacks matching question counts, answers, and notice observations.The two sources therefore cannot form one GMMM analysis.
  • 2 Setup: A substantive application must use an observational unit common to media construction, response modeling, and treatment-effect definition.The simulations use geographic markets crossed with product clusters as units.
  • 2 Setup: The primary GEO effect compares expected responses with and without the specified source modification while retaining the GEM sequence.The evaluation window may extend beyond active GEO periods to include carried-over effects; GEM has an analogous contrast.

3 Generative Marketing Mix Modeling

GMMM constructs expected GEO and GEM counts from generative-system observations and attention, then processes them through standard MMM transformations and response estimation. Its estimation framework supports plug-in, cut, joint, and regularized procedures.

  • 3 Generative Marketing Mix Modeling: GMMM constructs the two generative-media counts before applying carryover, Hill saturation, and the response model.The GEO count can be positive without the modification, while the GEM count reflects noticed sponsored placements.
  • 3.1 The GEO Input: The GEO input combines question counts, generative-system shares, occurrence probabilities, and notice probabilities under each source state.The source modification changes occurrence, and notice is assumed unchanged conditional on occurrence.
  • 3.2 The GEM Input: The GEM input models sponsored placements as a function of spending and predictors, then weights shown placements by notice probability.Observed delivery counts can be used directly; a Poisson model supports missing delivery counts and unobserved spending sequences.
  • 3.4 Estimation: A general penalized criterion combines measurement, response, experimental, and regularization terms.With negative log likelihoods and no penalties it yields maximum likelihood estimation; nonzero penalties permit regularization.
  • 3.5 Bayesian Computation: Plug-in, joint, and Bayesian procedures differ in how they propagate measurement uncertainty and allow response data to update measurement parameters.The cut posterior keeps measurement parameters determined by measurement data, whereas the joint posterior allows response data to revise them.
  • 3.5 Bayesian Computation: The Bayesian computation samples measurement parameters, constructs corresponding media sequences, weights them using response-data fit, and evaluates the treatment effect.The cut procedure uses equal weights for measurement samples, while joint weighting incorporates the response likelihood.

4 Identification of the Treatment Effect of GEO

Identification of GEO effects requires both identifiable response-model components and data supporting expected responses under complete treatment sequences. The analysis also characterizes scale equivalence and limits imposed by collinearity, modeling assumptions, and heterogeneous market composition.

  • 4 Identification of the Treatment Effect of GEO: GEO treatment-effect recovery requires response-model variation and observations supporting expected responses under each treatment sequence.Observed market inputs and treatment assignments determine whether the relevant sequence comparison can be evaluated.
  • 4.1 Identification Within the Response Model: GMMM reduces to standard MMM when expected GEO and GEM inputs are observed, with the added source-state and generative-channel terms.Setting the GEO and GEM coefficients to zero recovers the established-media additive model under the stated restrictions.
  • 4.1 Identification Within the Response Model: Conditional on fixed inputs and transformations, response coefficients are identified exactly when the residualized design has full rank.A linear effect can remain identified when individual coefficients are not, provided it vanishes on the design matrix’s null space.
  • 4.1 Identification Within the Response Model: An established-media input equal to the transformed GEO input can prevent its coefficient from being separated from the GEO coefficient.The simulations omit established media, so residualization against the remaining controls is sufficient there but not in the general model.
  • 4.2 Occurrence Probabilities and Market Counts: Under a common multiplicative scale, rescaling the Hill midpoint preserves the transformed GEO sequence and treatment effect.A single common midpoint cannot absorb scale factors that differ across markets, while heterogeneous question volume or notice probabilities affect treatment effects.
  • 4.3 Identification across Treatment Sequences: Identification across treatment sequences requires target-population occurrence probabilities, a correctly specified conditional response mean, and control of confounding.For nonlinear transformations, transforming an expected input generally differs from averaging transformed realized inputs.

5 Answer Collection and Simulation Evidence

The study combines collected generated-answer data with simulations to evaluate GMMM measurement and treatment-effect estimators. Across conditions, estimating carryover and saturation improves coefficient recovery, while no posterior method uniformly improves treatment-effect accuracy.

  • 5.1 Target Name in Generated Answers: 33.8% of GPT-5.6 Luna answers and 27.8% of GPT-4o answers contain Glasp, with a 5.70-point posterior difference and 95% interval [3.12, 8.27].The ordering reverses for English questions, and GPT-4o more often names Glasp first conditional on occurrence.
  • 5.3 Estimation of Carryover and Saturation: Estimating carryover and saturation reduces plug-in coefficient RMSE from 1.207 to 0.945, while the joint posterior reaches 0.950.The interval for the joint-minus-plug-in difference is [−0.006, 0.016], and the effect-RMSE difference also contains zero.
  • 5.4 Use of Response Data in the Measurement Model: For estimated occurrence probabilities, effect RMSE is 17.326% for the cut posterior and 17.480% for the joint posterior.Their difference is 0.154 percentage points with interval [0.046, 0.259].
  • 5.5 Accuracy of the Components of the Treatment Effect: The mean simulated total effect is 266.944 in the controlled design and 234.755 with estimated occurrence probabilities.The source-state component accounts for most of the simulated total effect despite its coefficient being unknown to estimators.
  • 5.5 Accuracy of the Components of the Treatment Effect: In the estimated-occurrence design, total-effect RMSE is 17.359% for plug-in estimation, versus 47.129% for the GEO-input component.Opposite-signed component errors can make the total more accurate than either component.

6 Discussion

The discussion separates media-input construction from causal comparison and clarifies how treatment design, identification, and uncertainty affect GMMM. It emphasizes that estimator performance is context-dependent and that applications must account for uncertainty and omitted channels.

  • 6 Discussion: GMMM constructs expected noticed GEO and GEM inputs before comparing complete treatment sequences through the response model.GEO changes occurrence probabilities, whereas GEM changes the spending sequence determining sponsored placements.
  • 6 Discussion: Occurrence rates vary by language and use case, so application weights should match the target population and business question.Equal weights estimate the fixed panel rate and only represent a population rate when that population has the same composition.
  • 6 Discussion: Identification requires GEO-input variation beyond proportional movement with the source-state variable after other regressors are removed.When regressors move together, the null-space condition determines whether ΔG remains identified.
  • 6 Discussion: Causal interpretation depends on treatment assignment design: randomized rollouts supply variation, while observational rollouts require controls, support, and contemporaneous comparison groups.If earlier treatment affects later assignment, the g-formula averages over covariate distributions under each treatment sequence.
  • 6 Discussion: No estimator ranks uniformly: the joint posterior improves coefficient RMSE in one low-noise condition but worsens effect RMSE under estimated occurrence probabilities.The cut posterior keeps measurement-parameter distributions fixed by measurement data, whereas the joint posterior lets response data revise them.
  • 6 Discussion: Numerical approximations require candidate-count and placement checks, while Laplace and Gaussian approximation errors remain unchanged when candidate counts increase.This separates importance-weight concentration from approximation errors in the measurement and posterior calculations.
  • 6 Discussion: The simulations hold GEO and GEM inputs fixed across estimators and omit established media channels, isolating estimation of a common GMMM response model.Input-construction comparisons would require holding response modeling and transformation estimation fixed.
  • 6 Discussion: In applications, estimated question counts and changing system shares can be added to the measurement model so their uncertainty propagates into ΔG.This changes input construction but not the response model or treatment-effect definition.

7 Conclusion

GMMM extends MMM to unrecorded generative-AI media inputs and defines GEO effects through complete treatment-sequence comparisons. Identification can hold even when component coefficients are not separately recoverable, while simulations highlight the importance of estimating carryover and saturation.

  • Framework: GMMM defines noticed GEO and GEM inputs for each market and period before applying carryover and Hill transformations.GEO uses expected noticed generated-answer occurrences; GEM uses expected noticed sponsored placements.
  • Identification and results: Treatment effects can be identified in some designs even when their component coefficients cannot be recovered separately.The simulations also show that the total effect can conceal larger errors in its two model components.
  • Identification and results: Estimating carryover and saturation explains most of the improvement in coefficient recovery over a plug-in method that fixes them.Allowing response data to revise measurement parameters provides no consistent benefit for estimating ∆G.
  • Causal effects: The GEO treatment effect compares business responses under two complete treatment sequences rather than relying on a single regression coefficient.This accounts for carryover from earlier periods and the resulting changes across the affected sequence.
  • Empirical scope: The primary empirical index describes a fixed panel of 56 questions and equals a market occurrence rate only when its weights reproduce the market question distribution.Its intervals exclude uncertainty from replacing that panel with a different question population.

A.2 Occurrence Results

Occurrence rates differ by model, language, and use case, while the simulation evaluates how GMMM estimates coefficients and GEO treatment effects from collected-answer probabilities. Joint GMMM has lower coefficient RMSE than Plug-in GMMM at five answers per cell, whereas paired treatment-effect RMSE comparisons include zero.

  • Occurrence rates: 5.70 percentage points: GPT-5.6 Luna’s aggregate occurrence rate exceeds GPT-4o’s, with posterior means of 34.5% and 28.8%.The 95% posterior interval for the difference is [3.12, 8.27] percentage points, with posterior probability above 0.999 that it is positive.
  • Occurrence rates: 6.79 and 18.75 percentage points: the model ordering reverses by language, with GPT-4o higher on English and GPT-5.6 Luna higher on Japanese.The corresponding 95% posterior intervals are [2.90, 10.02] and [14.11, 21.57] percentage points.
  • Occurrence rates: Web-highlighting questions have the highest occurrence rates for both models, while knowledge-management questions have the lowest.The reversal by language is especially pronounced for summarization with AI and YouTube learning, motivating system-specific question effects.
  • Simulation results: 0.942 versus 1.120 and 0.981: Joint GMMM has lower coefficient RMSE than Plug-in and similar RMSE to Two-stage GMMM with five answers per cell.The Joint-minus-Plug-in difference is −0.178 with a 95% bootstrap interval of [−0.261, −0.096], while the Joint-minus-Two-stage interval includes zero.
  • Simulation results: Adding a randomized estimate lowers GEO treatment-effect RMSE from 17.576% to 17.057%, while shifting that estimate raises signed bias from 0.380% to 3.813%.The paired RMSE difference for the aligned estimate is −0.519 percentage points with a 95% bootstrap interval of [−0.862, −0.182].

B.3 Controlled Simulation Results

Across controlled simulations, Joint GMMM generally improves recovery of response coefficients, while treatment-effect accuracy remains similar across estimators and depends strongly on response noise and transformation handling.

  • Controlled comparisons: 0.950 vector RMSE for Joint GMMM is below 1.207 for Plug-in GMMM and 1.032 for Two-stage GMMM with five answers per cell.The oracle vector RMSE is 0.880, indicating that response variation remains important even when media inputs and transformations are known.
  • Controlled comparisons: 0.256 lower vector RMSE versus Plug-in GMMM and 0.082 lower versus Two-stage GMMM favor Joint GMMM in paired comparisons.Both 95% bootstrap intervals exclude zero, while treatment-effect RMSE remains similar because coefficient errors can offset in the scalar estimand.
  • Controlled comparisons: Across answer counts, Joint GMMM has lower coefficient error than Plug-in GMMM, but their GEO treatment-effect errors remain close.Figures 4 and 5 show this pattern while varying the number of answers per cell.
  • Transformation estimation: 0.945 vector RMSE results when transformations are estimated with fixed measurement inputs, compared with 1.207 for the corresponding plug-in setup.The paired difference is −0.262 with a 95% interval of [−0.339, −0.182].
  • Transformation estimation: 17.541% and 17.525% are the relative GEO-effect RMSEs for the compared transformation-estimation procedures, with an interval containing zero.Joint GMMM has vector RMSE 0.950, only 0.005 above the plug-in estimator in this comparison.
  • Cut and joint posteriors: 17.359%, 17.326%, and 17.480% are the effect RMSEs for Plug-in, Cut, and Joint, and no paired comparison across five conditions favors Joint.The joint posterior can improve coefficient recovery without improving GEO treatment-effect estimation.

B.7 Alternative Data-Generating Processes

Alternative data-generating processes show that estimator performance depends on misspecification and response noise: Joint GMMM can reduce treatment-effect error under shifted transformations, but does not uniformly dominate.

  • Alternative data-generating processes: 19.133% relative GEO-effect RMSE for Joint GMMM under shifted transformations is below 23.131% for Plug-in and 22.550% for Two-stage GMMM.Under combined departures, Joint GMMM has RMSE 18.059%, bias −4.683%, and coverage 91.7%.
  • Randomized estimate: 17.430% changes to 17.424% in effect RMSE and 0.950 to 0.947 in vector RMSE when a randomized estimate is added with 20 answers per cell.When the randomized estimate is shifted, signed bias is 4.169% versus 1.243% without it.

C.3 Numerical Comparison on Common Samples

Common-sample diagnostics show how paired RMSE comparisons and effective sample size describe estimator differences and importance-weight concentration, while channel interactions require additive attribution rules.

  • Paired comparisons: Negative paired RMSE differences favor the first method, with differences in relative GEO-effect RMSE expressed in percentage points.Each comparison uses the same 120 replications and 5,000 paired bootstrap resamples.
  • Integration diagnostics: Measurement ESS and transformation ESS summarize weight concentration, while the Cut marginal is uniform by definition and has no measurement-ESS entry.Minimum ESS and conditional ESS agree for Cut and Joint because their conditional transformation weights are the same.
  • Integration diagnostics: 1.602% is the largest change in the posterior-mean treatment effect when candidate counts increase from M = 128, J = 700 to M = 256, J = 1,400.The largest interval-bound change is 4.357%, while the 120-replication comparisons remain unchanged.
  • Channel attribution: GEO and GEM effects from disabling one channel need not sum to the joint effect because their interaction belongs jointly to both channels.Shapley values assign the interaction once and satisfy ΦG+ΦP = V(1,1)−V(0,0).

E Consequences of Treatment Design in Simulation

Treatment-design simulations show that difference-in-differences can remove shared platform growth, while post-treatment adjustment recovers a direct effect only under the stated independence condition.

  • Platform growth: Bias falls from 0.081 with zero coverage to below 0.001 with 0.954 coverage when difference-in-differences removes shared platform growth.The corresponding RMSE is 0.008 across 500 replications.
  • Post-treatment adjustment: A response model omitting brand search estimates 1.579 for a true total effect of 1.58, combining direct GEO effect 0.70 and mediated effect 0.88.The additive design separates these components as a direct effect of 0.700 when conditioning on brand search.
  • Post-treatment adjustment: Conditioning on brand search identifies the direct GEO effect only because the simulation assumes independent disturbances for brand search and the response.The passage explicitly does not extend this interpretation to arbitrary variables affected by treatment.

F Public Referral Traffic Analysis

The referral analysis estimates an immediate treated-to-control ratio while accounting for trends, uncertain rollout timing, and a later differential measurement change. Results consistently indicate ratios above one, but inference is sensitive to the treatment date and measurement specification.

  • F Public Referral Traffic Analysis: The referral application cannot construct GEO or GEM inputs because the repository lacks question counts, generated answers, notice observations, and sponsored-placement data.It therefore analyzes referral indices alone, using December 30, 2025 as the central transition date and nearby dates as alternatives.
  • F Public Referral Traffic Analysis: The analysis uses log treated-to-control ratios with segmented level and slope terms, Newey–West inference, and alternative treatment sequences around the rollout window.The specification compares the observed post-treatment path with the pre-treatment relation extrapolated beyond the intervention.
  • F Public Referral Traffic Analysis: For December 30, the immediate ratio is 1.842, while the fitted final-four-week geometric mean is 2.301 versus an observed adjusted mean of 2.055.The placebo rank is 0.150, but pre-treatment changes at least as large as rollout mean this diagnostic does not establish a treatment effect.
  • F Public Referral Traffic Analysis: Moving the first post-treatment week from December 16, 2025 to January 6, 2026 changes the immediate ratio from 1.183 to 2.404.Intervals for December 16 and December 23 include one, whereas those for December 30 and January 6 exclude it because the transition window does not identify one breakpoint.
  • F Public Referral Traffic Analysis: Modeling a March 17 measurement change lowers the ratio from 1.713 under a level change to 1.387 when the slope also changes.The later measurement change reflects altered session composition after bot-filtering changes and may remain differentially expressed in the treated-to-control ratio.
  • F Public Referral Traffic Analysis: The immediate ratio exceeds one in every reported specification, but interval exclusion of one depends on the first post-treatment week and measurement-change treatment.Reported ratios range from 1.842 to 2.647 across response specifications, while the baseline interval remains above one under Newey–West lag choices from zero through 12.

H.2 Alternative Data-Generating Processes

The alternative-data-generating-process experiments vary measurement and response specifications while preserving panel dimensions and treatment sequences. They evaluate robustness across overdispersed inputs, altered transformations, nonlinear response departures, and alternative estimators.

  • H.2 Alternative Data-Generating Processes: Five designs alter measurement models, response equations, or both while retaining panel dimensions, treatment sequences, candidate counts, and conditional posterior samples.The designs include occurrence overdispersion, sponsored-placement overdispersion, alternative transformations, and an added nonlinear response term.
  • H.2 Alternative Data-Generating Processes: The paired comparisons report 95% bootstrap intervals from 5,000 resamples of 120 common panels, with effect RMSE differences measured in percentage points.A positive difference favors the plug-in estimator.
  • H.2 Alternative Data-Generating Processes: The simulation uses 80 clusters over 48 periods on three platforms, with half receiving GEO in period 24 and 500 replications comparing difference-in-differences estimators.The platform paths combine deterministic growth, seasonal variation, and a random walk; the GEO log-odds coefficient is 0.65.
  • H.2 Alternative Data-Generating Processes: A randomized encouragement simulation sets the direct effect at 0.70 and the brand-search-mediated effect at 0.88, producing a total effect of 1.58.Across 800 replications, it compares a response regression on treatment with one additionally conditioning on brand search.

I Computation for the Referral Analysis

The referral computation implements segmented regression robustness checks across treatment dates, measurement changes, response specifications, lag lengths, influential observations, and generated-answer configurations. It uses Newey–West inference, bootstrap intervals, placebo diagnostics, and fixed model instructions for the generated answers.

  • I Computation for the Referral Analysis: The computation varies treatment dates, trend specifications, lag lengths, and influential observations after fitting the segmented referral regression.The analysis uses 47 weekly observations and considers every Newey–West lag from zero through 12.
  • I Computation for the Referral Analysis: Candidate measurement changes add post-change level terms, or both level and slope terms, to the log-ratio regression while holding the central treatment week fixed.The analysis fits these specifications to all sessions and engaged sessions, with a truncated version using observations strictly before March 10.
  • I Computation for the Referral Analysis: The final-four-week fitted-ratio interval uses 5,000 residual moving-block bootstrap replications, distinct from the Newey–West t-test interval.Residuals are sampled in circular four-week blocks within pre- and post-treatment segments before refitting the regression.
  • I Computation for the Referral Analysis: The placebo rank is descriptive rather than a randomization-test p value because the treatment date was not randomly selected from the candidate dates.The placebo procedure evaluates every admissible pre-treatment split with at least four observations on each side.
  • I Computation for the Referral Analysis: Generated answers were collected under common system instructions requiring web search, concrete product recommendations, and no mention of the audit or target brand.Requests varied only the response language and recorded model identifier while retaining search sources and default sampling temperature.
Loading 2609.11915v1…