EconBase
← Back to paper

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

74,892 characters

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact





\maketitle




\begin{abstract}
Generative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm's name in generated answers. We develop Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization (GEO) and Generative Engine Marketing (GEM). For GEO, GMMM combines repeated generated answers with question counts, shares of use across generative systems, and notice probabilities. For GEM, it combines records of sponsored placements with notice probabilities. GMMM compares expected business responses under alternative treatment sequences and establishes sufficient conditions for identifying the resulting effects. We investigate the empirical performance of the proposed method using simulated answers to product recommendation in English and Japanese.
\end{abstract}

\noindent\textbf{Keywords:} marketing mix modeling; generative engine optimization; generative engine marketing; causal inference; measurement error; channel attribution; Bayesian inference

\section{Introduction}
Marketing mix modeling (MMM) relates an aggregate response, such as sales or conversions, to media inputs observed across markets and periods. Two features of advertising motivate the transformations used in MMM. An advertisement can affect the response after the period in which it is shown, so a carryover function combines current and past inputs. The marginal change in response can also become smaller when the accumulated input is already high, so a saturation function maps the carried-over input into a bounded or slowly increasing regressor. MMM applies these transformations before estimating the coefficients of the media channels \citep{Nerlove1962optialadvertizing,Clarke1976econometricmeasurement,Hanssens2001marketresponse}.

Generative artificial intelligence (AI) changes what must be measured before this model can be used. Through Generative Engine Optimization (GEO), a firm modifies source material that may alter answers about the firm, for example by adding a frequently asked questions section or revising product documentation. Through Generative Engine Marketing (GEM), a firm pays for sponsored placements in generated answers \citep{Aggarwal2024geogenerative,Feizi2026onlineadvertisements,Hu2026gembencha}. Conventional impression logs do not record nonsponsored occurrences of a firm's name in generated answers, and a platform record of sponsored placements does not show whether users noticed them.

The frequency with which collected answers contain a firm's name is not yet a media input for a market and period. To obtain an expected count for a market and period, that frequency must be combined with the number of relevant questions, the share handled by each generative system, and the probability that a user notices the name. Referral sessions measure a different event because users may read an answer without following a link. For sponsored placements, spending must likewise be related to the number of placements shown and the probability of notice. GEO and GEM require this additional measurement step before they can be analyzed alongside media channels whose impressions are recorded directly.

Estimating a treatment effect adds a second requirement. Carryover makes the response in period $t$ depend on inputs from earlier periods, so disabling GEO or GEM changes the entire sequence of media inputs over the affected periods. The relevant comparison must replace the full sequence under one treatment with the full sequence under another treatment before the response is evaluated. A regression coefficient by itself does not perform this comparison. Throughout this study, GEO denotes the specified source modification, and its treatment effect is the total effect of applying that modification. The response may change through the generated answers represented by the GEO input and through other consequences of the modification, such as conventional search traffic or conversion on a landing page.

\subsection{Contributions}
We develop GMMM to connect measurements of generated answers and sponsored placements to MMM. For GEO, GMMM models the probability that an answer contains a specified feature and combines that probability with market query counts and notice probabilities to obtain the expected number of noticed occurrences under each source state. This count can be positive without the source modification because GEO changes the occurrence probability rather than creating the entire input from zero. For GEM, GMMM relates spending to the number of sponsored placements and then accounts for notice. The resulting inputs receive the same carryover and Hill transformations as established media before they enter the response model.

Estimation need not be Bayesian. A likelihood or a penalized criterion can support plug-in, joint, and other regularized procedures. We use a Bayesian formulation to average over uncertainty in the occurrence and notice probabilities. The cut posterior leaves the distribution of the measurement parameters determined by the measurement data, whereas the joint posterior also allows the response data to update it. Comparing the two procedures shows how estimation of the media transformations and uncertainty about the constructed inputs affect the treatment effect of GEO.

The identification analysis gives conditions under which the treatment effect remains identifiable even when the coefficients on the GEO source state and the GEO input cannot be determined separately after the media transformations are fixed. It also characterizes when replacing the expected market count by a sequence of occurrence probabilities changes only the scale and when changes in market composition prevent that reduction.

The empirical collection contains 2,240 complete answers from GPT-5.6 Luna and GPT-4o. Each model answers the same set of 56 questions, comprising 28 English questions and 28 Japanese questions. Every call uses the same system instruction and requires web search. The target name is not included in the system instruction or in the questions. Glasp occurs in $33.8\%$ of the GPT-5.6 Luna answers and $27.8\%$ of the GPT-4o answers, although GPT-4o has the higher rate for the English questions. These observations estimate occurrence probabilities under the source state present during collection. One simulation design uses those estimates for the baseline occurrence probabilities, while the simulation itself generates the treatment effect. Across the controlled comparisons, estimating carryover and saturation explains most of the improvement in coefficient recovery over a plug-in method that fixes them, while updating the measurement parameters with response data has no uniform advantage for estimating the treatment effect.

\subsection{Related Work}
Classical MMM combines distributed lags with nonlinear response functions, and recent methods use regularization or Bayesian pooling across markets \citep{Jin2017bayesianmethods,Ng2024bayesiantime,Runge2025packagingup,Gong2024causalmmmlearning,Sun2017geolevelbayesian}. These methods estimate a response model once its media inputs have been defined. Causal interpretation requires further evidence because predictive fit alone does not identify an advertising effect. Work on incrementality and geographic experiments studies the assignment mechanisms and randomized comparisons that can support such an interpretation \citep{Chan2017challengesand,Lewis2015theunfavorable,Gordon2019acomparison,Vaver2011measuringad,Chen2022robustcausal}.

Research on GEO studies how changes in source material affect generated answers and their ranking, while work on GEM considers sponsored retrieval and placement \citep{Bagga2026egeoa,Kim2026sageoarena,Martinez2026optimizingvisibility,Hajiaghayi2024adauctions,Dubey2024auctionswith,Dutting2024mechanismdesign,Xu2026adinsertion}. Studies of GEO treat citation of a source and use of its content in an answer as different outcomes \citep{Zhang2026fromcitation}. Neither quantity alone gives the number of users who notice the target feature. GMMM addresses the next step by placing the occurrence probability on a market and time scale suitable for MMM.

The use of records on reach and frequency in MMM provides a precedent for replacing spending with quantities closer to exposure \citep{Zhang2023bayesianhierarchical}. Randomized estimates have likewise been incorporated through informative priors or likelihood terms \citep{Zhang2024mediamix}. Meridian and PyMC-Marketing are Bayesian implementations of MMM, and Meridian supports inputs based on reach and frequency.\footnote{See \url{https://developers.google.com/meridian} and \url{https://www.pymc-marketing.io/}.} Our contribution concerns the measurements required when the media input itself must be inferred from generated answers or sponsored placements.

Cut and joint posteriors arise in modular Bayesian inference \citep{Plummer2015cutsin,Jacob2017bettertogether,Carmona2020semimodularinference}. A cut posterior passes uncertainty from one module to another without allowing the later module to revise the first distribution. Its interpretation in GMMM depends on the assignment of parameters to modules: fixing the prior for the media transformations also prevents their estimation from response data, whereas cutting the update of the measurement parameters still permits those transformations to be estimated.

\section{Setup}
\label{sec:setup}
We use \emph{response} for the business variable analyzed by MMM and \emph{answer} for text returned by a generative system. A \emph{treatment} is a specified setting of GEO or GEM. When several periods are involved, a treatment sequence lists those settings from period $1$ through period $T$ and includes any earlier values needed to initialize carryover.

\subsection{Response and Media Inputs in MMM}
For every observational unit $i\in\{1,\ldots,n\}$ and period $t\in\{1,\ldots,T\}$, let $Y_{it}$ denote the response and let $W_{it}$ contain baseline variables and controls observed before the period-$t$ treatment. The unit may be a geographic market or another aggregation used consistently in the measurement and response models. Let $\mathcal M_0$ be the set of established media channels. For every $m\in\mathcal M_0$, let $E_{mit}\geq0$ denote the media input in period $t$ and define its sequence through period $t$ by $\boldsymbol E_{mi,1:t}=(E_{mi1},\ldots,E_{mit})$.

A standard MMM transforms each sequence of media inputs before including it in the response model. We write
\begin{align}
H_{mit}
&=
h_m\left(
A_m(\boldsymbol E_{mi,1:t};\alpha_m);\theta_m
\right),
\label{eq:general-media-transform}
\end{align}
where $A_m$ combines current and earlier inputs according to carryover parameters $\alpha_m$, and $h_m$ represents saturation with parameters $\theta_m$. Let $b(W;\gamma)$ be the baseline response with coefficient vector $\gamma$. If $\mathcal I_0$ is a specified set of distinct pairs of media channels, the conditional mean satisfies
\begin{align}
g_Y\left(\mu^{\mathrm{MMM}}_{it}\right)
&=
b(W_{it};\gamma)
+
\sum_{m\in\mathcal M_0}\beta_mH_{mit}
+
\sum_{(m,m')\in\mathcal I_0}
\beta_{mm'}H_{mit}H_{m'it},
\label{eq:general-mmm}
\end{align}
where $\mu^{\mathrm{MMM}}_{it}=\mathbb{E}[Y_{it}\mid\mathcal H_{it}]$. The information set $\mathcal H_{it}$ contains the controls and the sequences of media inputs used to model $Y_{it}$. We specify $Y_{it}\mid\mathcal H_{it}\sim\mathcal F(\mu^{\mathrm{MMM}}_{it},\psi)$, where $\psi$ contains the remaining parameters of the response distribution. The identity link gives the usual model for a continuous response with additive noise of mean zero, while other links accommodate counts or rates.

\subsection{Channels Through Generative AI}
GMMM enlarges the media set to $\mathcal M=\mathcal M_0\cup\{G,P\}$, where $G$ indexes the input constructed from generated answers and $P$ indexes the input constructed from sponsored placements. For each period, the first input is the expected number of generated answers in which the specified property occurs and is noticed under the current source state. The second is the expected number of sponsored placements that users notice. The first count need not be zero when the source modification is absent; GEO changes its value by changing generated answers. Inputs for established media are observed directly, while the two additional inputs are constructed from generated answers, platform records, market counts, and observations of user attention.

Let $Z_{it}\in\{0,1\}$ denote the GEO source state, where $Z_{it}=1$ means that the specified source modification is present and $Z_{it}=0$ means that it is absent. The response specification studied here is
\begin{align}
g_Y\left(\mu_{it}\right)
={}&
b(W_{it};\gamma)
+
\sum_{m\in\mathcal M_0}\beta_mH_{mit}
+
\beta_D Z_{it}
+
\beta_GH^G_{it}
+
\beta_PH^P_{it}
+
\beta_{GP}H^G_{it}H^P_{it}.
\label{eq:gmmm-response}
\end{align}
The term $\beta_D Z_{it}$ permits the source modification to affect the response through variables that are not represented by the constructed GEO input, such as conventional search traffic or conversion on a landing page. The coefficient $\beta_{GP}$ permits the effect associated with one generative channel to depend on the other. Either term may be omitted when the corresponding mechanism is excluded from the model.

\subsection{Data Sources}
Let ${\mathcal{D}}_M$ contain the observations used to construct the GEO and GEM inputs. For GEO, these observations include repeated generated answers, counts of relevant questions, shares of use across generative systems, and observed notice indicators. For GEM, they include spending, the number of sponsored placements shown, predictors of that number, and observed notice indicators. Let ${\mathcal{D}}_Y$ contain $Y_{it}$ together with the controls and media inputs in the response model. When a randomized experiment estimates the same treatment effect for a comparable population and evaluation period, its estimate and standard error form ${\mathcal{D}}_E$.

The collection described in Section~\ref{sec:answer-collection} contains 20 answers from each of two GPT models for each of 56 questions, comprising 28 English questions and 28 Japanese questions. The public referral series analyzed in Appendix~\ref{sec:empirical} comes from an earlier period and lacks the corresponding question counts, generated answers, and notice observations. These missing observations prevent the two sources from being combined in one GMMM analysis.

The simulations use units of the form $i=(g,j)$, where $g\in\{1,\ldots,5\}$ indexes geographic markets and $j\in\{1,\ldots,6\}$ indexes product clusters. Each cluster contains pages and questions that respond to the same GEO treatment. A substantive application must use an observational unit that is common to the construction of the media inputs, the response model, and the treatment effect.

\subsection{Treatment Effect}
Because carryover links the period-$t$ response to media inputs observed before $t$, we define the treatment effect from complete treatment sequences over a specified evaluation window.

Let $a_G\in\{0,1\}$ select one of two GEO sequences: $a_G=1$ selects the specified sequence with GEO, and $a_G=0$ selects the sequence with the source modification removed. Let $a_P\in\{0,1\}$ similarly select the specified GEM spending sequence or the sequence with GEM removed. The pair $(a_G,a_P)$ selects one complete GEO sequence and one complete GEM sequence across all periods. The comparison uses the same history before period $1$ unless an intervention begins earlier; in that case, the treatment and input values before period $1$ follow the selected sequence. The simulations set all media inputs before period $1$ to zero. For an evaluation set ${\mathcal{T}}_E\subseteq\{1,\ldots,T\}$, define
\begin{align}
V(a_G,a_P)
&=
\sum_{i=1}^{n}
\sum_{t\in{\mathcal{T}}_E}
\mathbb{E}\left[Y_{it}(a_G,a_P)\right],
\label{eq:business-value}
\end{align}
where $Y_{it}(a_G,a_P)$ is the potential response under the selected treatment sequences and the specified evolution of all other variables. The primary estimand is
\begin{align}
\Delta_G
&=
V(1,1)-V(0,1).
\label{eq:geo-treatment-effect}
\end{align}
We refer to $\Delta_G$ as the treatment effect of GEO. It is the total effect of the specified source modification, including the change transmitted through $H^G$ and any additional change represented by $\beta_DZ_{it}$. The contrast retains the specified GEM sequence. The evaluation window may extend beyond the periods in which GEO is active so that the effect includes responses associated with carried-over inputs. The analogous treatment effect of GEM is $\Delta_P=V(1,1)-V(1,0)$, which compares the specified GEM sequence with the sequence in which GEM is disabled while retaining the GEO sequence.

\section{Generative Marketing Mix Modeling}
\label{sec:method}
The data record the source state, spending, and established media inputs, but they omit the expected counts associated with GEO and GEM. GMMM constructs these two counts before applying the standard media transformations and estimating the response model.

\subsection{The GEO Input}
Let $q\in\{1,\ldots,Q\}$ index clusters of questions and let $p\in\mathcal P$ index generative systems. A system is defined by the model version and the settings used to obtain an answer. The set $\mathcal Q(i)$ contains the question clusters associated with unit $i$. For every $i$, $q$, and $t$, let $N_{iqt}\geq0$ be the number of times that a question in cluster $q$ is submitted in unit $i$ during period $t$. For every $i$, $p$, and $t$, let $\omega_{ipt}\in[0,1]$ be the share of those questions handled by system $p$, with $\sum_{p\in\mathcal P}\omega_{ipt}=1$.

Before collecting answers, the analyst specifies the property to be recorded. For repetition $r\in\{1,\ldots,R_{qpt}\}$ and source state $z\in\{0,1\}$, let $X_{qptr}(z)=1$ when the stored answer has that property and let it equal zero otherwise. In the empirical collection, the property is the occurrence of a spelling of the target name in a fixed dictionary. Let
$\pi_{qpt}(z)=\Prb(X_{qptr}(z)=1)$ be its occurrence probability. Conditional on $X_{qptr}(z)=1$, let $\lambda_{qpt}\in[0,1]$ be the probability that the user notices the recorded property. The input associated with GEO is
\begin{align}
E^G_{it}(z)
&=
\sum_{q\in\mathcal Q(i)}
N_{iqt}
\sum_{p\in\mathcal P}
\omega_{ipt}\pi_{qpt}(z)\lambda_{qpt}.
\label{eq:geo-exposure}
\end{align}
The quantity $E^G_{it}(z)$ is the expected number of generated answers in unit $i$ and period $t$ for which the specified property occurs and is noticed under source state $z$. It may be positive when $z=0$; the change in this input caused by the source modification is $E^G_{it}(1)-E^G_{it}(0)$. A user who encounters the property more than once contributes more than once to this count. The comparison between $z=1$ and $z=0$ holds $N_{iqt}$ and $\omega_{ipt}$ fixed. We assume that the source modification does not change notice conditional on occurrence, which is why $\lambda_{qpt}$ has no argument $z$.

Equation~\eqref{eq:geo-exposure} pools occurrence and notice probabilities across markets after conditioning on the question cluster, generative system, period, and source state. It also uses a system share that is common across question clusters within a market and period. When the data distinguish these cells, the same construction can use $\pi_{iqpt}(z)$, $\lambda_{iqpt}$, and $\omega_{iqpt}$ without changing the media transformations or the treatment contrasts.

We estimate the occurrence probabilities with the hierarchical model
\begin{align}
X_{qptr}\mid\pi_{qpt}
&\sim\operatorname{Bernoulli}(\pi_{qpt}),
\label{eq:occurrence-binomial}\\
\operatorname{logit}(\pi_{qpt})
&=
A_{qpt}^{\mathsf T}\alpha+u_q,
\qquad
\boldsymbol u=B_Q\sigma_u\widetilde{\boldsymbol u},
\qquad
\widetilde{\boldsymbol u}\sim\mathcal N(0,I_{Q-1}).
\label{eq:occurrence-logit}
\end{align}
The columns of $B_Q\in\mathbb{R}^{Q\times(Q-1)}$ form an orthonormal basis for vectors whose coordinates sum to zero, so $\sum_{q=1}^{Q}u_q=0$. The predictor vector $A_{qpt}$ may contain indicators for the generative system, the source state, and other observed determinants of occurrence. Its coefficient vector is $\alpha$. The effect for question cluster $q$ is represented in noncentered form by standard normal coordinates $\widetilde{\boldsymbol u}$ and scale $\sigma_u>0$. Replacing the observed sequence of source states by another treatment sequence yields probabilities under that sequence only when the data contain the required variation in source state. Answers collected under one source state identify only the occurrence probabilities for that state; the change caused by the source modification requires observations under both states.

Suppose that a user study presents generated answers and records $C_{qpt}$ cases in which participants notice the property among $n_{qpt}$ answers in which it occurs. We use
\begin{align}
C_{qpt}\mid\lambda_{qpt}
&\sim\operatorname{Binomial}(n_{qpt},\lambda_{qpt}),
\label{eq:notice-binomial}\\
\operatorname{logit}(\lambda_{qpt})
&=
(D^\lambda_{qpt})^{\mathsf T}\xi+v_q,
\qquad
\boldsymbol v=B_Q\sigma_v\widetilde{\boldsymbol v},
\qquad
\widetilde{\boldsymbol v}\sim\mathcal N(0,I_{Q-1}).
\label{eq:notice-logit}
\end{align}
Here, $D^\lambda_{qpt}$ is a vector of observed predictors with coefficient vector $\xi$. The effect for question cluster $q$, denoted by $v_q$, uses the same zero-sum basis and its own scale $\sigma_v>0$. Together, the two models estimate the probabilities needed in \eqref{eq:geo-exposure}.

\subsection{The GEM Input}
A sponsored placement contributes to the GEM input only if the platform shows it and the user notices it. Let $S^P_{it}\geq0$ denote the spending level selected by the firm for the GEM intervention in unit $i$ and period $t$, before the platform determines delivery. Let $L^P_{it}$ denote the number of sponsored placements shown. Conditional on a placement being shown, let $\rho^P_{it}\in[0,1]$ be the probability that the user notices it. For positive spending, we model the number shown by
\begin{align}
L^P_{it}\mid S^P_{it}>0
&\sim\operatorname{Poisson}(\mu^P_{it}),
\label{eq:paid-placement-distribution}\\
\log\mu^P_{it}
&=
\phi_0
+
\phi_S\log S^P_{it}
+
\phi_DD_{it}
+
\phi_RR_{it},
\qquad
\phi_S>0,
\label{eq:paid-placement-mean}
\end{align}
where $D_{it}$ is a demand predictor observed before the period-$t$ treatment and $R_{it}$ is a promotion indicator. We set $\mu^P_{it}=0$ when $S^P_{it}=0$. The input associated with GEM is
\begin{align}
E^P_{it}
&=
\mu^P_{it}\rho^P_{it}.
\label{eq:paid-exposure}
\end{align}
The restriction $\phi_S>0$ makes the expected number of sponsored placements increase with positive spending. A user study in which participants report whether they noticed each sponsored placement can estimate $\rho^P_{it}$ with a binomial model. The analysis below uses one notice probability for all sponsored placements, while \eqref{eq:paid-exposure} permits variation across units and periods when the data support it. When $L^P_{it}$ is observed and the analysis conditions on realized delivery, $L^P_{it}\rho^P_{it}$ can be used directly. The Poisson model is needed for missing delivery counts and for spending sequences that were not observed.

\subsection{Carryover and Saturation}
The two expected counts are transformed in the same way as inputs for established media. We use normalized geometric carryover,
\begin{align}
A_{mit}
&=
\frac{
\sum_{\ell=0}^{L_m}
\alpha_m^{\ell}E_{mi,t-\ell}/s_m
}{
\sum_{\ell=0}^{L_m}\alpha_m^{\ell}
},
\qquad
0\leq\alpha_m<1,
\label{eq:adstock}
\end{align}
where $L_m\in\{0,1,\ldots\}$ is the maximum lag and $s_m>0$ is a fixed scale used in the analysis. Values before period $1$ come from the earlier history assigned to the treatment sequence; the simulations set these values to zero. The transformed input is
\begin{align}
H_{mit}
&=
\frac{A_{mit}^{\kappa_m}}
{A_{mit}^{\kappa_m}+\theta_m^{\kappa_m}},
\qquad
\theta_m>0,
\qquad
\kappa_m>0.
\label{eq:hill}
\end{align}
The Hill function equals $1/2$ when $A_{mit}=\theta_m$, and $\kappa_m$ determines how sharply it changes near that point. Carryover is applied before the Hill transformation \citep{Jin2017bayesianmethods}. Substituting $H^G_{it}$ and $H^P_{it}$ into \eqref{eq:gmmm-response} completes the response specification.

\subsection{Estimation from Measurement and Response Data}
Let $\eta_M$ collect the parameters used in \eqref{eq:geo-exposure} and \eqref{eq:paid-exposure}, let $\eta_T$ collect the carryover and Hill parameters, and let $\vartheta$ collect the response parameters. We write $\mathcal J_M(\eta_M;{\mathcal{D}}_M)$ for a loss based on the data used to construct the two inputs and $\mathcal J_Y(\vartheta,\eta_M,\eta_T;{\mathcal{D}}_Y)$ for a loss based on the response data. If ${\mathcal{D}}_E$ contains a randomized estimate of the same treatment effect, its contribution is $\mathcal J_E(\vartheta,\eta_M,\eta_T;{\mathcal{D}}_E)$; otherwise, we set $\mathcal J_E=0$. A general estimator minimizes
\begin{align}
\mathcal J(\eta_M,\eta_T,\vartheta)
&=
\mathcal J_M
+
\mathcal J_Y
+
\mathcal J_E
+
\mathcal P_M(\eta_M)
+
\mathcal P_T(\eta_T)
+
\mathcal P_Y(\vartheta),
\label{eq:general-criterion}
\end{align}
where the penalty terms may be zero. With negative log likelihoods and no penalties, this criterion gives maximum likelihood estimation. Nonzero penalties permit regularization.

The criterion covers several ways of using the data. A plug-in method estimates $\eta_M$ from ${\mathcal{D}}_M$, substitutes that estimate into the response model, and either fixes $\eta_T$ or estimates it from ${\mathcal{D}}_Y$. A joint likelihood estimates $\eta_M$, $\eta_T$, and $\vartheta$ together. Resampling can be used with either procedure to assess uncertainty. A Bayesian version makes both the treatment of uncertainty and the direction of updating explicit.

\subsection{Bayesian Computation}
Minimizing the criterion with negative log likelihoods gives point estimates. In the Bayesian analysis, the cut posterior leaves the distribution of the measurement parameters determined by ${\mathcal{D}}_M$, whereas the joint posterior allows ${\mathcal{D}}_Y$ to update it. Both procedures average over uncertainty in the constructed inputs. The hierarchical occurrence and notice models are fitted in noncentered coordinates. Their posterior modes and analytic Hessians define a Laplace approximation $q_L(\eta_M\mid{\mathcal{D}}_M)$ to the posterior based only on ${\mathcal{D}}_M$. Sampling the transformation parameters from their prior gives
\begin{align}
q(\eta\mid{\mathcal{D}}_M)
&=
q_L(\eta_M\mid{\mathcal{D}}_M)p(\eta_T),
\qquad
\eta=(\eta_M,\eta_T).
\label{eq:measurement-distribution}
\end{align}
Let $\beta=(\beta_D,\beta_G,\beta_P,\beta_{GP})^{\mathsf T}$. After replacing the exact posterior for $\eta_M$ by the Laplace approximation, the joint posterior is
\begin{align}
\widetilde p(\eta,\beta,\sigma_Y^2\mid{\mathcal{D}}_M,{\mathcal{D}}_Y)
&\propto
p({\mathcal{D}}_Y\mid\eta,\beta,\sigma_Y^2)
p(\beta,\sigma_Y^2)
q(\eta\mid{\mathcal{D}}_M).
\label{eq:joint-posterior}
\end{align}
For each $k\in\{1,\ldots,K\}$, we sample $\eta_k\sim q(\eta\mid{\mathcal{D}}_M)$ and construct the corresponding sequences of media inputs. Conditional on $\eta_k$, the response parameters are integrated under the normal inverse-gamma model with the sign restrictions in Appendix~\ref{sec:response-posterior}. Let $m_k$ be the numerical approximation to $p({\mathcal{D}}_Y\mid\eta_k)$ obtained with a Gaussian approximation to the posterior probability of the sign restrictions. The joint calculation assigns weight
\begin{align}
w_k
&=
\frac{m_k}{\sum_{h=1}^{K}m_h}.
\label{eq:importance-weights}
\end{align}
After selecting $\eta_k$ according to these weights, we sample the response parameters from their conditional posterior and evaluate $\Delta_G$.

The two-stage procedure uses equal weights for the samples from \eqref{eq:measurement-distribution}, leaving both the Laplace approximation for $\eta_M$ and the prior for $\eta_T$ unchanged by ${\mathcal{D}}_Y$. A cut posterior keeps only the first of these distributions fixed:
\begin{align}
p_{\mathrm{cut}}(\eta_M,\eta_T,\beta,\sigma_Y^2\mid{\mathcal{D}}_M,{\mathcal{D}}_Y)
&=
q_L(\eta_M\mid{\mathcal{D}}_M)
p(\eta_T,\beta,\sigma_Y^2\mid{\mathcal{D}}_Y,\eta_M).
\label{eq:cut-posterior}
\end{align}
Because the conditional posterior on the right is normalized separately for every $\eta_M$, integrating over $(\eta_T,\beta,\sigma_Y^2)$ leaves $q_L(\eta_M\mid{\mathcal{D}}_M)$ unchanged \citep{Plummer2015cutsin,Jacob2017bettertogether}.

For computation, we use $M$ samples of the measurement parameters and a common set of $J$ samples from the transformation prior. Let $m_{mj}$ approximate $p({\mathcal{D}}_Y\mid\eta_{M,m},\eta_{T,j})$. On this common set, the cut and joint weights are
\begin{align}
w^{\mathrm{cut}}_{mj}
&=
\frac{1}{M}
\frac{m_{mj}}{\sum_{h=1}^{J}m_{mh}},
&
w^{\mathrm{joint}}_{mj}
&=
\frac{m_{mj}}
{\sum_{r=1}^{M}\sum_{h=1}^{J}m_{rh}}.
\label{eq:modular-weights}
\end{align}
The plug-in method that estimates the transformations uses the same $J$ samples at the point estimate of $\eta_M$, while the two-stage procedure assigns weight $1/(MJ)$ to every pair. For the independent samples in \eqref{eq:importance-weights}, we report
\begin{align}
\operatorname{ESS}
&=
\left(\sum_{k=1}^{K}w_k^2\right)^{-1}
\label{eq:ess}
\end{align}
as a numerical diagnostic. A small effective sample size means that a few samples carry most of the weight. For the common set of samples, Appendix~\ref{sec:modular-diagnostics} reports the effective sample size within each measurement sample and the effective sample size of the joint weights across measurement samples.

Figure~\ref{fig:framework} summarizes the order of the calculations.
\begin{figure}[!htbp]
\centering
\includegraphics[width=\textwidth,draft=false]{Figure/figure_framework-eps-converted-to.pdf}
\caption{GMMM constructs the expected counts associated with GEO and GEM before applying carryover, the Hill transformation, and the response model. The plug-in, cut, and joint procedures differ in whether they average over uncertainty in the measurement parameters and whether the response data update those parameters.}
\label{fig:framework}
\end{figure}

\subsection{Randomized Estimate of the GEO Treatment Effect}
A randomized experiment can inform GMMM when it estimates the same treatment effect for a population, response definition, and evaluation window that can be related to the GMMM target. Suppose that the experiment reports $\widehat\tau_E$ with standard error $s_E$, and let $\tau_E(\vartheta,\eta)$ be the corresponding effect implied by the GMMM parameters. We use
\begin{align}
\widehat\tau_E\mid\vartheta,\eta
&\sim
\mathcal N\left(
\tau_E(\vartheta,\eta),
s_E^2+\sigma_{\mathrm{tr}}^2
\right),
\label{eq:randomized-estimate-likelihood}
\end{align}
where $\sigma_{\mathrm{tr}}\geq0$ is fixed before estimation and represents residual differences between the experimental population and the target population after the measured design variables have been aligned. This likelihood contributes $\mathcal J_E$ to \eqref{eq:general-criterion}. In the Bayesian analysis, it changes the weights and the conditional posterior for the response coefficients. For every parameter sample, the experimental effect is computed from the complete treatment and control sequences before the media transformations are applied.

The aligned simulation uses the average treatment effect of GEO, including the term associated with the source state. A second simulation adds a nonzero difference between the experimental and target populations to the same effect and thereby examines misspecification of the transport model. If an additive attribution across interacting channels is also required, Appendix~\ref{sec:interaction-allocation} gives a Shapley allocation as an additional summary.

\section{Identification of the Treatment Effect of GEO}
\label{sec:theory}
Recovering the treatment effect of GEO requires both statistical variation in the response model and observations that support the expected response under each treatment sequence. Statistical variation determines whether the response coefficients or the linear combination that defines the effect can be identified, while the observed market inputs and treatment assignments determine whether the comparison can be evaluated.

\subsection{Identification Within the Response Model}
If the two constructed inputs were observed directly, GMMM would differ from a standard MMM only through the additional term for the GEO source state. The following reduction states this relation before the rank condition for the response coefficients.

\begin{proposition}[Reduction to standard MMM]
Suppose that $\boldsymbol E^G_{i,1:T}$ and $\boldsymbol E^P_{i,1:T}$ are observed for every $i\in\{1,\ldots,n\}$. If $\beta_D=0$, then \eqref{eq:gmmm-response} is an instance of \eqref{eq:general-mmm} with channel set $\mathcal M_0\cup\{G,P\}$ and interaction set $\{(G,P)\}$. If $\beta_{GP}=0$ also holds, the response is additive on this enlarged channel set. Setting all four coefficients associated with GEO and GEM to zero gives the additive model for established media with $\mathcal I_0=\varnothing$.
\end{proposition}
\begin{proof}
The observed sequences determine $H^G$ and $H^P$ through \eqref{eq:general-media-transform}. Substitution into \eqref{eq:gmmm-response}, together with the stated restrictions on the coefficients, gives the corresponding cases of \eqref{eq:general-mmm}.
\end{proof}

GMMM adds the construction of $E^G$ and $E^P$ and the definition of their values under each treatment sequence. Once these quantities are fixed, identification of the response coefficients is a rank problem.

\begin{proposition}[Identification of response coefficients and linear effects]
\label{prop:response-rank}
Fix the sequences of media inputs and all parameters of the media transformations. Suppose that the response model has an identity link and $b(W;\gamma)=W\gamma$. Let $H_0$ be the matrix whose columns are the transformed inputs for established media and let $W_*=(W,H_0)$. Let $X_\eta$ contain the variable for source state, $H^G$, $H^P$, and $H^GH^P$. If $M_{W_*}$ is the orthogonal projection onto the complement of the column space of $W_*$, define
\begin{align}
D_\eta
&=M_{W_*}X_\eta,
&
G_\eta
&=X_\eta^{\mathsf T}M_{W_*}X_\eta.
\label{eq:response-rank}
\end{align}
Conditional on the fixed sequences and transformations, the four response coefficients are identified from the conditional mean if and only if $G_\eta$ is nonsingular. Under a Gaussian conditional response with nonsingular variance, the same condition is equivalent to identification by the likelihood. For a fixed vector $d\in\mathbb{R}^4$, the linear effect $d^{\mathsf T}\beta$ is identified if and only if $d^{\mathsf T}v=0$ for every $v\in\mathbb{R}^4$ such that $D_\eta v=0$.
\end{proposition}
\begin{proof}
Two coefficient vectors $\beta$ and $\beta'$ give the same conditional mean after the nuisance coefficients are adjusted if and only if $X_\eta(\beta-\beta')$ belongs to the column space of $W_*$. Equivalently, $D_\eta(\beta-\beta')=0$. The four coefficients are identified exactly when $D_\eta$ has full column rank, which is equivalent to nonsingularity of $G_\eta$. The linear effect is constant over observationally equivalent coefficient vectors exactly when it vanishes on the null space of $D_\eta$. The necessity statement remains valid under the sign restrictions used in the simulations because their parameter space has nonempty interior.
\end{proof}

The simulations omit established media, so $H_0$ has no columns and residualization with respect to $W$ is sufficient. In the general model, $H_0$ must be included. For example, if a transformed input for an established channel equals $H^G$, its coefficient cannot be separated from $\beta_G$ even when the four columns of $X_\eta$ are linearly independent after removing $W$ alone.

A linear effect may remain identified when its component coefficients are not. This distinction is especially relevant when the variable for source state and $H^G$ are nearly proportional after the other regressors have been removed. Proposition~\ref{prop:response-rank} states the exact null-space condition; finite-sample stability still depends on the degree of collinearity.

The proposition treats the transformations as fixed. Joint identification of $(\beta_G,\alpha_G,\theta_G,\kappa_G)$ additionally requires the observed input sequences to provide independent information about response amplitude, carryover, and the Hill curve. With all other parameters fixed, full column rank of the Jacobian of the residualized mean with respect to these four parameters is sufficient for local identification at an interior point of a continuously differentiable model. If other parameters are unknown, full column rank of the Jacobian with respect to the complete free parameter vector after removal of the fixed linear controls is sufficient. A step treatment whose untransformed GEO input has one level before treatment and another after treatment cannot separate response amplitude from the Hill parameters using steady-state observations alone. Identification then depends on transition dynamics and on variation within source states, including variation in question counts, system shares, or occurrence probabilities. A prior can stabilize estimation without identifying a likelihood that lacks this variation. Related ambiguities arise when nonlinear response and coefficients that change over time produce similar observed sequences but different allocation decisions \citep{Dew2024yourmmm}.

\subsection{Occurrence Probabilities and Market Counts}
The market input in \eqref{eq:geo-exposure} weights an occurrence probability by the number and composition of questions, the shares of use across systems, and the probability of notice. In a special case, this difference amounts only to a change of scale. Let $A(\pi)$ denote linear carryover applied to a sequence of exact occurrence probabilities, with the same initialization as \eqref{eq:adstock}.

\begin{proposition}[Constant rescaling of the GEO input]
\label{prop:exposure-scale}
Suppose that $E^G_{it}(z)=c_0\pi_{it}(z)$ for the same constant $c_0>0$ in every market, period, and evaluated source state, including values used to initialize carryover. For the Hill function $h(a;\theta,\kappa)=a^\kappa/(a^\kappa+\theta^\kappa)$, it holds that
\begin{align}
h(A(E^G);\theta,\kappa)
&=
h(c_0A(\pi);\theta,\kappa)
=
h(A(\pi);\theta/c_0,\kappa).
\label{eq:constant-scale-equivalence}
\end{align}
Rescaling $\theta$ preserves the transformed sequence and the treatment effect computed from it. A single common value of $\theta$ does not generally absorb scale factors that vary across markets or periods.
\end{proposition}
\begin{proof}
Linearity gives $A(c_0\pi)=c_0A(\pi)$. Dividing the numerator and denominator of $h(c_0a;\theta,\kappa)$ by $c_0^\kappa$ proves \eqref{eq:constant-scale-equivalence}. For the final statement, consider two markets with the same positive occurrence-probability sequence and different constants $c_1$ and $c_2$. The transformed market counts differ because $h$ is strictly increasing, whereas a common transformation of the identical probability sequences is the same. One common value of $\theta$ cannot represent both.
\end{proof}

The same equivalence holds within each market if its scale factor $c_i$ is constant over time and the model permits a market-specific parameter $\theta_i$, rescaled to $\theta_i/c_i$. A Bayesian analysis must transform the prior and its support consistently. Under the common-scale condition in Proposition~\ref{prop:exposure-scale}, a common multiplicative level in question volume or notice probability is absorbed by the Hill midpoint. Question volume, system shares, and notice probabilities affect the treatment effect when their relative values vary across markets, periods, systems, question clusters, or source states. When that variation changes the proportionality between occurrence probabilities and expected counts, one occurrence-probability sequence cannot represent every market input. A finite collection of answers adds a separate source of uncertainty about the probabilities. Appendix~\ref{sec:uncertainty-integration} examines both issues in a static linear model and under a nonlinear transformation.

\subsection{Identification across Treatment Sequences}
The rank condition concerns the response model after the inputs have been fixed. Identification of $\Delta_G$ also requires the observed data to determine the expected response under each complete treatment sequence. For every $i\in\{1,\ldots,n\}$ and $t\in\{1,\ldots,T\}$, let $\mathcal H^-_{it}$ contain the variables observed immediately before the period-$t$ GEO and GEM treatments. It includes earlier responses, prior media inputs, earlier treatments, and any aggregate variables used to represent interference. Let $\mathcal H_{it}$ add the period-$t$ treatments and the values of $E^G_{it}$ and $E^P_{it}$ implied by them. For a pair $(a_G,a_P)$ of complete treatment sequences, let $F_{it}^{a_G,a_P}$ be the distribution of $\mathcal H_{it}$ obtained by recursively replacing the observed treatment assignment with those sequences. Finally, define $m_{it}(h)=\mathbb{E}[Y_{it}\mid\mathcal H_{it}=h]$. The following assumptions place the standard longitudinal g-formula on these GMMM treatment sequences.

\begin{assumption}[Consistency]
If the observed GEO and GEM treatments through period $t$ equal the values specified by $(a_G,a_P)$, then the observed variables through period $t$ equal their potential values under those treatment sequences.
\end{assumption}

\begin{assumption}[Sequential exchangeability]
Conditional on $\mathcal H^-_{it}$, the period-$t$ GEO and GEM treatments are independent of future potential information sets and potential responses under the treatment sequences evaluated in \eqref{eq:business-value}.
\end{assumption}

\begin{assumption}[Positivity]
For every value of $\mathcal H^-_{it}$ in the target support, each discrete treatment specified by the evaluated sequences has positive conditional probability. A positive continuous GEM spending level lies in the interior of its conditional support and has positive density in a neighborhood of that level. A sequence that sets spending to zero requires positive conditional mass at zero when the absence of a campaign is represented by a separate point mass.
\end{assumption}

\begin{assumption}[Observed-data laws in the target population]
\label{ass:measurement-population}
The measurement and response data identify $m_{it}$ and the conditional laws used to construct $F_{it}^{a_G,a_P}$ on the support of the evaluated treatment sequences. When the response model uses the expected counts in \eqref{eq:geo-exposure} and \eqref{eq:paid-exposure}, the measurement parameters and market inputs identify those counts for the target population. Identification of $m_{it}$ on the same support must also be established. If $\mathcal H_{it}$ instead contains unobserved realized impressions, their joint law with the observed variables and $Y_{it}$ must also be identified; a marginal distribution of impressions is insufficient.
\end{assumption}

\begin{assumption}[Interference through specified aggregates]
Spillovers within an observational unit are included in that unit's treatments and response. Treatments assigned to other units may affect its potential response or generative inputs only through aggregate variables included in $\mathcal H^-_{it}$. The distribution $F_{it}^{a_G,a_P}$ specifies how those aggregates evolve under the evaluated treatment sequences.
\end{assumption}

Under these conditions, the longitudinal g-formula applies to the complete GEO and GEM treatment sequences.

\begin{theorem}[Longitudinal g-formula for GMMM]
\label{thm:g-formula}
Under the preceding assumptions, the expected response in \eqref{eq:business-value} is identified by
\begin{align}
V(a_G,a_P)
&=
\sum_{i=1}^{n}
\sum_{t\in{\mathcal{T}}_E}
\int
m_{it}(h)
\,dF_{it}^{a_G,a_P}(h).
\label{eq:g-formula}
\end{align}
The difference between this expression for $(a_G,a_P)=(1,1)$ and $(a_G,a_P)=(0,1)$ identifies $\Delta_G$.
\end{theorem}
\begin{proof}
Consistency links observed variables to their potential values under the realized treatments. Sequential exchangeability permits the observed assignment mechanism to be replaced, conditional on $\mathcal H^-_{it}$, by the evaluated treatment sequences. Positivity ensures that the required conditional laws are defined on their support. The measurement assumption identifies the conditional laws and $m_{it}$, including their dependence on the two constructed inputs. The interference assumption makes the potential response for each unit well defined at the chosen aggregation. Iterated expectation yields the longitudinal g-formula in \eqref{eq:g-formula} \citep{Robins1987agraphical}.
\end{proof}

Assumption~\ref{ass:measurement-population} states the general observed-data requirement. The next corollary gives concrete sufficient conditions for the parametric GMMM used in the simulations.

\begin{corollary}[Parametric identification with fixed market inputs]
\label{cor:expected-exposure}
Consider \eqref{eq:business-value} conditional on measured market inputs, with every response input other than GEO fixed across the two GEO treatment sequences. Suppose that consistency, sequential exchangeability, positivity, and the interference assumption hold. Assume that $\eta_M$ is identified on the required support and that the occurrence probabilities apply to the target population. Suppose that the media transformations are fixed or identified and that the response model with an identity link is correctly specified. If $W_*$ has full column rank and $D_\eta$ in Proposition~\ref{prop:response-rank} has rank four, then $\Delta_G$ is identified. More generally, for a known vector $d_\eta$, the condition $d_\eta^{\mathsf T}v=0$ whenever $D_\eta v=0$ is sufficient to identify $d_\eta^{\mathsf T}\beta$ even when the four coefficients are not separately identified.
\end{corollary}
\begin{proof}
The identified measurement parameters and fixed market inputs determine $E^G$ and $E^P$ under the two treatment sequences. Fixed or identified transformations then determine the corresponding regressors and their differences. Proposition~\ref{prop:response-rank} identifies either the four coefficients or the specified linear effect. The causal assumptions permit these identified quantities to be evaluated under both treatment sequences. Because the response model is written in terms of expected counts, these identified quantities suffice without modeling unobserved realized impressions.
\end{proof}

The corollary relies on occurrence probabilities that apply to the target population and on a correctly specified conditional response mean. Under a nonlinear transformation, a model for an expected count differs from a model for an unobserved realized count: for a random input $M$, $\mathbb{E}[h(M)\mid{\mathcal{D}}_M]$ generally differs from $h(\mathbb{E}[M\mid{\mathcal{D}}_M])$. Either specification also requires control of confounding between the media input and the response.

The general formula permits treatments to change variables observed later. The simulations condition on generated question counts, shares of system use, notice probabilities, and all response inputs other than GEO. For every parameter sample, the calculation replaces the complete GEO sequence before evaluating the response. The rollout dates are fixed by design, so the simulation obtains counterfactual values by evaluating the specified parametric response model under both sequences rather than by using nonparametric overlap for every cluster history.

A randomized rollout can supply treatment variation by design. An observational rollout instead requires pre-treatment controls and support for both treatment sequences. A variable such as referral traffic or brand search may transmit part of the effect of a current treatment to revenue or conversion. Fixing that variable does not recover the total effect and is not sufficient by itself to identify a controlled direct effect \citep{Acharya2016explainingcausal}. If the variable is affected by an earlier treatment and also influences a later treatment, it may belong to $\mathcal H^-_{it}$; the longitudinal g-formula then averages over its distribution under the treatment sequence \citep{Robins1987agraphical}.

\subsection{Decomposition under the Linear Response Model}
Under the linear response used in the simulations, $\Delta_G$ can be written as a component associated with the source state and a component due to the change in the transformed GEO input.

\begin{proposition}[Decomposition under the linear response model]
\label{prop:decomposition}
Suppose that $g_Y$ is the identity link and that every input to the response model other than the GEO source state and the GEO input is the same under the two treatment sequences. Assume that the terms below are integrable. Let $\boldsymbol Z$ denote the sequence of GEO source states selected by $a_G=1$, and let $\boldsymbol 0$ denote the sequence selected by $a_G=0$. Under \eqref{eq:gmmm-response}, it holds that
\begin{align}
\Delta_G
={}&
\beta_D\sum_{i=1}^{n}\sum_{t\in{\mathcal{T}}_E}\mathbb{E}[Z_{it}]
+
\beta_G\sum_{i=1}^{n}\sum_{t\in{\mathcal{T}}_E}
\mathbb{E}\left[H^G_{it}(\boldsymbol Z)-H^G_{it}(\boldsymbol 0)\right]
\nonumber\\
&+
\beta_{GP}\sum_{i=1}^{n}\sum_{t\in{\mathcal{T}}_E}
\mathbb{E}\left[H^P_{it}
\left(H^G_{it}(\boldsymbol Z)-H^G_{it}(\boldsymbol 0)\right)\right].
\label{eq:decomposition}
\end{align}
\end{proposition}
\begin{proof}
Subtract the two conditional means. Every term that contains neither the GEO source state nor the GEO input cancels. The remaining terms are the term associated with the source state, the change in $H^G$, and the interaction between that change and $H^P$. Taking expectations and summing over ${\mathcal{T}}_E$ gives \eqref{eq:decomposition}.
\end{proof}

The component due to the GEO input contains $(\beta_G+\beta_{GP}H^P_{it})(H^G_{it}(\boldsymbol Z)-H^G_{it}(\boldsymbol 0))$, so the interaction makes this component depend on the GEM input. Any other regressor changed by the GEO treatment would add another term.

Equation~\eqref{eq:decomposition} is an algebraic decomposition of the total GEO effect under the specified response model. We use it to study estimation error in the term for the source state and in the terms that contain the constructed GEO input. A causal mediation interpretation would require interventions that can vary the source state and the generated exposure separately, together with exchangeability and support for both interventions \citep{Imai2010ageneral,Pearl1995causaldiagrams}. Our estimand remains the causal effect of the complete source modification.

\section{Answer Collection and Simulation Evidence}
\label{sec:experiments}
The answer collection and the simulations serve different roles. The collection estimates occurrence probabilities under one observed source state. The simulations introduce changes in source state and generate the corresponding market responses, which permits evaluation of estimators of $\Delta_G$. Appendices~\ref{sec:response-collection-details} and~\ref{sec:additional-simulation-results} give the collection procedure and additional numerical results.

\subsection{Target Name in Generated Answers}
\label{sec:answer-collection}
We submit a fixed list of 56 questions that ask for product recommendations in seven use cases to GPT-5.6 Luna and GPT-4o. The list contains 28 English questions and 28 Japanese questions. Each API call uses the same system instruction and requires web search. Its user message contains only the requested language and question, and Glasp appears in neither message. After storing the complete answer, we record one if it contains a spelling of Glasp in a fixed dictionary and zero otherwise; the indicator records occurrence regardless of sentiment.

Each model answers every question 20 times, giving 2,240 answers. The requests are randomized within four blocks collected during one window, and Table~\ref{tab:model-occurrence} summarizes the resulting occurrence rates. Because the source state is unchanged throughout collection, these observations estimate probabilities for that state. Identification of the GEO treatment effect also uses variation in source state, market counts of the questions, shares of use across systems, notice probabilities, and contemporaneous response data.
\begin{table}[!htbp]
\centering
\caption{Occurrence of Glasp in the Collected Answers}
\label{tab:model-occurrence}
\footnotesize
\setlength{\tabcolsep}{3.8pt}
\begin{tabular}{lrrrcrr}
\toprule
Model & English & Japanese & Panel Rate & Posterior Mean (95\% Interval) & Mean Rank & Rank 1 (\%) \\
\midrule
GPT-4o & 35.4 & 20.2 & 27.8 & $28.8\ [27.0,30.7]$ & 1.62 & 60.5 \\
GPT-5.6 Luna & 28.6 & 38.9 & 33.8 & $34.5\ [32.8,36.3]$ & 1.94 & 39.4 \\
\bottomrule
\end{tabular}
\medskip
\begin{minipage}{0.94\textwidth}
\footnotesize
English, Japanese, and Panel Rate are percentages of answers that contain Glasp. Panel Rate averages the 56 observed rates with equal weights. Posterior means and intervals average the distributions for the 56 questions in \eqref{eq:occurrence-jeffreys}. Mean Rank is the rank at the first occurrence of Glasp among 11 candidate brands, conditional on its occurrence.
\end{minipage}
\end{table}

Glasp occurs in $33.8\%$ of the GPT-5.6 Luna answers and $27.8\%$ of the GPT-4o answers. Averaging the posterior distributions for the individual questions gives a difference of $5.70$ percentage points with a 95\% interval of $[3.12,8.27]$. The ordering reverses for English questions, where GPT-4o has the higher rate. Conditional on occurrence, GPT-4o also names Glasp first among the candidate brands more often. These results show why the recorded property and the weights assigned to the questions must match the business application. Appendix~\ref{sec:response-collection-details} reports results by language and use case.

\subsection{Simulation Design}
\label{sec:simulation-design}
The simulations compare estimators on the same panels and with the same construction of $E^G$ and $E^P$. Five geographic markets each contain six product clusters and are observed for 36 periods. Two clusters begin the GEO treatment in period 19 and two begin in period 23, while two never receive it. Every product cluster contains two question clusters observed on two generative systems. The resulting data contain 1,080 observations across units and periods but only six units of treatment assignment within each market. The response follows \eqref{eq:gmmm-response} with
\begin{align}
(\beta_D,\beta_G,\beta_P,\beta_{GP})
&=
(0.55,4.00,2.20,0.50).
\label{eq:true-coefficients}
\end{align}
In the controlled design, baseline occurrence probabilities follow the hierarchical logit model. A second design uses probabilities estimated from the collected answers and preserves the pairing of the two models for each question, as specified in \eqref{eq:paired-occurrence-probabilities}. The treatment effect of GEO is generated in both designs, and the evaluation window is ${\mathcal{T}}_E=\{1,\ldots,36\}$. The simulations isolate uncertainty in the occurrence and notice models together with estimation of the media transformations: question counts and system shares are known, the demand variable is included among the controls, and every cell has 50 notice observations. A cell consists of one question cluster, one generative system, and one period. Established media channels are omitted. Appendix~\ref{sec:simulation-details} gives the complete data-generating process.

The principal comparisons use 120 paired replications in every condition. We evaluate the four response coefficients and $\Delta_G$ using root mean squared error (RMSE) and coverage of 95\% intervals. Appendix~\ref{sec:simulation-evaluation} defines the metrics and numerical settings. Paired intervals for differences between methods resample the 120 common replication indices 5,000 times. Differences between relative RMSE values are reported in percentage points. Appendix~\ref{sec:additional-simulation-results} gives additional results for one, five, and 20 answers per cell and for randomized estimates included in estimation.

\subsection{Estimation of Carryover and Saturation}
\label{sec:transformation-comparison}
Before comparing how uncertainty passes between the two models, we examine whether carryover and saturation are estimated or fixed. On the same 120 controlled panels with five answers per cell, one plug-in method fixes the carryover and Hill parameters at their prior means, while a second estimates them from the response data using the same candidates as the joint posterior. Three diagnostic benchmarks use the true transformation parameters. Table~\ref{tab:transformation-comparison} reports the results.
\begin{table}[!htbp]
\centering
\caption{Effect of Estimating the Transformation Parameters With Five Answers per Cell}
\label{tab:transformation-comparison}
\small
\begin{tabular}{llrrr}
\toprule
Estimator & Transformations & Effect RMSE (\%) & Vector RMSE & Coverage \\
\midrule
Plug-in & Prior mean & 17.642 & 1.207 & 0.975 \\
Plug-in & Estimated from response data & 17.525 & 0.945 & 0.975 \\
Two-stage & Prior distribution & 17.705 & 1.032 & 0.958 \\
Joint & Estimated from response data & 17.541 & 0.950 & 0.975 \\
Plug-in & Known & 17.135 & 0.882 & 0.950 \\
Two-stage & Known & 17.066 & 0.876 & 0.950 \\
Joint & Known & 17.191 & 0.884 & 0.958 \\
\bottomrule
\end{tabular}
\end{table}

Estimating the transformation parameters reduces the RMSE of the coefficient vector from $1.207$ to $0.945$ for the plug-in method. The corresponding value for the joint posterior is $0.950$. The 95\% paired interval for the joint value minus the plug-in value is $[-0.006,0.016]$, and the interval for the difference in effect RMSE also contains zero. Estimation of carryover and saturation explains the principal improvement in coefficient recovery over the plug-in method with fixed transformations. The designs based on the collected answers and the designs with misspecification give the same qualitative result (Appendix~\ref{sec:transformation-details}).

\subsection{Use of Response Data in the Measurement Model}
\label{sec:modular-comparison}
The cut posterior in \eqref{eq:cut-posterior} estimates the transformation parameters from the response data while keeping the distribution of $\eta_M$ fixed by ${\mathcal{D}}_M$. Comparing it with the joint posterior isolates the effect of allowing ${\mathcal{D}}_Y$ to revise $\eta_M$. We use $M=128$ samples of the measurement parameters, a common set of $J=700$ samples from the transformation prior, and 1,000 conditional samples of the response coefficients for each method. Appendix~\ref{sec:modular-diagnostics} describes the computation.

Let $J_A$ be the number of generated answers per cell. The controlled design combines $J_A\in\{1,5\}$ with standard deviations $\sigma_Y\in\{0.4,1.2\}$ for the response errors. A fifth condition uses occurrence probabilities estimated from the collected answers, with $J_A=5$ and $\sigma_Y=1.2$. The 120 replications in each condition give 600 panels. Methods use the same observations within a condition, and the controlled conditions use the same market inputs and standardized response innovations.
\begin{table}[!htbp]
\centering
\caption{Comparison on a Common Set of Transformation Draws}
\label{tab:modular-comparison}
\small
\begin{tabular}{lrrr}
\toprule
Method & Vector RMSE & Effect RMSE (\%) & Coverage \\
\midrule
\multicolumn{4}{l}{Controlled: $J_A=1$, $\sigma_Y=0.4$} \\
Plug-in (estimated) & 0.667 & 6.000 & 0.958 \\
Cut (estimated) & 0.687 & 6.010 & 0.950 \\
Joint & 0.657 & 6.046 & 0.950 \\
Two-stage (prior) & 0.918 & 7.075 & 0.975 \\
Oracle & 0.493 & 5.723 & 0.950 \\
\midrule
\multicolumn{4}{l}{Controlled: $J_A=1$, $\sigma_Y=1.2$} \\
Plug-in (estimated) & 0.975 & 17.450 & 0.967 \\
Cut (estimated) & 0.943 & 17.519 & 0.975 \\
Joint & 0.974 & 17.545 & 0.975 \\
Two-stage (prior) & 1.020 & 17.882 & 0.967 \\
Oracle & 0.886 & 17.236 & 0.950 \\
\midrule
\multicolumn{4}{l}{Controlled: $J_A=5$, $\sigma_Y=0.4$} \\
Plug-in (estimated) & 0.639 & 5.912 & 0.967 \\
Cut (estimated) & 0.642 & 5.930 & 0.967 \\
Joint & 0.636 & 5.923 & 0.975 \\
Two-stage (prior) & 0.897 & 7.070 & 0.958 \\
Oracle & 0.493 & 5.723 & 0.950 \\
\midrule
\multicolumn{4}{l}{Controlled: $J_A=5$, $\sigma_Y=1.2$} \\
Plug-in (estimated) & 0.956 & 17.513 & 0.975 \\
Cut (estimated) & 0.953 & 17.442 & 0.975 \\
Joint & 0.963 & 17.479 & 0.975 \\
Two-stage (prior) & 1.027 & 17.858 & 0.975 \\
Oracle & 0.886 & 17.236 & 0.950 \\
\midrule
\multicolumn{4}{l}{Estimated occurrence probabilities: $J_A=5$, $\sigma_Y=1.2$} \\
Plug-in (estimated) & 0.931 & 17.359 & 0.942 \\
Cut (estimated) & 0.912 & 17.326 & 0.942 \\
Joint & 0.939 & 17.480 & 0.950 \\
Two-stage (prior) & 0.974 & 18.099 & 0.950 \\
Oracle & 0.821 & 16.984 & 0.958 \\
\bottomrule
\end{tabular}
\end{table}
\noindent Each row uses 120 paired replications. Effect RMSE is the relative RMSE of $\Delta_G$. Coverage refers to its 95\% interval. The parenthetical labels indicate whether the transformation parameters are estimated from the response data or retained at their prior distribution. With 120 independent replications, coverage near $95\%$ has a binomial Monte Carlo standard error of about 2 percentage points.

A smaller standard deviation for the response errors substantially reduces effect RMSE, but the joint posterior has no uniform advantage over the cut posterior or the plug-in method that estimates the transformations. When $J_A=1$ and $\sigma_Y=0.4$, the joint posterior reduces coefficient RMSE relative to the cut posterior by $0.0300$; the 95\% paired interval for this reduction is $[0.0221,0.0383]$. The corresponding interval for the difference in effect RMSE contains zero. In the condition based on estimated occurrence probabilities, effect RMSE is $17.326\%$ for the cut posterior and $17.480\%$ for the joint posterior. Their difference is $0.154$ percentage points with an interval of $[0.046,0.259]$. The interval comparing the cut posterior with the plug-in method contains zero. Table~\ref{tab:modular-paired} reports all 15 paired comparisons.

To assess numerical sensitivity, we doubled both numbers of candidates for eight replications in each condition. The cut and joint estimates of $\Delta_G$ changed by at most $1.602\%$ of the true effect, although the effective sample size for the transformation candidates can approach one when the standard deviation of the response errors is small. Appendix~\ref{sec:modular-diagnostics} reports the full diagnostics and states which approximation errors remain when the candidate counts increase.

\subsection{Accuracy of the Components of the Treatment Effect}
\label{sec:effect-components}
The main simulation sets $\beta_D=0.55$ because a source modification can affect the response outside generated answers. The total effect therefore combines the term for the source state, the change in the transformed GEO input, and its interaction with GEM. Errors in the two components can partially cancel when the sum is estimated. Let $n_Z=\sum_{i=1}^{n}\sum_{t\in{\mathcal{T}}_E}Z_{it}$ be the number of treated observations across units and periods. Define the component for the source state by $C_D=n_Z\beta_D$ and the component that contains the GEO input by $C_E=\Delta_G-C_D$. Every method estimates them as $\widehat C_D=n_Z\widehat\beta_D$ and $\widehat C_E=\widehat\Delta_G-\widehat C_D$, using its own estimate of $\beta_D$. The second component includes the interaction in Proposition~\ref{prop:decomposition} and is used here as a decomposition of the specified response model.

The treatment sequences give $n_Z=320$ and $C_D=176$ in every replication. With five answers per cell and $\sigma_Y=1.2$, the mean total effect is $266.944$ in the controlled design and $234.755$ in the design based on estimated occurrence probabilities. The corresponding mean values of $C_D/\Delta_G$ are $66.34\%$ and $75.38\%$. The term associated with the source state accounts for most of the simulated total effect even though its coefficient is unknown to the estimators.
\begin{table}[!htbp]
\centering
\caption{Accuracy of the Treatment Effect and Its Model Components}
\label{tab:effect-components}
\small
\begin{tabular}{lrrrr}
\toprule
Method & Total (\%) & Source State (\%) & GEO Input (\%) & Error Correlation \\
\midrule
\multicolumn{5}{l}{Controlled: $J_A=5$, $\sigma_Y=1.2$} \\
Plug-in (estimated) & 17.513 & 32.638 & 48.351 & -0.627 \\
Cut (estimated) & 17.442 & 32.526 & 47.756 & -0.626 \\
Joint & 17.479 & 32.646 & 48.366 & -0.631 \\
Two-stage (prior) & 17.858 & 36.426 & 47.860 & -0.677 \\
Oracle & 17.236 & 29.351 & 34.825 & -0.523 \\
\midrule
\multicolumn{5}{l}{Estimated occurrence probabilities: $J_A=5$, $\sigma_Y=1.2$} \\
Plug-in (estimated) & 17.359 & 27.551 & 47.129 & -0.555 \\
Cut (estimated) & 17.326 & 27.483 & 45.929 & -0.556 \\
Joint & 17.480 & 27.872 & 47.295 & -0.561 \\
Two-stage (prior) & 18.099 & 29.810 & 48.439 & -0.614 \\
Oracle & 16.984 & 24.470 & 31.468 & -0.409 \\
\bottomrule
\end{tabular}
\end{table}
\noindent The first three numerical columns report relative RMSE for each component. For a component $C$ and $R=120$ replications, the reported value is $100(R^{-1}\sum_{r=1}^{R}((\widehat C_r-C_r)/C_r)^2)^{1/2}$. The columns have different denominators. The final column is the correlation between estimation errors in the component associated with the source state and the component from the GEO input. Section~\ref{sec:modular-comparison} reports all five conditions.

In the condition based on estimated occurrence probabilities, the plug-in method has relative RMSE $17.359\%$ for the total effect and $47.129\%$ for the component from the GEO input. The latter value is $45.929\%$ for the cut posterior, $47.295\%$ for the joint posterior, and $31.468\%$ for the oracle. The negative correlations in Table~\ref{tab:effect-components} explain why the total can be more accurate than either component. If $e_D$ and $e_E$ denote their errors, the squared error of the total contains $2e_De_E$ in addition to the two squared component errors. Errors with opposite signs can cancel through this term. We retain $\Delta_G$ as the primary estimand and report the decomposition alongside the total effect.

\section{Discussion}
\label{sec:discussion}
GMMM separates construction of the two media inputs from the causal comparison. Generated answers determine occurrence probabilities; question counts, system shares, and notice probabilities put those probabilities on the scale of an MMM; complete treatment sequences then define the effects of GEO and GEM. Plug-in, cut, and joint procedures differ only in how they estimate the transformations and use uncertainty in the constructed inputs.

\subsection{Interpretation of the Results}
The GEO input describes expected noticed occurrence under a source state, and it may be positive both with and without the source modification. GEO changes this input through the occurrence probabilities. The estimand $\Delta_G$ compares the complete source-modification sequences, so it includes the path through the constructed input and the additional path represented by $\beta_DZ_{it}$. For GEM, the corresponding comparison changes the spending sequence that determines sponsored placements.

The collected answers also show that one overall occurrence rate is inadequate. The ordering of GPT-5.6 Luna and GPT-4o changes with language, and rates vary substantially across use cases. A feature used in an application should be chosen for the business question, and the probabilities for individual questions should receive the weights of the target population. Equal weights estimate the rate for the fixed panel used here. They also estimate a population rate when that population has the same distribution of questions.

Identification of $\beta_D$ and $\beta_G$ requires variation in the transformed GEO input beyond a proportional change in the variable for source state after the other regressors have been removed. When the two regressors move together, the condition on the null space in Proposition~\ref{prop:response-rank} determines whether $\Delta_G$ remains identified. The component results show why this distinction matters in finite samples: errors in the component associated with the source state and the component from the GEO input often have opposite signs and partially cancel in the total effect.

The design that assigns the treatment supplies its causal interpretation. A randomized rollout provides treatment variation by design. An observational rollout requires pre-treatment controls, support for both treatment sequences, and a contemporaneous comparison group when treated and untreated units experience common changes in the generative systems. Variables affected by an earlier treatment may also influence later assignment, in which case the g-formula in \eqref{eq:g-formula} averages over their distribution under each treatment sequence.

For estimation, the plug-in method that estimates carryover and saturation is a useful reference when the measurement distribution is concentrated. The cut posterior averages over measurement uncertainty while keeping the distribution of the measurement parameters fixed by the data used to construct the inputs. The joint posterior also uses the response likelihood to revise those parameters. We find no universal ranking: the joint posterior improves coefficient RMSE in one condition with a small standard deviation for the response errors, but it does not improve effect RMSE there and performs worse than the cut posterior in the condition based on estimated occurrence probabilities.

Numerical approximation raises a different question. Concentrated importance weights require checks of the number and placement of candidates, while a Markov chain would require checks of mixing and convergence. Increasing the number of candidates leaves the errors from the Laplace approximation for the measurement parameters and the Gaussian approximation to the posterior probability of the sign restrictions unchanged.

The simulations hold the GEO and GEM inputs fixed across estimators and omit established media channels, which isolates estimation of a common GMMM response model. A separate comparison of input construction would keep the response model and transformation estimation fixed while comparing a source indicator, occurrence probabilities, and expected counts under changes in question volume, system shares, and attention. A randomized estimate included in the likelihood informs the same treatment effect; an independently evaluated treatment would address predictive transfer to a different intervention.

\subsection{Alternative Response Models}
The construction of $E^G$ and $E^P$ precedes the choice of response distribution, which allows an application to use a link and distribution suited to sales, conversions, counts, or rates. Hierarchical coefficients can share information across related units, and coefficients that vary over time can represent changes in the relation between a media input and the response. Other carryover or saturation functions can replace \eqref{eq:adstock} and \eqref{eq:hill} without changing the definitions of the two expected counts or $\Delta_G$.

\subsection{Alternative Estimation Procedures}
The distinction between measurement and response parameters also applies outside the Bayesian formulation. A procedure based on likelihood can estimate all parameters together and use resampling for uncertainty, while a regularized procedure can penalize selected terms in \eqref{eq:general-criterion}. In the Bayesian analysis used here, the cut and joint posteriors differ in whether ${\mathcal{D}}_Y$ updates $\eta_M$.

\subsection{Uncertainty in Market Inputs}
The simulations treat the question counts $N_{iqt}$ and shares $\omega_{ipt}$ as known. In an application, the counts may be estimated from a sample and the shares may change during the evaluation window. A probability model for either quantity can be added to $\eta_M$, after which its uncertainty passes through \eqref{eq:geo-exposure}, the media transformations, and the calculation of $\Delta_G$. This extension changes the construction of the input but not the response model or the definition of the treatment effect.

\section{Conclusion}
GMMM extends MMM to settings in which the media inputs associated with generative AI are not recorded directly. Under each source state, the GEO input is the expected number of generated answers in which a specified property occurs and is noticed; GEO changes this count but need not create it from zero. The GEM input is the expected number of sponsored placements that users notice. Both inputs are defined for each market and period before carryover and the Hill transformation are applied. The treatment effect of GEO is the total effect of the specified source modification under two complete treatment sequences.

The identification results allow the treatment effect to be determined in some designs even when its component coefficients cannot be recovered separately. They also identify the conditions under which a sequence of occurrence probabilities differs from a market count only by scale. In the simulations, estimating carryover and saturation explains most of the improvement in coefficient recovery over a plug-in method that fixes them. Allowing response data to revise the measurement parameters has no consistent benefit for estimating $\Delta_G$, and the total effect can conceal larger errors in its two model components.

The collected answers identify occurrence probabilities under the source state present during collection. Identification of the change caused by GEO uses observations under both source states together with contemporaneous market counts, notice observations, and business responses. GMMM combines these measurements and evaluates the causal effect from the complete sequence of media inputs.


\bibliography{arXiv2.bbl}

\bibliographystyle{tmlr}

\clearpage