EconBase
← Back to paper

Linear Estimation of Structural and Causal Effects for Nonseparable Panel Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,091 characters · 23 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Linear Estimation of Structural and Causal Effects for Nonseparable Panel Data

abstractThis paper develops linear estimators for structural and causal parameters of nonseparable models using panel data. These models incorporate unobserved, time-varying, individual heterogeneity, which may be correlated with the regressors. Estimation is based on an approximation of a conditional average potential outcome by a linear sieve specification with individual-specific parameters. Effects of interest are estimated by a bias corrected average of individual ridge regressions. We demonstrate how this approach can be applied to estimate causal effects, counterfactual consumer welfare, and averages of individual taxable income elasticities. We show that the proposed estimator has an empirical Bayes interpretation and possesses a number of other useful properties. We formulate Large-$T$ asymptotics that can accommodate discrete regressors and which bypass partial identification in this case. We employ the methods to estimate average equivalent variation and deadweight loss for potential price increases using data on grocery purchases.

Models and Parameters of Interest

We consider panel data made up of observations across $n$ individuals, indexed by $i$, and $T$ time periods, indexed by $t$. The individuals are drawn independently and identically from some population. Each observation consists of a scalar outcome variable $Y_{it}$ and a vector of variables of interest $X_{it}$ that can affect $Y_{it}$.

\ignorespaces We assume a non-separable structural/causal model for $Y_{it}$ of the form

equation[equation omitted — 98 chars of source]

where $\eta_{it}$ are individual- and time-specific unobserved variables representing preferences, technology, or potential outcomes that may be correlated with $X_{it}$. The $\eta_{it}$ may be serially correlated across $t$. The model is nonseparable in the sense that it allows for general interactions between the observed variables $X_{it}$ and the unobserved heterogeneity $\eta_{it}$.

We specify that the function $g$ has a causal/structural interpretation, that $g(x,\eta_{it})$ is the counterfactual value of the outcome $Y_{it}$ in a counterfactual world in which the vaiables of interest take the value $x$. Thus, $Y_{it}(x):=g(x,\eta_{it})$ is a potential outcome.

A simple and quite flexible nonseparable model is a linear random coefficients (LRC) model, where there is a known $J\times1$ vector $b(x)=(b_1(x),...,b_J(x))'$ of functions of $x$, with first component $b_1(x)=1$, such that

equation[equation omitted — 63 chars of source]

Here a potential outcome is $b(x)'\eta_{it}$. Let $X_{i}=(X_{i1}^{\prime},...,X_{iT}^{\prime})^{\prime}$ be the full history of of variables of interest for individual $i$. Also let $\beta_{it}=E[\eta_{it}|X_i]$ where we suppress dependence of $\beta_{it}$ on $X_i$ for notational convenience. Then

equation[equation omitted — 51 chars of source]

Here we see that $\beta_{it}$ are regression coefficients so we could try to estimate them in order to estimate counterfactual objects like $b(x)'\beta_{it}$. The problem is there is only one observation $Y_{it}$ available to do this, which is not enough data to identify $\beta_{it}$.

\ignorespaces

In this paper we use conditional time-homogeneity of $\eta_{it}$ to identify and estimate objects of interest. \theoremstyle{definition} \newtheorem*{A1}{Assumption 1 (Time-Homogeneity)} \begin{A1} The distribution of $\eta_{it}$ conditional on $X_{i}$ does not depend on $t$. \end{A1}

\ignorespaces

The identifying power of Assumption 1 for the LRC model is that $\beta_{it}=E[\eta_{it}|X_i]=E[\eta_{i1}|X_i]:=\beta_i$ does not vary over time, so that

equation*[equation* omitted — 64 chars of source]

Here all $T$ observations observations are available to estimate each $\beta_i$ making identification of $\beta_i$ possible when $T$ is greater than or equal to the number $J$ of regressors in the LRC model. \ignorespaces

Assumption 1 allows for endogeneity where the conditional distribution of $(\eta_{i1},...,\eta_{iT})'$ given $X_{i}$ may depend on $X_{i}$. Such endogeneity is present when $X_{it}$ includes choice or equilibrium values that are determined by the preferences or technology represented by $\eta_{it}.$ It is common to decompose $\eta_{it}$ into time constant components $\alpha_{i}$ and time-varying components $v_{it}$. The $\alpha_{i}$ trivially satisfies Assumption 1, so that Assumption 1 only restricts $v_{it}$. Assumption 1 requires that $v_{it}$ have the same distribution in each period conditional on the history of regressors $X_{i}$. In economic applications, it is always important to accommodate time-varying preferences and technology as represented by $v_{it}$. Models that do not incorporate time-varying unobservables may fail to fit the variation in outcomes over time.

Assumption 1 has been imposed in the many previous papers on nonseparable panel data that were referred to in the Introduction. Assumption 1 is also satisfied in discrete choice, multinomial logit demand models such as those of Dubois2020, where $v_{it}$ are logit disturbances in discrete choice demand models. In demand models $v_{it}$ represents variation in preferences over time for a given individual. Such variation could be due to a taste for variety as further discussed in Section 4. Assumption 1 imposes stability over time on such preferences or technology. This condition seems important for the nonparametric model, where identifying structural/causal effects while allowing unrestricted correlation between $X_{it}$ and $\eta_{it}$ seems difficult without time homogeneity. Assumption 1 does allow for systematic variation over time in the outcome $Y_{it}$ via the observable variables of interest $X_{it}$ whose magnitude varies with time. For example, $X_{it}$ may include time trends or seasonal indicators. It may also be possible to allow for time period specific effects, as was recently done by Jin2026 in a related model, though that is beyond the scope of this paper.

\ignorespaces

A main focus and innovation of this paper is estimation of causal and structural effects for the nonparametric, nonseparable model where the functional form of $g(X_{it},\eta_{it})$ is unknown, the dimension of $\eta_{it}$ is unspecified and may even be infinite, and Assumption 1 is satisfied. This model does not impose any functional form assumption for the outcome function $g$ and is therefore very general. For example, it allows for discrete outcomes where $g(X_{it},\eta_{it})$ may depend on indicator functions for unknown functions of $X_{it}$ and $\eta_{it}$.

Our object (i.e., parameter) of interest is an average difference in linear combinations of potential outcomes with coefficients specified by the researcher. Let $X_{it}^{+}$ and $X_{it}^{-}$ be counterfactual values of variables of interest that can differ from the observed $X_{it}$. Also let $H_{it}^{+}$ and $H_{it}^{-}$ be coefficients that are specified by the researcher. We estimate objects of the form

equation[equation omitted — 157 chars of source]

In addition, to objects of the form above, we discuss identification and estimation of partial effects in Appendix B.

We assume throughout that the following conditional independence assumption holds. \theoremstyle{definition} \newtheorem*{A2}{Assumption 2 (Counterfactuals)} \begin{A2} $((X_{it}^{+}$,$H_{it}^{+}, X_{it}^{-}, H_{it}^{-}),t=1,...,T)$ is independent of $(\eta_{i1},t=1,...,T)$, conditional on $X_i$. \end{A2}

\ignorespaces Because the variables $X_{it}^{+}$, $X_{it}^{-}$, $H_{it}^{+}$, and $H_{it}^{-}$ are chosen by the researcher, Assumption 2 can be made to hold by construction. For example, it holds if these variables are constructed as functions of $X_i$ and/or simulated variables that are independent of the data, as is the case in all of the examples we consider. However, if these objects are chosen to be functions of $X_i$ and some observables $Z_{it}$ not included in $X_i$, then veracity of Assumption 2 depends on the statistical relationship between $X_i$, $Z_{it}$, and $\eta_{it}$.

We note that our object of interest $\theta_0$ implicitly depends on the number of time periods $T$, because we do not restrict any of the observable variables in the definition of $\theta_0$ to be stationary over time. For notational convenience we suppress this dependence in what follows, while noting here that it is accounted for in the asymptotic analysis of Section 6.

\ignorespaces

A number of important policy-relevant objects may be written in the form of $\theta_0$. Three examples we consider are average causal effects of alternative treatment regimes, bounds on average equivalent variation and on deadweight loss in demand analysis, and taxable income effects with nonlinear budget sets. In Appendix C we discuss marginal effects in binary choice models.

Example 1: Average Effects of Alternative Treatment Regimes

\ignorespaces

To define treatment effects within our model, recall that the potential outcome for a counterfactual level $x$ of the regressors is $Y_{it}(x)=g(x,\eta_{it})$. Here, the heterogeneity $\eta_{it}$ determines potential outcomes and endogeneity of $\eta_{it}$ corresponds to correlation between potential outcomes and observed treatments $X_{it}$, similarly to Imbens.

\ignorespaces

Consider two counterfactual treatment regimes. In the first, an individual $i$ in period $t$ receives a random treatment ${X}_{it}^{-}$ and in the second they receive treatment ${X}_{it}^{+}$. These counterfactual treatments may depend on $X_i$. For example, we may wish to compare mean outcomes under the counterfactual assignments ${X}_{it}^{+}$ with the factual treatments, in which case we can set $X_{it}^{-}=X_{it}$. The expected average difference between potential outcomes is \[ \theta_{0}:=E[\frac{1}{T}\sum_{t=1}^{T}\{Y_{it}({X}_{it}^{+})-Y_{it}( {X}_{it}^{-})\}]=E[\frac{1}{T}\sum_{t=1}^{T}\{g({X}_{it}^{+},\eta_{it})-g( {X}_{it}^{-},\eta_{it})\}]. \] A major challenge for estimating such a causal effect is the possibility of unobserved confounding. That is, there may be latent factors that jointly determine the treatment $X_{it}$ and the outcomes. The nonparametric nonseparable model we consider allows for the possibility of unobserved confounding. Assumption 1 allows us to estimate average treatment effects like $\theta_0$, as explained in Section 3.2.

We can motivate Assumption 1 in this context using a nonparametric structural model for the treatment assignments and the heterogeneity in potential outcomes. Let $\alpha_{i}$ be a vector of time-invariant confounding factors and consider the following model where the time-varying innovations $\{u_{it}\}_{t=1}^{T}$ and $\{w_{it}\}_{t=1}^{T}$ are each jointly independent of $\alpha_{i}$:

align*[align* omitted — 80 chars of source]

The first equation decomposes the endogenous heterogeneity in potential outcomes into variation between individuals, captured in $\alpha_i$, and variation over time $u_{it}$. Let us suppose that the innovations $\{u_{it}\}_{t=1}^{T}$ are jointly independent of $\{w_{it}\}_{t=1}^{T}$ and $\alpha_i$. This implies $u_{it}$ is independent of the history of treatments $X_i$. In addition, suppose the marginal distribution of $u_{it}$ (but not necessarily $w_{it}$) is time-invariant. Under these conditions, Assumption 1 holds.

In this model, the time-invariant factors $\alpha_i$ are akin to fixed effects. They are individual-specific characteristics that explain the confounding between treatments and outcomes and do not vary over time. The condition that the temporal variation in potential outcomes, captured in $u_{it}$, is independent of the history of treatment assignments is akin to strict exogeneity. Unlike in the classic fixed effects model, $\alpha_i$ may enter non-separably into the possibly non-linear model for the outcome $Y_{it}$. Similar panel data treatment effect models were explicitly formulated in Chernozhukov2013 and Torgovitsky.

Example 2: Average Equivalent Variation and Deadweight Loss Bounds

The average equivalent variation and deadweight loss of a price change are important objects of interest in empirical demand analysis. Obtaining bounds on these quantities is a crucial step in assessing the welfare impact of a policy that may alter consumer prices, such as a sales tax. Suppose $Y_{it}$ is the expenditure share of some commodity and $X_{it}=(P_{it},Z_{it})$, where $P_{it}$ is the product price and $Z_{it}$ is a vector of covariates that includes total expenditure $M_{it}$ and the prices of other goods.

In order to define the welfare effects of a price change, we must choose an initial price paid by individual $i$ at time $t$, which we denote by ${P}_{it}^{-}$. This starting price may depend on $X_i$. For example, ${P}_{it}^{-}$ could simply be $P_{it}$, the price paid in period $t$ by individual $i$ for the good. Let $\Delta_{it}$ denote the change in the price of the good for individual $i$ in period $t$ that also may depend on $X_i$. Let $\omega_{t}(X_i)$ be some weighting that may depend upon $X_i$. Using Hausman, we obtain bounds on the weighted average equivalent variation from a price change from $P_{it}^{-}$ to $P_{it}^{-}+\Delta_{it}$. Let $\pi$ be an upper or lower bound on the income effect for every individual and let $U_{i}=(U_{i1},...,U_{iT})'$ be a vector of $T$ random variables that are uniformly distributed on $(0,1)$ and independent of the data, i.e. that are simulation draws from the standard uniform distribution. Taking $g(p,Z_{it},\eta_{it})$ to be the counterfactual expenditure share at price $p$, a bound on the weighted average equivalent variation is

align[align omitted — 276 chars of source]

If $\pi$ is a lower (upper) bound on the income effect for every individual then, by Hausman, $\theta_{EV}$ is an upper (lower) bound on the average over time and individuals of the equivalent variation for a change from ${P}_{it}^{-}$ to ${P}_{it}^{-}+\Delta_{it}$, weighted by $\omega_{t}(X_i)$. The weights $\omega_{t}(X_i)$ allow us to assess the welfare impact on particular sub-populations, such as those in a low income bracket or with a certain family size.

A corresponding deadweight loss bound can be obtained by subtracting the weighted average change in final demand as below.

align[align omitted — 229 chars of source]

If $\pi$ is a lower (upper) bound on the income effect then $\theta_{0}$ will be an upper (lower) bound for weighted deadweight loss averaged over all time periods and individuals.

Both $\theta_{EV}$ and $\theta_{DWL}$ are of the form in ((ref)). In particular, $\theta_{EV}$ corresponds to the case with $H_{it}^{+}$ defined by ((ref)), $H_{it}^{-}=0$, and $X_{it}^{+}=({P}_{it}^{-}+\Delta_{it}U_{it},Z_{it}')'$. The deadweight loss shares these choices of $H_{it}^{+}$ and $X_{it}^{+}$, but in this case, $H_{it}^{-}$ is set as in ((ref)) and $X_{it}^{-}=({P}_{it}^{-}+\Delta_{it},Z_{it}')'$. This example is discussed further in Section 5, which provides an application to consumer panel demand data.

Example 3: Average Heterogeneous Taxable Income Elasticities

In structural economic models, an object of interest can be the expectation of a coefficient of an LRC model. An example is the panel, budget set regression of Blomquist2024. There an individual specific isoelastic utility function, together with time-varying scale heterogeneity in preferences, leads to an outcome $Y_{it}$, equal to the log of taxable income, that is a nonlinear function of the budget set. By integrating out the time-varying scale factor, using linear approximation of certain integrals, and specifying $X_{it}$ to be a piecewise linear budget frontier, an LRC is obtained with $J=3$, $\beta_{i2}$ equal to the taxable income elasticity for individual $i$, and

equation[equation omitted — 85 chars of source]

where $\beta_{i2}$ is the taxable income elasticity for individual $i$, $b_2(x)$ is the natural logarithm of the slope of the last budget segment, and $b_3(x)$ is the ln of rato of slopes of the last and first budget segments. The panel data setting of this paper allows for endogeneity of piecewise linear budget sets, where budget sets may be correlated with preferences. The parameter of interest $\theta_0=E[\beta_{i2}]$ can be represented in the form of equation (4) by choosing $H^+_{it}=H^-_{it}=1$ and specifying $X_{it}=(1,b_2(X_{it}),b_3(X_{it}))',X_{it}^{+}=X_{it}+e_2,$ and $X_{it}^{-}=X_{it}$, where $e_2$ is a unit vector with $1$ in the second position and zeros elsewhere.

\ignorespaces

Linear Approximation and Estimation

In this Section we discuss approximation of nonparametric nonseparable models by LRC specifications. Thus we obtain a linear approximation to the object of interest $\theta_0$ which motivates a linear estimator. The approximation relies on the fact that $\theta_0$ is linear in the conditional average potential outcome, which is

equation*[equation* omitted — 71 chars of source]

where the second equality follows by Assumption 1. The following result gives the formula for $\theta_0$ in terms of $h(x,X_i)$ which motivates our linear approximation and estimation method.

\theoremstyle{plain} \newtheorem*{T1new}{Theorem 1} \begin{T1new} If Assumptions 1 and 2 are satisfied and if the moments $E[|g(X_{it}^+,\eta_{it})|]$, $E[|H_{it}^+g(X_{it}^+,\eta_{it})|]$, and $E[|H_{it}^-g(X_{it}^-,\eta_{it})|]$ are finite, then

equation[equation omitted — 170 chars of source]

\end{T1new}

Approximation

\ignorespaces Theorem 1 demonstrates that we can approximate $\theta_0$ by approximating $h(x,X_i)$. We can approximate $h(x,X_i)$ by the average potential outcome for the LRC model, which is

equation*[equation* omitted — 74 chars of source]

The idea is that if $b(x)$ is a rich enough vector of approximating functions and $h(x,X_i)$ is a smooth enough function of $x$, uniformly in $X_i$, then $b(x)'\beta_{i}$ will approximate $h(x,X_i)$ as a function of $x$ for some $\beta_i$.

\ignorespaces

In this approximation the function $h(x,X_i)$ is being approximated by $b(x)'\beta_i$ separately for each value of $X_i$. As $X_i$ varies the coefficients $\beta_i$ are allowed to vary so that $b(x)'\beta_i$ remains a good approximation of $h(x,X_i)$. Such approximations are known to exist when $x$ is bounded and $h(x,X_i)$ is continuously differentiable up to order $d$ with derivatives bounded, uniformly in $x$ and $X_i$. Such uniform approximation results are often referred to as Jackson theorems, following Jackson1911 where approximation rates were obtained for power series and scalar $x$. Jackson theorems for multivariate $x$, power series, splines, and other approximating functions are given, for example, in DeVore1993. We provide a more formal analysis of the approximation error in Section 6.

One could also consider the LRC model $b(x)'\eta_{it}$ as an approximation to the nonparametric model $g(x,\eta_{it})$, as was done in a previous version of this paper. Approximating $g(x,\eta_{it})$ is more difficult because in important examples of interest, such as those with discrete $Y_{it}$, $g(x,\eta_{it})$ will not be smooth in $x$. In contrast, the average potential outcome $h(x,X_i)$ can be very smooth in $x$ even when $Y_{it}$ is discrete, as long as some elements of $\eta_{it}$ are continuously distributed with smooth pdf, because $h(x,X_i)=E[g(x,\eta_{it})|X_i]$ integrates over $\eta_{it}$. For this reason, and because our objects of interest depend on just $h(x,X_i)$, we focus on approximating $h(x,X_i)$.

\ignorespaces

The approximation of $h(x,X_i)$ leads to a corresponding approximation for $\theta_0$. The approximation $h(x,X_i)\approx b(x)'\beta_i$ implies $h(X_{it}^+,X_i)\approx b(X_{it}^+)'\beta_i$ and $h(X_{it}^-,X_i)\approx b(X_{it}^-)'\beta_i$. Plugging these approximations into Theorem 1 gives

equation[equation omitted — 171 chars of source]

Thus, $\theta_0$ is approximately an expectation of the inner product of a known vector $a_i$ with $\beta_i$. Also, this approximation is exact in the LRC model where $h(x,X_i)=b(x)'\beta_i$. Consequently, an estimator of $\theta_0$ can be constructed by forming the inner product of the known $a_i$ with an estimator of $\beta_i$ and averaging over individuals, as described in the next subsection.

The approximating LRC model is a random parameters specification like that of newey2025. We innovate in approximating a nonparametric, nonseparable model by one that is linear in known functions of $x$.

\ignorespaces

Estimation

\ignorespaces To estimate $\theta_0$ using equation (13) we need an estimator $\hat\beta_i$ of $\beta_i$ for each individual $i$. To construct these estimators we use the approximate linearity of $E[Y_{it}|X_i]=h(X_{it},X_i)\approx b(X_{it})'\beta_i$ in $b(X_{it})$ by regressing $Y_{it}$ on $b(X_{it})$ over all time periods $(t=1,...,T)$ to form $\hat\beta_i$. From the literature on series estimation of conditional means (e.g. Gallant, 1981), it is known that such a linear regression can be used to estimate functions of a conditional mean under specific identification and regularity conditions. We apply this insight to the panel data setting by using linear regression for each individual over all $t$ and then averaging across individuals. In Section 6 we make this intuition precise by giving regularity conditions for consistency and asymptotic normality of the resulting $\hat\theta$ for large $T$. \ignorespaces

In practice, there could be high multicollinearity in this regression, particularly if $T$ is not much larger than the number $J$ of basis functions. Here we address this problem using individual specific ridge regression. To describe these ridge regressions let $Y_i := (Y_{i1},...,Y_{iT})'$, $B_i :=[b(X_{i1}),...,b(X_{iT})]'$, $Q_{i}=B_{i}^{\prime}B_{i}/T,$ $D_{i}$ be a diagonal matrix with $0$ as its upper left entry and all other diagonal entries strictly positive, and $\lambda$ a positive constant. A ridge regression estimator of ${\beta}_i$ is defined as

equation[equation omitted — 97 chars of source]

The zero in the top left entry of $D_i$ ensures that we do not penalize the intercept in the ridge regression. By allowing $D_i$ to be individual-specific we can accommodate individual-level re-scaling of the regressors. \ignorespacesIn our empirical application we simply set $D_i$ to be the identity matrix with its upper left entry set to zero.\ignorespaces

We note here that perfect multicollinearity in the regression for individual $i$, where $Q_i$ is singular, is an identification problem and not just a computational issue. If $Q_i$ is singular then the population least squares coefficients $\beta_i$ from regressing $E[Y_{it}|X_i]$ on $b(X_{it})$ are not unique, and hence not identified. We return to this important identification issue in Section 3.5 that follows.

These individual ridge estimators are biased, as usual for ridge regression. It is possible to mitigate this ridge bias in the estimation of $\theta_0=E[a_i'{\beta}_i]$. \ignorespaces Let $A_{i}$ denote a square $J$-dimensional matrix with $a_{i}^{\prime}$ as its first row and the remaining rows consisting of $J-1$ distinct rows of the identity matrix chosen so that $\frac{1}{n}\sum_{i=1}^{n}A_{i}$ is non-singular. In our application we simply use the last $J-1$ rows of the identity. \footnote{\ignorespacesOther choices of $A_i$ can also be used. What is necessary is that $A_i$ is square, a function of $a_i$, constructed so that there is a fixed vector $v$ so that $a_{i}'=v'A_{i}$, and $\frac{1}{n}\sum_{i=1}^{n}A_{i}$ is non-singular. Our theoretical results apply for any such $A_i$ and the asymptotic variance we derive is not affected by the choice of $A_i$.\ignorespaces} \ignorespaces Also, let

equation[equation omitted — 70 chars of source]

The debiased average ridge estimator of $\theta_{0}$ is then

equation[equation omitted — 275 chars of source]

An estimator $\hat{V}$ for the asymptotic variance of $\sqrt{n}(\hat{\theta}-\theta_0)$ can be obtained via the delta method as

equation[equation omitted — 270 chars of source]

Example 3 illustrates that parameters of interest may include elements of the vector $E[\beta_i]$. Each component of this vector has the form $E[a_i'\beta_i]$ where $a_i$ is a unit vector. When $a_i$ is constant, the formula for the debiased estimator simplifies so that $A_i$ cancels out. Thus we obtain an estimator $\hat{\theta}$ of $E[\beta_i]$ and a corresponding estimator $\hat{V}$ of the asymptotic variance, for $\overline{W}:= \sum_{i=1}^{n}W_{i}/n$,

equation[equation omitted — 257 chars of source]

This estimator has a straight-forward interpretation. The term $\sum_{i=1}^{n}\hat{\beta}_{i}/n$ is the sample average of individual ridge estimates that suffers from ridge bias. Multiplying by $\overline{W}^{-1}$ effectively undoes the ridge bias on average. Specifically, $\hat\theta$ has Property C mentioned in the Introduction, being unbiased in the LRC model when $\beta_{ij}$ does not vary with $i$ for $j>1$. In the next subsection we discuss this and other properties of $\hat\theta$.

\ignorespaces The debiased average ridge estimator belongs to a general class of regularized panel estimators that includes Graham2012. See the discussion of Property C in Section 6.1 for details.

Summary of Properties

The debiased panel ridge estimator has several interesting characteristics that help explain its form and how it may be used and interpreted. Here we provide a brief summary of these properties, with a more formal discussion deferred to Section 6 along with our asymptotic analysis.

Property A: Empirical Bayes Interpretation

The average coefficient estimator $\hat\theta$ of equation ((ref)) can be interpreted as an empirical Bayes estimator. Suppose we assume an LRC specification holds and the corresponding residuals are iid normally distributed, and that each $\beta_i$ has a normal prior with common nonzero mean $\bar{\beta}$. For each $i$, let $\beta_i^{\text{Post}}$ be the posterior mode for each individual parameter $\beta_i$ with $D_i$ being directly proportional to the prior variance-covariance matrix for $\beta_i$. A Bayesian estimator of $E[\beta_i]$ can then be constructed as $\sum_{i=1}^n \beta_i^{\text{Post}}/n$. This estimator depends on the prior mean $\bar{\beta}$ and the empirical Bayes approach is to use the data to determine this hyper-parameter. Suppose one chooses the unique $\bar{\beta}$ that ensures $\sum_{i=1}^n \beta_i^{\text{Post}}/n=\bar{\beta}$, which can be understood as a `self-consistency' restriction. The resulting estimator coincides with our estimator $\hat{\theta}$.

Property B: Convergence to Fixed Effects with Large Penalty $\lambda$

As the penalty parameter $\lambda$ diverges to infinity, the estimator $\hat{\theta}$ converges towards a fixed-effects estimator $\frac{1}{n}\sum_{i=1}^n a_i' \hat{\beta}_{FE,i}$ where $\hat{\beta}_{FE,i}$ is a fixed effects estimator with individually-varying intercept and constant slopes,when $D_i$ does not vary with $i$.

Property C: Unbiased When Slope Coefficients Do Not Vary with Individuals

For the LRC model the debiased ridge estimator of equation (18) is unbiased when $\beta_{ij}$ does not vary with $i$ for $j>1$. Note that this holds under `exogenous slopes', that is, if $\eta_{itj}$ is independent of $X_i$ for $j>1$. This property holds regardless of the choice of $\lambda$ and applies in cases in which $Q_i$ is singular for some individuals so that $\beta_i$ is not identified.

Property D: Convergence Under Small Penalty

As $\lambda$ shrinks to zero, $\hat\theta$ converges to the average of the product of $a_i$ and the individual OLS estimate of $\beta_i$ if $Q_i$ is non-singular for all individuals.\\ \\

Properties B and D together demonstrate that the choice of penalty parameter $\lambda$ allows for a smooth transition between the two extremes of fixed effects and average of individual OLS estimates.

\ignorespaces

Guidance for Applications

In order to construct the estimate $\hat{\theta}$, in the case of nonparametric $g(x,\eta)$ one must choose the vector of basis functions $b(\cdot)$. In our empirical application, we use polynomial basis functions (powers and interactions of the regressors) and report results for polynomial bases of various degrees. The number of regressors in our application is moderately large, so we use relatively low order polynomials. Our most flexible specification is log-linear in a subset of regressors, and for other regressors we include powers and interactions up to order three. In cases with fewer regressors, we suggest the use of cubic spline bases (DeVore1993).

The estimator $\hat{\theta}$ given in the formula ((ref)) depends on two regularization parameters. The first of these is the matrix $D_i$. We suggest researchers take $D_i$ to be the identity matrix with its upper-left entry set to zero, as we do in our empirical application.

For the penalty parameter $\lambda$, we suggest researchers report results for a range of values. One effect of the debiasing strategy is that, by counteracting the shrinkage of ridge towards zero, it tends to reduce the sensitivity of the final estimates to the choice of the penalty parameter. The problem of developing a data-driven method for selecting $\lambda$ is a problem for future research. Nonetheless, a version of Lepski's method (see e.g., Birge2001) may provide a useful heuristic. To apply the method, one selects $\lambda$ to be the largest value such that for any $\lambda'\leq\lambda$, the absolute difference in the corresponding estimates is no greater than four times the standard error under $\lambda'$.

Evaluating the Extent of Identification

Ridge regression and debiasing affect the contribution of each observation to $\hat\theta$. To see this in the LRC, note that by Assumption 2 and i.i.d. individual data, \[ E[\hat\theta|X_1,...,X_n]=\frac{1}{n}\sum_{i=1}^{n}\hat{a}_i'\beta_{i},\hat{a}_i'=\bar{a}'(\overline{AW})^{-1}A_iW_i. \] Here we see that $\hat\theta$ is (conditionally) unbiased for the regularized weighted average $\frac{1}{n}\sum_{i=1}^{n}\hat{a}_i'\beta_{i}$ rather than the correct $\frac{1}{n}\sum_{i=1}^{n}a_i'\beta_{i}$.

For example consider the average slope estimator $\hat\theta_2$ in the LRC for $J=2$, where $b(x)=(1,b_2(x))'$ and $a_i=(0,1)'$. Let $\tilde{Q}_i$ denote the sample variance over $t$ of $b_2(X_{it})$, $w_i=\tilde{Q_i}/ [\lambda+\tilde{Q_i}]$, and $\omega_i=w_i/{\sum_{k=1}^{n}w_k}$ be nonnegative weights that sum to $1$. Then it is straightforward to show that $\hat{a}_i=(0,\omega_i)'$, so that $E[\hat\theta_2|X_1,...,X_n]={\sum_{i=1}^{n}\omega_i\beta_{i2}}$. Here we see that the expectation of the debiased ridge estimator is a weighted average of individual specific coefficients $\beta_{i2}$, with larger weight $\omega_i$ given to observations where $\beta_{i2}$ is more strongly identified, in the sense that $\tilde{Q}_i$ is larger, and zero weight given to observations where $\beta_{i2}$ is not identified, i.e. where $\tilde{Q_i}=0$ (and so $Q_i$ is singular).

In general, $\hat{a}_i$ differs from the corresponding $a_i$ depending on the size of $\lambda$ and $Q_i$. If every $Q_i$ is nonsingular then as $\lambda$ shrinks each $\hat{a}_i$ will converge to $a_i$ for every $i$, as in Property D and shown in Section 6. If some $Q_i$ are singular, then each $\hat{a}_i$ need not converge to $a_i$ and there can be especially sharp differences when $Q_i$ is singular. In the example given in the previous paragraph, each $\hat{a}_i$ with $Q_i$ singular converges to $(0,0)'$ and each $\hat{a}_i$ with nonsingular $Q_i$ converges to $(0,1)'$. More generally, when $Q_i$ is singular and $a_i$ is not in the identified set of linear combinations of $\beta_i$, the limit of $\hat{a}_i$ will not be $a_i$ as $\lambda$ shrinks. Thus, the extent of nonidentification of $a_i'\beta_i$ across $i$ should be indicated by differences between $\hat{a}_i$ and $a_i$ across observations for small $\lambda$.

We can use the Euclidean norm $\|\hat{a}_i-a_i\|$ to evaluate the extent of identificatioin in applications. Note that \[ |\hat{a}_i'\beta_i-a_i'\beta_i|\leq\|\hat{a}_i-a_i\|\|\beta_i\|. \] Therefore, a quantity that measures the effect of possible singularity of $Q_i$ on the contribution of the $ith$ observation to the estimated effect $\hat\theta$ of interest is \[ \zeta_i=\|\hat{a}_i-a_i\| \|\hat\beta\|/n\hat\sigma, \] where $\hat\beta$ is the esimator of the average coefficient vector and $\hat\sigma$ is the standard error of $\hat\theta$. A quantile plot of this object may aid in assessing both the degree of nonsingularity of $Q_i$ over the all $i$ as well as the strength of identification for particular individuals.

One can also upper-bound the conditional bias due to regularization by\footnote{To derive this bound, first note that $\frac{1}{n}\sum_{i=1}^{n}\hat{a}_i'E[\beta_{i}]=\frac{1}{n}\sum_{i=1}^{n}a_i'E[\beta_{i}]$ and thus $\frac{1}{n}\sum_{i=1}^n \hat{a}_i'\beta_i-\frac{1}{n}\sum_{i=1}^n a_i'\beta_i=\frac{1}{n}\sum_{i=1}^n (\hat{a}_i-a_i')(\beta_i-E[\beta_i])$, then apply Cauchy-Schwarz.} \[\bigg|E[\hat\theta|X_1,...,X_n]-\frac{1}{n}\sum_{i=1}^n a_i'\beta_i\bigg|\leq \sqrt{\frac{1}{n}\sum_{i=1}^n \|\hat{a}_i-a_i\|^2}\sqrt{ \frac{1}{n}\sum_{i=1}^n \|\beta_i-E[\beta_i]\|^2}.\] The first square root on the RHS can be calculated directly, and the second may be bounded according to one's a priori beliefs about the variation in $\beta_i$. Indeed, in recent work, kwon2025 use a priori bounds on the variance of heterogeneous coefficients to obtain bias-aware confidence intervals in a related setting.

Nonparametric, Nonseparable Demand Models for Panel Data

The nonseparable, nonparametric model $Y_{it}=g(X_{it},\eta_{it})$ (of equation (ref)) provides a very general specification of individual demand. In the application, we take the outcome variable $Y_{it}$ to be the expenditure share on a class of goods, which has long been a useful specification, as in Deaton1980, Chaudhuri2006, and hsiao2021, but other choices of outcome variable will also do. Discrete choice is included as a special case where $Y_{it}$ is the number of units of a particular good purchased by an individual in time period $t$ and the outcome model is specified analogous to that in Section 3.1.

The model allows unobserved heterogeneity $\eta_{it}$ to affect demand in very general ways. The $\eta_{it}$ is allowed to be infinite-dimensional corresponding to stochastic revealed preference as in Richter1990, McFadden2005, and Kitamura, with demand restricted to be single valued. Such choice specifications have been considered by Lewbel2001, Blomquista, Blundell2014, Hoderlein2014, Bhattacharya, Dette2016, and Hausman. In addition, $\eta_{it}$ may include product specific unobserved characteristics as in Berry1994 and Berry1995. The presence of such could create correlation across individuals in $\eta_{it}$. Alternatively, if $\eta_{it}$ is tastes by an individual for unobserved product characteristics and preferences are independent across individuals, then correlation across individuals need not be present.

To help this model relate to existing demand models, it is helpful to decompose the heterogeneity $\eta_{it}$ into a component $\alpha_i$ that does not vary with $t$ and a time-varying component $v_{it}$. This decomposition is common in panel demand models, including discrete choice, as in Chamberlain1984. Here, $\alpha_i$ represents preference features that are stable over time for a given individual, while $v_{it}$ allows some time variation in demand. For example, $v_{it}$ could represent a taste for variety that is not observable to the econometrician. Tastes for variety could also be incorporated by including functions of $t$ in $b(X_{it})$. Generally, it is quite common to incorporate time-varying heterogeneity as represented by $v_{it}$ in nonlinear panel data models.

An important feature of panel demand data is that prices are common across consumers and are determined in market equilibrium. As a result, prices will generally be endogenous in being related to individual preferences. Restrictions on $v_{it}$ mitigate potential price endogeneity. If $v_{it}$ is i.i.d. over time and independent of unobserved supply shocks, and there are many consumers in the market, then bias from price endogeneity will be small, as shown by moon2026. Intuitively, price effects will be (nearly) identified from the movement of prices over time because supply shocks are independent of time variation in preferences. This independence seems plausible when $v_{it}$ is a stochastic taste for variety of an individual and variation in supply is due to cost shocks. Also, Hausmana found that the use of prices from other markets as instruments did not change demand estimates for scanner data, providing evidence that relying on time variation in prices for identification of price effects is consistent with scanner data.

The interpretation of $v_{it}$ as preference heterogeneity means that preferences are allowed to change over time, even being correlated over time, in order to represent a taste for variety. Demand specifications with time-varying, unobserved preference effects $v_{it}$ are common in panel data, discrete choice demand being a prime example. The presence of $v_{it}$ helps demand and other models fit the data better. It allows for departures from the weak axiom of revealed preference in the choice of an individual over time, as has been found in empirical work, for example, crawford2019.

Time-varying preferences have little effect on the interpretation of welfare calculations. The average equivalent variation and deadweight loss calculations just average over time. As such, they estimate the expected value of welfare integrated over as $v_{it}$, similar to welfare estimates for discrete choice panel data. Such time-average welfare measures are consistent with utility maximization over time if there are no dynamic linkages in goods. Of course, such preferences are not consistent with stockpiling models like that of Hendel2006. In the application, we take one month as the time unit and focus on goods with little potential for stockpiling to avoid this concern.

Allowing for zero demand is important in demand modeling. For example, consumer data that considers alcohol or tobacco consumption will have many individuals with zero consumption. Including zero demand observations in the data correctly accounts for zero consumption in calculations of average equivalent variation and deadweight loss, as shown in Hausman, Theorem 3. Intuitively, there is no effect of a price change on the welfare of a consumer who never purchases a product, and the average is also correct when the product is only purchased sometimes. The nonseparable, nonparametric specification gives the demand equations flexibility to allow for zeros while being consistent with utility maximization.

An observation arising from economic theory is that often, but not always, the policy question of interest depends on only one, or a very few, price effects. For example, estimation of individual welfare effects typically depends only on the own price effect when all other prices are held constant (Hausman1981). Also, small cross-price effects will mitigate market equilibrium effects of changing one price. Price changes for one good will shift demand for other goods by small amounts, so that the equilibrium welfare effect from changing only one price can be well approximated by the effect of just that price on average demand.

Computational simplicity is an important virtue of demand analysis in panel data based on linear approximation we give. Average equivalent variation and deadweight loss are estimated by a debiased average of individual specific linear combinations of ridge regressions. Simulation is used to approximate the integrals in the welfare estimates. Simple inference is based on the independence of estimates across individuals. All of these features make this approach to demand estimation simple to implement, even in very large data sets.

Application to Scanner Data

We apply our methods to estimate price elasticities for groceries and to analyze the impact of counterfactual tax changes on consumer welfare. In this context, the outcome variable $Y_{it}$ is the share of expenditure on a particular class of goods, and the regressors $X_{it}$ include the natural log of prices and total expenditure. Our specification generalizes the popular Almost Ideal Demand System (AIDs) of Deaton1980a to approximate a nonparametric, fully nonseparable demand model as we describe in Section 4.

Given this specification, our debiasing method has important implications for our elasticity estimates, and consequently, our estimates of counterfactual welfare. In the absence of debiasing, ridge regression tends to shrink parameters to zero. Therefore, in an AIDS-type specification, ridge would shrink the own price elasticity towards $-1$, the cross-price elasticities towards $0$, and the expenditure elasticity towards $1$. The debiasing mitigates the effects of shrinkage and results in estimates that are less sensitive to the choice of penalty parameter. However, we note that shrinkage of the cross-price elasticities may be appropriate in consumer demand panel datasets, where small cross-price effects re often found in the literature Burda2008 and Burda2012. Indeed, cross-sectional OLS regressions in our empirical setting recover small cross-price elasticities, as we report in Table IV in Appendix A.

Rather than estimate demand for particular products, we instead focus on the demand for classes of goods. In effect, we model demand at an intermediate level of multi-stage budgeting to estimate welfare effects of price changes for good types. Here, the consumer decides how much to spend on a class of goods based on individual- and type-specific second-order flexible price indices, and on the total expenditure on all included classes of goods.

Modeling demand for good types can be justified by certain conditions on the separability of preferences, as in, for example, Gorman1959, Gorman1981, Deaton1980a, and Blundell2000. An alternative motivation relies on statistical aggregation for the many prices into a price index which is independent of consumer preferences, as in Hoderlein2012a. For the intermediate level of commodities we consider (e.g., soda), it may be important to allow for more general substitution patterns across the dissimilar kinds of goods. The flexibility in allowing for general cross-price effects provided by the AIDS demand system of Deaton1980a, may be useful here as it is in Chaudhuri2006 and hsiao2021.

We use NielsenIQ retail scanner data to construct price indices, and the NielsenIQ Homescan Panel to track purchases and household characteristics.\footnote{The empirical work is researchers' own analyses calculated (or derived) based in part on data from Nielsen Consumer LLC and marketing databases provided through the NielsenIQ Datasets at the Kilts Center for Marketing Data Center at The University of Chicago Booth School of Business. The conclusions drawn from the NielsenIQ data are those of the researcher(s) and do not reflect the views of NielsenIQ. NielsenIQ is not responsible for, had no role in, and was not involved in analyzing and preparing the results reported herein.}

The data include $2585$ households with Houston-area ZIP codes in the years 2010-2014. The number of monthly observations for each household ranges from $1$ to $60$, and we restrict our analysis to the households included for at least $12$ months.\footnote{We checked for differences in results between using all households and the 2197 that were present for at least a year and found no statistically significant differences. The insensitivity to panel length suggests that attrition bias does not play a large role in this data.}

We construct the price indices for each consumer from data on the monthly total expenditures per good category, and on the quantity purchased per month. The original data contained time-stamps for purchases. The price indices span 15 aggregated groups of goods: soda, milk, soup, water, butter, cookies, eggs, orange juice, ice cream, bread, chips, salad, yogurt, coffee, and cereal. As in Burda2008 and Burda2012, we chose these groups because they made up a relatively large proportion of total grocery expenditure. The data also includes demographics such as race, marital status, household size and composition, and employment status.

The price index for each group of goods is computed as a weighted geometric average of the actual purchase prices (expenditure divided by quantity) over all purchases made by the household in the month, with weights equal to the proportion of expenditure on a specific item associated with a unique item code. The price index $P_{g,it}$ for household $i$ at time $t$, for the $g^{th}$ group of goods is specified by \[ \ln(P_{g,it})=\sum_{j=1}^{J_{g}}w_{gj,it}\ln(P_{gj,it}), \] where $j$ denotes a particular item code, $J_{g}$ is the number of codes for the $g^{th}$ commodity, $w_{gj,it}$ is the proportion of expenditure on commodity $g$ that is spent on code $j$, and $P_{gj,it}$ is expenditure by household $i$ on code $j$ divided by quantity of code $j$ in month $t$. This is a T\"ornqvist price index, which was shown by Diewert1976 to be exact for a quadratic utility specification, and a second-order approximation to the exact price index for any utility. Deaton1980 (pp. 132-133) showed that with weak separability, this price index appears in share equations for a Rotterdam demand specification (i.e., log quantity as a linear function of log prices and log expenditure) and suggest that it could lead to a good approximation when prices within a group tend to move together.

The price indices may be endogenous because the amount spent on a particular item in a group of goods is a choice of the consumer. Price endogeneity could be particularly important when a group of goods contains commodities of varying quality, such as organic and non-organic milk, or fresh and frozen orange juice. As we discuss in the previous section, our approach can accommodate such endogeneity, provided that the unobserved heterogeneity satisfies the time-invariance condition formalized in Assumption 1.

As stated above, we construct price indices using prices actually paid by each household. Including zero expenditures makes it necessary to impute price indices for time periods where an individual purchased none of a particular good. If a household had purchased the good before, then price indices are imputed as the most recent price faced by the household in a past purchase. Rarely, a good is never purchased prior to a given month, in which case its imputed price is the average price of the same good within a subset of stores similar to those at which the household shops.\footnote{Specifically, we group retailers in the Houston area into $4$ categories, and assign households to their most-visited retailer category each year. Then we construct monthly price indices for each retailer category and each good, which are used to fill in missing prices.} The frequency of household-month observations with zero total expenditures varies by good: for some goods, most households record purchases each month, while other goods, such as orange juice and ice cream, are purchased more infrequently. Our analysis focuses on estimating demand for the goods for which we have the most reliable data, namely, soda and milk.

The inclusion of prices for all $15$ categories of goods allows estimation of cross-price demand effects. This gives us $16$ price and expenditure regressors. This is too large a number of regressors for standard nonparametric estimation, such as kernel regression, where it is thought to be impractical to use more than five or six regressors. For panel estimation, $16$ regressors may also be excessively large. The large number of regressors with small coefficients for the many cross-price effects motivates our use of ridge regularization.

In total, our analysis uses $86,122$ observations across the households. As a baseline, we consider the log-linear AIDS-type specification below.

equation[equation omitted — 131 chars of source]

where $Y_{it}$ is the share of expenditure by household $i$ in month $t$ on a particular class of goods. $Exp_{it}$ is that household's total monthly expenditure over the $15$ categories of goods, and $P_{g,it}$ is the household's price index for good $g$ in that month. $\alpha_i$, $\gamma_i$, and $\beta_{g,i}$ are individual-specific coefficients, and $u_{it}$ a time-varying residual. We estimate separate models for soda and milk, with no restrictions that the coefficients in each case are the same.

In order to more precisely approximate a possibly non-linear and non-separable underlying demand model, in some of our analyses we enrich the specification ((ref)) by including some powers and interactions of log prices and total expenditure.

Table I contains elasticity estimates for both soda and milk. We employ the model ((ref)) and compare three methods for estimation. These are cross-sectional OLS, fixed-effects estimates, individual-specific ridge without debiasing, and estimates that employ our debiased individual-specific ridge method.

In order to perform individual-ridge, we must select the matrix $D_i$ in the formula ((ref)). We let $D_i$ be the identity matrix with its first diagonal entry set to zero. We carry out ridge using two alternative choices for the penalty parameter $\lambda$. As a robustness check, we carry out the analysis with and without the inclusion of seasonal dummy variables. Seasonal variation in both price and tastes could be problematic for our analysis, as it suggests that heterogeneity in preferences is time-varying given prices, which would contradict Assumption 1.

table[table omitted — 1,566 chars of source]

\ignorespaces The columns in Table I respectively contain own-price elasticity estimates obtained using cross-sectional OLS, Fixed Effects, individual-ridge without debiasing and penalty parameters, and our debiased ridge estimates. For the latter two methods, we present estimates for two different values of the penalty parameter, namely $0.05$ and $0.0005$. To account for dependence between the coefficient estimates and the mean expenditure share, standard errors are calculated by bootstrap.\ignorespaces

In all cases, the estimates are insensitive to the inclusion of seasonal dummies. The ridge estimates without debiasing are sensitive to the choice of penalty parameter, particularly in the case of milk. As we discuss above, in our specification, the shrinkage associated with ridge will tend to bias the elasticity estimates towards $-1$. Indeed, when we do not debias, the elasticity estimates from ridge are closer to $-1$ when we employ a higher penalty than with a smaller penalty. This is particularly striking for milk. By contrast, when we debias, which mitigates the shrinkage associated with ridge, we obtain elasticity estimates that are much less sensitive to the choice of penalty.

The elasticity estimates from our debiased ridge method roughly align with those found in the previous literature (see, for example, the meta-analysis of Andreyeva2010).

\ignorespaces Compared with cross-sectional OLS and Fixed Effects, the individual ridge estimates relax the assumption of homogeneous coefficients (and intercepts, in the case of OLS). For sufficiently small values of $\lambda$, this relaxation will typically result in estimates with greater standard errors. On the other hand, if $\lambda$ is sufficiently large, the slope estimates will be shrunk strongly towards zero, generally leading to smaller standard errors than those of the Fixed Effects estimates, and possibly the OLS estimates as well. Indeed, in Table 1 the individual ridge elasticity estimates with a small value of $\lambda$ have greater standard errors than OLS and Fixed Effects, but for the larger value, the milk elasticity standard errors are smaller for individual ridge. Shrinkage towards zero can be a major source of bias, and our debiasing method acts to counter the overall shrinkage towards zero. As such, for sufficiently large values of $\lambda$, the debiasing will typically increase standard errors. Indeed, in Table 1, for the larger value of $\lambda$, the debiasing is associated with substantially increased standard errors for the milk elasticities. However, for the smaller value of $\lambda$, this effect is less strong. In fact, in Table 1, the debiased ridge estimates with a small $\lambda$ have standard errors that are no greater, and in one case smaller, than without debiasing.

\ignorespaces We apply our methods to estimate an upper bound on the average equivalent variation consumer surplus and deadweight loss from a $10\%$ increase in price for both soda and milk, while excluding time periods for each individual where the larger price was outside the range of prices in the observed data. This increase is relative to the actual price faced by each household in a particular period. The bound follows the formula in Hausman, as detailed in Example 2 in Section 2. The formula requires that we impose a lower bound on the income effect. We take our lower bound to be $0$ which corresponds to the assumption that soda and milk are normal goods. This lower bound on the income effects corresponds to an upper bound on the welfare loss. \ignorespaces

In order to provide some distributional analysis, we estimate the welfare bounds separately for households in three different income groups. In particular, for those whose household income (averaged over all periods for which there is data on that household) is in the bottom quartile, top quartile, and for all households.

Tables II and III contain our estimation results for the welfare upper bounds. We apply our analysis for both the log-linear specification ((ref)) and a cubic specification, which supplements the regressors in the linear model with all powers and interactions of the log own-price and total expenditure up to order three. We provide results for various choices of the penalty parameter $\lambda$. The welfare estimates have been annualized; that is, the numbers represent the welfare change over the course of a year.

table[table omitted — 2,342 chars of source]
table[table omitted — 2,331 chars of source]

The estimates of average deadweight loss and consumer surplus for the full set of households are remarkably stable, both between the linear and cubic specifications, and for different values of the penalty parameter. In part, this may reflect the tendency of debiasing to mitigate the shrinkage induced by regularization, and thus to reduce sensitivity to the choice of penalty parameter $\lambda$.

We estimate that the deadweight loss from a price increase for soda is markedly higher than for milk. This is not surprising given that milk, unlike soda, is a staple food, and so demand for this product may be relatively inelastic. Indeed, this aligns with our elasticity estimates in Table I.

Harding2017 analyze the role of prices in determining food purchases and nutrition and estimate the impact of taxes on nutrition and individual welfare. Allcott2019 and Dubois2020 have also considered the welfare effects of taxing soda. Like Dubois2020, our panel approach estimates individual-specific demands. Our approach is simpler in that it is based on continuous demand modeling and individual ridge regression with total expenditure included in the demand function. Also, our application averages over on-the-go and larger store purchases as well as over individuals that purchase soda and those that do not. We obtain substantially larger estimates of average equivalent variation than their compensating variation which is to be expected because we model household demand and they model individual demand.

Figure 1 plots the quantiles of the nonidentification measure $\zeta_i$ from Section 3.5. \ignorespaces.

figure[figure omitted — 390 chars of source]

We see from the figures that for a large proportion of individuals, estimates the $\zeta_i$ is small. This is particularly clear in the case of our consumer surplus estimates, for which the quantiles below 90% are almost zero.

Theoretical Results

We now turn to a formal analysis of the properties of our estimation procedure. For this purpose, we will be explicit about allowing the number of periods for which we have observations to vary between individuals. In particular, we let $T_i$ denote the number of periods for which we observe data on individual $i$. In addition, we explicitly define the approximation error that results from the use of an LRC approximation. Recall that we employ an approximation $h(x,X_i)\approx b(x)'\beta_i$. To explicitly define $\beta_i$, we suppose that $\frac{1}{T_{i}}\sum_{t=1}^{T_{i}}E[b(X_{it})b(X_{it})']$ is non-singular. For each fixed value $\mathbb{X}$ in the support of $X_i$, define the function $\beta(\mathbb{X})$ as follows. \[ \beta(\mathbb{X})=E[\frac{1}{T_{i}}\sum_{t=1}^{T_{i}}b(X_{it})b(X_{it})']^{-1}E[\frac{1}{T_{i}}\sum_{t=1}^{T_{i}}b(X_{it})h(X_{it},\mathbb{X})]. \] That is, $\beta(\mathbb{X})$ is the vector of coefficients from a best approximation of $h(\cdot,\mathbb{X})$ by a linear combination of the basis functions $b(\cdot)$. Then we let $\beta_i:=\beta(X_i)$. If an LRC model holds (that is, if $g(x,\eta_{it})=b(x)'\eta_{it}$), then under Assumption 1, the definition above ensures that $\beta_i=E[\eta_{it}|X_i]$, which coincides with the discussion of the LRC model in Section 2.

We define the approximation error $r(x,X_i)$ as, \[ r(x,X_i)=h(x,X_i)-b(x)'\beta_i. \] We then let $r_{it}=r(X_{it},X_i)$. Note that if outcomes follow an LRC model, then $r(\cdot,X_i)$ is identically zero. We also define a residual $u_{it}$ by, \[u_{it}=Y_{it}-h(X_{it},X_i).\] Let $r_i$ and $u_i$ be the length-$T_i$ column vectors whose $t$-th entries are $r_{it}$ and $u_{it}$ respectively, then we obtain the model

equation[equation omitted — 80 chars of source]

where $r_i=0$ if the LRC model holds.

It is helpful to introduce some notation. \ignorespacesThroughout, for a vector $v$, $\|v\|$ is its Euclidean norm, and for a matrix $A$, $\|A\|$ is the operator norm defined by $\|A\|=\sup_{v\neq0}\|Av\|/\|v\|$. \ignorespaces Let $D_{1,i}$ be equal to $D_{i}$ but with the first row and column removed (recall that by definition, the first row and column of $D_i$ contain only zeros). $J$ is the length of the vector $b(X_{it})$, and let $B_{1,it}$ denote $b(X_{it})$ with its first entry (the constant) removed. Finally, define $\bar{Y}_{i}=\frac{1}{T_{i}}\sum_{t=1}^{T}Y_{it}$, $\bar{B}_{i}=\frac{1}{T_{i}}\sum_{t=1}^{T}B_{1,it}$, and $\tilde{Q}_{i}=\frac{1}{T_{i}}\sum_{t=1}^{T}B_{1,it}B_{1,it}'-\bar{B}_{i}\bar{B}_{i}'$.

Recall that Section 3 defines $D_i$ to be diagonal with the first diagonal entry zero and the others non-zero. In this section, we allow for more general choices of $D_i$: we retain the condition that the first row and column of $D_i$ contain only zeroes, but unless we state otherwise, $D_{1,i}$ can be any strictly positive-definite matrix. We implicitly assume throughout that the estimator ((ref)) is well defined, that is, $\overline{AW}$ is non-singular.

Properties of the Estimator

Before we turn to the asymptotic behavior of our estimation procedures, we elaborate on a number of notable properties of the estimator outlined in Section 3. In particular, we consider its motivation as an empirical Bayes estimator, its limiting behavior under large and small values of the penalty parameter, and the sense in which the estimator eliminates regularization bias.

Property A: Empirical Bayes Interpretation

The estimator $\hat{\theta}$ can be expressed as an empirical Bayes procedure. The empirical Bayes approach imposes a prior on the individual coefficient vectors $\{\beta_i\}_{i=1}^n$ in ((ref)) and the mean of this prior is estimated jointly with the individual-level coefficients.

To be more precise, the debiased ridge estimator is a Bayesian maximum a posteriori (MAP) estimate. The likelihood corresponds to the model ((ref)) in which $r_i=0$ (as in an LRC model) and the residuals $u_{it}$ are iid Gaussian. The prior for the individual-specific coefficients $\{\beta_i\}_{i=1}^n$ is Gaussian and independent across individuals.

It is well-known that standard ridge regression estimates can be expressed as MAP estimates in which the prior is Gaussian with mean zero. What distinguishes our approach is the manner in which the prior mean for $\beta_i$ is determined by the data. In particular, the prior mean $\bar{\beta}$, is pinned down by the restriction

equation[equation omitted — 132 chars of source]

where $\{{\beta}^{\text{Post}}_{i}\}_{i=1}^n$ is the posterior mode for $\{{\beta}_{i}\}_{i=1}^n$. Thus the prior mean for the individual coefficients is fixed by imposing that the prior mode and posterior modes of $\frac{1}{n}\sum_{i=1}^{n}A_{i}{\beta}_i$ are identical. Loosely speaking, it ensures that observing the data does not lead us to update (i.e., improve) upon our prior for $\frac{1}{n}\sum_{i=1}^{n}A_{i}{\beta}_{i}$.

To be yet more precise, consider a Gaussian conditional likelihood for the outcomes $Y_{it}|X_{i}\overset{iid}{\sim}N(b(X_{it})'\beta_i,\sigma^2)$ and prior for the individual slope parameters $\beta_i\overset{iid}{\sim}N(\bar{\beta},\Sigma)$, where $\bar{\beta}$ is the prior mean and $\Sigma$ is the prior variance-covariance matrix. Given this likelihood and prior, the posterior density $f^{\text{Post}}$ satisfies the expression below:

equation[equation omitted — 235 chars of source]

The parameters $\{\beta_{i}\}_{i=1}^{n}$ that maximize the above (given a fixed $\bar{\beta}$) are the MAP estimates. In order to obtain an empirical Bayes estimate, we estimate the prior mean $\bar{\beta}$ from the data by imposing equation ((ref)).

\theoremstyle{plain} \newtheorem*{P6}{Proposition 1} \begin{P6} The estimator $\hat{\theta}$ in ((ref)) can be written as $\hat{\theta}=\frac{1}{n}\sum_{i=1}^{n}a_{i}'{\beta}_{i}^{\text{Post}}$ where $\{{\beta}_{i}^{\text{Post}}\}_{i=1}^n$ and $\bar{\beta}$ jointly solve the equation ((ref)), and equation ((ref)) below:

align[align omitted — 236 chars of source]

\end{P6} The objective ((ref)) is a monotone transformation of a Bayesian posterior ((ref)) when $D_{i}=\frac{\sigma}{\lambda}\Sigma^{-1}$. Thus $\{{\beta}_{i}^{\text{Post}}\}_{i=1}^n$ are MAP estimates of the individual slopes. The prior is fixed by the second equation.\footnote{Note that we take $D_i$ to have first row and columns composed of zeroes, and so $D_i$ is singular. As such, strictly speaking we use flat `improper' prior for the individual intercepts.} In the special case in which $A_i$ is the identity, solving the two equations above is equivalent to maximizing the objective in ((ref)) jointly over both $\beta_i$ and $\bar{\beta}$. That is, $\bar{\beta}$ can be obtained by

align*[align* omitted — 244 chars of source]

Property B: Convergence to Fixed Effects with Large Penalty

As the penalty parameter grows to infinity, our estimator converges to a plug-in fixed effects or generalized fixed effects estimator. To state this formally, let us first note that the standard fixed effects estimate $\hat{\beta}_{FE,i}$ may be expressed as follows. The first component of this vector is an individual intercept given by $\bar{Y}_{i}-\bar{B}_{i}'\hat{\beta}_{FE,1}$, where $\hat{\beta}_{FE,1}$ is a vector of shared slope parameters and constitutes the remaining components of $\hat{\beta}_{FE,i}$. The slopes are given by \[ \hat{\beta}_{FE,1}=(\frac{1}{n}\sum_{i=1}^{n}\tilde{Q}_{i})^{-1}\frac{1}{n}\sum_{i=1}^{n}\frac{1}{T_{i}}\sum_{t=1}^{T_{i}}(B_{1,i,t}-\bar{B}_{i})Y_{it}. \]

$\hat{\beta}_{FE,i}$ is a special case of a generalized fixed-effects estimator $\hat{\beta}_{GFE,i}$. Again, the first component of $\hat{\beta}_{GFE,i}$ is an individual intercept, in this case $\bar{Y}_{i}-\bar{B}_{i}'\hat{\beta}_{GFE,1}$, where $\hat{\beta}_{GFE,1}$ is a vector of shared slopes. Let $G_{i}$ be a non-singular weighting matrix, then the corresponding vector of slope parameters $\hat{\beta}_{GFE,1}$ is defined as follows: \[ \hat{\beta}_{GFE,1}=(\frac{1}{n}\sum_{i=1}^{n}G_{i}\tilde{Q}_{i})^{-1}\frac{1}{n}\sum_{i=1}^{n}\frac{1}{T_{i}}\sum_{t=1}^{T_{i}}G_{i}(B_{1,i,t}-\bar{B}_{i})Y_{it}. \]

A plug-in fixed-effects estimate of $\theta_{0}$ is given by $\frac{1}{n}\sum_{i=1}^{n}a_{i}'\hat{\beta}_{FE,i}$, and plug-in generalized fixed-effects estimator by $\frac{1}{n}\sum_{i=1}^{n}a_{i}'\hat{\beta}_{GFE,i}$.

\theoremstyle{plain} \newtheorem*{P3}{Proposition 2} \begin{P3}

$\lim_{\lambda\to\infty}\hat{\theta}=\frac{1}{n}\sum_{i=1}^{n}a_{i}'\hat{\beta}_{GFE,i}$ where $G_{i}=D_{1,i}^{-1}$. If $D_{i}$ does not vary with $i$, then $\lim_{\lambda\to\infty}\hat{\theta}=\frac{1}{n}\sum_{i=1}^{n}a_{i}'\hat{\beta}_{FE,i}$. \end{P3} The proposition states that as the penalty parameter grows towards infinity, the debiased panel ridge estimator converges to the plug-in generalized fixed-effects estimator with $G_{i}$ equal to the inverse of $D_{1,i}$. In the special case in which $D_{i}$ does not depend on $i$, this is identical to the standard plug-in fixed effects estimator.

Property C: No Regularization Bias Under Exogenous Effects

Our estimator is based on individual-level ridge regressions. Ridge estimates typically suffer from `regularization' bias. The form of our estimator ((ref)) is designed to mitigate, and in some cases entirely eliminate, regularization bias. In particular, in the case in which $\beta_i$ is constant, apart from the intercept, and there is no approximation error.

For some insight into the bias properties of the estimator, it is helpful to compare our method with a plug-in estimator based on individual-specific OLS. An individual OLS estimate of $\beta_{i}$ is given below, where $Q_{i}^{\dagger}$ is the pseudo-inverse of $Q_{i}$ and is well-defined even if $Q_{i}$ is singular: \[ \tilde{\beta}_{i}=Q_{i}^{\dagger}B_{i}'Y_{i}/T_i \]

A plug-in OLS estimator of $\theta_{0}$ is then $\frac{1}{n}\sum_{i=1}^{n}a_{i}'\tilde{\beta}_{i}$. Suppose Assumptions 1 and 2 hold. If $Q_{i}$ is non-singular for all $i$, and the mean of $a_{i}'\tilde{\beta}_{i}$ is finite, then the plug-in OLS estimator is unbiased (up to approximation error) for $\theta_{0}$. This is because, in the absence of approximation error, $\tilde{\beta}_{i}$ is a conditionally (on $X_i$) unbiased estimate of $\beta_{i}$. Moreover, by Assumption 2, $\tilde{\beta}_{i}$ is independent of $a_i$ conditional on $X_i$. Therefore, if $E[|a_{i}'\tilde{\beta}_{i}|]<\infty$ then we can apply the law of iterated expectations and obtain: \[ E[a_{i}'\tilde{\beta}_{i}]=E\big[a_{i}'E[\tilde{\beta}_{i}|X_{i},a_i]\big]=E\big[a_{i}'E[\tilde{\beta}_{i}|X_{i}]\big]=E[a_{i}'\beta_{i}]. \]

If individuals are drawn independently and identically from the population, then consistency of the plug-in OLS estimator follows by the law of large numbers. However, if $Q_{i}$ is singular with positive probability, this argument fails because an $\tilde{\beta}_{i}$ is (in general) biased when $Q_{i}$ is singular. Moreover, the assumption that $E[|a_{i}'\tilde{\beta}_{i}|]$ is finite is crucial. If this moment is infinite, then one cannot apply the law of iterated expectations, nor the law of large numbers. The first moment may be infinite even if $Q_{i}$ is non-singular almost surely. Graham2012 acknowledge that the finite mean condition may fail, particularly if the number of regressors is close to the number of time periods. In the case of $a_{i}$ constant, this situation coincides with the case in which the information bound derived in Chamberlain1992 is infinite, and thus regular estimation is impossible with the number of time periods fixed.

In contrast to OLS, individual-specific ridge estimates have finite expectation under weak conditions. Proposition 3 provides conditions under which the expectation of $a_{i}'\hat{\beta}_{i}$ is finite, where $\hat{\beta}_i$ is the individual ridge regression estimate defined in Section 3. The proposition applies even if $Q_{i}$ is singular with positive probability.

\theoremstyle{plain} \newtheorem*{P1}{Proposition 3} \begin{P1} Suppose $\lambda>0$, $D_{1,i}$ has eigenvalues bounded below by $c>0$, $\|B_{i}\|$ and $\|a_{i}\|$ are uniformly bounded, and $E[|Y_{it}|]$ is finite. Then $E[|a_{i}'\hat{\beta}_{i}|]<\infty$. \end{P1}

As we discuss in Section 3, individual ridge estimates are conditionally biased, even in the absence of approximation error. This in turn suggests that the sample average $\frac{1}{n}\sum_{i=1}^{n}a_{i}'\hat{\beta}_{i}$ is biased for $E[a_{i}'\beta_{i}]$, even if $r_{it}$ in ((ref)) is identically zero (i.e., if an LRC holds). This motivates our debiasing strategy. Proposition 4 shows that if effects are exogenous, then the debiased estimator is exactly unbiased up to approximation error. Note that the theorem holds for any $\lambda>0$ and also applies when $Q_{i}$ is singular with positive probability. This contrasts with the average of plug-in OLS, which is in general conditionally biased when $Q_i$ is singular with positive probability. By `effects are exogenous', we mean that $\beta_{i,2}$ is mean independent of $X_i$, where $\beta_{i,2}$ is the subvector of $\beta_i$ formed by removing its first component.

\theoremstyle{plain} \newtheorem*{P2}{Proposition 4} \begin{P2} Suppose Assumptions 1 and 2 hold, $r_{it}=0$ almost surely, and $\beta_{i,2}$ is mean independent of $X_i$. In addition, suppose $D_i$ is a function of $X_i$. If $\frac{1}{n}\sum_{i=1}^{n}A_{i}W_{i}$ is non-singular then \[ E[\hat{\theta}|X_{1},X_{2},...,X_{n}]=\frac{1}{n}\sum_{i=1}^{n}E[a_{i}'\beta_{i}|X_{1},X_{2},...,X_{n}]. \] If the above holds and $E[|\hat{\theta}|]<\infty$ then $E[\hat{\theta}]=E[a_{i}'\beta_{i}]$. \end{P2}

Proposition 4 requires that $\beta_{i,2}$ is mean independent of $X_i$. The first component of $\beta_i$ is unrestricted. We do not need to restrict this component because we do not penalize the intercept in our individual-ridge regressions (this is why the first row and column of $D_i$ are composed of zeros). A sufficient (but not necessary) condition for the whole vector $\beta_i$ to be mean independent of $X_i$ is that $\eta_{it}$ is independent of $X_i$.

If we strengthen the condition in Proposition 4 so that the entire vector $\beta_i$ is mean independent of $X_i$, then the proposition applies to a general class of estimators. Consider that we can rewrite the estimator ((ref)) as follows:

equation[equation omitted — 289 chars of source]

If we replace $W_{i}:=(Q_{i}+\lambda D_{i})^{-1}Q_{i}$ with some other conformable matrix that depends only on the regressors, then we obtain an alternative estimate of $\theta_{0}$. Consider the special case in which $a_{i}$ is constant, and let $W_{i}$ be an indicator that the determinant of $Q_{i}$ exceeds a cut-off $h$, then the formula yields the estimator of Graham and Powell (absent adjustment for time-effects). If $\beta_i$ is mean independent of $X_i$, Proposition 2 applies for any estimator of the form above so long as $W_{i}Q_{i}^{\dagger}Q_{i}=W_{i}$, which holds both for our choice of $W_{i}$ as well as that of Graham and Powell.

Property D: Convergence Under Small Penalty

Suppose that for each $i$, the matrix $Q_{i}$ is non-singular. Then, as $\lambda$ goes to zero, the matrix $(Q_{i}+\lambda D_{i})^{-1}$ converges to $Q_{i}^{-1}$, for every $i$. As such, our estimate converges to the plug-in average of individual OLS estimates $\frac{1}{n}\sum_{i=1}^{n}a_{i}'\tilde{\beta}_{i}$. Note the contrast with Property A. As $\lambda\to\infty$ our estimator converges to a plug-in average of (generalized) fixed effects estimates, with individual-specific intercepts and shared slope parameters. If $Q_{i}$ is non-singular for all $i$, then as $\lambda\to0$, the estimator converges instead to the plug-in average of OLS estimates, which have both individual-specific intercepts and slopes. The choice of $\lambda$ thus allows us to smoothly transition between these two estimators.

When $Q_{i}$ is singular for some individuals, our estimator generally does not converge to plug-in individual OLS, but nonetheless, it has an interpretable limit. Proposition 5 considers the special case in which one of the regressors, denoted by $Z_{it}$, is discretely distributed. Because $Z_i$ is discrete, it may be constant over time for some individual $i$, in which case $Q_i$ is singular.

\theoremstyle{plain} \newtheorem*{P4}{Proposition 5} \begin{P4} Suppose $b(X_{it})=(1,Z_{it},B_{2,it}')'$, where $Z_{it}$ is a discrete scalar and the vector $B_{2,it}$ is continuously distributed. Let $C_{i}$ be equal to $1$ if $Z_{i1}=Z_{i2}=...=Z_{iT_{i}}$ and zero otherwise, and let $\hat{p}=\frac{1}{n}\sum_{i=1}^{n}(1-C_{i})$.

Suppose i. $D_i$ is diagonal, ii. if $C_i= 0$ then $Q_{i}$ is non-singular, and iii. the submatrix of $\tilde{Q}_{i}$ formed by removing its first row and column is non-singular. Define estimates

align*[align* omitted — 246 chars of source]

where $\tilde{\beta}_{i,2}$ and $\tilde{\beta}_{i,3}$ are respectively the individual OLS coefficients on $Z_{it}$ and $B_{2,it}$. Let $\tilde{\beta}_{i}^{*}=(\tilde{\beta}_{i,1}^{*},\tilde{\beta}_{i,2}^{*},\tilde{\beta}_{i,3}')'$. Then $\lim_{\lambda\to0}\hat{\theta}=\frac{1}{n}\sum_{i=1}^{n}a_{i}'\tilde{\beta}_{i}^{*}$. \end{P4}

The limit in Proposition 5 differs from plug-in individual OLS in that the OLS coefficient on $Z_{it}$ is replaced with the alternative estimate $\tilde{\beta}_{i,2}^{*}$, and the intercept is adjusted accordingly. If $Z_{it}$ varies for individual $i$, then $\tilde{\beta}_{i,2}^{*}$ is equal to the individual OLS estimate $\tilde{\beta}_{i,2}$. However, if $Z_{it}$ does not vary, then $\tilde{\beta}_{i,2}^{*}$ is equal to $\frac{1}{\hat{p}n}\sum_{i=1}^{n}(1-C_{i})\tilde{\beta}_{i,2}$, which is the average of the OLS coefficients among the individuals for whom $Z_{it}$ does vary. In other words, for individuals without variation in $Z_{it}$, we impute the value of this coefficient as the average among individuals for whom $Z_{it}$ varies.

\theoremstyle{plain} \newtheorem*{P5}{Proposition 6} \begin{P5} Define $\tilde{B}_{it}=D_{1,i}^{-1/2}(B_{1,it}-\bar{B}_{i})$ and let $\tilde{B}_{i}=(\tilde{B}_{i1},\tilde{B}_{i2},...,\tilde{B}_{iT_{i}})'$. Define the projection matrix $P_{i}=D_{1,i}^{-1/2}\tilde{B}_{i}^{\dagger}\tilde{B}_{i}D_{1,i}^{1/2}$ and the following vectors of coefficients

align*[align* omitted — 250 chars of source]

In addition, let $\tilde{\beta}_{i,1}^{*}=(\bar{Y}_{i}-\bar{B}_{i}'\tilde{\beta}_{i,2}^{*})$ and define $\tilde{\beta}_{i}^{*}=(\tilde{\beta}_{i,1}^{*},{\tilde{\beta}_{i,2}^{*}}{}')'$/ Then $\lim_{\lambda\to0}\hat{\theta}=\frac{1}{n}\sum_{i=1}^{n}a_{i}'\tilde{\beta}_{i}^{*}$. \end{P5}

Proposition 6 considers the small $\lambda$ limit of our estimator in the general case. If $Q_{i}$ has full rank, then $\tilde{\beta}_{i}^{*}=\tilde{\beta}_{i}$, the individual OLS estimate. To interpret $\tilde{\beta}_{i}^{*}$ when $Q_{i}$ is singular, first note that $P_{i}$ is the orthogonal (with respect to the inner-product $\langle a,b\rangle:=a'D_{i}b$) projection onto the range of $\tilde{B}_{i}'$. Therefore, $I-P_{i}$ is the orthogonal projection onto the null space of $\tilde{B}_{i}$. If $Q_{i}$ is singular, this null space is non-trivial. This is problematic because, by construction, $(I-P_{i})\tilde{\beta}_{i,2}^{\circ}=0$. That is, the projection of the coefficients $\tilde{\beta}_{i,2}^{\circ}$ onto this subspace is zero, and so, loosely speaking, this part of $\tilde{\beta}_{i,2}^{\circ}$ is missing. The second term in the definition of $\tilde{\beta}_{i,2}^{*}$ adjusts for this by, in effect, replacing the missing part of $\tilde{\beta}_{i,2}^{\circ}$ with the projection of $(\frac{1}{n}\sum_{i=1}^{n}P_{i})^{-1}\frac{1}{n}\sum_{i=1}^{n}\tilde{\beta}_{i,2}^{\circ}$ onto this subspace. In the special case in Proposition 5, when $Q_{i}$ is singular, the null space consists simply of the vectors proportional to $(1,0,0,...,0)'$, and projecting onto this space yields the coefficients on $Z_{it}$.

Consistency and Asymptotic Normality

In order to establish the statistical properties of our estimator, we impose some additional conditions. In the assumptions below, inequalities involving random variables are understood to hold almost surely. Throughout, we define $\delta:=E[\|(Q_{i}+\lambda D_{i})^{-1}\|]$. \ignorespacesNote that $\delta$ may change with the sample size. This is because it depends directly on $\lambda$, which we let shrink as the sample grows, and $Q_i$ which changes with the dimension of the vector of basis functions $b(\cdot)$ and the number of time series observations $T_i$. Similarly, $a_i$, $A_i$, $B_i$, $r(\cdot,\cdot)$, $u_i$, and $\beta_i$ all depend on the dimension of the series approximation, and therefore, they too may change with the sample size.\ignorespaces

\theoremstyle{definition} \newtheorem*{A3}{Assumption 3 (Consistency)} \begin{A3} For some scalar $0<c<\infty$, i. $\|A_{i}\|,\|a_{i}\|,\|B_i\|\leq c$, ii. $\|(\frac{1}{n}\sum_{i=1}^{n}A_{i})^{-1}\|\leq c$, iii. $\|\beta_{i}\|\leq c$, iv. $\|E[u_{i}u_{i}'|X_{i}]\|\leq c$, v. $H_{t}^{-}(\cdot)$ and $H_{t}^{+}(\cdot)$ are uniformly bounded, vi. $\sup_{x,\mathbb{X}}|r(x,\mathbb{X})|\leq\ell$ with $\ell\to0$, vii. $T_{i}\geq T$, viii. The first row and column of $D_{i}$ contains only zeros, and the eigenvalues of $D_{1,i}$ are bounded above by $c$ and below by $1/c$. \end{A3}

\theoremstyle{definition} \newtheorem*{A4}{Assumption 4 (Asymptotic Normality)} \begin{A4} For some finite constants $c,\xi,v,q>0$ such that $(v-2)(q-2)>4$, i. $\frac{1}{T}\sum_{t=1}^{T}E[u_{it}^{v}]\leq c$, ii. $1/c\leq Var(a_{i}'\beta_{i})$, and iii. we have \[ n^{\frac{1}{v}+\frac{1}{q}-\frac{1}{2}}E[\|(Q_{i}+\lambda D_{i})^{-1}\|^{q/2}]^{1/q}(\delta/T)^{\xi}=o(1). \] \end{A4}

\theoremstyle{definition} \newtheorem*{A5}{Assumption 5 (Remainders)} \begin{A5} $\lambda\delta,\ell\sqrt{\delta},\ell=o(\sqrt{\frac{1}{n}})$, $\frac{\delta}{T}=O(1)$, and $\frac{J\lambda^{2}\delta^{3}}{T}=o(1)$. \end{A5}

Assumption 3 i. and ii. impose conditions on $a_{i}$, $A_{i}$, and $B_i$ which are chosen directly by the researcher. 3.iii imposes that the individual-specific mean parameter $\beta_{i}$ is bounded in norm. iv. concerns the conditional second moments of $u_{i}$. The condition allows for serial correlation in $u_{it}$, but it restricts the correlation so that the operator norm of $E[u_{i}u_{i}'|X_{i}]$ remains bounded as $T$ grows. In the special case in which $u_{i}$ is serially uncorrelated and homoskedastic, $\|E[u_{i}u_{i}'|X_{i}]\|$ equals the constant variance of $u_{it}$, and so, the condition holds. 3.v restricts $H_{t}^{-}(\cdot)$ and $H_{t}^{+}(\cdot)$. 3.vi is a condition on the sieve approximation error. Conditions of this form hold for many choices of sieve space used in practice under smoothness conditions on $h(\cdot,\mathbb{X})$, see e.g., DeVore1993 for examples. 3.vii imposes that the number of time periods $T_{i}$, which can vary between individuals, is bounded below by some $T$. 3.viii is a weak condition on $D_{i}$ which is chosen by the researcher.

Assumption 4 ensures a normal limiting distribution via a Lyapunov condition. The assumption restricts the $v$-th moment of one random variable and the $q$-th moment of another. $q$ and $v$ must be strictly positive and satisfy $(v-2)(q-2)>4$, which implies that $v,q>2$. The conditions trade-off in the sense that if $v$ is large, then $q$ need not be much larger than $2$ and vice versa. 4.i bounds the average $v$-th moment of $u_{it}$. 4.ii states that the variance of the individual-specific approximation $a_{i}'\beta_{i}$, is bounded below \ignorespaces which requires that either $a_i$ or $\beta_i$ varies between individuals. If the condition fails, then the estimator may achieve faster than $\sqrt{n}$-convergence, in which case the asymptotic variance that we derive is degenerate. This issue is studied in a related setting by fernandezval2025, who show that it results in conservative inference for plug-in methods. A similar phenomenon is also considered in Pesaran\ignorespaces. 4.iii requires that a particular sequence is $o(1)$. Each entry in the sequence is a product of three terms. The first is $n^{\frac{1}{v}+\frac{1}{q}-\frac{1}{2}}$. The conditions on $v$ and $q$ imply that this term goes to zero with $n$. The second term is the $q/2$-th moment of $\|(Q_{i}+\lambda D_{i})^{-1}\|$ raised to the power $1/q$, the third is $\delta/T$ raised to the power $\xi$. Note that $\xi$ can be any strictly positive constant. As such, in the case of $\delta/T\to\infty$, 4.iii holds for a sufficiently large choice of $\xi$ so long as the $q$-th moment of $\|(Q_{i}+\lambda D_{i})^{-1}\|$ grows (at most) at a polynomial rate with $T$.

Assumption 5 restricts the rates at which various sequences converge to zero. It ensures some terms in the asymptotic expansion of the estimation error are second order.\\

Theorem 2 provides general asymptotic theory for our estimator ((ref)). It applies both when regressors are continuous or discrete or a combination of the two. The result follows from the more general result in Lemma 1 in the appendix, which applies to any estimator of the form ((ref)) such that $W_{i}Q_{i}^{\dagger}Q_{i}=W_{i}$. \theoremstyle{plain} \newtheorem*{T1}{Theorem 2 ( Asymptotics)} \begin{T1} Suppose Assumptions 1, 2, and 3 hold and $\lambda\delta\to0$.

a. (Consistency)

align*[align* omitted — 177 chars of source]

b. (Asymptotic Normality)

In addition, if Assumptions 4 and 5 hold, then $\hat{\theta}-\theta_{0}=O_{p}(n^{-1/2})$ and $\sqrt{n}\sigma^{-1}(\hat{\theta}-\theta_{0})\sim^{a}N(0,1)$, where $\sigma^{2}$ is given by:

equation[equation omitted — 122 chars of source]

\end{T1}

Theorem 2 shows the crucial role of $\delta$ in the asymptotic behavior of the estimator. To understand what $\delta$ represents, suppose for simplicity that $D_{i}$ is the identity matrix. Then $\|(Q_{i}+\lambda D_{i})^{-1}\|$ is equal to $(\mu_{\text{min}}(Q_{i})+\lambda)^{-1}$, where $\mu_{\text{min}}(Q_{i})$ is the smallest eigenvalue of $Q_{i}$. As such, if $Q_{i}$ is close to singular, then $\|(Q_{i}+\lambda D_{i})^{-1}\|$ is close to $1/\lambda$. In the extreme case, if $Q_{i}$ is singular with probability $p$, then $p/\lambda\leq\delta$. Therefore, the condition $\lambda\delta\to0$ is only possible if $p$ shrinks to zero with the sample size, which generally requires that $T$ grows with $n$. On the other hand, if $E[1/\mu_{\text{min}}(Q_{i})]$ is bounded by a finite constant, then $\delta\leq E[1/\mu_{\text{min}}(Q_{i})]$, uniformly over $\lambda$. It follows that $\lambda\delta=o(n^{-1/2})$, so long as $\lambda$ shrinks sufficiently quickly to zero.

To examine this in more detail, we consider two extreme cases below. In the first case, captured in Corollary 1, we suppose that $E[\mu_{\min}(\tilde{Q}_{i})^{-1}]$ is bounded above, where $\mu_{\min}(\tilde{Q}_{i})$ is the smallest eigenvalue of $\tilde{Q}_{i}$. This is only possible if all regressors are continuously distributed and $T>J$. In the absence of approximation error (so that $\ell=0$ with $J<T$ fixed), root-$n$ consistency and asymptotic normality do not require that $T$ grows with the sample size. Indeed, if $E[\mu_{\min}(\tilde{Q}_{i})^{-1}]$ is finite, then the efficiency bound in Chamberlain1992 is finite and root-$n$ regular estimation is possible.

\theoremstyle{plain} \newtheorem*{C1}{Corollary 1 (Continuous Case)} \begin{C1} Let $\|Q_{i}\|\leq c$. Suppose Assumptions 1, 2, and 3 hold. If $E[\mu_{\min}(\tilde{Q}_{i})^{-1}]$ is bounded above and $\lambda\to0$, then:

align*[align* omitted — 117 chars of source]

If in addition Assumption 4 holds, $\ell,\lambda=o(\sqrt{\frac{1}{n}})$ and $\frac{J\lambda^{2}}{T}=o(1)$, then $\hat{\theta}-\theta_{0}=O_{p}(n^{-1/2})$ and $\sqrt{n}\sigma^{-1}(\hat{\theta}-\theta_{0})\sim^{a}N(0,1)$, where $\sigma^{2}$ is given by ((ref)). \end{C1}

At the other extreme, Corollary 2 applies in the case of a single binary regressor. We assume that for a given individual $i$, the probability $X_{it}=1$ is given by $\pi_{i}\in(0,1)$ and that, conditional on $\pi_{i}$, the regressor is independent over time. In this case, $\tilde{Q}_{i}$ must be singular with positive probability because for an individual $i$, the regressor is constant with positive probability. However, as $T_{i}$ grows, this probability shrinks to zero at a rate that depends on the distribution of $\pi_i$.

\theoremstyle{plain} \newtheorem*{C2}{Corollary 2 (Binary Case)} \begin{C2} Suppose Assumptions 1, 2, and 3 hold, and $b(X_{it})=(1,X_{it})'$ where $X_{it}$ is binary. Suppose $P(X_{it}=1|\pi_{i})=\pi_{i}$ and the entries of the sequence $\{X_{it}\}_{t=1}^{T_i}$ are jointly independent conditional on $\pi_{i}$. Let $\pi_{i}$ admit a probability density $f_{\pi}$ so that $f_{\pi}(\pi)\leq C(1-\pi)^{\omega}\pi{}^{\omega}$ where $\omega>0$. Then if $\lambda\to 0$ and $T\to\infty$ we have:

align*[align* omitted — 162 chars of source]

In addition, if $\frac{T^{-(1+\omega)}}{\lambda}=O(1)$, $T^{-(1+\omega)},\lambda=o(\sqrt{1/n})$, and Assumptions 4.i and 4.ii hold with $q/2<\omega$, then $\hat{\theta}-\theta_{0}=O_{p}(n^{-1/2})$ and $\sqrt{n}\sigma^{-1}(\hat{\theta}-\theta_{0})\sim^{a}N(0,1)$, where $\sigma^{2}$ is given by ((ref)). \end{C2}

Corollary 2 establishes root-$n$ consistency only under the condition that $T^{-(1+\omega)}=o(\sqrt{1/n})$. Thus the rate at which $T$ must grow with $n$ depends on the rate at which $f_{\pi}(\pi)$ goes to zero as $\pi$ goes to zero or one. If $f_{\pi}(\pi)$ goes quickly to zero, then the probability $X_{it}$ is constant over time goes to zero quickly as $T$ grows, and so $T$ need not increase rapidly with $n$. This phenomenon is considered in Chernozhukov2013 and tied to the rate at which the identified set shrinks with $T$.

Results for other cases, for example, with both discrete and continuous regressors, may also be obtained from Theorem 2. As in the proofs of Corollaries 1 and 2, it would suffice to derive a convergence rate for $\delta$ under suitable assumptions.