Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,043 characters · 17 sections · 96 citation commands
Welfare Analysis in Dynamic Models
\setcounter{page}{1}
\makeatletter \def\thanks#1{\protected@xdef\@thanks{\@thanks \footnotetext{1}}} \makeatother
\makeatletter \def\thanks#1{\protected@xdef\@thanks{\@thanks \footnotetext{1}}} \makeatother
Dynamic considerations are important in applied work Miller1984, Wolpin, Pakes:1986, Rust, HotzMiller, HotzMillerSandersSmith, keane1994solution, AMira2002,AMira2007,BBL. These considerations are captured by the value function, defined as the present discounted value of agents' expected per-period utility. Prior work has focused on identifying utility parameters from the optimality of agents' behavior and conducting inference on them. Building on this foundation, this paper targets a different yet complementary class of parameters that depend on the value function which we introduce below.
Welfare analysis is a central objective in economics. This paper develops identification, estimation, and inference results for welfare metrics in dynamic models. Examples of such metrics include the average value function and average marginal welfare effects, e.g., due to the change in observable time-invariant characteristics such as initial wealth. Most importantly, we decompose welfare effects into the direct effect through changes in utilities holding agents' behavior fixed and the indirect effect from changes in agents' behavior. This decomposition is analogous in spirit to Oaxaca-Blinder decomposition of outcome distributions yet different since welfare metrics are based on latent utilities rather than observed outcomes.
The economic motivation for the welfare decomposition we propose is drawn from various fields in applied economics, including labor economics, public finance, and health economics. In labor economics, kitagawa55, oaxaca73 and blinder73 decompose the gender wage gap into structural and composition effects. In public finance, it is common to distinguish between mechanical and behavioral effects of taxation policies, see e.g. Chetty2009. In health economics, the work by EinavFinkelsteinCullen2010 differentiates the direct effect of a health insurance price change from the indirect effect arising from changes in selection. For an example of cancer screening EinavEtAl2020, the total welfare effect corresponds to the introduction of a new screening option, and the direct effect is its counterpart as if the frequency of screening were mandated. One of the paper's contribution is to give identification, estimation, and inferential results for direct and indirect effects in dynamic models.
The next contribution of the paper is the dynamic dual representation of welfare metrics, which directly maps the per-period utility to the welfare metric of interest. For the case of average welfare, the dynamic dual representation reduces to a known expression of expected per-period utility. Furthermore, neither value function nor other dynamic object needs to be estimated for certain discrete choice models, such as those of Rust and HotzMiller. For cases beyond average welfare, the dynamic dual representation involves backward discounting. In these cases, we derive a doubly robust representation that facilitates consistent estimation if at least one of the two nuisance components, the value function or the dynamic dual representation, is correct. Appendix (ref) gives large sample properties of the proposed estimators based on Neyman-orthogonal moment equations, in which the dynamic dual representation plays a central role. This paper is the first result in the literature to leverage duality in dynamic models whose state distribution is strictly stationary.
Another contribution of the paper is to introduce a Lasso estimator for the value function, the dynamic dual representation, and associated mean square convergence rates in the setting of high-dimensional state variables. Specifically, the number of state variables may exceed the sample size, provided that only a small subset is relevant. We derive a novel least squares criterion that distinguishes our approach from previous IV-based methods. Adding an $\ell_1$ penalty to the sample analog of this criterion yields a Lasso estimator of the value function. In addition, we propose a neural network estimator and derive corresponding mean square convergence rates in the low-dimensional case. Although the results are presented for the value function, the methods apply to any fixed-point solution of an integral equation of the second kind -- such as the Q-function in reinforcement learning -- and do not require strict stationarity of the state distribution.
We revisit the study of teacher absenteeism in the nonformal education centers (NFEs) of Rajasthan, India, initially analyzed by DHR (henceforth, DHR). Their approach emphasizes unobserved heterogeneity in teachers' leisure preferences—shaped by factors such as past effort, illness, fatigue, and informal obligations. We conjecture that this heterogeneity can be effectively proxied by a sufficiently long window of prior work history, even if the history itself may not have a structural or causal interpretation. DHR's structural estimates remain robust to this exercise. In addition, we focus on teacher welfare and find that the failure to properly account for teacher heterogeneity results in the average teacher welfare overestimated by 13–20%%. As a side contribution, we demonstrate that our methodological framework remains valid in finite-horizon settings, provided the panel is sufficiently long relative to standard discount factors.
A large body of work is dedicated to welfare analysis in discrete choice models with unobserved heterogeneity, as studied, e.g., in Bhattacharya2015. As Bhattacharya2024 discusses, welfare calculations are based on latent utilities rather than observed outcomes. For dynamic choice -- the focus of this paper -- important early references include Miller1984, Wolpin, and Pakes:1986. Within this group, a notable subclass of models includes those with the terminal action property, as in HotzMiller and HotzMillerSandersSmith, and, more broadly, models with finite dependence, as in ArcidiaconoMillerFD. In HotzMiller, structural parameters are identified via a regression problem that depends on conditional choice probabilities. The dynamic dual representation we derive extends this insight to the average value function and other welfare metrics.
Estimation of dynamic discrete choice models has received substantial attention; see, e.g., AMira2002, AMira2007, BBL, PSM, ArcidiaconoMiller, Arcidiacono:2013, ArcidiaconoEtAl2013, ChenAck, AMagesan, BShum. Recent work by AdsmEck2022 relaxes the terminal action property while allowing for an infinitely supported state space. Appendix (ref) of the present paper outlines an automatic debiasing approach that targets nonlinear functionals of the value function and thus could be applicable to the model in AdsmEck2022.
Last but not least, we contribute to the literature on estimating fixed points of integral equations of the second kind; see, e.g., SLinton. Prior work by chen2022wellposedness has shown that this is a special case of a well-posed NPIV problem and has developed value function estimators that achieve minimax-optimal rates in low-dimensional settings; see also xia2022krylovbellmanboostingsuperlinearpolicy for related IV-based approaches. However, these methods may not readily extend to penalized estimators that remain consistent in high-dimensional settings. The least squares criterion we derive circumvents this limitation. We also give a least squares criterion for estimating the dynamic dual representation, that only depends on the welfare metric of interest, and so enables automatic debiasing like chernozhukov2021automatic and chernozhukov2024qm.
The paper is organized as follows. Section (ref) presents the general framework and examples of welfare metrics. Section (ref) decomposes the differences in average welfare in the spirit of Oaxaca and Blinder. Section (ref) revisits DHR study of teacher absenteeism. Section (ref) gives the dynamic dual and doubly robust representations for welfare metrics. Section (ref) describes orthogonal estimating equations for average welfare and related averages. Least squares estimators of the value function and dynamic dual representation are presented in Section (ref). Section (ref) concludes. Appendix (ref) considers the example of dynamic discrete choice. Appendix (ref) gives proofs of the results in main text. Appendix (ref) gives large sample properties of the estimators. Appendices (ref)--(ref) give mean square rates for first-stage estimators. Appendix (ref) extends the proposed method to nonlinear functionals of the value function. Appendix (ref) demonstrates the method of Appendix (ref) for dynamic binary choice.
We consider estimation and inference on welfare metrics that depend on the value function. To define the value function, let $X_t, (t=0,1,...)$, denote a time series of observed state variables, which we assume to be a time-homogeneous, first-order Markov process with initial element $X:=X_0$. The value function is determined by a per-period reward $\zeta_0(X)$, or, in other words, expected utility in a single period conditional on the state $X$, which we assume to be identifiable. The value function $V_0(X)$ is the present discounted value of per-period rewards given the current state, satisfying
where $\beta \in [0,1)$ is a known discount factor. The value function satisfies the integral equation:
where $X$ is the current and $X_{+}$ is the next period element of the first-order Markov process (see Lemma (ref)). In what follows, we assume that the time series is strictly stationary. The welfare metrics we consider are linear functionals of the value function $V_0(X)$ having the form
where $w_0(X)$ is some function of the state.
To give an example of per-period utility, we consider the dynamic discrete choice problem Rust,HotzMiller,AMira2002. In each period, $(t=0,1,...),$ the agent chooses an action $j$ from a finite choice set $\mathcal{A}$. The utility of choice $j$ in period $t$ is additively separable in a function of the current state $u(X_t,j)$ and private shock $\varepsilon_t(j)$ and is given by $$\bar{u} (X_t, j, \epsilon_t) = u(X_t,j) + \epsilon_t(j), \quad j \in \mathcal{A}.$$ Here the sequence ${(X_{t}, J_t)}$ is strictly stationary, so the time index $t$ can be dropped. The per-period reward is the expected utility $\zeta_0(x)$, obtained by taking the expectation over choices: $$ \zeta_0(x) = \sum_{j \in \mathcal{A}} (u(x,j) + \mathrm{E} [ \epsilon (j) \mid X=x, J=j]) {\mathrm{P}} (J=j \mid X=x), $$ where ${\mathrm{P}} (J=j \mid X=x)$ is the probability that an agent chooses $j$ when $X=x$, and $\mathrm{E} [ \epsilon (j) \mid X=x, J=j]$ is the expectation of $\epsilon (j)$ under the choice $J=j$ and given $X=x$. For example, in a special case where the choice is binary and the private shocks are distributed as Gumbel,
where $p_0(x) = {\mathrm{P}} (J=1 \mid X=x)$, $\mathcal{A}=\{1, 0\}$, and $H(t) = \gamma_e - t\ln t -(1-t) \ln (1-t),$ with $\gamma_e = 0.5227$ denoting the Euler constant. Here and generally for dynamic discrete choice the per-period utility will be the sum of the expected value of the observable part of the utility plus the expected value over optimal choices of the private shock part. We assume that the utility components $u(x,1)$ and $u(x,0)$ are known up to a structural parameter that is identified.
Within the discrete choice problems, our target parameter $\delta_0$ represents a welfare metric since $V_0(\cdot)$ is the expected value of an agent making optimal dynamic choices conditional on state $X$. Our first example is the expected value function where $w_0(X)=1$.
We focus particularly on settings where the state variable include observable time-invariant heterogeneity in individual, per-period expected utility denoted by a vector $K$. Examples of $K$ include occupation type KeaneWolpin1997, gender and class grade of a child ToddWolpin and teacher test score DHR. Time invariant state variables could also represent observable heterogeneity in individual, per-period expected utility. In this case, the state vector $X_t$ can be decomposed as $X_t = (S_t, K),$ where $K_t= K$ does not vary over time. Here first-order time homogeneity of $X_t$ implies that $(S_t)_{t > 0 }$ is a first-order time-homogeneous Markov chain conditional on $K$.
Extending this example to differences in average welfare across groups is straightforward by differencing the parameter of interest in Example (ref) across different values of $K$. In this example and the others, the weight $w_0(X)$ is unknown and will need to be estimated. The identification and estimation of $w_0(X)$ will be accounted for in the results that follow.
When $K$ represents an endowment of some resource it may be of interest to consider the welfare effect of changing the distribution of that endowment.
For continuously distributed $K$ an effect of interest could be the average effect of changing $K$ on the value function.
All of the above examples of target parameters fall in the following general theoretical framework. Let $Z$ denote a data vector that includes $X$ and $X_+$, and let $V$ denote a possible value function. Also let $m(Z,V)$ denote a function of $Z$ and the function $V(\cdot)$ (i.e. $m(Z,V)$ is a functional of $V$.) We consider parameters of the form
where $\mathrm{E}[m(Z, V)]$ is linear in $V$. We will impose throughout that the expectation $\mathrm{E}[m(Z, V)]$ is mean square continuous as a function of $V$, meaning that there is a constant $C>0$ such that for all $V(X)$ with $\mathrm{E}[V(X)^2]<\infty$,
By the Riesz representation theorem mean square continuity of $\mathrm{E}[m(Z,V)]$ is equivalent to existence of a function $w_0(X)$ with $\mathrm{E}[w_0(X)^2]<\infty$ such that
for all $V(\cdot)$ with $\mathrm{E}[V(X)^2]<\infty$. Here, we see that under mean square continuity, any parameter as in equation (ref) can be represented as a linear function of the value function. There are many other potentially interesting examples of such welfare metrics. In the next section we consider decomposing differences in average welfare into direct and indirect components.
To motivate our analysis, we present a simple running example in the context of a randomized controlled trial, in which a one-shot treatment assigned at time $t = 0$ generates dynamic incentives. Suppose we aim to analyze welfare differences between the treated and control populations, denoted by $1$ and $0$, respectively. In each group, the time-varying state variable is denoted by $S$. The full state vector is $X = (S, K)$ where $K$ is the indicator of the treatment status. We adopt the notation in chernozhukovmelly.
Let $V_0^1(s) = V_0(s,1)$ and $V_0^0(s)=V_0(s,0)$ denote the treated and control value functions. Define the treated average welfare as
and the control average welfare as
Both quantities are special cases of the group average welfare defined in Example (ref). For $k \in \{1, 0\}$, we have
where $\pi^k(s)$ denotes the probability distribution function of $S$ conditional on $K = k$. The counterfactual welfare metric
does not correspond to the group average welfare of any observable subpopulation. Instead, it is constructed by integrating the treated value function with respect to the stationary distribution of the control population. Provided that $\pi^1(s)$ and $\pi^0(s)$ share the same support, this parameter is well-defined.
The treatment-control welfare difference can be decomposed in the spirit of kitagawa55, oaxaca73, and blinder73:
If treatment is randomly assigned, this welfare difference admits a causal interpretation.
The proposed decomposition has an intuitive interpretation when the treatment affects only per-period utilities and does not impact the state transition. We describe such empirical settings in Examples (ref) and (ref) below. In this case, the stationary distributions $\pi^1(s)$ and $\pi^0(s)$ differ solely due to agents making different optimal choices in the treated and control states, respectively. The first summand,
captures the difference in per-period utilities holding the distribution of states fixed at the control level. This term can be interpreted as the direct or mechanical effect. The second summand,
reflects changes in agents' behavior that alter the distribution of states, holding the utilities fixed. This term can be interpreted as the indirect or behavioral effect. The sum of the direct and behavioral effects gives the total welfare difference.
Proposition (ref) expresses the counterfactual welfare $\delta_{\langle 1 \mid 0 \rangle}$ as a linear functional of the treated value function.
Recent work has extended classical decomposition methods to modern settings. chernozhukov2021automatic derive the Riesz representer for the Average Treatment Effect on the Treated (ATET). vafa2024estimatingwagedisparitiesusing provides an Oaxaca-Blinder decomposition of wage differences. Proposition (ref) departs from these approaches by offering a decomposition of average welfare in dynamic models where welfare is based on latent utilities rather than observed outcomes.
We include the average counterfactual welfare $\delta_{\langle 1 \mid 0 \rangle}$ as Example (ref) and discuss its estimation in Section (ref). Notice that the parameters $ \delta_{\langle 1 \mid 1 \rangle}$ and $ \delta_{\langle 0 \mid 0 \rangle}$ are special cases of Example (ref) with $k=1$ and $k=0$, respectively. Therefore, it is straightforward to extend this example to accommodate direct and indirect effects. We discuss the estimation of the counterfactual welfare measure and related effects further in Section (ref).
We study how daily financial incentives affect teacher attendance in single-teacher nonformal education centers (NFEs) operated by the NGO Seva Mandir in tribal villages of Udaipur, Rajasthan, India. From 2003 to 2005, DHR conducted a randomized trial in which tamper-proof cameras recorded photographs at school opening and closing. A school day was deemed valid if the two images were at least five hours apart and at least eight students were present. At the end of each month, teachers earned a base salary of 500 Rupees (Rs) if they worked fewer than 10 days, plus a 50 Rs bonus for each additional day of work beyond that threshold. The 10-day cutoff thus created a nonlinear dynamic incentive, which we focus on in this application.\footnote{We abstract from the firing threat, which appears negligible: no teacher was fired during the study period, even in cases of near-total absence. According to DHR, Seva Mandir adopts a long-term view in assessing teacher performance, which may explain the lack of dismissals.} Our dataset comprises daily attendance records for 57 teachers over an 18-month period (January 1, 2004 to June 30, 2005), along with a test score administered prior to the start of teaching.\footnote{Following DHR, the estimation sample includes only weekdays when teachers actively choose between working and taking leisure. Holidays and weekends, though counted toward pay, are excluded from analysis.}
We revisit the dynamic behavioral model of DHR, henceforth DHR. Let $t$ denote the day of the month, ranging from $1$ to $T=30$. On each day $t$, a teacher chooses between working ($j_t = 1$) and taking leisure ($j_t = 0$). On the final day of the month, the consumption utility is determined by the monthly paycheck
where $500$ is the base salary, $d_T$ is the total number of days worked by day the final day $T$, and $50$ is the bonus. For all days $t < T$, there is no consumption; utility accumulates only through leisure and is modeled as
Here, $ \bar{x}_t \in \mathrm{R}^{p_X}$ is the state vector including a constant and possibly other observable characteristics, and $\mu_0 \in \mathrm{R}^{p_X}$ is a parameter to be estimated. The state vector is $X_t = (\bar{X}_t, d_t)$. For example, in Model I of DHR, the leisure utility is assumed to be the same for all teachers, which corresponds to $u(x_t, 0) = \mu_0$ and $X_t =d_t$. The per-period utilities are not discounted. DHR includes only a handful of observables into $\bar{x}_t$ so as to leverage standard maximum likelihood estimators.
We consider a stylized dynamic binary choice model as described in Section (ref). Since teachers cannot be fired, the base salary is assumed not to affect their choice between work and leisure. We decompose the total monthly bonus into daily payments. Specifically, we assume the utility of working ($j_t = 1$) is given by:
where the indicator function, referred to as “In the money” by DHR, captures the bonus structure in the stylized model. The parameter $\mu_1$ converts monetary rewards (in Rupees) into utility units. The stylized model in equations (ref)--(ref) preserves the monetary incentives of the exact model (ref)--(ref), aside from discounting of the bonus. To abstract from the finite-horizon considerations, we restrict attention to calendar days 15, 16, and 17 of each month. If this model is relevant, the methods developed in this paper permit the state vector $\bar{X}_t$ to include high-dimensional covariates.
Table (ref) compares the exact estimates reported by DHR (Columns (1)-(2)) to our stylized infinite-horizon replications (Columns (3)-(4)) for selected coefficients. The results are encouraging. First, the estimated bonus coefficient under the stylized model falls within the range of the exact estimates. The standard errors of the replicated coefficients are, on average, 2.5 to 3 times larger, as expected given the smaller sample (only three days per month). Second, the coefficient on teacher test scores remains negative across all specifications, consistent with the original findings. Other coefficients (not reported here) also closely match their exact counterparts. These results suggest that the stylized infinite-horizon framework is appropriate for this dataset.
We further investigate the role of prior work history in explaining teachers work decisions, a possibility raised by DHR. DHR included the first lag of work history in Model VIII, and we extend this idea by incorporating 89 more lags. Additionally, we summarize work history using a work streak variable, defined as the number of consecutive days a teacher has worked without taking leisure:
We investigate the role of prior work history both in the predictive and structural settings.
Table (ref) shows the out-of-sample mean squared error (MSEs) for predicting a teacher’s decision to work, using models that sequentially expand the covariate set from Model I to Model IV. Adding the teacher's test score on top of days worked yields minimal improvement. Including month dummies (Model II) results in a modest reduction in MSE. Incorporating the work streak (Model III) leads to a substantial improvement: relative to Model II, the MSE falls by roughly 32% for Logit, 32% for Probit, and 35% for Random Forest. Finally, adding the full 90-day work history dummies (Model IV) further improves prediction, halving the MSE relative to the baseline. These findings show that prior work history has high predictive power even after other observables have been taken into account.
As a next exercise, we revisit the stylized structural model with the aim of flexibly modeling the utility of leisure. As noted by DHR, decisions to skip work may be influenced by factors such as social norms, informal requests and commitments, accumulated effort, or fatigue. These factors are unlikely to be fully captured by basic observables like test scores. A longer spell of prior work history may be more granular and thus may better represent teacher's heterogeneity even if the history itself has no structural or causal interpretation. Assuming only a small number of lags suffices to capture the history, we include a 90-day window of lagged attendance indicators. This sparsity assumption calls for the use of Lasso estimators of the value function and the dynamic dual representation, developed in Section (ref).
Table (ref) reports selected coefficients for the structural parameter estimates (Panel A) as well as welfare metrics (Panels B, C and D). Columns (1)-(2) correspond to a simple model of (ref) whose only observable is teacher's test score. Columns (3)-(4) correspond to a more sophisticated model of (ref) where observables include prior work history. In both cases, the model is estimated using Algorithm (ref) described in Appendix G based on Logit and Random Forest estimators of conditional choice probability. The welfare metrics are estimated using the dual estimator described in Section (ref). Instead of using the value function, it combines the structural estimates of Panel A with the choice probabilities.
Our findings are as follows. First and foremost, the structural parameters—particularly the bonus coefficient ($\mu_1$) and the effect of teacher test score —are robust to the inclusion of the work history as a state component as well as to the choice of CCP estimator. In contrast, the welfare metrics are more sensitive to model specification. In particular, failure to account for the prior work history results in overestimating teacher welfare by 13-20$\%\%$. This overstatement persists across both the full sample and the subgroups defined by test scores values. Finally, the average welfare is lower for teachers with test scores at or below 30 than for those with scores above 40, as shown in Panels C and D, which is consistent with we include a 90-day window of lagged attendance indicators. Similar to DHR's interpretation, more skilled teachers are more committed to work and receive higher utility from teaching.
In this Section, we give a dual representation of the parameter of interest. This representation is important for several purposes. When $w_0(X)$ depends only on $K$, and so is time-invariant, the dual representation gives a simplified formula for $\delta_0$ that does not require solving any dynamic problem. Otherwise, the dual representation leads to a doubly robust moment condition for identification and estimation of the parameter of interest. The dual representation\footnote{See equations (16)-(17) in the first version of the paper CNS. } was derived in the previous version of this paper \citet*{CNS}.
A key part of the dual representation is a function of the state variable that is a backward discounted value of $w_0(X)$, given by
where $X_{-t}$ is the state variable in period $-t$ in the extended stochastic process ${X_t}$ where $t$ ranges over all the integers. Alternatively, $\alpha_0(X)$ is a fixed point of the backward dynamic operator
The following result gives the dynamic dual representation of weighted average welfare.
An interesting implication of this dual representation is that if $w_0(X)$ is time-invariant, then $\delta_0$ depends only on the per-period expected utility $\zeta_0(X)$.
In each of our first three examples, the weight was time-invariant so that Corollary (ref) applies and the parameter of interest $\delta_0$ depends only on $\zeta_0(X)$. Here are expressions for $\delta_0$ for Examples (ref)-(ref).
When the weight is not time-invariant, $\alpha_0$ will not generally have a closed form or explicit expression because it depends in a complicated way on the dynamic distribution of the state vector $X_t$. To help understand better the nature of $\alpha_0$ we revisit Example (ref) where the state variable follows an autoregressive process of order 1 with a Gaussian innovation. While this example may not correspond to a state distribution under a dynamic discrete choice model, we include it for pedagogic purposes to help explain the nature of $\alpha_0$.
In this section, we give an identifying moment condition for the parameter of interest that is doubly robust in the sense that it holds if just one of $V(\cdot)$ or $\alpha(\cdot)$ is the true function. This moment condition uses the identifying conditional moment restriction for $V_0$ in equation (ref). Let $Z$ denote a data observation which includes $(X,X_{+})$, $V$ denote a possible value function, and $\lambda(Z, V):= \beta V(X_{+}) - V(X) + \zeta_0(X)$. Equation (ref) is equivalent to the conditional moment restriction
This is a nonparametric conditional moment restriction like those of NeweyPowell and AiChen2003 where $X_+$ is an "endogenous" variable, $X$ is an "instrument", and $\lambda(Z,V)$ is a nonparametric residual as considered in CNS and chen2022wellposedness. Here we take $\zeta_0(X)$ to be a known function and will consider estimation of $\zeta_0(X)$ in the next Section.
Let $\alpha$ denote a possible function $\alpha_0$. A doubly robust moment function can then be formed as
Given a function $\xi$ of $X$ define
Lemma (ref) establishes double robustness of the moment function $g(Z,V,\alpha,\delta)$ which has zero expectation at $\delta=\delta_0$ if either $V=V_0$ or $\alpha=\alpha_0$ by equation ((ref)). A doubly robust estimator of the average treatment effect was given in (Robins) and (LRSP) characterize doubly robust moment functions as being linear in both non-parameric components.
In this Section, we give estimators of the welfare metrics we introduced in Sections (ref) and (ref). These estimators will account for the estimation of $\zeta_0(\cdot)$ and of $w_0(\cdot)$ or $m(Z,V)$ by including influence functions for their effect on identifying moments. The inclusion of these influence functions debiases for model selection and/or regularization in the estimation of unknown functions and corrects resulting standard errors for their estimation, as in LRSP.
For simplicity of exposition, we focus on panels with $T=2$ time periods where the pairs of consequent states $(X_{i1},X_{i2})_{i=1}^n$ are i.i.d. We use standard cross-fitting for i.i.d data (schick1986asymptotically) as common in work on debiased machine learning, chernozhukov2016double. For a weakly dependent time series with $T\geq 3$ periods, cross-fitting along both unit and time dimension is possible by leaving out neighboring folds, as discussed in CGST. Related work on conditional moment restrictions with weak dependence includes ChenSieveRiesz,ChenLiao2015,ChenLiaoWang2024.
We will first consider a weighted average value function parameter with time-invariant weight that is possibly estimated. This case includes average welfare and related averages in Examples (ref)--(ref). Let $F$ denote an unrestricted distribution for $Z$ and $w(K,F)$ and $\zeta(X,F)$ denote the probability limit (plim) of an estimated weight $\widehat{w}(K)$ and an estimator $\widehat\zeta(X)$ respectively. Let $\phi_{w}(Z)$ and $\phi_{\zeta}(Z)$ be the influence functions of $(1-\beta)^{-1}\mathrm{E}[ w(K,F)\zeta_0(X)]$ and $(1-\beta)^{-1}E[w_0(K)\zeta(X,F)]$ respectively. To nonparametrically debias for the estimation of $w(K)$ and $\zeta(X)$ and so construct a Neyman orthogonal moment function we add $\phi_{w}(Z)$ and $\phi_{\zeta}(Z)$ to the identifying moment function, as in LRSP, to obtain
where the true parameter $\delta_0$ solves
at the true value $\gamma_0$ of $\gamma$ and $\phi_0$ of $\phi$.
For cross-fitting purposes, we partition the set of data indices ${1, \ldots, n}$ into $L$ disjoint subsets $I_\ell$ of about equal size, $\ell = 1, \ldots, L$. Let $\widehat\gamma_\ell=(\widehat w_{\ell}, \widehat \zeta_{\ell})$ and $\widehat\phi_\ell =(\widehat\phi_{w\ell},\widehat\phi_{\zeta\ell})$ be estimators of the weight, per-period utility, and influence functions constructed using all observations not in $I_\ell$. Also let $\psi(Z,\widehat\gamma_\ell,\widehat\phi_\ell,\delta)$ be as in equation (ref) with $\widehat\gamma_\ell$ and $\widehat\phi_\ell$ in place of $\gamma$ and $\phi$. A cross-fit estimator of $\delta_0$ can be obtained from solving $\sum_{\ell = 1}^L \sum_{i \in I_\ell} \psi(Z_i,\widehat\gamma_\ell,\widehat\phi_\ell,\delta)/n =0$ for $\delta$ giving
An example of estimated per-period utility $\zeta$ and its correction term for dynamic binary choice is given in Appendix (ref). The Lemma (ref) in Appendix (ref) provides sufficient conditions for the validity of asymptotic inference.
We next continue with the description of group average welfare in Example (ref). A key difference from Example (ref) is that the time-invariant weighting function $w$ depends on the group probability ${\mathrm{P}} (K=k)$ which needs to be estimated.
In this Section we descrite the estimator of $\delta_0=\mathrm{E}[m(Z,V_0)]$ when $w_0(\cdot)$ varies with time. Let $m(Z,V,F)$ denote the plim of the estimated $m(Z,V)$ function, $\phi_m(Z)$ the influence function of $\mathrm{E}[m(Z,V_0,F)]$, and $\phi_{\zeta}(Z)$ the influence function of $\mathrm{E}[\alpha_0(X)\zeta(X,F)]$. The orthogonal moment function is
where $\phi_m$ corrects for the estimation of $m(\cdot)$ and $\phi_{\zeta}$ corrects for the estimation of $\zeta$. Algorithm (ref) below gives the proposed estimator of the parameter of interest. Lemma (ref) in Appendix (ref) establishes the validity of asymptotic inference.
Here there is no correction $\phi_m$ since the functional $m(Z,V)=\partial_K V(X)$ of $V$ does not involve any unknown components.
Algorithm (ref) summarizes the estimation steps of welfare metrics.
This section introduces novel least squares estimators for both the value function and the dynamic dual representation. In contrast to the welfare metrics, defined in Section (ref), these objects do not require strict stationarity to be well-defined. Thus, the estimators of the value function and the dynamic dual representation delivered here do not require the time series to be strictly stationary. Consequently, these estimators apply to any fixed point of a second-kind integral operator.
Sections (ref) and (ref) develop a new least squares criterion for the value function that accommodates high-dimensional covariates through penalization, enabling consistent estimation in a high-dimensional state space. Section (ref) proposes a distinct least squares criterion for the dynamic dual representation, which depends only on the welfare metric of interest. This innovation permits automatic debiasing in the style of chernozhukov2021automatic and chernozhukov2024qm. Section (ref) discusses the results.
The starting point of our analysis is the expectation operator $\mathrm{A}_0$ defined as
Rewriting (ref) in terms of $\mathrm{A}_0$ gives
or, equivalently, $V_0 = (\mathrm{I} - \mathrm{A}_0)^{-1} \zeta_0$. The operator $\mathrm{A}_0$ is akin to the integral equation operator in SLinton.
The value function can be represented as a minimizer of a criterion function that depends on $\mathrm{A}_0$. $V_0$ will minimize the expected squared difference of the left and right-hand sides of equation (ref), that is
where the second equality follows by squaring and dropping the term that does not depend on $V$ and the third and fourth equalities by iterated expectations. The expression minimized following the first equality is the nonparametric two-stage least squares criterion for NeweyPowell, Newey1991, and AiChen2003. The expression following the third equality is a hybrid that uses iterated expectations to remove the conditional expectation $\mathrm{A}_0$ from all but one term.
Proposition (ref) gives two least squares criterion functions. We use (ref) and (ref) to construct a Lasso and a Neural Network estimator, respectively. We describe the Lasso estimator in Section (ref) and Neural Network estimator in Appendix D.
To describe the Lasso estimator of the value function let $$b(x) = (b_1(x), \dots, b_p(x)) \in \mathrm{R}^p,$$ be a vector of basis functions. We approximate the value function using a linear form $$ V(x) \approx \sum_{j=1}^{p} b_j(x) \rho_{Vj} = b(x)^{\prime} \rho_V, $$ where $\rho_V = (\rho_{V1}, \rho_{V2}, \dots, \rho_{Vp}) \in \mathrm{R}^p$ is a $p$-vector of coefficients. The vector is chosen to minimize an approximate least squares criterion $$ \rho_V = \arg \min_{\rho \in \mathrm{R}^p} \rho^{\prime} G^V \rho - 2 M^V \rho $$ where $G^V$ is a symmetric $p \times p$ matrix
and $M^V$ is a linear term
The FOC reduces to $$ G^V \rho_V = M^V. $$ We choose the criterion (ref) as opposed to (ref) so that the sample version of matrix $G^V$ is symmetric and positive-definite.
Given an i.i.d sample $(X_i, X_{i+})_{i=1}^n$, we construct a sample estimate of $\rho_V$ in the regime where $\dim (\rho_V) = p_V \gg n$. For simplicity of exposition, we abstract away from subsequent estimation steps and drop respective cross-fitting indices. Given a plug-in estimate $\widehat \mathrm{A} b$ of $\mathrm{A}_0 b$ and $\widehat \zeta$ and $\zeta_0$ estimated on a hold-out sample, define
Given a radius $\rho_V$, an $\ell_1$-regularized estimator of the value function takes the form
Assumption (ref) requires that $V_0$ belongs to the mean square closure $\Gamma$ of linear combinations $b(x)^{\prime} \bar{\rho}$, as well as that the approximating coefficients $\bar{\rho}$ are sufficiently sparse.
Assumption (ref)(1) requires value function to be approximately sparse in the chosen basis. Assumptions (ref)(2)-(4) are standard regularity conditions. Assumption (ref)(5) reduces to a rate condition on the first-stage estimators.
Theorem (ref) gives a mean square convergence rate for the value function. The rate is determined by the sparsity parameter $\xi_V$ of the value function and the first-stage rate parameter $d_V \in (0, 1/2)$.
In this Section, we derive dynamic dual criterion function. Define
From the dual representation of (ref) we know that $\alpha_0$ satisfies
and $\mathrm{I} - \mathrm{A}^{*}_0$ is invertible.
We show that $\alpha_0$ can be represented as a minimizer of a criterion function that depends on $\mathrm{A}^{*}_0$. $\alpha_0$ will minimize the expected squared difference of the left and right-hand sides of equation (ref), that is
where the second equality follows by squaring and dropping the term that does not depend on $\alpha$, the third equality follows the Riesz representation in (ref), and the third equality by iterated expectations.
Given a vector of basis functions $$b(x) = (b_1(x), b_2(x), \dots, b_j(x), \dots, b_p(x)) \in \mathrm{R}^p,$$ where $p$ can differ from the one Section (ref), we approximate the dynamic dual representation via a linear form $$ \alpha(x) \approx \sum_{j=1}^{p} b_j(x) \rho_{\alpha j} = b(x)^{\prime} \rho_{\alpha} $$ where $\rho_{\alpha} = (\rho_{\alpha 1}, \rho_{\alpha 1}, \dots, \rho_{\alpha p}) \in \mathrm{R}^p$ is a $p$-vector. The vector is chosen to minimize an approximate least squares criterion $$ \rho_{\alpha} = \arg \min_{\rho \in \mathrm{R}^p} \rho^{\prime} G_{\alpha} \rho - 2 M_{\alpha} \rho $$ where the $p \times p$ matrix is an outer product of
and the free term is
We choose the criterion (ref) as opposed to (ref) so that the sample version of matrix $G_{\alpha}$ is symmetric and positive-definite. Given a radius $r_{\alpha}$, an $\ell_1$-regularized minimum distance estimator of the dynamic dual representation
where
Theorem (ref) gives a mean square convergence rate for dynamic dual representation. The rate is determined by the sparsity parameter $\xi_{\alpha}$ of the dynamic dual representation and the first-stage rate parameter $d_{\alpha}$. The estimator is automatic in the sense that it only requires knowledge of the linear functional $m(Z, (\mathrm{I} - \mathrm{A}^{*})b)$.
In this Section, we discuss related results. Remark (ref) discusses cross-fitting. Remark (ref) introduces a neural network estimator of the value function. Remark (ref) introduces a neural network estimator of the dynamic dual representation. Remark (ref) describes the automatic property of the dynamic dual criterion function. Remark (ref) verifies rate conditions for asymptotic theory.
In this paper we introduce welfare metrics -- including welfare decompositions into direct and indirect effects -- and give a complete set of estimation and inference results for them in the presence of high-dimensional state space. The results are presented for the dynamic binary choice model of Rust,HotzMiller,AMira2002 but are applicable to other dynamic models (e.g., dynamic games AMira2007,BBL). For the case of average welfare and related metrics, the proposed estimator is a known function of choice probabilities and the structural parameter. In particular, if the model has “terminal action” property, value function or any other dynamic object does not have to be estimated at all. We have applied these methods to estimate the average teachers welfare in an application to teachers absenteeism as in DHR.