EconBase
← Back to paper

Welfare Analysis in Dynamic Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,043 characters · 17 sections · 96 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Welfare Analysis in Dynamic Models

\setcounter{page}{1}

\makeatletter \def\thanks#1{\protected@xdef\@thanks{\@thanks \footnotetext{1}}} \makeatother

\makeatletter \def\thanks#1{\protected@xdef\@thanks{\@thanks \footnotetext{1}}} \makeatother

abstractThis paper introduces metrics for welfare analysis in dynamic models. We develop estimation and inference for these parameters even in the presence of a high-dimensional state space. Examples of welfare metrics include average welfare, average marginal welfare effects, and welfare decompositions into direct and indirect effects similar to oaxaca73 and blinder73. We derive dual and doubly robust representations of welfare metrics that facilitate debiased inference. For average welfare, the value function does not have to be estimated. In general, debiasing can be applied to any estimator of the value function, including neural nets, random forests, Lasso, boosting, and other high-dimensional methods. In particular, we derive Lasso and Neural Network estimators of the value function and associated dynamic dual representation and establish associated mean square convergence rates for these functions. Debiasing is automatic in the sense that it only requires knowledge of the welfare metric of interest, not the form of bias correction.  The proposed methods are applied to estimate a dynamic behavioral model of teacher absenteeism in DHR and associated average teacher welfare. Keywords: Oaxaca-Blinder decomposition, value function, dynamic discrete choice, dynamic dual representation, average derivative, stationarity, finite dependence, double robustness

Introduction

Dynamic considerations are important in applied work Miller1984, Wolpin, Pakes:1986, Rust, HotzMiller, HotzMillerSandersSmith, keane1994solution, AMira2002,AMira2007,BBL. These considerations are captured by the value function, defined as the present discounted value of agents' expected per-period utility. Prior work has focused on identifying utility parameters from the optimality of agents' behavior and conducting inference on them. Building on this foundation, this paper targets a different yet complementary class of parameters that depend on the value function which we introduce below.

Welfare analysis is a central objective in economics. This paper develops identification, estimation, and inference results for welfare metrics in dynamic models. Examples of such metrics include the average value function and average marginal welfare effects, e.g., due to the change in observable time-invariant characteristics such as initial wealth. Most importantly, we decompose welfare effects into the direct effect through changes in utilities holding agents' behavior fixed and the indirect effect from changes in agents' behavior. This decomposition is analogous in spirit to Oaxaca-Blinder decomposition of outcome distributions yet different since welfare metrics are based on latent utilities rather than observed outcomes.

The economic motivation for the welfare decomposition we propose is drawn from various fields in applied economics, including labor economics, public finance, and health economics. In labor economics, kitagawa55, oaxaca73 and blinder73 decompose the gender wage gap into structural and composition effects. In public finance, it is common to distinguish between mechanical and behavioral  effects of taxation policies, see e.g. Chetty2009. In health economics, the work by EinavFinkelsteinCullen2010  differentiates the direct effect of a health insurance price change from the indirect effect arising from changes in selection. For an example of cancer screening EinavEtAl2020, the total welfare effect corresponds to the introduction of a new screening option, and the direct effect is its counterpart as if the frequency of screening were mandated. One of the paper's contribution is to give identification, estimation, and inferential results for direct and indirect effects in dynamic models.

The next contribution of the paper is the dynamic dual representation of welfare metrics, which directly maps the per-period utility to the welfare metric of interest. For the case of average welfare, the dynamic dual representation reduces to a known expression of expected per-period utility. Furthermore, neither value function nor other dynamic object needs to be estimated for certain discrete choice models, such as those of   Rust and HotzMiller.  For cases beyond average welfare, the dynamic dual representation involves backward discounting. In these cases,  we derive a doubly robust representation that facilitates consistent estimation if at least one of the two nuisance components, the value function or the dynamic dual representation,  is correct. Appendix (ref) gives large sample properties of the proposed estimators based on Neyman-orthogonal moment equations, in which the dynamic dual representation plays a central role. This paper is the first result in the literature to leverage duality in dynamic models whose state distribution is strictly stationary.

Another contribution of the paper is to introduce a Lasso estimator for the value function, the dynamic dual representation, and associated mean square convergence rates in the setting of high-dimensional state variables. Specifically, the number of state variables may exceed the sample size, provided that only a small subset is relevant. We derive a novel least squares criterion that distinguishes our approach from previous IV-based methods. Adding an $\ell_1$ penalty to the sample analog of this criterion yields a Lasso estimator of the value function. In addition, we propose a neural network estimator and derive corresponding mean square convergence rates in the low-dimensional case. Although the results are presented for the value function, the methods apply to any fixed-point solution of an integral equation of the second kind -- such as the Q-function in reinforcement learning -- and do not require strict stationarity of the state distribution.

We revisit the study of teacher absenteeism in the nonformal education centers (NFEs) of Rajasthan, India, initially analyzed by DHR (henceforth, DHR). Their approach emphasizes unobserved heterogeneity in teachers' leisure preferences—shaped by factors such as past effort, illness, fatigue, and informal obligations. We conjecture that this heterogeneity can be effectively proxied by a sufficiently long window of prior work history, even if the history itself may not have a structural or causal interpretation. DHR's structural estimates remain robust to this exercise. In addition, we focus on teacher welfare and find that the failure to properly account for teacher heterogeneity results in the average teacher welfare overestimated by 13–20%%. As a side contribution, we demonstrate that our methodological framework remains valid in finite-horizon settings, provided the panel is sufficiently long relative to standard discount factors.

Literature Review

A large body of work is dedicated to welfare analysis in discrete choice models with unobserved heterogeneity, as studied, e.g., in Bhattacharya2015. As Bhattacharya2024 discusses, welfare calculations are based on latent utilities rather than observed outcomes. For dynamic choice -- the focus of this paper -- important early references include Miller1984, Wolpin, and Pakes:1986. Within this group, a notable subclass of models includes those with the terminal action property, as in HotzMiller and HotzMillerSandersSmith, and, more broadly, models with finite dependence, as in ArcidiaconoMillerFD. In HotzMiller, structural parameters are identified via a regression problem that depends on conditional choice probabilities. The dynamic dual representation we derive extends this insight to the average value function and other welfare metrics.

Estimation of dynamic discrete choice models has received substantial attention; see, e.g., AMira2002, AMira2007, BBL, PSM, ArcidiaconoMiller, Arcidiacono:2013, ArcidiaconoEtAl2013, ChenAck, AMagesan, BShum. Recent work by AdsmEck2022 relaxes the terminal action property while allowing for an infinitely supported state space. Appendix (ref) of the present paper outlines an automatic debiasing approach that targets nonlinear functionals of the value function and thus could be applicable to the model in AdsmEck2022.

Last but not least, we contribute to the literature on estimating fixed points of integral equations of the second kind; see, e.g., SLinton. Prior work by chen2022wellposedness has shown that this is a special case of a well-posed NPIV problem and has developed value function estimators that achieve minimax-optimal rates in low-dimensional settings; see also xia2022krylovbellmanboostingsuperlinearpolicy for related IV-based approaches. However, these methods may not readily extend to penalized estimators that remain consistent in high-dimensional settings. The least squares criterion we derive circumvents this limitation. We also give a least squares criterion for estimating the dynamic dual representation, that only depends on the welfare metric of interest, and so enables automatic debiasing like chernozhukov2021automatic and chernozhukov2024qm.

The paper is organized as follows. Section (ref) presents the general framework and examples of welfare metrics. Section (ref) decomposes the differences in average welfare in the spirit of Oaxaca and Blinder. Section (ref) revisits DHR study of teacher absenteeism. Section (ref) gives the dynamic dual and doubly robust representations for welfare metrics. Section (ref) describes orthogonal estimating equations for average welfare and related averages. Least squares estimators of the value function and dynamic dual representation are presented in Section (ref). Section (ref) concludes. Appendix (ref) considers the example of dynamic discrete choice. Appendix (ref) gives proofs of the results in main text. Appendix (ref) gives large sample properties of the estimators. Appendices (ref)--(ref) give mean square rates for first-stage estimators. Appendix (ref) extends the proposed method to nonlinear functionals of the value function. Appendix (ref) demonstrates the method of Appendix (ref) for dynamic binary choice.

Setup

We consider estimation and inference on welfare metrics that depend on the value function.  To define the value function, let $X_t, (t=0,1,...)$, denote a time series of observed state variables, which we assume to be a time-homogeneous, first-order Markov process with initial element $X:=X_0$. The value function is determined by a per-period reward $\zeta_0(X)$, or, in other words,  expected utility in a single period conditional on the state $X$, which we assume to be identifiable. The value function $V_0(X)$ is the present discounted value of per-period rewards given the current state, satisfying

align[align omitted — 99 chars of source]

where $\beta \in [0,1)$ is a known discount factor. The value function satisfies the integral equation:

align[align omitted — 85 chars of source]

where $X$ is the current and $X_{+}$ is the next period element of the first-order Markov process (see Lemma (ref)). In what follows, we assume that the time series is strictly stationary. The welfare metrics we consider are linear functionals of the value function $V_0(X)$ having the form

align[align omitted — 81 chars of source]

where $w_0(X)$ is some function of the state.

To give an example of per-period utility, we consider the dynamic discrete choice problem Rust,HotzMiller,AMira2002. In each period, $(t=0,1,...),$ the agent chooses an action $j$ from a finite choice set $\mathcal{A}$. The utility of choice $j$ in period $t$ is additively separable in a function of the current state $u(X_t,j)$ and private shock $\varepsilon_t(j)$ and is given by $$\bar{u} (X_t, j, \epsilon_t) = u(X_t,j) + \epsilon_t(j), \quad j \in \mathcal{A}.$$ Here the sequence ${(X_{t}, J_t)}$ is strictly stationary, so the time index $t$ can be dropped. The per-period reward is the expected utility $\zeta_0(x)$, obtained by taking the expectation over choices: $$ \zeta_0(x) =   \sum_{j \in \mathcal{A}} (u(x,j) + \mathrm{E} [ \epsilon (j) \mid X=x, J=j]) {\mathrm{P}} (J=j \mid X=x), $$ where ${\mathrm{P}} (J=j \mid X=x)$ is the probability that an agent chooses $j$ when $X=x$, and  $\mathrm{E} [ \epsilon (j) \mid X=x, J=j]$ is the expectation of $\epsilon (j)$ under the choice $J=j$ and given $X=x$. For example, in a special case where the choice is binary and the private shocks are distributed as Gumbel,

align[align omitted — 89 chars of source]

where $p_0(x) = {\mathrm{P}} (J=1 \mid X=x)$,   $\mathcal{A}=\{1, 0\}$, and $H(t) = \gamma_e - t\ln t -(1-t) \ln (1-t),$ with  $\gamma_e =  0.5227$ denoting the Euler constant. Here and generally for dynamic discrete choice the per-period utility will be the sum of the expected value of the observable part of the utility plus the expected value over optimal choices of the private shock part. We assume that the utility components $u(x,1)$ and $u(x,0)$ are known up to a structural parameter that is identified.

Within the discrete choice problems, our target parameter $\delta_0$ represents a welfare metric since $V_0(\cdot)$ is the expected value of an agent making optimal dynamic choices conditional on state $X$. Our first example is the expected value function where $w_0(X)=1$.

example[Average Welfare] If $w_0(X)=1$, the parameter \begin{align} \delta_0 = \mathrm{E} [V_0(X)] \end{align} represents the unconditional expected value or welfare of making optimal dynamic choices.

We focus particularly on settings where the state variable include observable time-invariant heterogeneity in individual, per-period expected utility denoted by a vector $K$. Examples of $K$ include occupation type KeaneWolpin1997, gender and class grade of a child ToddWolpin and teacher test score DHR. Time invariant state variables could also represent observable heterogeneity in individual, per-period expected utility.     In this case, the state vector $X_t$ can be decomposed as $X_t = (S_t, K),$  where $K_t= K$ does not vary over time. Here first-order time homogeneity of $X_t$ implies that $(S_t)_{t > 0 }$ is a first-order time-homogeneous Markov chain conditional on $K$.

example[Group Average Welfare] When $K$ is discrete, taking on a finite number of values, the group average welfare is \begin{align} \delta_0 = \mathrm{E}[1(K=k)V_0(X)]/{\mathrm{P}}(K=k) = \mathrm{E}[V_0(X)\mid K=k]. \end{align} In this case $w_0(X)=1(K=k)/{\mathrm{P}}(K=k)$.

Extending this example to differences in average welfare across groups is straightforward by differencing the parameter of interest in Example (ref) across different values of $K$. In this example and the others, the weight $w_0(X)$ is unknown and will need to be estimated. The identification and estimation of $w_0(X)$ will be accounted for in the results that follow.

When $K$ represents an endowment of some resource it may be of interest to consider the welfare effect of changing the distribution of that endowment.

example[Average Policy Effect] Let  $\pi(k)$ and $\pi^{*}(k)$ be the probability density (or mass) function of $K$ with respect to a base measure, corresponding to the actual data and a proposed policy shift. The average policy effect from this shift is \begin{align} \delta_0 &=\mathrm{E}[w_0(K)V_0(X)] , \quad w_0(K) = [\pi^{*} (K) - \pi (K)]/\pi(K).\nonumber \end{align} This object differs from the policy effect of Stock in being the average effect of a policy on dynamic welfare rather than the average effect on some outcome variable.

For continuously distributed $K$ an effect of interest could be the average effect of changing $K$ on the value function.

example[Average Marginal Effect] For continuously distributed $K$, the average derivative of the value function with respect to $K$, \begin{align} \delta_0 = \mathrm{E}[\partial_k V_0(X)]=\mathrm{E}[\partial_k V_0(S,K)], \end{align} measures the average change in welfare due to varying $K$.  Letting $f(K|S)$ denote the conditional PDF of $K$ given $S$, integration by parts gives equation (ref) with  $w_0(X)=- \partial_k \ln f(K|S)$ as long as  $f(K|S)$ is equal to zero at the boundary of the support of $K$ conditional on $S$. This parameter differs from the average derivative\footnote{The average marginal welfare effects herein are different from the average marginal effects of ACarro in panel data setup, where marginal effects are taken with respect to unobserved unit heterogeneity.} of Stoker1986 in being a dynamic welfare effect rather than an outcome effect.
exampleLet $\mathcal{X}$ be a set and $X$ be a state vector. Define \begin{align} \delta_0 = \mathrm{E}[1(X \in \mathcal{X}) V_0(X)]/{\mathrm{P}}(X \in  \mathcal{X}) = \mathrm{E}[V_0(X)\mid X \in \mathcal{X}]. \end{align} In this case the weighting function $w_0(X)=1(X \in \mathcal{X})/{\mathrm{P}}(X \in  \mathcal{X})$ is time varying.

All of the above examples of target parameters fall in the following general theoretical  framework. Let $Z$ denote a data vector that includes $X$ and $X_+$, and let $V$ denote a possible value function. Also let $m(Z,V)$ denote a function of $Z$ and the function $V(\cdot)$ (i.e. $m(Z,V)$ is a functional of $V$.) We consider parameters of the form

align[align omitted — 62 chars of source]

where $\mathrm{E}[m(Z, V)]$ is linear in $V$. We will impose throughout that the expectation $\mathrm{E}[m(Z, V)]$ is mean square continuous as a function of $V$, meaning that there is a constant $C>0$ such that for all $V(X)$ with $\mathrm{E}[V(X)^2]<\infty$,

align[align omitted — 68 chars of source]

By the Riesz representation theorem mean square continuity of $\mathrm{E}[m(Z,V)]$ is equivalent to existence of a function $w_0(X)$ with $\mathrm{E}[w_0(X)^2]<\infty$ such that

align[align omitted — 78 chars of source]

for all $V(\cdot)$ with $\mathrm{E}[V(X)^2]<\infty$. Here, we see that under mean square continuity, any parameter as in equation (ref) can be represented as a linear function of the value function. There are many other potentially interesting examples of such welfare metrics. In the next section we consider decomposing differences in average welfare into direct and indirect components.

Decomposition of Differences in Average Welfare

To motivate our analysis, we present a simple running example in the context of a randomized controlled trial, in which a one-shot treatment assigned at time $t = 0$ generates dynamic incentives. Suppose we aim to analyze welfare differences between the treated and control populations, denoted by $1$ and $0$, respectively. In each group, the time-varying state variable is denoted by $S$. The full state vector is $X = (S, K)$ where $K$ is the indicator of the treatment status. We adopt the notation in chernozhukovmelly.

Let $V_0^1(s) = V_0(s,1)$ and $V_0^0(s)=V_0(s,0)$ denote the treated and control value functions. Define the treated average welfare as

align[align omitted — 79 chars of source]

and the control average welfare as

align[align omitted — 80 chars of source]

Both quantities are special cases of the group average welfare defined in Example (ref). For $k \in \{1, 0\}$, we have

align[align omitted — 115 chars of source]

where $\pi^k(s)$ denotes the probability distribution function of $S$ conditional on $K = k$. The counterfactual welfare metric

align[align omitted — 81 chars of source]

does not correspond to the group average welfare of any observable subpopulation. Instead, it is constructed by integrating the treated value function with respect to the stationary distribution of the control population. Provided that $\pi^1(s)$ and $\pi^0(s)$ share the same support, this parameter is well-defined.

The treatment-control welfare difference can be decomposed in the spirit of kitagawa55, oaxaca73, and blinder73:

align[align omitted — 275 chars of source]

If treatment is randomly assigned, this welfare difference admits a causal interpretation.

The proposed decomposition has an intuitive interpretation when the treatment affects only per-period utilities and does not impact the state transition. We describe such empirical settings in Examples (ref) and (ref) below. In this case, the stationary distributions $\pi^1(s)$ and $\pi^0(s)$ differ solely due to agents making different optimal choices in the treated and control states, respectively. The first summand,

align[align omitted — 178 chars of source]

captures the difference in per-period utilities holding the distribution of states fixed at the control level. This term can be interpreted as the direct or mechanical effect. The second summand,

align[align omitted — 179 chars of source]

reflects changes in agents' behavior that alter the distribution of states, holding the utilities fixed. This term can be interpreted as the indirect or behavioral effect. The sum of the direct and behavioral effects gives the total welfare difference.

example[School attendance] DHR estimates a dynamic behavioral structural model in which teachers choose between working and taking leisure. DHR focuses the student achievements as the primary target. In this paper, we take a complementary perspective and focus on teacher welfare. Let $K = 1$ indicate the treatment status, where treatment constitutes a cash bonus for each additional day of work once the count of days worked exceeds 10 in a given month. The state variable $S$ denotes the number of days worked since the beginning of the month. The full state vector is $X = (S, K)$. The direct effect, $\delta_{\langle 1 \mid 1 \rangle} - \delta_{\langle 1 \mid 0 \rangle}$, captures the impact of the bonus on welfare through per-period utilities, holding teachers work decisions fixed. The indirect effect, $\delta_{\langle 1 \mid 0 \rangle} - \delta_{\langle 0 \mid 0 \rangle}$, arises from the change in teachers work decisions, holding per-period utilities fixed. A similar decomposition applies to the model in ToddWolpin.
example[Breast cancer screening] A standard approach to evaluating welfare in the context of cancer screening focuses on mortality. Yet, a broader perspective considers the costs of screening—such as time, resource use, and the psychological burden of false positives— which cannot be captured by observable outcomes. Let $K = 1$ indicate assignment to a novel breast cancer screening technology. The state variable $S$ denotes the time elapsed since the most recent screening, and the full state vector is $X = (S, K)$. The direct effect, $\delta_{\langle 1 \mid 1 \rangle} - \delta_{\langle 1 \mid 0 \rangle}$, captures the impact of the new technology on welfare through per-period utilities, holding screening behavior fixed. The indirect effect, $\delta_{\langle 1 \mid 0 \rangle} - \delta_{\langle 0 \mid 0 \rangle}$, reflects changes in screening behavior induced by the treatment, holding utilities fixed.

Proposition (ref) expresses the counterfactual welfare $\delta_{\langle 1 \mid 0 \rangle}$ as a linear functional of the treated value function.

proposition[Decomposition of Differences in Average Welfare] Suppose both density functions $\pi^0(s)$ and $\pi^1(s)$ have the same support of the state variable $S$. Then the counterfactual welfare $ \delta_{\langle 1 \mid 0 \rangle}$ is a special case of equation (ref) with \begin{align} m(Z, V) &= \dfrac{V(S, 1) 1\{ K = 0\}}{{\mathrm{P}} (K=0)} \end{align} whose Riesz representation (ref) holds with \begin{align} w_0(X) = w_0(S,K) =\dfrac{ 1\{ K=1\} } {{\mathrm{P}} (K=0)} \dfrac{{\mathrm{P}} (K=0 \mid S)}{{\mathrm{P}} (K=1 \mid S)}. \end{align}

Recent work has extended classical decomposition methods to modern settings. chernozhukov2021automatic derive the Riesz representer for the Average Treatment Effect on the Treated (ATET). vafa2024estimatingwagedisparitiesusing provides an Oaxaca-Blinder decomposition of wage differences. Proposition (ref) departs from these approaches by offering a decomposition of average welfare in dynamic models where welfare is based on latent utilities rather than observed outcomes.

We include the average counterfactual welfare $\delta_{\langle 1 \mid 0 \rangle}$ as Example (ref) and discuss its estimation in Section (ref). Notice that the parameters $ \delta_{\langle 1 \mid 1 \rangle}$ and $ \delta_{\langle 0 \mid 0 \rangle}$ are special cases of Example (ref) with $k=1$ and $k=0$, respectively. Therefore, it is straightforward to extend this example to accommodate direct and indirect effects. We discuss the estimation of the counterfactual welfare measure and related effects further in Section (ref).

remark[Overview of Related Literature] kitagawa55,oaxaca73,blinder73 pioneered the use of least squares methods for decomposing differences in average outcomes, such as wages. This approach was later extended to the distributional setting in dfl96,mm05,flf11,chernozhukovmelly, with chernozhukovmelly also providing inference methods. Other related contributions include OaxacaRansom,KlineOaxaca2011,chernozhukovmelly,KlineOaxaca,GuoBasse2021,vafa2024estimatingwagedisparitiesusing. In the canonical wage example, groups $1$ and $0$ correspond to men and women, respectively. The direct effect captures the difference in wage schedules faced by men and women and is often interpreted as a measure of discrimination or preferential treatment. The indirect effect reflects differences in job-related characteristics, such as skills or experience, and is typically interpreted as a composition or selection effect.

DHR revisited

We study how daily financial incentives affect teacher attendance in single-teacher nonformal education centers (NFEs) operated by the NGO Seva Mandir in tribal villages of Udaipur, Rajasthan, India. From 2003 to 2005, DHR conducted a randomized trial in which tamper-proof cameras recorded photographs at school opening and closing. A school day was deemed valid if the two images were at least five hours apart and at least eight students were present. At the end of each month, teachers earned a base salary of 500 Rupees (Rs) if they worked fewer than 10 days, plus a 50 Rs bonus for each additional day of work beyond that threshold. The 10-day cutoff thus created a nonlinear dynamic incentive, which we focus on in this application.\footnote{We abstract from the firing threat, which appears negligible: no teacher was fired during the study period, even in cases of near-total absence. According to DHR, Seva Mandir adopts a long-term view in assessing teacher performance, which may explain the lack of dismissals.} Our dataset comprises daily attendance records for 57 teachers over an 18-month period (January 1, 2004 to June 30, 2005), along with a test score administered prior to the start of teaching.\footnote{Following DHR, the estimation sample includes only weekdays when teachers actively choose between working and taking leisure. Holidays and weekends, though counted toward pay, are excluded from analysis.}

We revisit the dynamic behavioral model of DHR, henceforth DHR. Let $t$ denote the day of the month, ranging from $1$ to $T=30$. On each day $t$, a teacher chooses between working ($j_t = 1$) and taking leisure ($j_t = 0$). On the final day of the month, the consumption utility is determined by the monthly paycheck

align[align omitted — 74 chars of source]

where $500$ is the base salary, $d_T$ is the total number of days worked by day the final day $T$, and $50$ is the bonus.   For all days $t < T$, there is no consumption; utility accumulates only through leisure and is modeled as

align[align omitted — 65 chars of source]

Here, $ \bar{x}_t \in \mathrm{R}^{p_X}$ is the state vector including a constant and possibly other observable characteristics, and $\mu_0 \in \mathrm{R}^{p_X}$ is a parameter to be estimated. The state vector is $X_t = (\bar{X}_t, d_t)$.   For example, in Model I of DHR, the leisure utility is assumed to be the same for all teachers, which corresponds to $u(x_t, 0) = \mu_0$ and $X_t =d_t$.  The per-period utilities are not discounted. DHR includes only a handful of observables into $\bar{x}_t$ so as to leverage standard maximum likelihood estimators.

We consider a stylized dynamic binary choice model as described in Section (ref). Since teachers cannot be fired, the base salary is assumed not to affect their choice between work and leisure.  We decompose the total monthly bonus into daily payments.  Specifically, we assume the utility of working ($j_t = 1$) is given by:

align[align omitted — 83 chars of source]

where the indicator function, referred to as “In the money” by DHR, captures the bonus structure in the stylized model.   The parameter $\mu_1$ converts monetary rewards (in Rupees) into utility units. The stylized model in equations (ref)--(ref) preserves the monetary incentives of the exact model (ref)--(ref), aside from discounting of the bonus. To abstract from the finite-horizon considerations, we restrict attention to calendar days 15, 16, and 17 of each month. If this model is relevant, the methods developed in this paper permit the state vector $\bar{X}_t$ to include high-dimensional covariates.

Table (ref) compares the exact estimates reported by DHR (Columns (1)-(2)) to our stylized infinite-horizon replications (Columns (3)-(4)) for selected coefficients. The results are encouraging. First, the estimated bonus coefficient under the stylized model falls within the range of the exact estimates. The standard errors of the replicated coefficients are, on average, 2.5 to 3 times larger, as expected given the smaller sample (only three days per month). Second, the coefficient on teacher test scores remains negative across all specifications, consistent with the original findings. Other coefficients (not reported here) also closely match their exact counterparts. These results suggest that the stylized infinite-horizon framework is appropriate for this dataset.

We further investigate the role of prior work history in explaining teachers work decisions, a possibility raised by DHR. DHR included the first lag of work history in Model VIII, and we extend this idea by incorporating 89 more lags. Additionally, we summarize work history using a work streak variable, defined as the number of consecutive days a teacher has worked without taking leisure:

align[align omitted — 83 chars of source]

We investigate the role of prior work history both in the predictive and structural settings.

Table (ref) shows the out-of-sample mean squared error (MSEs) for predicting a teacher’s decision to work, using models that sequentially expand the covariate set from Model I to Model IV. Adding the teacher's test score on top of days worked yields minimal improvement. Including month dummies (Model II) results in a modest reduction in MSE. Incorporating the work streak (Model III) leads to a substantial improvement: relative to Model II, the MSE falls by roughly 32% for Logit, 32% for Probit, and 35% for Random Forest. Finally, adding the full 90-day work history dummies (Model IV) further improves prediction, halving the MSE relative to the baseline. These findings show that prior work history has high predictive power even after other observables have been taken into account.

As a next exercise, we revisit the stylized structural model with the aim of flexibly modeling the utility of leisure. As noted by DHR, decisions to skip work may be influenced by factors such as social norms, informal requests and commitments, accumulated effort, or fatigue. These factors are unlikely to be fully captured by basic observables like test scores. A longer spell of prior work history may be more granular and thus may better represent teacher's heterogeneity even if the history itself has no structural or causal interpretation. Assuming only a small number of lags suffices to capture the history, we include a 90-day window of lagged attendance indicators. This sparsity assumption calls for the use of Lasso estimators of the value function and the dynamic dual representation, developed in Section (ref).

table[table omitted — 1,161 chars of source]
table[table omitted — 1,107 chars of source]

Table (ref) reports selected coefficients for the structural parameter estimates (Panel A) as well as welfare metrics (Panels B, C and D). Columns (1)-(2) correspond to a simple model of (ref) whose only observable is teacher's test score. Columns (3)-(4) correspond to a more sophisticated model of (ref) where observables include prior work history. In both cases, the model is estimated using Algorithm (ref) described in Appendix G based on Logit and Random Forest estimators of conditional choice probability. The welfare metrics are estimated using the dual estimator described in Section (ref). Instead of using the value function, it combines the structural estimates of Panel A with the choice probabilities.

Our findings are as follows. First and foremost, the structural parameters—particularly the bonus coefficient ($\mu_1$) and the effect of teacher test score —are robust to the inclusion of the work history as a state component as well as to the choice of CCP estimator. In contrast, the welfare metrics are more sensitive to model specification. In particular, failure to account for the prior work history results in overestimating teacher welfare by 13-20$\%\%$. This overstatement persists across both the full sample and the subgroups defined by test scores values. Finally, the average welfare is lower for teachers with test scores at or below 30 than for those with scores above 40, as shown in Panels C and D, which is consistent with we include a 90-day window of lagged attendance indicators. Similar to DHR's interpretation, more skilled teachers are more committed to work and receive higher utility from teaching.

table[table omitted — 2,026 chars of source]

Dynamic Dual and Doubly Robust Representations of Welfare Metrics

Dynamic Dual Representation

In this Section, we give a dual representation of the parameter of interest.   This representation is important for several purposes. When $w_0(X)$ depends only on $K$, and so is time-invariant, the dual representation gives a simplified formula for $\delta_0$ that does not require solving any dynamic problem. Otherwise,  the dual representation leads to a doubly robust moment condition for identification and estimation of the parameter of interest. The dual representation\footnote{See equations (16)-(17) in the first version of the paper CNS. } was derived in the previous version of this paper \citet*{CNS}.

A key part of the dual representation is a function of the state variable that is a backward discounted value of $w_0(X)$, given by

align[align omitted — 103 chars of source]

where $X_{-t}$ is the state variable in period $-t$ in the extended stochastic process ${X_t}$ where $t$ ranges over all the integers. Alternatively, $\alpha_0(X)$ is a fixed point of the backward  dynamic operator

align[align omitted — 103 chars of source]

The following result gives the dynamic dual representation of weighted average welfare.

proposition[Dynamic Dual Representation] Let $V_{\zeta}$ be a net present discounted value of per-period utility $\zeta(x)$ as in (ref) with $\zeta(\cdot)$ replacing $\zeta_0(\cdot)$. The function $\alpha_0(X)$ in equation (ref) is the unique function such that \begin{align} \mathrm{E} [w_0(X) V_{\zeta}(X)] = \mathrm{E}[ \alpha_0(X) \zeta(X)] \quad \end{align}  for any $\zeta(X)$ with finite second moment.

An interesting implication of this dual representation is that if $w_0(X)$ is time-invariant, then $\delta_0$ depends only on the per-period expected utility $\zeta_0(X)$.

corollary[Dynamic Dual Representation With Time-Invariant Weight] If $m(Z,V)=w_0(X)V(X)$ and $w_0(X)=w_0(K)$ depends only on a time-invariant variable $K$ then $\alpha_0(X)=(1-\beta)^{-1}w_0(K)$ and \begin{align} \delta_0 = \mathrm{E} [w_0(X)V_{\zeta}(X)] = (1-\beta)^{-1}\mathrm{E} [w_0(K)\zeta(X)]. \end{align}

In each of our first three examples, the weight was time-invariant so that Corollary (ref) applies and the parameter of interest $\delta_0$ depends only on $\zeta_0(X)$. Here are expressions for $\delta_0$ for Examples (ref)-(ref).

example[continues=ex:average] The average welfare is given by \begin{align} \delta_0=\mathrm{E}[V_0(X)]=(1-\beta)^{-1}\mathrm{E}[\zeta_0(X)] \end{align}
example[continues=ex:group] The group average welfare is given by \begin{align} \delta_0 = (1-\beta)^{-1} \mathrm{E}[1(K=k)\zeta_0(X)]/{\mathrm{P}}(K=k). \end{align}
example[continues=ex:stock] The policy effect is given by \begin{align} \delta_0 = (1-\beta)^{-1}\mathrm{E}[w_0(K)\zeta_0(X)], w_0(K) = [\pi^{*} (K) - \pi (K)]/\pi(K). \end{align}
remark[Implications for Rust and HotzMiller] Consider a dynamic binary choice model where one of the actions has a terminal property similar to Rust. Furthermore, suppose the deterministic utilities take a linear index form \begin{align} u(x,1)=D_1(x)^{\prime} \theta_{11}, \quad u(x,0) =D_0(x)^{\prime} \theta_{10}, \end{align} where $\theta_0= (\theta_{11}, \theta_{10})$. This parameter can be identified via a semiparametric moment condition whose only nuisance parameter is conditional choice probability HotzMiller,LRSP. Stacking this moment condition with (ref) in Example (ref) gives a semiparametric moment condition for $(\theta_0, \delta_0)$ where neither value function nor any other fixed point of Bellman equation needs to be estimated.
remark[Implications for dynamic discrete choice models as in AMira2002] In a broad class of dynamic discrete choice models, including the one in DHR, neither action has a terminal choice property. In that case, the structural parameter can be identified by a PMLE moment condition described in AMira2002 whose debiased analog is proposed in AdsmEck2022. When the state space is continuously supported, the nuisance components includes choice-specific value functions that are fixed points of a nonparametric IV problem. Appendix (ref) describes a related yet different moment condition that we utilize in Section (ref).

When the weight is not time-invariant, $\alpha_0$ will not generally have a closed form or explicit expression because it depends in a complicated way on the dynamic distribution of the state vector $X_t$. To help understand better the nature of $\alpha_0$ we revisit Example (ref) where the state variable follows an autoregressive process of order 1 with a Gaussian innovation. While this example may not correspond to a state distribution under a dynamic discrete choice model, we include it for pedagogic purposes to help explain the nature of $\alpha_0$.

proposition[Dynamic Dual Representation for Average Marginal Effects] Consider an AR(1) model with a Gaussian innovation \begin{align} S_{t+1} =  \rho(K) S_t + U_t, \quad U_t \sim IID N(0, 1), \end{align} where $\rho(K) \in (-1,1) \text{ a.s.}$  is an autoregressive coefficient that may depend on $K$. Then the weighting function in Example (ref) is linear in  $S^2$ \begin{align} w_0(X) = \gamma_1(K) S^2 + \gamma_0(K) \end{align} whose intercept $\gamma_0(K)$ and the slope $\gamma_1(K)$ are functions of the time-invariant type $K$ given in (ref).   The dynamic dual representation is \begin{align} \alpha_0(X) =  \gamma_1(K) \dfrac{S^2}{1 - \beta \rho^2(K)}  +  (1-\beta)^{-1}  \gamma_1(K) \dfrac{\beta }{(1 - \beta \rho^2(K))} +   (1-\beta)^{-1} \gamma_0(K). \end{align}
corollaryConsider a white noise model with i.i.d states $S_t$, which is a special case of  (ref) with  $\rho(K)=0$. Then the dynamic dual representation  is time-invariant \begin{align*} w_0(X) = w_0(K) = \gamma_0(K) =  - \partial_K \ln f_K(K), \qquad \alpha_0(X) = (1-\beta)^{-1} \gamma_0(K). \end{align*}

Doubly Robust Representation

In this section, we give an identifying moment condition for the parameter of interest that is doubly robust in the sense that it holds if just one of $V(\cdot)$ or $\alpha(\cdot)$ is the true function. This moment condition uses the identifying conditional moment restriction for $V_0$ in equation (ref). Let $Z$ denote a data observation which includes $(X,X_{+})$, $V$ denote a possible value function, and $\lambda(Z, V):= \beta V(X_{+}) - V(X) + \zeta_0(X)$. Equation (ref) is equivalent to the conditional moment restriction

align[align omitted — 66 chars of source]

This is a nonparametric conditional moment restriction like those of NeweyPowell and AiChen2003 where $X_+$ is an "endogenous" variable, $X$ is an "instrument", and  $\lambda(Z,V)$ is a nonparametric residual as considered in CNS and chen2022wellposedness. Here we take $\zeta_0(X)$ to be a known function and will consider estimation of $\zeta_0(X)$ in the next Section.

Let $\alpha$ denote a possible function $\alpha_0$. A doubly robust moment function can then be formed as

align[align omitted — 91 chars of source]

Given a function $\xi$ of $X$ define

align[align omitted — 55 chars of source]
lemma[Double Robustness of Moment Function  (ref)] The moment function satisfies \begin{align} \mathrm{E} [g(Z,V,\alpha, \delta_0)] = \mathrm{E}[(\alpha(X)-\alpha_0(X))(\lambda(Z,V)-\lambda(Z,V_0))]. \end{align} Also \begin{align} | \mathrm{E} [g(Z, V,\alpha, \delta_0)] | \leq (1+\beta) \| \alpha - \alpha_0 \| \| V - V_0 \|. \end{align}

Lemma (ref) establishes double robustness of the moment function $g(Z,V,\alpha,\delta)$ which has zero expectation at $\delta=\delta_0$ if either $V=V_0$ or $\alpha=\alpha_0$ by equation  ((ref)). A doubly robust estimator of the average treatment effect was given in (Robins) and (LRSP) characterize doubly robust moment functions as being linear in both non-parameric components.

example[continues=ex:avder] The doubly robust representation for the average derivative is \begin{align} \mathrm{E} [g(Z, V, \alpha,\delta_0)] &= \mathrm{E} [ \partial_K V(X) + \alpha(X)  (\beta V(X_{+}) - V(X) + \zeta_0(X))  - \delta_0 ] \end{align} where the true value of $\alpha$ is the backward discounted value (ref) based on $w_0(X) = - \partial_k \ln f(K| S)$.

Overview of Estimation and Inference

In this Section, we give estimators of the welfare metrics we introduced in Sections (ref) and (ref). These estimators will account for the estimation of $\zeta_0(\cdot)$ and of $w_0(\cdot)$ or $m(Z,V)$ by including influence functions for their effect on identifying moments. The inclusion of these influence functions debiases for model selection and/or regularization in the estimation of unknown functions and corrects resulting standard errors for their estimation, as in LRSP.

For simplicity of exposition, we focus on panels with $T=2$ time periods where the pairs of consequent states $(X_{i1},X_{i2})_{i=1}^n$ are i.i.d. We use standard cross-fitting for i.i.d data (schick1986asymptotically) as common in work on debiased machine learning, chernozhukov2016double. For a weakly dependent time series with $T\geq 3$ periods, cross-fitting along both unit and time dimension is possible by leaving out neighboring folds, as discussed in CGST. Related work on conditional moment restrictions with weak dependence includes ChenSieveRiesz,ChenLiao2015,ChenLiaoWang2024.

Average Welfare and Related Averages

We will first consider a weighted average value function parameter with time-invariant weight that is possibly estimated. This case includes average welfare and related averages in Examples (ref)--(ref). Let $F$ denote an unrestricted distribution for $Z$ and $w(K,F)$ and $\zeta(X,F)$ denote the probability limit (plim) of an estimated weight $\widehat{w}(K)$ and an estimator $\widehat\zeta(X)$ respectively. Let $\phi_{w}(Z)$ and  $\phi_{\zeta}(Z)$ be the influence functions of $(1-\beta)^{-1}\mathrm{E}[ w(K,F)\zeta_0(X)]$ and $(1-\beta)^{-1}E[w_0(K)\zeta(X,F)]$ respectively. To nonparametrically debias for the estimation of $w(K)$ and $\zeta(X)$ and so construct a Neyman orthogonal moment function we add $\phi_{w}(Z)$ and $\phi_{\zeta}(Z)$ to the identifying moment function, as in LRSP, to obtain

align[align omitted — 185 chars of source]

where the true parameter $\delta_0$ solves

align[align omitted — 65 chars of source]

at the true value $\gamma_0$ of $\gamma$ and $\phi_0$ of $\phi$.

For cross-fitting purposes, we partition the set of data indices ${1, \ldots, n}$ into $L$ disjoint subsets $I_\ell$ of about equal size, $\ell = 1, \ldots, L$. Let $\widehat\gamma_\ell=(\widehat w_{\ell}, \widehat \zeta_{\ell})$ and $\widehat\phi_\ell =(\widehat\phi_{w\ell},\widehat\phi_{\zeta\ell})$ be estimators of the weight, per-period utility, and influence functions constructed using all observations not in  $I_\ell$. Also let $\psi(Z,\widehat\gamma_\ell,\widehat\phi_\ell,\delta)$ be as in equation (ref) with $\widehat\gamma_\ell$ and $\widehat\phi_\ell$ in place of $\gamma$ and $\phi$. A cross-fit estimator of $\delta_0$ can be obtained from solving $\sum_{\ell = 1}^L \sum_{i \in I_\ell} \psi(Z_i,\widehat\gamma_\ell,\widehat\phi_\ell,\delta)/n =0$ for $\delta$ giving

align[align omitted — 380 chars of source]

An example of estimated per-period utility $\zeta$ and its correction term for dynamic binary choice is given in Appendix (ref). The Lemma (ref) in Appendix (ref) provides sufficient conditions for the validity of asymptotic inference.

example[continues=ex:average] The estimate of the average welfare is \begin{align*} \widehat \delta &= \frac{1}{n}\sum_{\ell = 1}^L \sum_{i \in I_\ell} [ (1-\beta)^{-1} \widehat{\zeta}_{\ell}(X_i)+ \widehat \phi_{\zeta\ell} (Z_i)] \end{align*}

We next continue with the description of group average welfare in Example (ref). A key difference from Example (ref) is that the time-invariant weighting function $w$ depends on the  group probability ${\mathrm{P}} (K=k)$ which needs to be estimated.

example[continues=ex:group] Let $\phi_{\zeta}$ be the influence function of $\mathrm{E}[ (1-\beta)^{-1} 1\{ K=k\} \zeta(X,F)/{\mathrm{P}} (K=k)]$.  The weighting function $w(K) = 1\{K=k\}/{\mathrm{P}} (K=k)$. The influence function for ${\mathrm{P}} (K=k)$ is \begin{align*} \phi_w(Z) =  -\dfrac{ \delta_0}{{\mathrm{P}}(K=k)}(1\{K=k\} - {\mathrm{P}} (K=k)). \end{align*} Thus, the estimator of $\delta_0$ reduces to \begin{align*} \widehat \delta &=  \frac{1}{n}\sum_{\ell = 1}^L \sum_{i \in I_\ell} [[ (1-\beta)^{-1}  1\{ K_i = k\}/ \widehat {\mathrm{P}} (K=k) ]\widehat{\zeta}_{\ell}(X_i)+\widehat{\phi}_{\zeta\ell}(Z_i)+\widehat{\phi}_{w\ell}(Z_i)], \end{align*} whose standard error in (ref) accounts for estimation of ${\mathrm{P}} (K=k)$ by including the term $\phi_w(\cdot)$.

General case.

In this Section we descrite the estimator of $\delta_0=\mathrm{E}[m(Z,V_0)]$ when $w_0(\cdot)$ varies with time. Let $m(Z,V,F)$ denote the plim of the estimated $m(Z,V)$ function, $\phi_m(Z)$ the influence function of $\mathrm{E}[m(Z,V_0,F)]$, and $\phi_{\zeta}(Z)$ the influence function of $\mathrm{E}[\alpha_0(X)\zeta(X,F)]$. The orthogonal moment function is

align[align omitted — 220 chars of source]

where $\phi_m$ corrects for the estimation of $m(\cdot)$ and $\phi_{\zeta}$ corrects for the estimation of $\zeta$. Algorithm (ref) below gives the proposed estimator of the parameter of interest. Lemma (ref) in Appendix (ref) establishes the validity of asymptotic inference.

example[continues=ex:avder] The estimate of the average derivative is \begin{align} \widehat \delta = \frac{1}{n}\sum_{\ell = 1}^L \sum_{i \in I_\ell} [\partial_k \widehat V_{\ell} (X_i) + \widehat \alpha_{\ell}(X_i)(\beta\widehat V_{\ell} (X_{+i})-\widehat V_{\ell} (X_i)+\widehat \zeta_{\ell}(X_i)) +\widehat{\phi}_{\zeta\ell}(Z_i)]. \end{align}

Here there is no correction $\phi_m$ since the functional $m(Z,V)=\partial_K V(X)$ of $V$ does not involve any unknown components.

example[Counterfactual Welfare] The counterfactual welfare $ \delta_{\langle 1 \mid 0 \rangle}$ in equation (ref) is a special case of (ref) with $m(z,V)$ in (ref) and $w_0$ in (ref). Let $\phi_{\zeta}(Z)$ be the first step influence function (FSIF) of $\mathrm{E}[ \alpha_0(X) \zeta(X,F) ]$ where $\alpha_0$ is given in (ref) based on the weighting function $w_0$ in (ref). The influence function for ${\mathrm{P}} (K=0)$ is $$ \phi_w(Z) = - \dfrac{\delta_{\langle 1 \mid 0 \rangle}}{{\mathrm{P}} (K=0)} \left (1\{K=0\} -{\mathrm{P}} (K=0)\right). $$ Thus the estimator of $\delta_{ 1 \mid 0}$ reduces to $$ \widehat \delta_{\langle 1 \mid 0 \rangle}=\frac{1}{n}\sum_{\ell = 1}^L \sum_{i \in I_\ell} \dfrac{1\{ K_i = 0 \}}{\widehat {\mathrm{P}} (K=0)} \widehat V_{\ell}(S_i, 1) + \widehat \alpha_{\ell}(X_i) \lambda(Z_i, \widehat V_{\ell}) +\widehat{\phi}_{\zeta\ell}(Z_i)+\widehat{\phi}_{w\ell}(Z_i), $$ whose standard error in (ref) accounts for estimation of ${\mathrm{P}} (K=0)$ by including the term $\phi_w(\cdot)$.

Algorithm (ref) summarizes the estimation steps of welfare metrics.

algorithm[algorithm omitted — 1,955 chars of source]

Estimation of Value Function and Dynamic Dual Representation

This section introduces novel least squares estimators for both the value function and the dynamic dual representation. In contrast to the welfare metrics, defined in Section (ref), these objects do not require strict stationarity to be well-defined. Thus, the estimators of the value function and the dynamic dual representation delivered here do not require the time series to be strictly stationary. Consequently, these estimators apply to any fixed point of a second-kind integral operator.

Sections (ref) and (ref) develop a new least squares criterion for the value function that accommodates high-dimensional covariates through penalization, enabling consistent estimation in a high-dimensional state space. Section (ref) proposes a distinct least squares criterion for the dynamic dual representation, which depends only on the welfare metric of interest. This innovation permits automatic debiasing in the style of chernozhukov2021automatic and chernozhukov2024qm. Section (ref) discusses the results.

Least Squares Criterion for Value Function

The starting point of our analysis is the expectation operator $\mathrm{A}_0$ defined as

align[align omitted — 104 chars of source]

Rewriting (ref) in terms of $\mathrm{A}_0$ gives

align[align omitted — 71 chars of source]

or, equivalently, $V_0 = (\mathrm{I} - \mathrm{A}_0)^{-1} \zeta_0$. The operator $\mathrm{A}_0$ is akin to the integral equation operator in SLinton.

The value function can be represented as a minimizer of a criterion function that depends on $\mathrm{A}_0$. $V_0$ will minimize the expected squared difference of the left and right-hand sides of equation (ref), that is

align[align omitted — 485 chars of source]

where the second equality follows by squaring and dropping the term that does not depend on $V$ and the third and fourth equalities by iterated expectations. The expression minimized following the first equality is the nonparametric two-stage least squares criterion for NeweyPowell, Newey1991, and AiChen2003. The expression following the third equality is a hybrid that uses iterated expectations to remove the conditional expectation $\mathrm{A}_0$ from all but one term.

proposition[Least Squares Criterion for Value Function] The value function $V_0$ in (ref) is the unique minimizer of least squares criterion functions \begin{align} V_0&= \arg \min_{V} \mathrm{E} [((\mathrm{I} - \mathrm{A}_0) V)(X)^2- 2(V(X)-\beta V(X_+))(X)\zeta_0(X)] \\ &= \arg \min_{V} \mathrm{E} [(V(X)-\beta V(X_+))((\mathrm{I} - \mathrm{A}_0) V)(X)- 2\zeta_0(X))]. \end{align}

Proposition (ref) gives two least squares criterion functions. We use (ref) and (ref) to construct a Lasso and a Neural Network estimator, respectively. We describe the Lasso estimator in Section (ref) and Neural Network estimator in Appendix D.

To describe the Lasso estimator of the value function let $$b(x) = (b_1(x), \dots, b_p(x)) \in \mathrm{R}^p,$$ be a vector of basis functions. We approximate the value function using a linear form $$ V(x) \approx \sum_{j=1}^{p} b_j(x) \rho_{Vj} = b(x)^{\prime} \rho_V, $$ where $\rho_V = (\rho_{V1}, \rho_{V2}, \dots, \rho_{Vp}) \in \mathrm{R}^p$ is a $p$-vector of coefficients. The vector is chosen to minimize an approximate least squares criterion $$ \rho_V = \arg \min_{\rho \in \mathrm{R}^p} \rho^{\prime} G^V \rho - 2 M^V \rho $$ where $G^V$ is a symmetric $p \times p$ matrix

align[align omitted — 131 chars of source]

and $M^V$ is a linear term

align*[align* omitted — 69 chars of source]

The FOC reduces to $$ G^V \rho_V = M^V. $$ We choose the criterion (ref) as opposed to (ref) so that the sample version of matrix $G^V$ is symmetric and positive-definite.

Given an i.i.d sample $(X_i, X_{i+})_{i=1}^n$, we construct a sample estimate of $\rho_V$ in the regime where $\dim (\rho_V) = p_V \gg n$. For simplicity of exposition, we abstract away from subsequent estimation steps and drop respective cross-fitting indices. Given a plug-in estimate $\widehat \mathrm{A} b$ of $\mathrm{A}_0 b$ and $\widehat \zeta$ and $\zeta_0$ estimated on a hold-out sample, define

align*[align* omitted — 238 chars of source]

Given a radius $\rho_V$, an $\ell_1$-regularized estimator of the value function takes the form

align[align omitted — 195 chars of source]

Mean Square Convergence for Lasso Estimator

Assumption (ref) requires that $V_0$ belongs to the mean square closure $\Gamma$ of linear combinations $b(x)^{\prime} \bar{\rho}$, as well as that the approximating coefficients $\bar{\rho}$ are sufficiently sparse.

assumption(1) There exist constants $C>1$ and $\xi_V>1/2$ and $d_V \in (0, 1/2)$ such that for each positive integer $s_V\leq Cn^{-2(d_V)/(2\xi_V+1)}$ there is $ \bar{\rho}$ with $s_V$ nonzero elements such that \begin{equation*} \mathrm{E}[ (V_0(X)- b^{\prime}(X) \bar{\rho} )^{2}]\leq Cs_{V}^{-2\xi_V}. \end{equation*} (2) The matrix $\mathrm{E} [ b(X) b(X)^{\prime}]$ is positive definite with eigenvalues bounded from above by $\bar{\lambda}$ and below by $\underline{\lambda}>0$. (3) $\sup_j |b_j(X)|$ are bounded a.s. (4) The radius $r_V$ is chosen such that $\varepsilon_{n}=o(r_V),$ $r_V=o(n^{c}\varepsilon_{n})$ for all $c>0$, and there exists $C>0$ such that $p_V\leq Cn^{C}.$ (5) The first-stage estimators $\widehat \zeta$ of $\zeta_0$ and $\widehat \mathrm{A} b$ of $\mathrm{A}_0 b$ converge as $\| \widehat \zeta - \zeta_0 \| = o_P (\zeta_n)$ and $\| \widehat \mathrm{A} - \mathrm{A}_0 \| = o_P(a_n)$ where $\zeta_n + a_n = o (n^{-d_{V}})$ for some positive constant $d_{V} \in (0, 1/2)$.

Assumption (ref)(1) requires value function to be approximately sparse in the chosen basis. Assumptions (ref)(2)-(4) are standard regularity conditions. Assumption (ref)(5) reduces to a rate condition on the first-stage estimators.

theorem[Mean Square Rate for Lasso estimator of Value Function] If Assumption (ref) holds, then, for any $c>0$, \begin{equation} \Vert \widehat V- V_0\Vert =o_{p}(n^{c} n^{-2 d_{V} \xi_V/(2\xi_V+1)}). \end{equation}

Theorem (ref) gives a mean square convergence rate for the value function. The rate is determined by the sparsity parameter $\xi_V$ of the value function and the first-stage rate parameter $d_V \in (0, 1/2)$.

remark[Verification of Rate Condition (5) in Assumption (ref)] To give an example of operator $\mathrm{A}_0$, we consider the dynamic discrete choice problem discussed in Section (ref) where the state transition is deterministic conditional on action. Then the estimator of $\mathrm{A}_0$ reduces to an estimator of choice probabilities, and the rate $a_n = O (p_n)$ where $p_n$ is the $\ell_{\infty}$-rate for CCPs.
remark[Lasso estimator of $\mathrm{A}_0$] The following example of $\mathrm{A}_0$ is based on a Lasso estimator. For $j=1,2, \dots, p$, define the estimators as \begin{align} \widehat G &= n^{-1} \sum_{i=1}^n b(X_i) b^{\prime}(X_i), \quad \widehat M_j = n^{-1} \sum_{i=1}^n b(X_i) b_j(X_{+i}). \\ \widehat \rho_j & = \arg \min_{\rho} \rho^{\prime} \widehat G \rho - 2 \widehat M_j \rho + r_j \| \rho \|_1 \end{align} where \begin{align*} (\widehat \mathrm{A} b)_j = b(x)^{\prime} \widehat \rho_j, \quad j=1,2,\dots, p. \end{align*} Then, the rate condition (5) in Assumption (ref) on $a_n$ can be verified by the first-case rate of $\sup_{ 1 \leq j \leq p }\| \widehat \rho_j - \rho_j \|$ and can be established using the tools of CGST.

Least Squares Criterion for Dynamic Dual Representation

In this Section, we derive dynamic dual criterion function. Define

align[align omitted — 111 chars of source]

From the dual representation of (ref) we know that $\alpha_0$ satisfies

align[align omitted — 75 chars of source]

and $\mathrm{I} - \mathrm{A}^{*}_0$ is invertible.

We show that $\alpha_0$ can be represented as a minimizer of a criterion function that depends on $\mathrm{A}^{*}_0$. $\alpha_0$ will minimize the expected squared difference of the left and right-hand sides of equation (ref), that is

align[align omitted — 581 chars of source]

where the second equality follows by squaring and dropping the term that does not depend on $\alpha$, the third equality follows the Riesz representation in (ref), and the third equality by iterated expectations.

proposition[Least Squares Criterion for Dynamic Dual Representation] The dynamic dual representation $\alpha_0$ in (ref) is the unique minimizer of the quadratic criterion function \begin{align} \alpha_0 &=\arg \min_{\alpha} \mathrm{E} [((\mathrm{I} - \mathrm{A}^{*}_0) \alpha)(X)^2- 2m(Z,(\mathrm{I} - \mathrm{A}^{*}_0)\alpha)] \\ &=\arg \min_{\alpha} \mathrm{E} [((\mathrm{I} - \mathrm{A}^{*}_0) \alpha)((\mathrm{I} - \mathrm{A}^{*}_0) \alpha)(X)- 2m(Z,(\mathrm{I} - \mathrm{A}^{*}_0)\alpha)]. \\ &=\arg \min_{\alpha} \mathrm{E} [(\alpha(X)-\beta \alpha(X_-))((\mathrm{I} - \mathrm{A}^{*}_0) \alpha)(X)- 2m(Z,(\mathrm{I} - \mathrm{A}^{*}_0)\alpha)]. \end{align}

Given a vector of basis functions $$b(x) = (b_1(x), b_2(x), \dots, b_j(x), \dots, b_p(x)) \in \mathrm{R}^p,$$ where $p$ can differ from the one Section (ref), we approximate the dynamic dual representation via a linear form $$ \alpha(x) \approx \sum_{j=1}^{p} b_j(x) \rho_{\alpha j} = b(x)^{\prime} \rho_{\alpha} $$ where $\rho_{\alpha} = (\rho_{\alpha 1}, \rho_{\alpha 1}, \dots, \rho_{\alpha p}) \in \mathrm{R}^p$ is a $p$-vector. The vector is chosen to minimize an approximate least squares criterion $$ \rho_{\alpha} = \arg \min_{\rho \in \mathrm{R}^p} \rho^{\prime} G_{\alpha} \rho - 2 M_{\alpha} \rho $$ where the $p \times p$ matrix is an outer product of

align*[align* omitted — 127 chars of source]

and the free term is

align*[align* omitted — 84 chars of source]

We choose the criterion (ref) as opposed to (ref) so that the sample version of matrix $G_{\alpha}$ is symmetric and positive-definite. Given a radius $r_{\alpha}$, an $\ell_1$-regularized minimum distance estimator of the dynamic dual representation

align[align omitted — 233 chars of source]

where

align[align omitted — 298 chars of source]
theorem[Mean Square Rate for Lasso estimator of Dynamic Dual Representation] Suppose Assumption (ref) holds with $\alpha_0$ in place of $V_0$ and $\xi_{\alpha}$ in place of $\xi_V$, and suppose $\xi_{\alpha}>1/2$ and $a^{*}_n = o(n^{-d_{\alpha}})$. Then \begin{equation} \Vert\widehat \alpha- \alpha_0\Vert=o_{p}(n^{c}n^{ (-2d_{\alpha}\xi_{\alpha})/(2\xi_{\alpha}+1)}). \end{equation}

Theorem (ref) gives a mean square convergence rate for dynamic dual representation. The rate is determined by the sparsity parameter $\xi_{\alpha}$ of the dynamic dual representation and the first-stage rate parameter $d_{\alpha}$. The estimator is automatic in the sense that it only requires knowledge of the linear functional $m(Z, (\mathrm{I} - \mathrm{A}^{*})b)$.

Discussion and Related Results

In this Section, we discuss related results. Remark (ref) discusses cross-fitting. Remark (ref) introduces a neural network estimator of the value function. Remark (ref) introduces a neural network estimator of the dynamic dual representation. Remark (ref) describes the automatic property of the dynamic dual criterion function. Remark (ref) verifies rate conditions for asymptotic theory.

remark[Cross-fitting] To ensure that the nuisance components $\zeta_0, \mathrm{A}_0, \mathrm{A}^{*}_0$ and their respective criterion functions are estimated on different samples, the standard cross-fitting procedure is modified as follows. Let $L \geq 3$ be the number of partitions. For each partition $\ell \in \{1,2,\dots, L\}$, let $I^c_{\ell} = (Z_i)_{i \notin I_{\ell}}$ denote the set observations not in $I_{\ell}$. Partition $I^c_{\ell} = I^{c1}_{\ell} \sqcup I^{c2}_{\ell}$ into two halves. For each nuisance parameter $ \gamma \in \{ \zeta, \mathrm{A}, \mathrm{A}^{*} \}$, let $\widehat \gamma^{1}_{\ell}, \gamma^{2}_{\ell}$ denote the estimator computed on $I^{c1}_{\ell}$ and $ I^{c2}_{\ell}$, respectively. A cross-fit criterion for the value function is \begin{align*} L^V_n (V) = \sum_{i \in I^{c1}_\ell} \ell_V (Z_i, V, \widehat\omega^{2}_{\ell} ) + \sum_{i \in I^{c2}_\ell} \ell_V (Z_i, V, \widehat\omega^{1}_{\ell}), \quad \omega = (\zeta, \mathrm{A}). \end{align*} In a special case of Lasso estimator, the criterion $L^V_n (V)$ reduces to \begin{align*} L^V_n (\rho) &= \rho^{\prime} \widehat G^V \rho - 2 \widehat M^V \rho + r_V \| \rho \|_1 \end{align*} where \begin{align*} \widehat G^V &= \sum_{i \in I^{c1}_\ell} (\mathrm{I} - \widehat \mathrm{A}^{2}_{\ell}) b (X_i) (\mathrm{I} - \widehat \mathrm{A}^{2}_{\ell}) b^{\prime} (X_i) + \sum_{i \in I^{c2}_\ell} (\mathrm{I} - \widehat \mathrm{A}^{1}_{\ell}) b (X_i) (\mathrm{I} - \widehat \mathrm{A}^{1}_{\ell}) b^{\prime} (X_i) \\ \widehat M^V &= \sum_{i \in I^{c1}_\ell} (b (X_i) - \beta b(X_{i+})) \widehat \zeta^{2}_{\ell} (X_i) + \sum_{i \in I^{c2}_\ell} (b (X_i) - \beta b(X_{i+})) \widehat \zeta^{1}_{\ell} (X_i). \end{align*}
remark[Neural network estimator of value function] Proposition (ref) facilitates a general plug-in estimator of the value function with an arbitrary function class. Given a first-stage estimator of the per-period utility $\widehat \zeta$ of $\zeta_0$ and the expectation operator $\widehat \mathrm{A}$ of $\mathrm{A}_0$, define \begin{equation} \widehat{V}:=\arg\min_{V \in\mathcal{V}_{n}} n^{-1} \sum_{i=1}^n \ell_V (Z_i, V, \widehat \zeta, \widehat \mathrm{A}), \end{equation} where $\ell_V(Z, V, \zeta_0, \mathrm{A}_0)$ is taken to as in (ref). Theorem (ref) in Appendix (ref) establishes mean square consistency of $\widehat{V}$ to $V_0$ for an arbitrary function class. Corollary (ref) establishes mean square consistency of the neural network estimator of $V_0$.
remark[Neural network estimator of dynamic dual representation] Proposition (ref) facilitates a general plug-in estimator of dynamic dual representation with an arbitrary function class. Given a first-stage estimator of the operator $\widehat {\mathrm{A}}^{*}$ of $\mathrm{A}^{*}_0$, define \begin{equation} \widehat{\alpha}:=\arg\min_{\alpha \in\mathcal{A}_{n}} n^{-1} \sum_{i=1}^n \ell_{\alpha} (Z_i, \alpha, \widehat \mathrm{A}^{*}). \end{equation} Theorem (ref) in Appendix (ref) establishes mean square consistency of $\widehat{\alpha}$ to $\alpha_0$ for an arbitrary function class. It can be used to derive mean square consistency of neural network estimator of $\alpha_0$.
remark[Automatic property of criterion function (ref)] The criterion function (ref) have a convenient property that it only depends on the parameter of interest through $m(Z, (\mathrm{I} - \mathrm{A}_0) \alpha)$ and does not require an explicit formula for the weighting function $w_0$. A similar property has been establishes for the results in chernozhukov2021automatic and chernozhukov2024qm. When the model is static, that is $\beta=0$, Theorem (ref) recovers Theorem 1 of chernozhukov2021automatic. Likewise, the criterion function (ref) reduces to $$ \alpha_0 =\arg \min_{\alpha} \mathrm{E} [\alpha^2(X)- 2m(Z,\alpha)] $$ proposed in chernozhukov2024qm.
remark[Verification of rate conditions for asymptotic theory.] Suppose the conditions of Theorems (ref) and (ref) hold with $\xi_V, \xi_{\alpha}$ and $d_V, d_{\alpha}$. Furthermore, suppose $$ 2(d_{\alpha}\xi_{\alpha})/(2\xi_{\alpha}+1) + 2(d_{V}\xi_{V})/(2\xi_{V}+1) > 1/2. $$ For example, if $\min (\xi_{\alpha}, \xi_V) > 1 + 0.01$ and $\min (d_V, d_{\alpha}) \in (3/8, 1/2)$, the product of mean square rates can be upper bounded as $$ n^{1/2} \| \widehat V_{\ell} - V_0 \| \| \alpha_{\ell} - \alpha_0 \| = o_P (n^{1/2} n^{-1/4} \cdot n^{-1/4}) = o_P(1) $$ which suffices for the product rate condition $$ n^{1/2} \| \widehat V_{\ell} - V_0 \| \| \alpha_{\ell} - \alpha_0 \| = o_P(1) $$ of Assumption (ref) in Appendix (ref).

Conclusions

In this paper we introduce welfare metrics -- including welfare decompositions into direct and indirect effects -- and give a complete set of estimation and inference results for them in the presence of high-dimensional state space. The results are presented for the dynamic binary choice model of Rust,HotzMiller,AMira2002 but are applicable to other dynamic models (e.g., dynamic games AMira2007,BBL). For the case of average welfare and related metrics, the proposed estimator is a known function of choice probabilities and the structural parameter. In particular, if the model has “terminal action” property, value function or any other dynamic object does not have to be estimated at all. We have applied these methods to estimate the average teachers welfare in an application to teachers absenteeism as in DHR.