Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,123 characters · 12 sections · 75 citation commands
Structural models for policy-making
\setcounter{page}{1} \thispagestyle{empty}
\setcounter{page}{1} \FloatBarrier
Structural microeconometricians use highly parameterized computational models to investigate economic mechanisms, predict the impact of proposed policies, and inform optimal policy-making Wolpin.2013. These models represent deep structural relationships of theoretical economic models invariant to policy changes Hood.1953. The sources of uncertainty in such an analysis are ubiquitous Saltelli.2020. For example, models are often misspecified, there are numerical approximation errors in their implementation, and model parameters are uncertain. Therefore, most disciplines require a proper account of uncertainty before using computational models to inform decision-making Council.2012,SAPEA.2019.\\
The following study focuses on parametric uncertainty in structural microeconometric models that are estimated on observed data. Researchers often do not account for parametric uncertainty and conduct an as-if analysis in which the point estimates serve as a stand-in for the true model parameters. They then continue to study the implications of their models at the point estimates Adda.2017,Blundell.2016,Eckstein.2019,Eisenhauer.2015b and rank competing policy proposals based on the point predictions alone Blundell.2012,Cunha.2010,Gayle.2019,Todd.2006. In fact, Keane.2011d states in their handbook article that they are unaware of any applied work that reports the distribution of policy predictions under parametric uncertainty. To the best of our knowledge, this statement remains true more than a decade later. Consequently, economists risk accepting fragile findings as facts, ignoring the trade-off between model complexity and prediction uncertainty, and neglecting to frame policy advice as a decision problem under uncertainty.\\
To mitigate these shortcomings, we develop an approach that copes with parametric uncertainty in structural microeconometric models and embeds model-informed policy-making in a decision-theoretic framework. Ideally, policy-makers fix the parameter space ex-ante and then evaluate the policy options according to decision rules. However, this approach is often computationally intractable. We, therefore, follow Manski.2021's suggestion and, instead of using the parameter estimates as-if they were true, incorporate uncertainty in the analysis by treating the estimated confidence set as-if it is correct. We use the confidence set to construct an uncertainty set that is anchored in empirical estimates, statistically meaningful, and computationally tractable Ben-Tal.2013. Instead of just focusing on the point estimates, we evaluate counterfactual policies based on all parametrizations within the uncertainty set.\\
We draw on statistical decision theory Manski.2013 to deal with the uncertainty in counterfactual predictions. This approach promotes a well-reasoned and transparent policy process. Before a decision, it clarifies trade-offs between choices Gilboa.2018. Afterward, decision-theoretic principles allow constituents to scrutinize the coherence of choices Gilboa.2020, ease the ex-post justification Berger.2021, and facilitate the communication of uncertainty Manski.2019.\\
We tailor our approach to the class of Eckstein-Keane-Wolpin (EKW) models Aguirregabiria.2010. Labor economists often use EKW models to learn about human capital investment and consumption-saving decisions and predict the impact of proposed reforms to education policy and welfare programs Keane.2011d,Low.2017,Blundell.2017. The analysis of these models poses serious computational challenges. During estimation, EKW models are solved thousands of times and even a single solution often takes several minutes. Thus, a decision-theoretic ex-ante analysis of alternative decision rules across the whole parameter space, as intended by Wald.1950, is infeasible. Instead we construct an uncertainty set, a subset of the whole parameter space, and deal with the ex-post uncertainty after estimating the model. This compromise allows us to garner the benefits of using statistical decision theory to shape policy-making under uncertainty while ensuring the computational tractability of our analysis.\\
As an example of our approach, we analyze the seminal human capital investment model by Keane.1997 as a well-known, empirically grounded, and computationally demanding test case. We follow the authors and estimate the model on the National Longitudinal Survey of Youth 1979 (NLSY79) NLSY.2019 using the original dataset and reproduce all core results. We revisit their predictions for the impact of a tuition subsidy on completed years of schooling. The economics of the model implies that the nonlinear mapping between the model parameters and predictions is truncated at zero, and we thus use the Confidence Set (CS) bootstrap Woutersen.2019 to estimate the confidence set for the counterfactuals. We document considerable uncertainty in the policy predictions and highlight the resulting policy recommendations from different formal rules on decision-making under uncertainty.\\
Our work extends existing research exploring the sensitivity of implications and predictions to parametric uncertainty in macroeconomics and climate economics. For example, Harenberg.2019 study uncertainty propagation and sensitivity analysis for a standard real business cycle model. Cai.2019 examine how uncertainties and risks in economic and climate systems affect the social cost of carbon. However, neither of them estimates their model on data. Instead, they rely on expert judgments to inform the degree of parametric uncertainty. They do not investigate the consequences of uncertainty for policy decisions in a decision-theoretic framework.\\
We complement a burgeoning literature on the sensitivity analysis of policy predictions in light of model or moment misspecification. For example, Andrews.2017 and Andrews.2020 treat the model specification as given and then analyze the sensitivity of the parameter estimates to the misspecification of the moments used for estimation. Christensen.2019 study global sensitivity of the model predictions to misspecification of the distribution of unobservables. Jorgensen.2021 provides a local measure for the sensitivity of counterfactuals to model parameters that are fixed before the estimation of the model.\footnote{For other examples, see Armstrong.2021, Bonhomme.2020, Bugni.2019, and Mukhin.2018.} This literature does not embed the counterfactual predictions in a decision-theoretic setting. Recent work by Kalouptsidi.2020, Kalouptsidi.2021, and Norets.2014 studies (partial) identification and inference on counterfactuals. However, they all adopt the setup outlined in Rust.1987 and exploit the additive separability of the immediate utility function between observed and unobserved state variables, which does not apply to EKW models. In related work, Blesch.2021 conduct a decision-theoretic ex-ante analysis to determine optimal decision rules in Rust.1987's stochastic dynamic investment model where the decision-maker directly accounts for uncertainty in the model's transition dynamics. They only consider uncertainty in a subset of the model's parameters which are estimated outside the model and remain fixed to their point estimates during the analysis.\\
In Section (ref), we describe the decision-theoretic framework for making model-informed decisions under parametric uncertainty using an illustrative example. After summarizing the empirical setting of Keane.1997 in Section (ref), we present our results in Section (ref). We complete our analysis in Section (ref) with a brief conclusion and outlook.
\FloatBarrier
In the following section, we discuss uncertainty propagation and the common practice of using estimated parameters as a plug-in replacement for the true model parameters. We then explore the limitations of this strategy and introduce our alternative approach, in which we implement estimated confidence sets to construct uncertainty sets. In so doing, we are able to cope with uncertain policy predictions in a proper decision-theoretic framework. \\
At a high level, a structural microeconometric model provides a mapping $\mathcal{M}(\bm{\theta})$ between the $l$ model parameters $\bm{\theta} \in \bm{\Theta}$ and a quantity $y$ that is of interest to policy-makers.
A policy $g \in \mathcal{G}$ changes the mapping to $\mathcal{M}_g(\bm{\theta})$ and produces a counterfactual $y_g$.\\
Estimation of a baseline model $\mathcal{M}(\bm{\theta})$ describing the status-quo on observed data allows researchers to learn about the true parameters. Frequentist estimation procedures such as maximum likelihood estimation and the method of simulated moments produce a point estimate $\hat{\bm{\theta}}$. However, uncertainty about the true parameters remains.\\
Previewing our empirical analysis of Keane.1997, our $\mathcal{M}$ is provided by a dynamic model of human capital accumulation, which we estimate on observed schooling and labor market decisions using simulated maximum likelihood estimation. The policy $g$ is the implementation of a college tuition subsidy, and the counterfactual is the level of completed schooling in the population. Example parameters that drive the economics of the model are time preferences of individuals, the return to schooling, and the transferability of work experience across occupations.\\
The following illustrative example highlights our key points. We consider two policies $g \in\{1, 2\}$ that result in two different mappings $ (\mathcal{M}_1, \mathcal{M}_2)$ of the same scalar $\theta$ to a counterfactual $y_g$. Higher values of $y_g$ are more desirable for a policy-maker. The point estimate $\hat{\theta}$ is determined by estimating a baseline model on an observed dataset. We denote the probability density function of its sampling distribution by $f_{\hat{\theta}}$.\\
Under the first policy, the counterfactual is an increasing nonlinear function of $\theta$. In the case of the second policy, the relationship is decreasing and linear.
\FloatBarrier Figure (ref) traces the counterfactual from both models over a range of the parameter. At the point estimate, both models yield the same value for the counterfactual. Once we account for uncertainty in our estimates of the true parameter, deciding which policy to adopt becomes less straightforward: for higher values of $\theta$, the first policy is preferred, while the opposite is true for lower values.
Manski.2021 suggests acknowledging parametric uncertainty by working with estimated confidence sets instead of point estimates. A confidence set $\bm{\Theta}(\alpha) \subset \bm{\Theta}$ covers the true parameters, from an ex-ante point of view, with a predetermined coverage probability of $(1 - \alpha)$. Proceeding with our analysis, we refine the status quo procedure, in which estimated parameter values serve as a stand-in for the model's true parametrization. Instead, we assume the estimated confidence set for the parameters $\hat{\bm{\Theta}}(\alpha)$ and the counterfactual $\hat{\bm{\Theta}}_{y_g}(\alpha)$ are correct and analyze policy decisions accordingly. \\
Based on the estimated confidence sets, we construct so-called uncertainty sets for the parameters $\mathcal{U}(\alpha)$ and the prediction $\mathcal{U}_{y_g}(\alpha)$ by only considering parameterizations that we cannot reject based on a hypothesis test with confidence level $1 - \alpha$. This approach ensures the tractability of our decision-theoretic analysis, as the uncertainty set of the parameters is much smaller than the whole parameter space of the model. We adopt this procedure from the literature on data-driven robust optimization in operations research Ben-Tal.2013,Bertsimas.2018.
In our setting, a policy-maker relies on a structural model with an uncertain parametrization to map alternative policies to counterfactual predictions. In most cases, the preferred policy depends on the model's uncertain true parameters. We, therefore, draw on statistical decision theory to organize the decision-making process Gilboa.2009,Marinacci.2015.\\
Returning to our example, we rank the two policies according to alternative statistical decision rules using an uncertainty set derived from a confidence set with a $90\%$ coverage probability. In what follows, we postulate a simple linear utility function $U(y_g)$ to describe the policy-maker's preferences.\footnote{We assume that the sampling distribution of the point estimate is normal with a mean of three and a standard deviation of three-fourths. We can derive the uncertainty sets directly and simply consider realizations of $\theta\in[1.76, 4.23]$.}\\
Figure (ref) shows the implied sampling distribution of the predictions for the two alternative policies and the corresponding uncertainty sets $\mathcal{U}_{y_g}(0.1)$. The mapping $\mathcal{M}_1$ is highly nonlinear, while the mapping $\mathcal{M}_{2}$ is linear. When evaluated at the point estimate, the counterfactual is the same under both policies, so a policy-maker is indifferent. However, the spread of the uncertainty set differs considerably.\\
\FloatBarrier
Decision theory proposes a variety of different rules for reasonable decisions in this setting. We explore the following four: (1) as-if optimization, (2) maximin criterion, (3) minimax regret rule, and (4) subjective Bayes.\\
As-if optimization describes the predominant practice. The estimation of the model produces point estimates that serve as a plug-in for the true parameters. The decision maximizes the utility at the point estimate. More formally,
Given our example, an as-if policy-maker is indifferent between the two policies, since both policies result in the same counterfactual at the point estimates as indicated by the dashed line in Figure (ref).\\
The maximin criterion and minimax regret rule are two common alternatives that favor actions that work uniformly well over all possible parameters in the uncertainty set. This approach departs from as-if optimization, which only considers a policy's performance at a single point in the uncertainty set. The maximin decision Gilboa.1989, Wald.1950 is determined by computing the minimum utility for each policy within the uncertainty set and choosing the one with the highest worst-case outcome. Stated concisely,
Returning to Figure (ref), a maximin policy-maker prefers $g_2$ as the worst-case outcome. Within the uncertainty set, $\underline{y}_2$ is better than under the alternative policy, $g_1$.\\
The minimax regret rule Manski.2004, Niehans.1948 computes the maximum regret for each policy over the whole uncertainty set and chooses the policy that minimizes the maximum regret. The regret of choosing a policy $g$ for a given parameterization of the model is the difference between the maximum possible utility achieved from adopting $\tilde{g} \in \mathcal{G}$ and the actual utility obtained. The decision maximizes:
Figure (ref) compares our two policy examples over the uncertainty sets. A policy-maker adopting policy $g_1$ regrets his choice for small values of the model parameter, while the opposite is true for larger values. The regret of each policy is maximized at the boundaries of the uncertainty set. Maximum regret is minimized when a policy-maker chooses $g_1$. It corresponds to the difference in the counterfactual at the lower boundary of the uncertainty set instead of the larger difference at its upper bound. This outcome contradicts the maximin decision in which policy $g_2$ is preferred.
\FloatBarrier
Each decision rule presented so far focuses on a single point in the uncertainty set as the policy's relevant performance measure. Bayesian approaches aggregate a policy's performance over the complete uncertainty set.\\
Maximization of the subjective expected utility Savage.1954 requires the policy-maker to place a subjective probability distribution $f_{\bm{\theta}}$ over the parameters in the uncertainty set. A policy-maker then selects the alternative with the highest expected subjective utility. Formally,
Applying a uniform distribution to our example, a policy-maker chooses $g_1$, which performs well for high values of $\bm{\theta}$ and still reasonably well for low values.
\FloatBarrier
We now present the general structure of Eckstein-Keane-Wolpin (EKW) models Aguirregabiria.2010 and their solution approach. We then turn to the customized version used by Keane.1997 to study the career decisions of young men and investigate the consequences of parametric uncertainty in this empirically-grounded and computationally demanding setting. We outline their model's basic setup, provide some descriptive statistics of the empirical data used in our estimation, and then discuss the core findings.
EKW models describe sequential decision-making under uncertainty Gilboa.2009, Machina.2014. At time $t = 1, \hdots, T$ each individual observes the state of their choice environment $s_t\in S$ and chooses an action $a_t$ from the set of admissible actions $\mathcal{A}$. The decision has two consequences: an individual receives an immediate utility $u_t(s_t, a_t)$ and their environment evolves to a new state $s_{t + 1}$. The transition from $s_t$ to $s_{t + 1}$ is affected by the action but remains uncertain. Since individuals are forward-looking, they do not simply choose the alternative with the highest immediate utility. Instead, they take the future consequences of their actions into account.\\
A policy $\pi =(d^\pi_1, \hdots, d^\pi_T)$ provides the individual with instructions for choosing an action in any possible future state. It is a sequence of decision rules $d^\pi_t$ that specify the action $d^\pi_t(s_t) \in \mathcal{A}$ at a particular time $t$ for any possible state $s_t$ under $\pi$. The implementation of a policy generates a sequence of utilities that depends on the objective transition probability distribution $p_t(s_t, a_t)$ for the evolution from state $s_t$ to $s_{t + 1}$ induced by the model.\\
Figure (ref) depicts the timing of events for two generic periods. At the beginning of period $t$, an individual fully learns about each action's immediate utility, selects one of the alternatives, and receives its immediate utility. Then, the state evolves from $s_t$ to $s_{t + 1}$, and the process repeats itself in $t + 1$.\\
Individuals make their decisions facing uncertainty about the future and seek to maximize their expected total discounted utilities over all decision periods given all available information. They have rational expectations Muth.1961, so their subjective beliefs about the future agree with the objective probabilities for all possible future events provided by the model. Immediate utilities are separable between periods Kahneman.1997, and a discount factor $\delta$ parameterizes a preference for immediate over future utilities Samuelson.1937.\\
Equation ((ref)) formally describes the individual's objective. Given an initial state $s_1$, they implement a policy $\pi$ that maximizes the expected total discounted utilities over all decision periods given the information available at the time.
EKW models are set up as a standard Markov decision process (MDP) Puterman.1994,Rust.1994,White.1993 that can be solved by a simple backward induction procedure. In the final period $T$, there is no future to consider, and the optimal action is choosing the alternative with the highest immediate utility in each state. With the decision rule for the final period, we can determine all other optimal decisions recursively. We use our group's open-source research code \verb+respy+ Gabler.2020b, which allows for the flexible specification, simulation, and estimation of EKW models. Detailed documentation of the software and its numerical components is available at \url{http://respy.readthedocs.io}.
Keane.1997 specialize the model above to explore the career decisions of young men regarding their schooling, work, and occupational choices using the National Longitudinal Survey of Youth 1979 (NLSY79) NLSY.2019 for the estimation of the model. We restrict ourselves to a basic summary of their setup. Further documentation of the model specification and the observed dataset is available in the Appendix.\\
Keane.1997 follows individuals over their working life from young adulthood at age 16 to retirement at age 65. Each decision period $t = 16, \dots, 65$ represents a school year. Figure (ref) illustrates the initial decision problem as individuals select one of five alternatives from the set of admissible actions $a\in\mathcal{A}$. They can decide to either work in a blue-collar or a white-collar occupation ($a = 1, 2$), serve in the military $(a = 3)$, attend school $(a = 4)$, or stay at home $(a = 5)$.\\
\FloatBarrier
Individuals are already heterogeneous when entering the model. They differ with respect to their level of initial schooling $h_{16}$, and have one of four different $\mathcal{J} = \{1, \hdots, 4\}$ alternative-specific skill endowment types $\bm{e} = \left(e_{j,a}\right)_{\mathcal{J} \times \mathcal{A}}$.\\
The immediate utility $u_a(\cdot)$ of each alternative consists of a non-pecuniary utility $\zeta_a(\cdot)$ and, at least for the working alternatives, an additional wage component $w_a(\cdot)$. Both depend on the level of human capital as measured by their alternative-specific skill endowment $\bm{e}$, their years of completed schooling $h_t$, and their occupation-specific work experience $\bm{k_t} = \left(k_{a,t}\right)_{a\in\{1, 2, 3\}}$. The immediate utilities are influenced by last-period choices $a_{t -1}$ and alternative-specific productivity shocks $\bm{\epsilon_t} = \left(\epsilon_{a,t}\right)_{a\in\mathcal{A}}$ as well. Their general form is given by:
Work experience $\bm{k_t}$ and years of completed schooling $h_t$ evolve deterministically. There is no uncertainty about grade completion Altonji.1993 and no part-time enrollment. Schooling is defined by time spent in school, not by formal credentials acquired. Once individuals reach a certain amount of schooling, they acquire a degree.
The productivity shocks $\bm{\epsilon_t}$ are uncorrelated across time and follow a multivariate normal distribution with mean $\bm{0}$ and covariance matrix $\bm{\Sigma}$. Given the structure of the utility functions and the distribution of the shocks, the state at time $t$ is $s_t = \{\bm{k_t}, h_t, t, a_{t -1}, \bm{e},\bm{\epsilon_t}\}$.\\
Skill endowments $\bm{e}$ and initial schooling $h_{16}$ are the only sources of persistent heterogeneity in the model. All remaining differences in life-cycle decisions result from different transitory shocks $\bm{\epsilon_t}$ that occur over time.\\
Theoretical and empirical research from specialized disciplines within economics informs the specification of each $u_a(\cdot)$. As an example, we provide the exact functional form of the non-pecuniary utility from schooling in Equation ((ref)). Further details on the specification of the utility functions are available in the Appendix.
There is a direct cost in the form of tuition for continuing education after high school $\beta_{tc_1}$ and college $\beta_{tc_2}$. The decision to leave school is reversible, but entails re-enrollment costs that differ by schooling category ($\beta_{rc_1}, \beta_{rc_2}$).\\
We analyze the original dataset used by Keane.1997. We only provide a brief description and relegate further details to the Appendix. The authors construct their sample based on the NLSY79, a nationally representative sample of young men and women living in the United States in 1979 and born between 1957 and 1964. Individuals were followed from 1979 onwards and repeatedly interviewed about their schooling decisions and labor market experiences. Based on this information, individuals are assigned to either working in one of the three occupations, attending school, or simply staying at home.\\
Keane.1997 restrict attention to white men, who turned 16 between 1977 and 1981, and exploit information collected between 1979 and 1987. Thus, individuals in the sample range in age between 16 and 26 years old. While the sample initially consists of 1,373 individuals at age 16, this number drops to 256 at the age of 26 due to sample attrition and missing data. Overall, the final sample consists of 12,359 person-period observations.\\
Figure (ref) summarizes the evolution of choices and wages over the sample period. Roughly 86% of individuals initially enroll in school, but this share steadily declines with age. Nevertheless, about 39% pursue some form of higher education and obtain more than a high school degree. As individuals leave school, most of them initially pursue a blue-collar occupation. However, the relative share of white-collar workers increases as individuals entering the labor market later gain access to higher levels of schooling. At age 26, about 48% work in a blue-collar occupation and 34% in a white-collar occupation. The share of individuals in the military peaks around age 20 at 8%. At its maximum around age 18, approximately 20% of individuals stay at home.\\
\FloatBarrier
For an individual, the average wage starts at about \$10,000 at age 16 and increases considerably up to about \$25,000 by the age of 26. While starting wages for blue-collar workers are about \$10,286, wages in white-collar occupations and the military start around \$9,000. However, wages for white-collar occupations increase sharply over time, overtaking blue-collar wages around age 21. By the end of the observation period, wages for white-collar occupations are about 50% higher than blue-collar wages at \$32,756 compared to only \$20,739. Military wages remain lowest throughout.\\
We consider observations for $i = 1, \hdots, N$ individuals in each time period $t = 1, \dots, T_i$. For every observation $(i, t)$ in the data, we observe the action $a_{it}$, some components $\bar{u}_{it}$ of the utility, and a subset $\bar{s}_{it}$ of the state $s_{it}$. Therefore, from an economist's point of view, we must distinguish between two types of state variables $s_{it} = \{\bar{s}_{it}, \bm{e},\bm{\epsilon_t}\}$. At time $t$, the economist and individual both observe $\bar{s}_{it}$, while $\{ \bm{e},\bm{\epsilon_t}\}$ is only observed by the individual.\\
We use simulated maximum likelihood Fisher.1922,Manski.1977 estimation and determine the $88$ model parameters $\hat{\bm{\theta}}$ that maximize the likelihood function $\mathcal{L}(\bm{\theta}\mid\mathcal{D})$. As we only observe a subset $\bar{s}_t = \{\bm{k_t}, h_t, t, a_{t -1}\}$ of the state, we can determine the probability $p_{it}(a_{it}, \bar{u}_{it} \mid \bar{s}_{it}, \bm{\theta})$ of individual $i$ at time $t$ in $\bar{s}_{it}$ choosing $a_{it}$ and receiving $\bar{u}_{it}$ given parametric assumptions about the distribution of $\bm{\epsilon_t}$. The objective function takes the following form:
Overall, our parameter estimates are in broad agreement with the results reported in the original paper and the related literature. For example, individuals discount future utilities by $6\%$ per year. The returns to schooling vary according to occupation. While wages for white-collar occupations increase by about $6\%$ with each additional year of schooling, they only increase by $2\%$ for those working blue collar jobs. Skills are transferable across occupations as work experience increases wages in both blue and white-collar occupations.\\
Figure (ref) shows the overall agreement between the empirical data and a dataset simulated using the estimated model parameters. We show average wages and the share of individuals choosing a blue-collar occupation over time. The results are based on a simulated sample of $10,000$ individuals. Additional model fit statistics are available in the Appendix.
\FloatBarrier
We adhere to the procedure outlined by the authors of the original paper and use the estimated model to conduct the ex-ante evaluation of a $\$2,000$ tuition subsidy on educational attainment. We simulate a sample of $10,000$ individuals using the point estimates and compare completed schooling to a sample of the same size, but with a reduction of $\hat{\beta}_{tc_1}$ by $\$2,000$. The subsidy increases average final schooling by 0.65 years. College graduation increases by 13 percentage points and high school graduation rates improve by 4 percentage points.
The construction of confidence sets for counterfactuals in many structural models poses two distinct challenges. First, the computational burden of even a single estimation of the model is considerable. This makes the application of a standard bootstrap approach Efron.1979 infeasible. Second, the nonlinear mapping from the parameters of the model to the counterfactual predictions often has kinks or is truncated. For example, in our case, the predicted impact of a tuition subsidy is bounded from below by zero. This violates the smoothness requirements of the delta method.\\
We use the Confidence Set (CS) bootstrap to construct the confidence set of the counterfactual. Although the CS bootstrap was originally proposed in Rao.1973, it has only recently been formalized by Woutersen.2019. Its application does not require repeated estimations of the model, as it uses the asymptotic normal distribution of the estimator for $\hat{\bm{\theta}}$. Furthermore, its validity does not depend on the differentiability of the prediction function.\footnote{See Reich.2020 for a critical assessment of confidence sets based on asymptotic arguments. They advocate the use of likelihood-ratio confidence intervals instead and set up their computation as a constraint optimization problem.}\\
Algorithm (ref) provides a concise description of the steps involved, where $\chi_l^2(1 - \alpha)$ is the quantile function for probability $1 - \alpha$ of the chi-square distribution with $l$ degrees of freedom.\\
\floatname{algorithm}{ Algorithm}
\FloatBarrier
To summarize, we draw a large sample of $M$ parameters from the estimated asymptotic normal distribution of our estimator with mean $\hat{\bm{\theta}}$ and covariance matrix $\hat{\boldsymbol{\Sigma}}$, accepting only those draws that are elements of the confidence set of the model parameters. We then compute the counterfactual for all remaining draws and calculate the confidence set for the counterfactual based on its lowest and highest value.\\
The CS bootstrap poses a considerable computational challenge. In many applications, including our own, a single prediction of a counterfactual takes several minutes. At the same time, the number of parameter samples must be large to ensure that the minimum and maximum values for the counterfactual prediction are reliable. However, the algorithm is amenable to parallelization using modern high-performance computational resources by processing each of the $M$ parameter draws independently.\\
Our uncertainty sets then take the following form:
\FloatBarrier
Turning to the presentation of our results, we focus on the impact of a $\$2,000$ tuition subsidy on completed schooling and use the 90% uncertainty set to measure the degree of uncertainty. All our results potentially depend on the size of the uncertainty set. In practice, policy-makers choose the uncertainty set's size in line with their underlying preferences - the more desirable protection against unfavorable outcomes is, the larger the uncertainty set will be.\footnote{In a different setting, Blesch.2021 conduct an ex-ante performance evaluation of the statistical decision functions over the whole parameter space Wald.1950,Manski.2021.}\\
All results are based on $30,000$ draws from the asymptotic normal distribution of our parameter estimates. We follow Keane.1997 and start by analyzing the prediction for a general subsidy. Then we turn to the situation where we use endowment types for policy targeting. Throughout our analysis, we postulate a linear utility function for the policy-maker.
Figure (ref) explores the impact prediction for a general tuition subsidy. We show the point prediction, its sampling distribution, and the uncertainty set. At the point estimate, average schooling increases by $0.65$ years. However, there is considerable uncertainty about the prediction, as the uncertainty set ranges from $0.15$ to $1.10$ years.\\
\FloatBarrier In Figure (ref), we trace the effect of the discount rate $\delta$ on the subsidy's impact over the uncertainty set, while keeping all other parameters at their point estimate. Initially, as $\delta$ increases, so does the policy's impact as individuals value the long-term benefits from increasing their level of schooling more and more. However, for high levels of the discount factor, the policy's impact starts to decrease as most individuals already complete a high school or college degree even without the subsidy.
\FloatBarrier
So far, we restricted the analysis to a general subsidy available to the whole population and the average predicted impact. We now examine the setting in which a policy-maker can target individuals based on the type of their initial endowment. The importance of early endowment heterogeneity in shaping economic outcomes over the life-cycle is the most important finding from Keane.1997. It served as motivation for a host of subsequent research on the determinants of skill heterogeneity among adolescents Caucutt.2020,Erosa.2010,Todd.2007.\\
To ease the exposition, we initially focus our discussion of results on Type 1 and Type 3 individuals. We later rank policies targeting either of the four types based on the different decision-theoretic criteria. Additional results are available in our Appendix.\\
Figure (ref) confirms that life-cycle choices differ considerably by initial endowment type. On the left, we show the number of periods the two types spend on average in each of the five alternatives. Those characterized as Type 1 individuals spend more than six years on their education even after entering the model. Type 3 individuals, on the other hand, extend their academic pursuits for only an additional two years. This difference translates into very different labor market experiences. While Type 1 individuals work for about 35 years in a white-collar occupation, Type 3 workers switch more frequently between white and blue-collar occupations and spend a comparable amount of time working in either occupation -- approximately 44 years split equally among white and blue-collar occupations. Both types only spend a short time at home.
\FloatBarrier On the right, we show the distribution of final schooling for both types. Years of schooling are considerably higher for Type 1 individuals with an average of more than 16 years compared to only 12 years for those identified as Type 3 individuals. Nearly all Type 1 individuals enroll in college and most graduate with a degree.\\
Figure (ref) provides a visualization of our core results for a targeted subsidy. At the point estimates, the predicted impact is considerably lower for Type 1 than Type 3. However, the prediction uncertainty is much larger for Type 3 compared to Type 1. The uncertainty set for Type 3 ranges all the way from $0$ to $1.2$ years, while the prediction for Type 1 is between $0.18$ and $0.75$.\\
\FloatBarrier This heterogeneity in impact and prediction uncertainty follows directly from the underlying economics of the model. Type 1 individuals are already more likely to have a college degree before the subsidy, and thus, the predicted impact is smaller. Alternatively, Type 1 individuals affected by the subsidy are in the middle of pursuing a college education and thus directly benefit from it. Since Type 3 individuals are at the lower end of the schooling distribution, a tuition subsidy can considerably increase their level of schooling. Whether the subsidy succeeds in doing so, however, remains uncertain.\\
We now consider the policy option to target Type 2 and Type 4 as well. Their point predictions are actually highest with an additional $0.81$ years on average for Type 2 and $0.75$ years for Type 4. However, both predictions are fraught with uncertainty. For Type 2 the uncertainty set ranges from $0.17$ to $1.3$, while for Type 4 it starts at zero and spans all the way to $1.18$.\\
Figure (ref) shows the policy alternative's ranking by the decision-theoretic criteria we discussed in Section (ref). Ranking alternatives using as-if optimization is straightforward. A policy targeting Type 2 is the most preferred alternative, while a focus on Type 1 is the least attractive. However, once we account for the presence of uncertainty in the predictions, a more nuanced picture emerges. Moving from as-if optimization to a subjective Bayes criterion using a uniform distribution over the uncertainty set does not change the ordering. However, once a decision-maker is concerned with performance across the whole range of values in the uncertainty set -- we move to the minimax regret or maximin criterion -- a policy targeting Type 1 becomes more and more attractive despite its low point prediction because its worst-case utility is highest.\\
\FloatBarrier In general, framing policy advice as a decision problem under uncertainty shows that there are many different ways of making reasonable decisions. The ranking of policies varies depending on the decision criteria. Not only that, but due to the necessary ex-post nature of our implementation, the ranking for a given criteria also depends on the choice of $\alpha$. The selection of $\alpha$ is part of the decision problem: the more a policy-maker is concerned about worst-case scenarios, the smaller the appropriate value for $\alpha$ will be. After deciding on a preferred decision rule, we suggest performing a sensitivity analysis around the selected $\alpha$ value by checking how much the policy ranking varies within a neighborhood.
\FloatBarrier
We develop a generic approach that addresses parametric uncertainty when using models to inform policy-making. We propose a decision-theoretic analysis of computationally demanding structural models based on uncertainty sets. We construct the uncertainty sets from empirical estimates and ensure their computational tractability by using the confidence set bootstrap. We revisit the seminal work by Keane.1997 to document the empirical relevance of prediction uncertainty and showcase our analysis. Focusing on their ex-ante evaluation of a tuition subsidy, we report considerable uncertainty in the policy's impact on completed schooling. We show how a policy-maker's preferred policy depends on the choice of alternative formal rules for decision-making under uncertainty.\\
In our ongoing research, we pursue three avenues for further improvements. First, we link our work with the literature on inference under (local) model misspecification to refine the construction of our uncertainty sets. For example, Armstrong.2021 and Bonhomme.2020 propose different methods for taking misspecification into account when constructing confidence sets. Second, we incorporate ideas from the literature on global sensitivity analysis Razavi.2021 to identify the parameters most responsible for uncertainty in predictions. The attribution of importance based on Shapely values, familiar to economists from game theory, appears promising Owen.2014, Shapley.1953 as well. Third, we address our analysis's computational burden using surrogate modeling Forrester.2008, which emulates the full model's behavior at a negligible cost per run and allows us to determine prediction uncertainty using a nonparametric bootstrap procedure.