Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,228 characters · 17 sections · 73 citation commands
Group Shapley Value and Counterfactual Simulations in a Structural Model
\doublespacing
The Shapley value, first introduced by ShapleyValue, is a foundational solution concept in cooperative game theory, with desirable properties supported by axiomatic approaches. It is also sustainable in non-cooperative settings gul1989bargaining. Shorrocks:2013 popularized its use in empirical applications, primarily for decomposing $R^2$ within linear regression models. Many studies have used the Shapley-Shorrocks decomposition: see, e.g., allen2014information, allen2014trade, dustmann2014out and henderson2018global among others. Most notably since the work of LundbergLee, Shapley values have become one of the main tools of explainable artificial intelligence (AI). They are likewise recognized as useful interpretational tools in other disciplines, such as global sensitivity analysis Owen:2014.
We propose to make a use of a particular variant which we call the group Shapley value. Specifically, we use it to understand the outcomes of counterfactual simulations in a structural economic model by quantifying the importance of various elements embedded within the model. In a structural model, the crucial task involves estimating or calibrating the underlying model parameters. Consequently, any changes to these parameter values will directly influence the outcomes of the structural model. Our motivating example is from DV2019, who developed a method to disentangle sources of capital misallocation using data from both China and the United States. Their model includes various components such as shocks to productivity, investment adjustment costs, uncertainty, and so on. They are specified via a number of parameters that are estimated or calibrated: Table (ref) reproduces the parameter estimates from DV2019.
In general, it would be challenging to compare the relative importance across different components of the parameters because they tend to be defined on totally non-comparable scales. For example, in DV2019, the changes in the parameter estimates directly influence the main outcome variable, namely, dispersion in average revenue products of capital ($\Delta \sigma^2_{arpk}$ using their notation); however, they are non-comparable across different components.
One common practice in economics is to apply the principle of ceteris paribus, measuring the effect of a component of interest by turning it on (or changing its value), while turning off all the others (or fixing all the others at constant values). For example, Table 3 in DV2019, which is reproduced here as Table (ref), reports the relative contributions of adjustment costs ($\xi$), uncertainty ($V$), and other factors ($\gamma$, $\sigma_\varepsilon^2$, and $\sigma_\chi^2$) under the assumption that only the factor of interest is operational. This type of evaluation has been popular but the relative contributions do not necessarily sum to one, as acknowledged by DV2019.
The ceteris paribus exercise in Table (ref) is applied to each country, thus explaining how much individual factors affect each country's capital misallocation. Their method does not directly compare the United States with China; nor does it account for the impact of replacing a particular factor from China to the United States. As an alternative, we propose a different decomposition method based on the group Shapley value. Our framework needs two sets of parameters to compare (e.g., China and the United States in Table (ref)). Our application of the group Shapley value decomposition generates unique additive contributions for the changes between the two sets of parameters, thereby providing a coherent method for quantitatively evaluating the importance of different components of the model. Decomposition results are summarized in a table, where their relative shares sum to one, allowing for ranking of importance by magnitude.
Generally speaking, there are two modes of sensitivity analysis: global sensitivity analysis Owen:2014 and local sensitivity analysis AGS2017. Our application of the Shapley value is more closely related to the concept of global sensitivity analysis. Although our decomposition may resemble a “derivative” to some readers, unlike a regular derivative, the Shapley value can be defined for a non-differentiable “utility” function and our setting is not limited to variance decomposition or the GMM framework.
Although Shapley values are widely used across different disciplines, to the best of our knowledge, they have not been adopted to interpret counterfactual simulations in structural economic models. One might wonder in what sense it would be useful to adopt the framework of the Shapley value in this context. It is well known that the Shapley value is the unique solution that satisfies the four cooperative game theory axioms: Linearity, Dummy, Symmetry, and Efficiency moulin2004fair. When analyzing the impacts of changes in the underlying parameters of a highly nonlinear and complex structural model, where the parameters can be partitioned into multiple groups, it is useful to assign importance of each group to explain the changes in the outcome of the counterfactual simulation. To evaluate these contributions, we may prefer the following properties: separable additive contributions of each group (i.e., Linearity); zero contribution if a change in one group of parameters has no impact at all (i.e., Dummy); identical contributions if two different groups produce identical marginal impacts (i.e., Symmetry); and a sum of all additive contributions that equals the total changes in the outcome (i.e., Efficiency). In other words, if we wish to retain all four axiomatic properties in a decomposition exercise, our proposed method is the only solution that satisfies them. Thus, we regard our approach as an attractive decomposition method based on a strong theoretical foundation. Furthermore, from a practical perspective, we believe our method will help practitioners produce an importance table that is as easily interpretable as a regression table.
The remainder of the paper is organized as follows. In Section (ref), we provide the definition of the group Shapley value; in Section (ref), we show that the group Shapley value can be characterized as a solution to a constrained weighted least squares problem. Additionally, we consider scenarios where some input for the group Shapley value is missing and propose robust decomposition methods to address this issue. In Section (ref), we describe a canonical use of the Shapley value in the machine learning community. In Section (ref), we provide a novel use of the group Shapley value in order to quantitatively evaluate the different components in counterfactual simulations that are generated by structural models, using a simple Roy model \`{a} la honore2017poor. In Section (ref), we demonstrate the usefulness of our framework by revisiting DV2019 and offering new perspectives on contributing factors for capital misallocation. Specifically, we offer an output table for Shapley value decomposition where we can rank the importance of parameter changes by magnitude. In Section (ref), we revisit CGT2016 and use its setting to illustrate the value of our decomposition methods when some input for the group Shapley value is missing. We show that our method is effective in providing robust estimates, even when utility computations are costly. Section (ref) provides the concluding remarks and Appendix (ref) gives the proof omitted in the main text.
We start with gul1989bargaining's generalized Shapley value. Let $P := \{ 1, 2, \ldots, p \}$ denote a set with $p$ elements. Consider a function $g: 2^P \mapsto \mathbb{R}$ that satisfies $g(\emptyset) = 0$. Let $\mathcal{P}$ denote the set of all possible partitions of $P$. For each $\Pi \in \mathcal{P}$ and $M \in \Pi$, gul1989bargaining defines the generalized Shapley value by
Here, we use Greek capital letters (e.g., $\Pi$ and $\Psi$) to denote a partition of $P$ and Roman capital letters (e.g., $M$ and $A$) to denote an element in a partition. Furthermore, $|\Psi|$ denotes the number of elements in partition $\Psi$. For ease of notation, we drop the subscript $g$ and use $\phi (M, \Pi)$ to denote $\phi_g (M, \Pi)$ when the context is clear. Note that the Shapley value is $\phi (\{i\}, \{ \{j\}: j \in P \})$ with $M = \{i\}$ and $\Pi = \{ \{j\}: j \in P \}$. The generalized Shapley value can also be written as
Two expressions (ref) and (ref) are identical; (ref) is expressed in terms of the marginal changes in the utility from subtracting $M$, whereas (ref) is given via the marginal changes in the utility from adding $M$. Instead of calling $\{ \phi (M, \Pi): M \in \Pi \}$ the generalized Shapley value, we may term it the group Shapley value because it resembles group lasso in variable selection yuan2006model and it may better describe our application of the generalized Shapley value. In many economic applications, it would be reasonable to assume a known group structure. See Section (ref) for an example of the group structure in a simple Roy model.
The original Shapley value provides a principled approach to allocating the total utility to individual elements by delivering a unique valuation that satisfies the fairness axioms in cooperative game theory ShapleyValue. The fairness axioms mathematically describe desired conditions for a valuation as a functional of a function $g$, namely Linearity, Dummy, Symmetry, and Efficiency. Extending this original axiomatic characterization, we present a set of axioms for the generalized Shapley value.
With these axioms, the following proposition is directly derived from Section 5.5 of moulin2004fair.
A generalized Shapley value gives a unique way to attribute $g(P)$ to individual groups $M \in \Pi$ while satisfying the four axioms above. In other words, no other evaluation function satisfies one or more of these axioms.
The Shapley value has multiple equivalent formulations. For example, charnes1988extremal, LundbergLee, and aas2021explaining express the Shapley value as the optimal solution to a constrained quadratic minimization problem, namely a constrained weighted least squares estimator. The following proposition is a straightforward application of Theorem 3 in charnes1988extremal to our setting.
The proof follows steps similar to those used in charnes1988extremal. However, they did not specifically consider the case of the group Shapley value. For the purpose of self-containment, we provide details of the proof in Appendix (ref).
While the group Shapley value can be expressed as closed-forms as shown in (ref) and (ref), Proposition (ref) yields a new expression because the optimization problem is a weighted linear regression problem with linear constraints. To be more specific, we consider a vector $(\Psi_1, \dots, \Psi_{2^{|\Pi|}-2})$ of subsets of $P$ where $\Psi_i$'s are possible unions of $|\Pi|$'s subset except $\emptyset$ and $P$. We set $\bm{g} = (g(\Psi_1), \dots, g(\Psi_{2^{|\Pi|}-2}))$ and denote a diagonal matrix whose $i$-th diagonal element is $k(\Pi, \Psi_i)$ by $\bm{K}$. Each $\Psi_i$ can be uniquely represented by a $|\Pi|$-dimensional binary vector, and we denote a binary matrix whose $i$-th row represents $\Psi_i$ by $\bm{D}$. Then, the optimization problem in Proposition (ref) can be re-expressed as:
subject to
where $\mathds{1}_{|\Pi|}$ is a $|\Pi|$ dimensional vector of ones. Moreover, the solution of this constraint linear regression problem is given as follows:
where $A = \bm{D}^\top \bm{K} \bm{D}$ and $b = \bm{D}^\top \bm{K} \bm{g}$.
Suppose that $P = \{1,2,\ldots,5 \}$ and $\Pi = \{ \{1 \}, \{2,3\}, \{4,5\} \}$. As an example, the group Shapley value with $M = \{1\}$ is written as
To write the least squares problem, define
In this example, $k ( \Pi, \Psi ) = 1$ for any $\Psi \subsetneqq \Pi$, implying that $\bm{K}$ is the identity matrix. Thus, $\bm{\phi}$ minimizes
subject to $$ \phi_{ \{1 \} } + \phi_{ \{2, 3 \} } + \phi_{ \{4, 5 \} } = g(P). $$ This is simply a constrained least squares problem, which can be solved easily as shown in (ref). However, when $|\Pi|$ gets large, the number of rows of $\bm{D}$ increases very rapidly: $2^{|\Pi|}-2$. In the literature LundbergLee, this difficulty is avoided by sampling the rows of $\bm{D}$ according to the probability distribution induced by the Shapley kernel weights $k ( \Pi, \Psi )$. See Section (ref) for details.
In practice, it is often challenging to observe every element of the utility vector $\bm{g} = (g(\Psi_1), \dots, g(\Psi_{2^{|\Pi|}-2}))$, especially when the computational cost of each $g$ is expensive. For instance, CGT2016, who analyze the contribution of firing cost, tariff rate, and iceberg trade cost to aggregate statistics in a ceteris paribus manner, do not provide the entire utility values. We will revisit this example in Section (ref). In such situations, the Shapely value in (ref) cannot be calculated. To address this problem, we propose Shapley bounds and Shapley Minimum Norm Solution that infer the Shapley values under user-specified linear constraints.
For $r \in \mathbb{N}$, $A_{\mathrm{const}} \in \mathbb{R}^{r \times |\bm{g}|}$, and $b_{\mathrm{const}} \in \mathbb{R}^r$, we suppose the utility vector $\bm{g}$ satisfies the linear constraints $A_{\mathrm{const}} \bm{g} \leq b_{\mathrm{const}}$. These linear constraints allow researchers to formulate their domain expertise and infer a missing part of $\bm{g}$. Since the Shapley value $\phi$ can be seen as a linear function of $\bm{g}$, we denote it by $\phi(\bm{g})$. Then, for $j \in |\Pi|$, Shapley Upper Bounds (SUB) for the $j$-th element of $\phi$ is defined as follows.
where $c_j \in \{0,1\}^{|\Pi|}$ is the $j$-th canonical basis in $\mathbb{R}^{|\Pi|}$ and $\mathcal{G} := \{ \bm{h} \in \mathbb{R}^{|\bm{g}|} : \bm{h}_i = \bm{g}_i $ if $\bm{g}_i$ is observed $\}$. That is, SUB is the maximum possible Shapely value that satisfies given linear constraints. Similarly, we define Shapley Lower Bounds (SLB) as follows.
While SUB and SLB can provide informative bounds for the Shapley value, it does not satisfy the efficiency axiom in general, and its value does not sum to $g(P)$. To address this problem, we propose the Shapley Minimum Norm Solution (SMNS) that solves the following optimization problem.
where $\mathds{1}_{N}$ is the $N$-dimensional one vector and $\lVert v \rVert_2$ is the $\ell^2$-norm of a vector $v$. Here, $\frac{g(P)}{N} \mathds{1}_{N}$ can be seen as the best guess of the Shapley value when the practitioner believes the contribution of every factor is identical, and SMNS infers the Shapely value that is not very different from $\frac{g(P)}{N} \mathds{1}_{N}$ while satisfying the linear constraints.
Note that the solution ($\tilde{g}$) to any of the optimization problems above is a utility. The resulting Shapley value is then obtained by plugging solution $\tilde{g}$ into equation (ref). Either reporting the bounds [SLB, SUB] or using SMNS can be viewed as a reasonable alternative in the presence of missing input. We use both approaches in Section (ref).
The Shapley value and its variations have been deployed in various machine learning applications. As mentioned in the introduction, LundbergLee advocate its use in the model interpretation problem (also known as explainable artificial intelligence), where it is employed to attribute a model's prediction to features. ghorbani2019data extend its utility to the data valuation problem, introducing a method to quantify the impact of individual data. pmlr-v151-kwon22a relax the efficiency axiom and use a semivalue in evaluating data values. wang2020principled leverage the Shapley value to quantify the contribution of a local model to a global model in federated learning settings. In addition, the Shapley value has been proposed as a fair valuation method in data marketplaces tian2022private. For an in-depth exploration of the literature on the Shapley value and its applications in machine learning, we refer the readers to the comprehensive review by rozemberczki2022shapley. In this section, we provide a brief description of an application of Shapley value in the context of explainable artificial intelligence.
For a vector of covariates $x$, let $f(x)$ denote a prediction model of a real-valued outcome of interest using $x$. For $A \subset P$, decompose $x = (x_A, x_{P \setminus A})$. Here, $x_A$ denotes a subvector of $x$ whose indices are in $A$. For example, if $x = (x_1, x_2, x_3)$ and $A = \{ 1, 2 \}$, $x_A = (x_1, x_2)$. For each $x$ and $A$, define
where the upper case refers to a random vector. Note that $g_{x} ( A )$ is a function of $x_A$.
We end this section by commenting that there have been some debates on causal interpretations of the Shapley values for explainable AI, in particular, when the covariates are dependent on each other. See, e.g., LundbergLee, aas2021explaining, janzing2020feature, and eskes2020causal among others. This debate is irrelevant for our application of Shapley values to counterfactual simulations.
We propose to use Shapley values for evaluating the different components in counterfactual simulations that are generated by structural models.
To be explicit, we consider a simple Roy model used in honore2017poor. In their setup, there are two sectors: $s \in \{1, 2\}$. Worker $i$ earns sector-specific income $w_{si1}$ in period 1 in the following form:
where $x_{si}$ is a vector of sector-specific human capital, $\beta_s$ is a vector of sector-specific parameters, and $\varepsilon_{si1}$ is a sector-specific unobserved random variable in period 1. Sector-specific income $w_{si2}$ in period 2 is generated by
where $d_{i1}$ is the sector chosen in period 1, $\gamma_s$ is a sector-specific parameter that represents the premium of staying in the same sector in period 2, and $\varepsilon_{si1}$ is a sector-specific unobserved random variable in period 2. The unobserved random variables $(\varepsilon_{1it}, \varepsilon_{2it})$ are assumed to be bivariate normally distributed with mean zeros, variances $(\sigma_1^2, \sigma_2^2)$ and correlation $\tau$, and i.i.d. over time.
Workers maximize discounted income. In time period 2, $d_{i2} = 1$ and the resulting income is $w_{i2} = w_{1i2}$ if
and $d_{i2} = 2$ and $w_{i2} = w_{2i2}$ otherwise. In time period 1, $d_{i1} = 1$ if and only if
where $\rho$ is discount factor.
Let $\theta = (\beta_1, \beta_2, \gamma_1, \gamma_2, \sigma_1^2, \sigma_2^2, \tau, \rho)$ denote the vector of all structural parameters. Let $\theta^b$ denote the benchmark vector of parameter values, for example, estimated values, and $\theta^c$ a counterfactual vector of parameter values. If we write $(\theta^{c}_{A}, \theta^b_{P \setminus A})$, we mean the combination of parameter values such that the $A$ subset of $\theta$ is set at the counterfactual values, while setting the $P \setminus A$ subset of $\theta$ at the benchmark values. Let $W = W(\theta)$ denote the observed random variables generated by the model given $\theta$. In the Roy model example, we have that $W = (w_{i1}, w_{i2}, d_{i1}, d_{i2}, x_{i1}, x_{i2})$. Let $f(\theta)$ be the real-valued quantity of interest in a counterfactual simulation that can be computed by simulating $W$ given $\theta$. For example, $f(\theta)$ be a measure of the changes in income inequality from period 1 to period 2. Finally, we let
Suppose that $\tau = 0$ and $\rho = 0.95$ are fixed, as in honore2017poor. Then, $P = \{ \beta_1, \beta_2, \gamma_1, \gamma_2, \sigma_1^2, \sigma_2^2 \}$. Suppose that $$ \Pi = \{ \{\beta_1, \beta_2\}, \{\gamma_1, \gamma_2\}, \{\sigma_1^2, \sigma_2^2\} \}. $$ That is, we consider the three groups: the coefficients for sector-specific human capital, the sector-specific benefits of staying in the same sector, and sector-specific variances.
As an example, we set the benchmark vector of parameter values at $\theta^b$ at the values used in honore2017poor. That is, \[ \beta_1^b = (1,1)^\top, \beta_2^b = (0.5, 1)^\top, \gamma_1^b = 0, \gamma_2^b = 1, (\sigma_1^2)^b = 2, (\sigma_2^2)^b = 3. \] Suppose that we set the counterfactual vector $\theta^c$ of parameter values at \[ \beta_1^c = (1,2)^\top, \beta_2^c = (0.5, 2)^\top, \gamma_1^c = 0, \gamma_2^c = 2, (\sigma_1^2)^c = 2, (\sigma_2^2)^c = 6. \] If we focus on income inequality, we may consider:
As an example, we consider the change in overall inequality, measured by the $0.9-0.1$ quantile difference, from period 1 to period 2, that is: $$ f(\theta) = h_{\mathrm{overall-ineq}, 2}(\theta; \tau_1 = 0.1, \tau_2 = 0.9) - h_{\mathrm{overall-ineq}, 1}(\theta; \tau_1 = 0.1, \tau_2 = 0.9). $$ This example of $f(\theta)$ is not generally differentiable with respect to $\theta$. In consequence, the resulting $g(A)$ can not be viewed as the standard derivative; however, our approach provides how to quantify the importance of changes in the parameters. Additionally, we offer a global sensitivity analysis, as opposed to a local sensitivity analysis.
As in Section (ref), we have that
and $k ( \Pi, \Psi ) = 1$ for any $\Psi \subsetneqq \Pi$. Thus, $\bm{\phi}$ minimizes
subject to $$ \phi_{\{\beta_1, \beta_2\} } + \phi_{ \{\gamma_1, \gamma_2\} } + \phi_{ \{\sigma_1^2, \sigma_2^2\} } = g( \{ \beta_1, \beta_2, \gamma_1, \gamma_2, \sigma_1^2, \sigma_2^2 \}). $$ Again, this is a constrained least squares problem, which can be solved easily. The values of $\bm{g}$ are obtained by Monte Carlo simulation with the sample size of $10^7$.
Table (ref) gives the counterfactual simulation results. It can be seen that all three components contribute to the increase in inequality. However, $\phi_{ \{\gamma_1, \gamma_2\} }$ matters most in this example. Recall that $(\gamma_1^b = 0, \gamma_2^b = 1)$ and $(\gamma_1^c = 0, \gamma_2^c = 2)$. In other words, sector 2 specific return to the log wages doubled in period 2, which explains the largest increase in the change in the overall inequality.
In short, we demonstrate that although a measure of income equality from a Roy model can be highly nonlinear in inputs, the group Shapley value generates unique additive contributions from $\Pi$. From a practical perspective, we find that Table (ref) is just as effortless to interpret as a regression table. Generally speaking, the group Shapley value provides a straightforward method to interpret counterfactual simulation results when a structural model simulates the output.
In this section, we revisit the work of DV2019, who developed a method to disentangle sources of capital misallocation---specifically, the dispersion in average revenue products of capital (arpk) using observable data on value-added and inputs from both China and United States.
For firm $i$ in period $t$, DV2019 define the average revenue product of capital (arpk) as $arpk_{it} = va_{it} - k_{it}$, where $va_{it}$ is the log value added and $k_{it}$ is the capital stock. The parameter of interest is the variance $\sigma_{aprk}^2$ of $arpk_{it}$. To simulate this quantity, we need to specify a variety of parameters. First of all, we need to choose the value of $\alpha$, which determines the curvature of operating profits. There are two important dynamics in DV2019. One is associated with the log productivity (denoted by $a_{it}$):
where $\rho$ is the autoregressive parameter and $\sigma_\mu^2$ is the variance of the innovation term for productivity (see equation (5) in DV2019). The other dynamics involves log distortion (denoted by $\tau_{it}$):
where $\gamma$ determines how the distortion co-moves with productivity, $\sigma_\varepsilon^2$ is the variance of i.i.d. shocks and $\sigma_\chi^2$ is the variance of time-invariant firm-specific shocks (see equation (6) in DV2019). The distribution of future productivity ($a_{it+1}$) conditional on the firm’s information set ($\mathcal{I}_{it}$) in period $t$ follows a normal distribution with the posterior mean $E_{it}[a_{it+1}]$ and the posterior variance $V$. Finally, investment is subject to quadratic adjustment costs and the severity of the adjustment cost is parameterized via $\xi$. In summary, there are 8 key parameters: $\alpha$, $\rho$, $\sigma_\mu^2$, $\xi$, $V$, $\gamma$, $\sigma_\varepsilon^2$, and $\sigma_\chi^2$ (see Table 1 in DV2019).
Recall that Table (ref) reproduces the parameter estimates from DV2019. The value of $\alpha$ is slightly larger for China. The level of persistence for productivity is similar between the two countries, whereas the variance of productivity is larger in China than in U.S. The adjustment costs are much higher in the U.S. but uncertainty is larger in China. Note that $\gamma$ is more negative for China, indicating that the extent to which the distortion discourages investment by firms with higher productivity is larger in China than in U.S. Finally, in China, the variance of the transitory shocks is smaller but that of the permanent shocks is larger.
DV2019 report the relative contributions of adjustment costs, uncertainty, and other factors to the arpk dispersion under the assumption that only the factor of interest is operational. To be more specific, we let $\theta = (\alpha, \rho, \sigma_\mu^2, \xi, V, \gamma, \sigma_\varepsilon^2,\sigma_\chi^2)$ and $f(\theta) = \sigma_{aprk}^2$. Assuming $f(\mathbf{0})=0$, DV2019 estimate the contribution of each factor by $\phi_{\mathrm{DV}} (A) := f(\theta_A, \mathbf{0}_{P \backslash A}) - f(\mathbf{0})$ where $\theta$ is the parameter estimates in Table (ref). The contribution estimation is based on the principle of ceteris paribus, measuring the impact of each factor from a hypothetical void setting $f(\mathbf{0})$. However, while it intends to quantify the change when a factor of interest is added, $\phi_{\mathrm{DV}}$ can yield an unrealistic or imprecise assessment of individual impacts because setting some entries of $\theta$ at zeros results in degenerate dynamics. For instance, $V$ is the posterior variance, which needs to be strictly greater than zero. Furthermore, $\phi_{\mathrm{DV}}$ does not satisfy the Efficiency axiom as it is a weighted sum of utility changes, which is referred to as a semivalue semivalue_without_efficiency. This is why the relative contributions do not necessarily sum to $1$, as acknowledged by DV2019. We address these two issues in the following subsection.
In this subsection, we carry out a new counterfactual exercise that is different from that of DV2019. We set the benchmark vector ($\theta_b$) of parameter values using estimates from U.S. and the counterfactual vector ($\theta_c$) using estimates from China. As each parameter represents a distinct component of the model developed in DV2019, we consider singleton groups:
In the data, $\sigma_{aprk}^2$ was 0.92 for China and 0.45 for U.S (see Table 2 in DV2019). Thus, the degrees of capital misallocation double by changing the U.S. parameter estimates to those based on Chinese data. Our Shapley value decomposition will indicate the extent to which each parameter contributes to this increase.
Table (ref) reports the resulting Shapley value decomposition. First of all, note that there are positive and negative values for $\phi_{ \{ \cdot\} }$ because some changes in the parameter values imply increases in the arpk dispersion but other changes indicate decreases. However, the sum of the Shapley shares is 1 by design.
We now examine each row. The increase of $\alpha$ from $0.62$ to $0.71$ is associated with $\phi_{ \{\alpha\} } = 0.0073$, implying the share of $0.016$, which is quite small. The persistence parameter $\rho$ gets slightly smaller moving from the U.S. to China (that is, from $0.93$ to $0.91$). However, the resulting Shapley value is $-0.0599$ with a relatively large share of $-0.129$. We can interpret that a small decrease in the autoregressive parameter in the log productivity is associated with a relatively large reduction in the capital misallocation. The variance of the innovation term for productivity increases more substantially from 0.08 to 0.15, resulting in the second largest share in the Shapley value decomposition. It is interesting to note that the change in the adjustment costs looks more visible in the sense that $\xi$ decreases by a factor of 10, while the implied share is only $-0.126$. This shows that the Shapley value decomposition is useful to compare changes in parameters on the same scale. The increase in uncertainty is associated with the share of $0.095$. The largest Shapley share of $0.542$ comes from $\gamma$, which is negative for both countries. A more negative $\gamma$ means that the distortion plays a larger role in disincentivizing investment by more productive firms. The reduction in the variance of the transitory shocks is small and as a result, the Shapley share is small as well. The increase in the variance of the permanent shocks is quite large, ranking as the third most important factor in terms of the Shapley share. In short, the two components ($\gamma$ and $\sigma_\chi^2$) of the distortion as well as the variance of the productivity shocks are the three most important factors affecting the increase in the arpk dispersion. Overall, our Shapley value analysis reconfirms the quantitative findings in DV2019, while providing novel perspectives.
CGT2016 examine decreases in trade barriers, tariffs, and firing costs in an open economy, using establishment-level data from Colombia. Their counterfactual experiments suggest that Colombia's integration into global product markets raised its national income but also increased unemployment, wage inequality, and firm-level volatility. Table 4 in CGT2016 provides the results of their counterfactual experiments. Suppose that we are interested in comparing the last column “Reforms and globalization” with the first column “Baseline” and construct Shapley value decomposition among three elements: firing cost ($c_f$), tariff rate ($\tau_a$), and iceberg trade cost ($\tau_c$). Then, we have the same structure as the toy example in Section (ref). That is, we have that
For self-containment, we re-produce the parameter values in Panel A of Table (ref). In Panel A, the benchmark (“Baseline”) parameters are $(c_f^b, \tau_a^b, \tau_c^b) = (0.60, 1.21, 2.50)$, whereas the counterfactual (“Reforms and globalization”) parameters are $(c_f^c, \tau_a^c, \tau_c^c) = (0.30, 1.11, 2.19)$. Regarding $g(\cdot)$, we consider the aggregates reported in Table 4 in CGT2016, which are reproduced in Panel B of Table (ref). As before, $g(\cdot)$ is obtained by taking the difference between the counterfactual (“Reforms and globalization”) quantities and the benchmark (“Baseline”) quantities, where the latter quantities are normalized to be one.
It can be seen from Table (ref) that the four middle columns are the input for optimization (i.e. elements of $\bm{g}$): using the notation above, we have values for $g ( \{c_f\} )$, $g ( \{\tau_a\} )$, $g ( \{\tau_c\} )$, and $g (\{c_f\} \cup \{\tau_a\} )$. They are called “Labor”, “Tariff”, “Iceberg”, and “Reforms” in the table. However, the last two elements of $\bm{g}$ are missing: namely, $g (\{c_f\} \cup \{\tau_c\} )$ and $g ( \{\tau_a\} \cup \{\tau_c\} )$. It would be ideal to complete the missing elements in order to construct Shapley value decomposition. However, one interesting research question is what one can do if it is expensive or impossible to carry out further experiments to obtain the missing elements for $\bm{g}$.
In view of the fact that each of the parameters is reduced to favor globalization, we consider the following set of restrictions on the missing elements: for $g (\{c_f\} \cup \{\tau_c\} )$,
and for $g (\{d_a\} \cup \{\tau_c\} )$,
In addition, suppose that we have known upper and lower bounds on $g (\{c_f\} \cup \{\tau_c\} )$ and $g (\{d_a\} \cup \{\tau_c\} )$:
where $g_{\min}$ and $g_{\max}$ are predetermined constants such that $g_{\min} = 0.5$ and $g_{\max} = 1.5$. Combining all the constraints above in (ref)-(ref) forms the linear constraints $A_{\mathrm{const}} \bm{g} \leq b_{\mathrm{const}}$ in Section (ref).
Table (ref) and Figure (ref) present the empirical results for the Shapley bounds and the Shapley minimum norm solutions using the methodology developed in Section (ref). As can be seen from Table (ref), the revenue share of exports increases by 1.5 from the benchmark quantity to the counterfactual quantity. The largest portion of this increase is explained by the decrease in the iceberg trade cost (with a lower bound of 0.98). The reduction in the tariff rate appears to explain slightly more than the reduction in firing costs, although their bounds overlap, making it difficult to definitively rank these two factors. Whether considering the Shapley bounds or the Shapley minimum norm solutions, firing cost reductions are linked to decreases in the exit rate, vacancy filling rate, and unemployment rate. Conversely, reductions in the tariff rate and iceberg trade cost are associated with increases in these rates. The decomposition result for the mass of firms is ambiguous: for all three factors, the minimum norm solutions are negative, but the bounds include zero. All the inequality measures---Std. wages (firms), Std. wages (workers), Std. $J$ (firms), and Std. $J$ (workers)---increase with the reduction in each factor. However, there is no definite ranking among the three factors in terms of their contribution to increases in inequality. Finally, real income increases by 0.12 from the benchmark quantity to the counterfactual quantity, primarily explained by the reduction in the iceberg trade cost (with a lower bound of 0.1). The Shapley bounds are rather wide for both the firing cost and the tariff rate.
Overall, the empirical results suggest that the reduction in iceberg trade costs is the most significant factor for real income growth. However, it is less clear which factor is the primary driver of increases in inequality. Returning to our research question, this example demonstrates that it is feasible to perform a coherent Shapley value decomposition even when it is impossible to conduct further experiments to obtain missing elements for $\bm{g}$. Thus, our methodology holds potential for real-world applications that involve costly counterfactual simulations.
Although Shapley values are widely used across various disciplines, their application to interpreting counterfactual simulations in structural economic models has not yet been explored. We have demonstrated that the group Shapley value provides a natural framework for quantifying the importance of parameter changes in complex models. By preserving the axiomatic properties of the Shapley value, our method offers a well-grounded decomposition approach for explaining outcomes in counterfactual simulations. Practically, it enables researchers to generate importance tables that are as intuitive and accessible as regression tables. In this sense, our use of the Shapley value could pave the way for “interpretable structural economics,” much like LundbergLee's introduction of Shapley values to the machine learning community has advanced explainable artificial intelligence.
However, our method has some limitations. First, we focus on comparing two sets of parameters. Extending our approach to cases involving more than two sets of parameter values (e.g., the United States, China, and South Korea, using the example in the introduction) would be valuable. Second, our numerical examples are limited to small-scale optimization problems, where computational complexity is less of a concern. For larger-scale problems, computational challenges become more significant, and developing strategies to address them—such as devising effective sampling techniques and accounting for sampling errors—will be crucial. These are exciting directions for future research.