EconBase
← Back to paper

Group Shapley Value and Counterfactual Simulations in a Structural Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

66,228 characters · 17 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Group Shapley Value and Counterfactual Simulations in a Structural Model

abstractWe propose a variant of the Shapley value, the group Shapley value, to interpret counterfactual simulations in structural economic models by quantifying the importance of different components. Our framework compares two sets of parameters, partitioned into multiple groups, and applying group Shapley value decomposition yields unique additive contributions to the changes between these sets. The relative contributions sum to one, enabling us to generate an importance table that is as easily interpretable as a regression table. The group Shapley value can be characterized as the solution to a constrained weighted least squares problem. Using this property, we develop robust decomposition methods to address scenarios where inputs for the group Shapley value are missing. We first apply our methodology to a simple Roy model and then illustrate its usefulness by revisiting two published papers. \\ \\ Keywords: Shapley value, explainable artificial intelligence, structural model

\doublespacing

Introduction

The Shapley value, first introduced by ShapleyValue, is a foundational solution concept in cooperative game theory, with desirable properties supported by axiomatic approaches. It is also sustainable in non-cooperative settings gul1989bargaining. Shorrocks:2013 popularized its use in empirical applications, primarily for decomposing $R^2$ within linear regression models. Many studies have used the Shapley-Shorrocks decomposition: see, e.g., allen2014information, allen2014trade, dustmann2014out and henderson2018global among others. Most notably since the work of LundbergLee, Shapley values have become one of the main tools of explainable artificial intelligence (AI). They are likewise recognized as useful interpretational tools in other disciplines, such as global sensitivity analysis Owen:2014.

We propose to make a use of a particular variant which we call the group Shapley value. Specifically, we use it to understand the outcomes of counterfactual simulations in a structural economic model by quantifying the importance of various elements embedded within the model. In a structural model, the crucial task involves estimating or calibrating the underlying model parameters. Consequently, any changes to these parameter values will directly influence the outcomes of the structural model. Our motivating example is from DV2019, who developed a method to disentangle sources of capital misallocation using data from both China and the United States. Their model includes various components such as shocks to productivity, investment adjustment costs, uncertainty, and so on. They are specified via a number of parameters that are estimated or calibrated: Table (ref) reproduces the parameter estimates from DV2019.

table[table omitted — 819 chars of source]

In general, it would be challenging to compare the relative importance across different components of the parameters because they tend to be defined on totally non-comparable scales. For example, in DV2019, the changes in the parameter estimates directly influence the main outcome variable, namely, dispersion in average revenue products of capital ($\Delta \sigma^2_{arpk}$ using their notation); however, they are non-comparable across different components.

table[table omitted — 457 chars of source]

One common practice in economics is to apply the principle of ceteris paribus, measuring the effect of a component of interest by turning it on (or changing its value), while turning off all the others (or fixing all the others at constant values). For example, Table 3 in DV2019, which is reproduced here as Table (ref), reports the relative contributions of adjustment costs ($\xi$), uncertainty ($V$), and other factors ($\gamma$, $\sigma_\varepsilon^2$, and $\sigma_\chi^2$) under the assumption that only the factor of interest is operational. This type of evaluation has been popular but the relative contributions do not necessarily sum to one, as acknowledged by DV2019.

The ceteris paribus exercise in Table (ref) is applied to each country, thus explaining how much individual factors affect each country's capital misallocation. Their method does not directly compare the United States with China; nor does it account for the impact of replacing a particular factor from China to the United States. As an alternative, we propose a different decomposition method based on the group Shapley value. Our framework needs two sets of parameters to compare (e.g., China and the United States in Table (ref)). Our application of the group Shapley value decomposition generates unique additive contributions for the changes between the two sets of parameters, thereby providing a coherent method for quantitatively evaluating the importance of different components of the model. Decomposition results are summarized in a table, where their relative shares sum to one, allowing for ranking of importance by magnitude.

Generally speaking, there are two modes of sensitivity analysis: global sensitivity analysis Owen:2014 and local sensitivity analysis AGS2017. Our application of the Shapley value is more closely related to the concept of global sensitivity analysis. Although our decomposition may resemble a “derivative” to some readers, unlike a regular derivative, the Shapley value can be defined for a non-differentiable “utility” function and our setting is not limited to variance decomposition or the GMM framework.

Although Shapley values are widely used across different disciplines, to the best of our knowledge, they have not been adopted to interpret counterfactual simulations in structural economic models. One might wonder in what sense it would be useful to adopt the framework of the Shapley value in this context. It is well known that the Shapley value is the unique solution that satisfies the four cooperative game theory axioms: Linearity, Dummy, Symmetry, and Efficiency moulin2004fair. When analyzing the impacts of changes in the underlying parameters of a highly nonlinear and complex structural model, where the parameters can be partitioned into multiple groups, it is useful to assign importance of each group to explain the changes in the outcome of the counterfactual simulation. To evaluate these contributions, we may prefer the following properties: separable additive contributions of each group (i.e., Linearity); zero contribution if a change in one group of parameters has no impact at all (i.e., Dummy); identical contributions if two different groups produce identical marginal impacts (i.e., Symmetry); and a sum of all additive contributions that equals the total changes in the outcome (i.e., Efficiency). In other words, if we wish to retain all four axiomatic properties in a decomposition exercise, our proposed method is the only solution that satisfies them. Thus, we regard our approach as an attractive decomposition method based on a strong theoretical foundation. Furthermore, from a practical perspective, we believe our method will help practitioners produce an importance table that is as easily interpretable as a regression table.

The remainder of the paper is organized as follows. In Section (ref), we provide the definition of the group Shapley value; in Section (ref), we show that the group Shapley value can be characterized as a solution to a constrained weighted least squares problem. Additionally, we consider scenarios where some input for the group Shapley value is missing and propose robust decomposition methods to address this issue. In Section (ref), we describe a canonical use of the Shapley value in the machine learning community. In Section (ref), we provide a novel use of the group Shapley value in order to quantitatively evaluate the different components in counterfactual simulations that are generated by structural models, using a simple Roy model \`{a} la honore2017poor. In Section (ref), we demonstrate the usefulness of our framework by revisiting DV2019 and offering new perspectives on contributing factors for capital misallocation. Specifically, we offer an output table for Shapley value decomposition where we can rank the importance of parameter changes by magnitude. In Section (ref), we revisit CGT2016 and use its setting to illustrate the value of our decomposition methods when some input for the group Shapley value is missing. We show that our method is effective in providing robust estimates, even when utility computations are costly. Section (ref) provides the concluding remarks and Appendix (ref) gives the proof omitted in the main text.

Group Shapley Value

We start with gul1989bargaining's generalized Shapley value. Let $P := \{ 1, 2, \ldots, p \}$ denote a set with $p$ elements. Consider a function $g: 2^P \mapsto \mathbb{R}$ that satisfies $g(\emptyset) = 0$. Let $\mathcal{P}$ denote the set of all possible partitions of $P$. For each $\Pi \in \mathcal{P}$ and $M \in \Pi$, gul1989bargaining defines the generalized Shapley value by

align[align omitted — 572 chars of source]

Here, we use Greek capital letters (e.g., $\Pi$ and $\Psi$) to denote a partition of $P$ and Roman capital letters (e.g., $M$ and $A$) to denote an element in a partition. Furthermore, $|\Psi|$ denotes the number of elements in partition $\Psi$. For ease of notation, we drop the subscript $g$ and use $\phi (M, \Pi)$ to denote $\phi_g (M, \Pi)$ when the context is clear. Note that the Shapley value is $\phi (\{i\}, \{ \{j\}: j \in P \})$ with $M = \{i\}$ and $\Pi = \{ \{j\}: j \in P \}$. The generalized Shapley value can also be written as

align[align omitted — 239 chars of source]

Two expressions (ref) and (ref) are identical; (ref) is expressed in terms of the marginal changes in the utility from subtracting $M$, whereas (ref) is given via the marginal changes in the utility from adding $M$. Instead of calling $\{ \phi (M, \Pi): M \in \Pi \}$ the generalized Shapley value, we may term it the group Shapley value because it resembles group lasso in variable selection yuan2006model and it may better describe our application of the generalized Shapley value. In many economic applications, it would be reasonable to assume a known group structure. See Section (ref) for an example of the group structure in a simple Roy model.

The original Shapley value provides a principled approach to allocating the total utility to individual elements by delivering a unique valuation that satisfies the fairness axioms in cooperative game theory ShapleyValue. The fairness axioms mathematically describe desired conditions for a valuation as a functional of a function $g$, namely Linearity, Dummy, Symmetry, and Efficiency. Extending this original axiomatic characterization, we present a set of axioms for the generalized Shapley value.

itemize• Linearity: $\phi_g$ satisfies the linearity axiom if $\phi_{g_1+ g_2} = \phi_{g_1} + \phi_{g_2}$ for any set functions $g_1$ and $g_2$. • Dummy: $\phi_g$ satisfies the dummy axiom if $g(S\cup A) = g(S)$ for any $S \subseteq \Pi$ implies $\phi_g (A) = 0$ for any set function $g$. • Symmetry: $\phi_g$ satisfies the symmetry axiom if $g(S\cup A) = g(S \cup B)$ for $S \subseteq \Pi \backslash \{A, B\}$ implies $\phi_g (A) = \phi_g (B)$ for any set function $g$. • Efficiency: $\phi_g$ satisfies the efficiency axiom if $\sum_{M \in \Pi} \phi(M, \Pi) = g(\bigcup_{A \in \Pi} A)$ for any set function $g$.

With these axioms, the following proposition is directly derived from Section 5.5 of moulin2004fair.

propA generalized Shapley value is a unique function that satisfies Linearity, Dummy, Symmetry, and Efficiency axioms.

A generalized Shapley value gives a unique way to attribute $g(P)$ to individual groups $M \in \Pi$ while satisfying the four axioms above. In other words, no other evaluation function satisfies one or more of these axioms.

Group Shapley Value via Constrained Weighted Least Squares

The Shapley value has multiple equivalent formulations. For example, charnes1988extremal, LundbergLee, and aas2021explaining express the Shapley value as the optimal solution to a constrained quadratic minimization problem, namely a constrained weighted least squares estimator. The following proposition is a straightforward application of Theorem 3 in charnes1988extremal to our setting.

prop[charnes1988extremal] The group Shapley value $\{ \phi (M, \Pi): M \in \Psi \subsetneqq \Pi, M \neq \emptyset \}$ solves the following optimization problem: \begin{align*} \min_{ \{\phi_M : M \in \Psi \subsetneqq \Pi, M \neq \emptyset \} } \sum_{\Psi \subsetneqq \Pi} \left\{ g\left( \bigcup_{M \in \Psi} M \right) - \sum_{M \in \Psi} \phi_M \right\}^2 k ( \Pi, \Psi ) \; subject to \; \sum_{M \in \Pi} \phi_M = g \left( P \right), \end{align*} where \begin{align*} k ( \Pi, \Psi ) := \begin{pmatrix} |\Pi| - 2 \\ |\Psi| - 1 \end{pmatrix}^{-1}. \end{align*}

The proof follows steps similar to those used in charnes1988extremal. However, they did not specifically consider the case of the group Shapley value. For the purpose of self-containment, we provide details of the proof in Appendix (ref).

While the group Shapley value can be expressed as closed-forms as shown in (ref) and (ref), Proposition (ref) yields a new expression because the optimization problem is a weighted linear regression problem with linear constraints. To be more specific, we consider a vector $(\Psi_1, \dots, \Psi_{2^{|\Pi|}-2})$ of subsets of $P$ where $\Psi_i$'s are possible unions of $|\Pi|$'s subset except $\emptyset$ and $P$. We set $\bm{g} = (g(\Psi_1), \dots, g(\Psi_{2^{|\Pi|}-2}))$ and denote a diagonal matrix whose $i$-th diagonal element is $k(\Pi, \Psi_i)$ by $\bm{K}$. Each $\Psi_i$ can be uniquely represented by a $|\Pi|$-dimensional binary vector, and we denote a binary matrix whose $i$-th row represents $\Psi_i$ by $\bm{D}$. Then, the optimization problem in Proposition (ref) can be re-expressed as:

align*[align* omitted — 104 chars of source]

subject to

align*[align* omitted — 59 chars of source]

where $\mathds{1}_{|\Pi|}$ is a $|\Pi|$ dimensional vector of ones. Moreover, the solution of this constraint linear regression problem is given as follows:

align[align omitted — 177 chars of source]

where $A = \bm{D}^\top \bm{K} \bm{D}$ and $b = \bm{D}^\top \bm{K} \bm{g}$.

Toy Example

Suppose that $P = \{1,2,\ldots,5 \}$ and $\Pi = \{ \{1 \}, \{2,3\}, \{4,5\} \}$. As an example, the group Shapley value with $M = \{1\}$ is written as

align*[align* omitted — 630 chars of source]

To write the least squares problem, define

align*[align* omitted — 426 chars of source]

In this example, $k ( \Pi, \Psi ) = 1$ for any $\Psi \subsetneqq \Pi$, implying that $\bm{K}$ is the identity matrix. Thus, $\bm{\phi}$ minimizes

align[align omitted — 91 chars of source]

subject to $$ \phi_{ \{1 \} } + \phi_{ \{2, 3 \} } + \phi_{ \{4, 5 \} } = g(P). $$ This is simply a constrained least squares problem, which can be solved easily as shown in (ref). However, when $|\Pi|$ gets large, the number of rows of $\bm{D}$ increases very rapidly: $2^{|\Pi|}-2$. In the literature LundbergLee, this difficulty is avoided by sampling the rows of $\bm{D}$ according to the probability distribution induced by the Shapley kernel weights $k ( \Pi, \Psi )$. See Section (ref) for details.

Shapley Bounds and Shapley Minimum Norm Solutions

In practice, it is often challenging to observe every element of the utility vector $\bm{g} = (g(\Psi_1), \dots, g(\Psi_{2^{|\Pi|}-2}))$, especially when the computational cost of each $g$ is expensive. For instance, CGT2016, who analyze the contribution of firing cost, tariff rate, and iceberg trade cost to aggregate statistics in a ceteris paribus manner, do not provide the entire utility values. We will revisit this example in Section (ref). In such situations, the Shapely value in (ref) cannot be calculated. To address this problem, we propose Shapley bounds and Shapley Minimum Norm Solution that infer the Shapley values under user-specified linear constraints.

For $r \in \mathbb{N}$, $A_{\mathrm{const}} \in \mathbb{R}^{r \times |\bm{g}|}$, and $b_{\mathrm{const}} \in \mathbb{R}^r$, we suppose the utility vector $\bm{g}$ satisfies the linear constraints $A_{\mathrm{const}} \bm{g} \leq b_{\mathrm{const}}$. These linear constraints allow researchers to formulate their domain expertise and infer a missing part of $\bm{g}$. Since the Shapley value $\phi$ can be seen as a linear function of $\bm{g}$, we denote it by $\phi(\bm{g})$. Then, for $j \in |\Pi|$, Shapley Upper Bounds (SUB) for the $j$-th element of $\phi$ is defined as follows.

align*[align* omitted — 182 chars of source]

where $c_j \in \{0,1\}^{|\Pi|}$ is the $j$-th canonical basis in $\mathbb{R}^{|\Pi|}$ and $\mathcal{G} := \{ \bm{h} \in \mathbb{R}^{|\bm{g}|} : \bm{h}_i = \bm{g}_i $ if $\bm{g}_i$ is observed $\}$. That is, SUB is the maximum possible Shapely value that satisfies given linear constraints. Similarly, we define Shapley Lower Bounds (SLB) as follows.

align*[align* omitted — 181 chars of source]

While SUB and SLB can provide informative bounds for the Shapley value, it does not satisfy the efficiency axiom in general, and its value does not sum to $g(P)$. To address this problem, we propose the Shapley Minimum Norm Solution (SMNS) that solves the following optimization problem.

align*[align* omitted — 224 chars of source]

where $\mathds{1}_{N}$ is the $N$-dimensional one vector and $\lVert v \rVert_2$ is the $\ell^2$-norm of a vector $v$. Here, $\frac{g(P)}{N} \mathds{1}_{N}$ can be seen as the best guess of the Shapley value when the practitioner believes the contribution of every factor is identical, and SMNS infers the Shapely value that is not very different from $\frac{g(P)}{N} \mathds{1}_{N}$ while satisfying the linear constraints.

Note that the solution ($\tilde{g}$) to any of the optimization problems above is a utility. The resulting Shapley value is then obtained by plugging solution $\tilde{g}$ into equation (ref). Either reporting the bounds [SLB, SUB] or using SMNS can be viewed as a reasonable alternative in the presence of missing input. We use both approaches in Section (ref).

Explainable Artificial Intelligence

The Shapley value and its variations have been deployed in various machine learning applications. As mentioned in the introduction, LundbergLee advocate its use in the model interpretation problem (also known as explainable artificial intelligence), where it is employed to attribute a model's prediction to features. ghorbani2019data extend its utility to the data valuation problem, introducing a method to quantify the impact of individual data. pmlr-v151-kwon22a relax the efficiency axiom and use a semivalue in evaluating data values. wang2020principled leverage the Shapley value to quantify the contribution of a local model to a global model in federated learning settings. In addition, the Shapley value has been proposed as a fair valuation method in data marketplaces tian2022private. For an in-depth exploration of the literature on the Shapley value and its applications in machine learning, we refer the readers to the comprehensive review by rozemberczki2022shapley. In this section, we provide a brief description of an application of Shapley value in the context of explainable artificial intelligence.

For a vector of covariates $x$, let $f(x)$ denote a prediction model of a real-valued outcome of interest using $x$. For $A \subset P$, decompose $x = (x_A, x_{P \setminus A})$. Here, $x_A$ denotes a subvector of $x$ whose indices are in $A$. For example, if $x = (x_1, x_2, x_3)$ and $A = \{ 1, 2 \}$, $x_A = (x_1, x_2)$. For each $x$ and $A$, define

align[align omitted — 133 chars of source]

where the upper case refers to a random vector. Note that $g_{x} ( A )$ is a function of $x_A$.

Algorithm for Explainable Artificial Intelligence

enumerate[(i)] • Choose $x=x^\ast$, where $x^\ast$ could be an observed vector of an individual we would like to study. • Predict $f(x^\ast)$ using a machine-learning prediction method. • Specify $\Pi$. • Sample $q$ rows of $\bm{D}$ according to the probability distribution induced by the Shapley kernel weights $k ( \Pi, \Psi )$, where $q \gg |\Pi|$. • For each row of $\bm{D}$, compute a sample analog of (ref). • Solve the constrained least squares problem in (ref).

We end this section by commenting that there have been some debates on causal interpretations of the Shapley values for explainable AI, in particular, when the covariates are dependent on each other. See, e.g., LundbergLee, aas2021explaining, janzing2020feature, and eskes2020causal among others. This debate is irrelevant for our application of Shapley values to counterfactual simulations.

Counterfactual Simulations

We propose to use Shapley values for evaluating the different components in counterfactual simulations that are generated by structural models.

A Simple Roy Model \'{a} la honore2017poor

To be explicit, we consider a simple Roy model used in honore2017poor. In their setup, there are two sectors: $s \in \{1, 2\}$. Worker $i$ earns sector-specific income $w_{si1}$ in period 1 in the following form:

align*[align* omitted — 71 chars of source]

where $x_{si}$ is a vector of sector-specific human capital, $\beta_s$ is a vector of sector-specific parameters, and $\varepsilon_{si1}$ is a sector-specific unobserved random variable in period 1. Sector-specific income $w_{si2}$ in period 2 is generated by

align*[align* omitted — 101 chars of source]

where $d_{i1}$ is the sector chosen in period 1, $\gamma_s$ is a sector-specific parameter that represents the premium of staying in the same sector in period 2, and $\varepsilon_{si1}$ is a sector-specific unobserved random variable in period 2. The unobserved random variables $(\varepsilon_{1it}, \varepsilon_{2it})$ are assumed to be bivariate normally distributed with mean zeros, variances $(\sigma_1^2, \sigma_2^2)$ and correlation $\tau$, and i.i.d. over time.

Workers maximize discounted income. In time period 2, $d_{i2} = 1$ and the resulting income is $w_{i2} = w_{1i2}$ if

align*[align* omitted — 156 chars of source]

and $d_{i2} = 2$ and $w_{i2} = w_{2i2}$ otherwise. In time period 1, $d_{i1} = 1$ if and only if

align*[align* omitted — 212 chars of source]

where $\rho$ is discount factor.

Let $\theta = (\beta_1, \beta_2, \gamma_1, \gamma_2, \sigma_1^2, \sigma_2^2, \tau, \rho)$ denote the vector of all structural parameters. Let $\theta^b$ denote the benchmark vector of parameter values, for example, estimated values, and $\theta^c$ a counterfactual vector of parameter values. If we write $(\theta^{c}_{A}, \theta^b_{P \setminus A})$, we mean the combination of parameter values such that the $A$ subset of $\theta$ is set at the counterfactual values, while setting the $P \setminus A$ subset of $\theta$ at the benchmark values. Let $W = W(\theta)$ denote the observed random variables generated by the model given $\theta$. In the Roy model example, we have that $W = (w_{i1}, w_{i2}, d_{i1}, d_{i2}, x_{i1}, x_{i2})$. Let $f(\theta)$ be the real-valued quantity of interest in a counterfactual simulation that can be computed by simulating $W$ given $\theta$. For example, $f(\theta)$ be a measure of the changes in income inequality from period 1 to period 2. Finally, we let

align[align omitted — 108 chars of source]

Algorithm for Counterfactual Simulations

enumerate• Choose $\theta = \theta^b$, where $\theta_b$ could be estimated parameter values using the dataset. • Simulate $W$ given $\theta^b$ and evaluate $f(\theta^{b})$. • Specify $\Pi$. • Sample $q$ rows of $\bm{D}$ according to the probability distribution induced by the Shapley kernel weights $k ( \Pi, \Psi )$, where $q \gg |\Pi|$. • For each row of $\bm{D}$, simulate $W$ given $(\theta^{c}_{A}, \theta^b_{P \setminus A})$ and evaluate $g(A) = f(\theta^{c}_{A}, \theta^b_{P \setminus A})$. • Solve the constrained least squares problem in (ref).

An Example: $\Pi$, $\theta^b$, $\theta^c$ and $f(\theta)$

Suppose that $\tau = 0$ and $\rho = 0.95$ are fixed, as in honore2017poor. Then, $P = \{ \beta_1, \beta_2, \gamma_1, \gamma_2, \sigma_1^2, \sigma_2^2 \}$. Suppose that $$ \Pi = \{ \{\beta_1, \beta_2\}, \{\gamma_1, \gamma_2\}, \{\sigma_1^2, \sigma_2^2\} \}. $$ That is, we consider the three groups: the coefficients for sector-specific human capital, the sector-specific benefits of staying in the same sector, and sector-specific variances.

As an example, we set the benchmark vector of parameter values at $\theta^b$ at the values used in honore2017poor. That is, \[ \beta_1^b = (1,1)^\top, \beta_2^b = (0.5, 1)^\top, \gamma_1^b = 0, \gamma_2^b = 1, (\sigma_1^2)^b = 2, (\sigma_2^2)^b = 3. \] Suppose that we set the counterfactual vector $\theta^c$ of parameter values at \[ \beta_1^c = (1,2)^\top, \beta_2^c = (0.5, 2)^\top, \gamma_1^c = 0, \gamma_2^c = 2, (\sigma_1^2)^c = 2, (\sigma_2^2)^c = 6. \] If we focus on income inequality, we may consider:

itemize• between-sector inequality: $h_{\mathrm{bs-ineq}, t}(\theta) := \mathbb{E}[ w_{it} | d_{it} = 1] - \mathbb{E}[ w_{it} | d_{it} = 2]$; • within-sector inequality: $h_{\mathrm{ws-ineq}, t}(\theta; \tau_1, \tau_2) := \mathbb{Q}_{w_{it}} (\tau_1| d_{it} = s ) - \mathbb{Q}_{w_{it}} (\tau_2| d_{it} = s )$, where $\mathbb{Q}_{w_{it}} (\tau| d_{it} = s )$ is the $\tau$-quantile of $w_{it}$ conditional on $d_{it} = s$; • overall inequality: $h_{\mathrm{overall-ineq}, t}(\theta; \tau_1, \tau_2) := \mathbb{Q}_{w_{it}} (\tau_1 ) - \mathbb{Q}_{w_{it}} (\tau_2 )$, where $\mathbb{Q}_{w_{it}} (\tau )$ is the $\tau$-quantile of $w_{it}$.

As an example, we consider the change in overall inequality, measured by the $0.9-0.1$ quantile difference, from period 1 to period 2, that is: $$ f(\theta) = h_{\mathrm{overall-ineq}, 2}(\theta; \tau_1 = 0.1, \tau_2 = 0.9) - h_{\mathrm{overall-ineq}, 1}(\theta; \tau_1 = 0.1, \tau_2 = 0.9). $$ This example of $f(\theta)$ is not generally differentiable with respect to $\theta$. In consequence, the resulting $g(A)$ can not be viewed as the standard derivative; however, our approach provides how to quantify the importance of changes in the parameters. Additionally, we offer a global sensitivity analysis, as opposed to a local sensitivity analysis.

As in Section (ref), we have that

align*[align* omitted — 606 chars of source]

and $k ( \Pi, \Psi ) = 1$ for any $\Psi \subsetneqq \Pi$. Thus, $\bm{\phi}$ minimizes

align*[align* omitted — 77 chars of source]

subject to $$ \phi_{\{\beta_1, \beta_2\} } + \phi_{ \{\gamma_1, \gamma_2\} } + \phi_{ \{\sigma_1^2, \sigma_2^2\} } = g( \{ \beta_1, \beta_2, \gamma_1, \gamma_2, \sigma_1^2, \sigma_2^2 \}). $$ Again, this is a constrained least squares problem, which can be solved easily. The values of $\bm{g}$ are obtained by Monte Carlo simulation with the sample size of $10^7$.

Table (ref) gives the counterfactual simulation results. It can be seen that all three components contribute to the increase in inequality. However, $\phi_{ \{\gamma_1, \gamma_2\} }$ matters most in this example. Recall that $(\gamma_1^b = 0, \gamma_2^b = 1)$ and $(\gamma_1^c = 0, \gamma_2^c = 2)$. In other words, sector 2 specific return to the log wages doubled in period 2, which explains the largest increase in the change in the overall inequality.

table[table omitted — 443 chars of source]

In short, we demonstrate that although a measure of income equality from a Roy model can be highly nonlinear in inputs, the group Shapley value generates unique additive contributions from $\Pi$. From a practical perspective, we find that Table (ref) is just as effortless to interpret as a regression table. Generally speaking, the group Shapley value provides a straightforward method to interpret counterfactual simulation results when a structural model simulates the output.

An Application to Capital Misallocation

In this section, we revisit the work of DV2019, who developed a method to disentangle sources of capital misallocation---specifically, the dispersion in average revenue products of capital (arpk) using observable data on value-added and inputs from both China and United States.

The Quantitative Framework of DV2019

For firm $i$ in period $t$, DV2019 define the average revenue product of capital (arpk) as $arpk_{it} = va_{it} - k_{it}$, where $va_{it}$ is the log value added and $k_{it}$ is the capital stock. The parameter of interest is the variance $\sigma_{aprk}^2$ of $arpk_{it}$. To simulate this quantity, we need to specify a variety of parameters. First of all, we need to choose the value of $\alpha$, which determines the curvature of operating profits. There are two important dynamics in DV2019. One is associated with the log productivity (denoted by $a_{it}$):

align[align omitted — 116 chars of source]

where $\rho$ is the autoregressive parameter and $\sigma_\mu^2$ is the variance of the innovation term for productivity (see equation (5) in DV2019). The other dynamics involves log distortion (denoted by $\tau_{it}$):

align[align omitted — 204 chars of source]

where $\gamma$ determines how the distortion co-moves with productivity, $\sigma_\varepsilon^2$ is the variance of i.i.d. shocks and $\sigma_\chi^2$ is the variance of time-invariant firm-specific shocks (see equation (6) in DV2019). The distribution of future productivity ($a_{it+1}$) conditional on the firm’s information set ($\mathcal{I}_{it}$) in period $t$ follows a normal distribution with the posterior mean $E_{it}[a_{it+1}]$ and the posterior variance $V$. Finally, investment is subject to quadratic adjustment costs and the severity of the adjustment cost is parameterized via $\xi$. In summary, there are 8 key parameters: $\alpha$, $\rho$, $\sigma_\mu^2$, $\xi$, $V$, $\gamma$, $\sigma_\varepsilon^2$, and $\sigma_\chi^2$ (see Table 1 in DV2019).

Recall that Table (ref) reproduces the parameter estimates from DV2019. The value of $\alpha$ is slightly larger for China. The level of persistence for productivity is similar between the two countries, whereas the variance of productivity is larger in China than in U.S. The adjustment costs are much higher in the U.S. but uncertainty is larger in China. Note that $\gamma$ is more negative for China, indicating that the extent to which the distortion discourages investment by firms with higher productivity is larger in China than in U.S. Finally, in China, the variance of the transitory shocks is smaller but that of the permanent shocks is larger.

DV2019's Decomposition

DV2019 report the relative contributions of adjustment costs, uncertainty, and other factors to the arpk dispersion under the assumption that only the factor of interest is operational. To be more specific, we let $\theta = (\alpha, \rho, \sigma_\mu^2, \xi, V, \gamma, \sigma_\varepsilon^2,\sigma_\chi^2)$ and $f(\theta) = \sigma_{aprk}^2$. Assuming $f(\mathbf{0})=0$, DV2019 estimate the contribution of each factor by $\phi_{\mathrm{DV}} (A) := f(\theta_A, \mathbf{0}_{P \backslash A}) - f(\mathbf{0})$ where $\theta$ is the parameter estimates in Table (ref). The contribution estimation is based on the principle of ceteris paribus, measuring the impact of each factor from a hypothetical void setting $f(\mathbf{0})$. However, while it intends to quantify the change when a factor of interest is added, $\phi_{\mathrm{DV}}$ can yield an unrealistic or imprecise assessment of individual impacts because setting some entries of $\theta$ at zeros results in degenerate dynamics. For instance, $V$ is the posterior variance, which needs to be strictly greater than zero. Furthermore, $\phi_{\mathrm{DV}}$ does not satisfy the Efficiency axiom as it is a weighted sum of utility changes, which is referred to as a semivalue semivalue_without_efficiency. This is why the relative contributions do not necessarily sum to $1$, as acknowledged by DV2019. We address these two issues in the following subsection.

Shapley Value Decomposition

In this subsection, we carry out a new counterfactual exercise that is different from that of DV2019. We set the benchmark vector ($\theta_b$) of parameter values using estimates from U.S. and the counterfactual vector ($\theta_c$) using estimates from China. As each parameter represents a distinct component of the model developed in DV2019, we consider singleton groups:

align*[align* omitted — 138 chars of source]

In the data, $\sigma_{aprk}^2$ was 0.92 for China and 0.45 for U.S (see Table 2 in DV2019). Thus, the degrees of capital misallocation double by changing the U.S. parameter estimates to those based on Chinese data. Our Shapley value decomposition will indicate the extent to which each parameter contributes to this increase.

table[table omitted — 646 chars of source]

Table (ref) reports the resulting Shapley value decomposition. First of all, note that there are positive and negative values for $\phi_{ \{ \cdot\} }$ because some changes in the parameter values imply increases in the arpk dispersion but other changes indicate decreases. However, the sum of the Shapley shares is 1 by design.

We now examine each row. The increase of $\alpha$ from $0.62$ to $0.71$ is associated with $\phi_{ \{\alpha\} } = 0.0073$, implying the share of $0.016$, which is quite small. The persistence parameter $\rho$ gets slightly smaller moving from the U.S. to China (that is, from $0.93$ to $0.91$). However, the resulting Shapley value is $-0.0599$ with a relatively large share of $-0.129$. We can interpret that a small decrease in the autoregressive parameter in the log productivity is associated with a relatively large reduction in the capital misallocation. The variance of the innovation term for productivity increases more substantially from 0.08 to 0.15, resulting in the second largest share in the Shapley value decomposition. It is interesting to note that the change in the adjustment costs looks more visible in the sense that $\xi$ decreases by a factor of 10, while the implied share is only $-0.126$. This shows that the Shapley value decomposition is useful to compare changes in parameters on the same scale. The increase in uncertainty is associated with the share of $0.095$. The largest Shapley share of $0.542$ comes from $\gamma$, which is negative for both countries. A more negative $\gamma$ means that the distortion plays a larger role in disincentivizing investment by more productive firms. The reduction in the variance of the transitory shocks is small and as a result, the Shapley share is small as well. The increase in the variance of the permanent shocks is quite large, ranking as the third most important factor in terms of the Shapley share. In short, the two components ($\gamma$ and $\sigma_\chi^2$) of the distortion as well as the variance of the productivity shocks are the three most important factors affecting the increase in the arpk dispersion. Overall, our Shapley value analysis reconfirms the quantitative findings in DV2019, while providing novel perspectives.

comment\section{Decomposition Studies in Applied Microeconomics} In the previous sections, we have described applications of the Shapley value to the counterfactual study of structural models. In this section, we consider an application to decomposition studies in applied microeconomics. \textcolor{red}{I am wondering why we want to decompose the difference. How is this section relevant to other sections?} The starting point for decomposition methods in applied microeconomics is the Oaxaca-Blinder decomposition. Consequently, a natural requirement for our method would be to have OB as a special case when the model is linear. We may then ask how our method compares with --in terms of, say, interpretability, identification conditions, invariances, etc-- to competing modern decomposition methods for nonlinear models. The OB decomposition relies on the model \[ Y_{gi}=\beta_{g0}+\sum_{k=1}^{K}X_{ik}\beta_{gk}+v_{gi},\ g=A,B, \] where $A$ and $B$ are the groups to be compared. \textcolor{red}{Definition of $v_{gi}$ is not given.} The overall difference between the groups is \[ \hat{\triangle}_{O,A\rightarrow B}^{\mu}=\bar{Y}_{B}-\bar{Y}_{A}. \] and is decomposed as \[ \hat{\triangle}_{O,A\rightarrow B}^{\mu}=\underset{\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{B})}{\underbrace{\left(\hat{\beta}_{B0}-\hat{\beta}_{A0}\right)+\sum_{k=1}^{K}\bar{X}_{Bk}\left(\hat{\beta}_{Bk}-\hat{\beta}_{Ak}\right)}}+\underset{\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})}{\underbrace{\sum_{k=1}^{K}\left(\bar{X}_{Bk}-\bar{X}_{Ak}\right)\hat{\beta}_{Ak}}}, \] where $\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{B})$ is the change attributed to the change in the wage structure given the composition at period/group $B$, and $\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})$ is the compositional change given the structure at period/group $A$. Note that the choice of fixing the “structure” at $\hat{\beta}_{Ak}$ to consider the impact of the compositional change $\bar{X}_{Bk}-\bar{X}_{Ak}$ matters for the decomposition estimates. The above gives the aggregate decomposition. The detailed decomposition is the covariate-wise equivalent. Specifically, $\bar{X}_{Bk}\left(\hat{\beta}_{Bk}-\hat{\beta}_{Ak}\right)$ and $\left(\bar{X}_{Bk}-\bar{X}_{Ak}\right)\hat{\beta}_{Ak}$. We would like to accommodate something similar while leveraging the Shapley construction. Let's say the covariate is a $p$-tuple \[ X=(X_{1},...,X_{p}), \] with a data set observed for both groups $\mathbf{X}^{A}$ and $\mathbf{X}^{B}$. Let's say the coefficient is a $K$-tuple, \[ \theta=(\theta_{1},...,\theta_{K}), \] with estimate on both groups $\hat{\theta}^{A}$ and $\hat{\theta}^{B}$. Somehow, we want to disentangle the change due to $X$ and the change due to $\theta$, and then further decompose those changes by individual covariates and coefficient --but note that we do not care to decompose the change in covariate by change in individual coefficient, and vice versa. We want to decompose \[ f\left(\mathbf{X}^{B};\hat{\theta}^{B}\right)-f\left(\mathbf{X}^{A};\hat{\theta}^{A}\right). \] If we treat $X$ and $\theta$ as the only two --vector-- components for which we want values, then we can get \[ S_{\theta}=\frac{1}{2}\left(\left(f\left(\mathbf{X}^{B};\hat{\theta}^{B}\right)-f\left(\mathbf{X}^{B};\hat{\theta}^{A}\right)\right)-\left(f\left(\mathbf{X}^{A};\hat{\theta}^{B}\right)-f\left(\mathbf{X}^{A};\hat{\theta}^{A}\right)\right)\right), \] which we may call the structural Shapley value, and \[ S_{X}=\frac{1}{2}\left(\left(f\left(\mathbf{X}^{B};\hat{\theta}^{B}\right)-f\left(\mathbf{X}^{A};\hat{\theta}^{B}\right)\right)-\left(f\left(\mathbf{X}^{B};\hat{\theta}^{A}\right)-f\left(\mathbf{X}^{A};\hat{\theta}^{A}\right)\right)\right), \] which we may call the compositional value. The detailed decomposition of the structural change could then be given as \[ S_{\theta_{i}}\propto\sum_{\mathcal{S}\in\{1,...,p\}\backslash i}\frac{1}{2}\left(\left(f\left(\mathbf{X}^{B};\hat{\theta}^{B}\right)-f\left(\mathbf{X}^{B};\left(\hat{\theta}_{S}^{A},\hat{\theta}_{-\mathcal{S}}^{B}\right)\right)\right)-\left(f\left(\mathbf{X}^{A};\hat{\theta}^{B}\right)-f\left(\hat{\theta}_{S\cup\{i\}}^{A},\hat{\theta}_{-\left(\mathcal{S\cup}\{i\}\right)}^{B}\right)\right)\right). \] The detailed decomposition of the compositional change could then be \[ S_{X_{i}}\propto\sum_{\mathcal{S}\in\{1,...,K\}\backslash i}\frac{1}{2}\left(\left(f\left(\mathbf{X}^{B};\hat{\theta}^{B}\right)-f\left(\left(\mathbf{X}_{S}^{A},\mathbf{X}_{-\mathcal{S}}^{B}\right);\hat{\theta}^{B}\right)\right)-\left(f\left(\mathbf{X}^{B};\hat{\theta}^{A}\right)-f\left(\left(\mathbf{X}_{S\cup\{i\}}^{A},\mathbf{X}_{-\left(\mathcal{S\cup}\{i\}\right)}^{B}\right);\hat{\theta}^{A}\right)\right)\right). \] Need to get clear connection with the OB decomposition in the linear case, and check if some modification gets the OB decomposition as its linear special case. \subsubsection*{Relating the Shapley decomposition to the Oaxaca-Blinder decomposition} If we use use the average fitted value as the value to decompose, then we can relate Shapley to Oaxaca-Blinder. Given that the Shapley decomposition --as described for structural models in the note-- in one parameter is invariant to the order of the other parameters, the best we can hope for is some average over the ordered Oaxaca-Blinder decompositions. {[}we get weird sign{]} Indeed \begin{align*} \hat{\triangle}_{O,A\rightarrow B}^{\mu} & =S_{\theta}+S_{X}\\ \\ \end{align*} where \begin{align*} S_{\theta} & =\frac{1}{2}\left(\left(\hat{\beta}_{B}-\hat{\beta}_{A}+\sum_{k}\bar{X}_{B}\left(\hat{\beta}_{B}-\hat{\beta}_{A}\right)\right)+\left(\hat{\beta}_{B}-\hat{\beta}_{A}+\sum_{k}\bar{X}_{A}\left(\hat{\beta}_{B}-\hat{\beta}_{A}\right)\right)\right)\\ & =\frac{1}{2}\left(\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{B})+\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{A})\right)\\ & =\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\frac{\bar{X}_{A}+\bar{X}_{B}}{2}) \end{align*} which is the impact of the structural change on the average composition, and \begin{align*} S_{X} & =\frac{1}{2}\left(\left(\sum_{k}\bar{X}_{B}\hat{\beta}_{B}-\sum_{k}\bar{X}_{A}\hat{\beta}_{B}\right)+\left(\sum_{k}\bar{X}_{B}\hat{\beta}_{A}-\sum_{k}\bar{X}_{A}\hat{\beta}_{A}\right)\right)\\ & =\frac{1}{2}\left(\left(\sum_{k}\left(\bar{X}_{B}-\bar{X}_{A}\right)\hat{\beta}_{B}\right)+\left(\sum_{k}\left(\bar{X}_{B}-\bar{X}_{A}\right)\hat{\beta}_{A}\right)\right)\\ & =\frac{1}{2}\left(\hat{\triangle}_{X,B\rightarrow A}^{\mu}(\hat{\beta}_{B})+\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})\right)\\ & =\hat{\triangle}_{X,B\rightarrow A}^{\mu}(\frac{\hat{\beta}_{A}+\hat{\beta}_{B}}{2}) \end{align*} which is the impact of the compositional change on the average structure. This is one of the many seemingly arbitrary choices of counterfactuals that can be picked to construct a decomposition. See Fortin, Lemieux and Firpo's handbook chapter, Section 3.3. We wonder if we argue that the decomposition above, because it is special case of of the Shapley decomposition, inherits its axiomatic motivation and is in some sense canonical. \subsubsection*{Symmetry, not satisfied} Here we have only two “players”, so the symmetry axioms requires that if both of them have the same payoff when added to the null set, they must have the same Shapley value. The is eminently reasonable in the case of a decomposition; if changing only the structure or only the composition has the same impact, then it should have be attributed the same value in a decomposition. Specifically, we ask if $f\left(\mathbf{X}^{B};\hat{\theta}^{A}\right)=f\left(\mathbf{X}^{A};\hat{\theta}^{B}\right)$ implies $\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})=\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{B})$. we work it out \[ f\left(\mathbf{X}^{B};\hat{\theta}^{A}\right)=f\left(\mathbf{X}^{A};\hat{\theta}^{B}\right) \] is \[ \hat{\beta}_{A0}+\sum_{k=1}^{K}\bar{X}_{Bk}\hat{\beta}_{Ak}=\hat{\beta}_{B0}+\sum_{k=1}^{K}\bar{X}_{Ak}\hat{\beta}_{Bk} \] \[ -\sum_{k=1}^{K}\bar{X}_{Ak}\hat{\beta}_{Bk}+\sum_{k=1}^{K}\bar{X}_{Bk}\hat{\beta}_{Bk}=\hat{\beta}_{B0}-\hat{\beta}_{A0}-\sum_{k=1}^{K}\bar{X}_{Bk}\hat{\beta}_{Ak}+\sum_{k=1}^{K}\bar{X}_{Bk}\hat{\beta}_{Bk} \] \[ \sum_{k=1}^{K}\left(\bar{X}_{Bk}-\bar{X}_{Ak}\right)\hat{\beta}_{Bk}=\hat{\beta}_{B0}-\hat{\beta}_{A0}+\sum_{k=1}^{K}\bar{X}_{Bk}\left(\hat{\beta}_{Bk}-\hat{\beta}_{Ak}\right) \] where the RHS is $\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{B})$, but the LHS is $\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{B})\neq\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})$. So the standard OB fails to satisfy the symmetry requirement. \subsubsection*{Null player property, not satisfied} Need that if$f\left(\mathbf{X}^{A};\hat{\theta}^{A}\right)=f\left(\mathbf{X}^{B};\hat{\theta}^{A}\right)$, then $\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})=0$. Let's inspect. We do the single covariate case without intercept for simplicity. \[ \bar{X}_{A}\hat{\beta}_{A}=\bar{X}_{B}\hat{\beta}_{A} \] \[ 0=\left(\bar{X}_{B}-\bar{X}_{A}\right)\hat{\beta}_{A} \] which is to say that, indeed, $\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})=0$ obtains. However, if $f\left(\mathbf{X}^{A};\hat{\theta}^{A}\right)=f\left(\mathbf{X}^{A};\hat{\theta}^{B}\right)$, then \[ \bar{X}_{A}\hat{\beta}_{A}=\bar{X}_{A}\hat{\beta}_{B} \] \[ 0=\bar{X}_{A}\left(\hat{\beta}_{B}-\hat{\beta}_{A}\right) \] means $\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{A})=0$, but does not mean $\hat{\triangle}_{\theta,A\rightarrow B}^{\mu}(\bar{X}_{B})$ is zero. Hence the null player property likewise fails. \subsubsection*{Additivity, satisfied} Additivity requires there being two games. So we should agree on what exactly the “game” is. The two players are the covariates and coefficients. So --not great-- notation could be $f(\mathbf{X}_{1}^{B};\hat{\theta}_{1}^{B})-f(\mathbf{X}_{1}^{A};\hat{\theta}_{1}^{A})$ to refer to the total payoff in game 1, and $f(\mathbf{X}_{2}^{B};\hat{\theta}_{2}^{B})-f(\mathbf{X}_{2}^{A};\hat{\theta}_{2}^{A})$ for game 2. Then \[ \hat{\triangle}_{\theta,A\rightarrow B}^{1+2}(\bar{X}_{B})=\hat{\triangle}_{\theta_{1},A\rightarrow B}^{1}(\bar{X}_{1,B})+\hat{\triangle}_{\theta_{2},A\rightarrow B}^{2}(\bar{X}_{2,B}) \] by construction. Likewise for $\hat{\triangle}_{X,A\rightarrow B}^{\mu}(\hat{\beta}_{A})$. \subsubsection*{Efficiency, satisfied} The components sum to the game total by construction.

An Application to Globalization

CGT2016 examine decreases in trade barriers, tariffs, and firing costs in an open economy, using establishment-level data from Colombia. Their counterfactual experiments suggest that Colombia's integration into global product markets raised its national income but also increased unemployment, wage inequality, and firm-level volatility. Table 4 in CGT2016 provides the results of their counterfactual experiments. Suppose that we are interested in comparing the last column “Reforms and globalization” with the first column “Baseline” and construct Shapley value decomposition among three elements: firing cost ($c_f$), tariff rate ($\tau_a$), and iceberg trade cost ($\tau_c$). Then, we have the same structure as the toy example in Section (ref). That is, we have that

align*[align* omitted — 443 chars of source]

For self-containment, we re-produce the parameter values in Panel A of Table (ref). In Panel A, the benchmark (“Baseline”) parameters are $(c_f^b, \tau_a^b, \tau_c^b) = (0.60, 1.21, 2.50)$, whereas the counterfactual (“Reforms and globalization”) parameters are $(c_f^c, \tau_a^c, \tau_c^c) = (0.30, 1.11, 2.19)$. Regarding $g(\cdot)$, we consider the aggregates reported in Table 4 in CGT2016, which are reproduced in Panel B of Table (ref). As before, $g(\cdot)$ is obtained by taking the difference between the counterfactual (“Reforms and globalization”) quantities and the benchmark (“Baseline”) quantities, where the latter quantities are normalized to be one.

It can be seen from Table (ref) that the four middle columns are the input for optimization (i.e. elements of $\bm{g}$): using the notation above, we have values for $g ( \{c_f\} )$, $g ( \{\tau_a\} )$, $g ( \{\tau_c\} )$, and $g (\{c_f\} \cup \{\tau_a\} )$. They are called “Labor”, “Tariff”, “Iceberg”, and “Reforms” in the table. However, the last two elements of $\bm{g}$ are missing: namely, $g (\{c_f\} \cup \{\tau_c\} )$ and $g ( \{\tau_a\} \cup \{\tau_c\} )$. It would be ideal to complete the missing elements in order to construct Shapley value decomposition. However, one interesting research question is what one can do if it is expensive or impossible to carry out further experiments to obtain the missing elements for $\bm{g}$.

In view of the fact that each of the parameters is reduced to favor globalization, we consider the following set of restrictions on the missing elements: for $g (\{c_f\} \cup \{\tau_c\} )$,

align[align omitted — 356 chars of source]

and for $g (\{d_a\} \cup \{\tau_c\} )$,

align[align omitted — 355 chars of source]

In addition, suppose that we have known upper and lower bounds on $g (\{c_f\} \cup \{\tau_c\} )$ and $g (\{d_a\} \cup \{\tau_c\} )$:

align[align omitted — 164 chars of source]

where $g_{\min}$ and $g_{\max}$ are predetermined constants such that $g_{\min} = 0.5$ and $g_{\max} = 1.5$. Combining all the constraints above in (ref)-(ref) forms the linear constraints $A_{\mathrm{const}} \bm{g} \leq b_{\mathrm{const}}$ in Section (ref).

table[table omitted — 1,588 chars of source]
table[table omitted — 1,977 chars of source]
figure[figure omitted — 296 chars of source]

Table (ref) and Figure (ref) present the empirical results for the Shapley bounds and the Shapley minimum norm solutions using the methodology developed in Section (ref). As can be seen from Table (ref), the revenue share of exports increases by 1.5 from the benchmark quantity to the counterfactual quantity. The largest portion of this increase is explained by the decrease in the iceberg trade cost (with a lower bound of 0.98). The reduction in the tariff rate appears to explain slightly more than the reduction in firing costs, although their bounds overlap, making it difficult to definitively rank these two factors. Whether considering the Shapley bounds or the Shapley minimum norm solutions, firing cost reductions are linked to decreases in the exit rate, vacancy filling rate, and unemployment rate. Conversely, reductions in the tariff rate and iceberg trade cost are associated with increases in these rates. The decomposition result for the mass of firms is ambiguous: for all three factors, the minimum norm solutions are negative, but the bounds include zero. All the inequality measures---Std. wages (firms), Std. wages (workers), Std. $J$ (firms), and Std. $J$ (workers)---increase with the reduction in each factor. However, there is no definite ranking among the three factors in terms of their contribution to increases in inequality. Finally, real income increases by 0.12 from the benchmark quantity to the counterfactual quantity, primarily explained by the reduction in the iceberg trade cost (with a lower bound of 0.1). The Shapley bounds are rather wide for both the firing cost and the tariff rate.

Overall, the empirical results suggest that the reduction in iceberg trade costs is the most significant factor for real income growth. However, it is less clear which factor is the primary driver of increases in inequality. Returning to our research question, this example demonstrates that it is feasible to perform a coherent Shapley value decomposition even when it is impossible to conduct further experiments to obtain missing elements for $\bm{g}$. Thus, our methodology holds potential for real-world applications that involve costly counterfactual simulations.

Conclusions

Although Shapley values are widely used across various disciplines, their application to interpreting counterfactual simulations in structural economic models has not yet been explored. We have demonstrated that the group Shapley value provides a natural framework for quantifying the importance of parameter changes in complex models. By preserving the axiomatic properties of the Shapley value, our method offers a well-grounded decomposition approach for explaining outcomes in counterfactual simulations. Practically, it enables researchers to generate importance tables that are as intuitive and accessible as regression tables. In this sense, our use of the Shapley value could pave the way for “interpretable structural economics,” much like LundbergLee's introduction of Shapley values to the machine learning community has advanced explainable artificial intelligence.

However, our method has some limitations. First, we focus on comparing two sets of parameters. Extending our approach to cases involving more than two sets of parameter values (e.g., the United States, China, and South Korea, using the example in the introduction) would be valuable. Second, our numerical examples are limited to small-scale optimization problems, where computational complexity is less of a concern. For larger-scale problems, computational challenges become more significant, and developing strategies to address them—such as devising effective sampling techniques and accounting for sampling errors—will be crucial. These are exciting directions for future research.