Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
128,367 characters · 18 sections · 55 citation commands
sensitivity of regular estimators
Balancing simplicity of statistical methodology with complexity of economic modeling is a challenge in empirical work. Structural models lead to estimators with nontransparent dependence on data. Both structural and predictive models are subject to specification choices that have nontransparent influence on inferences. However, regular estimators of parameters in these models have simple asymptotic behavior and can be understood well locally. Regularity allows to draw local comparisons (approximations) between two estimators and obtain local counterfactuals of their values. For example, it may be useful to know that a structural estimator is locally well approximated with a simple \{mean, variance, quantile, etc\}. Or that two alternative specifications provide similar results not only at the sampling distribution but in a neighborhood around it. Sensitivity measures formalize local comparisons and counterfactuals, and add transparency to inferences made with structural models.
We examine geometric foundations of estimator sensitivity and highlight the role of the information metric in asymptotics of regular estimation. Covariance of joint asymptotic distribution is the information inner-product that measures alignment of first-order approximations to regular parameters. This is a natural measure of local approximation quality between two estimators. Differentiability has a prominent role and a long history in regular asymptotics from von Mises (1947) mises1947 to van der Vaart (1991) vaart1991differentiable and Newey (1994) newey1994asymptotic, we go a step further and develop complete differential calculus on the model. We define sensitivity as a directional derivative and propose it as a general tool for local counterfactual analysis as in Stock (1989) stock1989nonparametric and Chernozhukov, Fern\'andez-Val and Melly (2013) chernozhukov2013inference. Instead of specifying a counterfactual distribution of control variables, we think of policy as shifting the value of a control parameter. Sensitivity measures the effect of policy on the value of a target parameter. For example, the local effect of changing the \{mean, variance, quantile, etc\} of a distribution on the \{mean, variance, quantile, etc\} of the distribution. Both the implicit counterfactual distribution and the sensitivity (directional derivative) depend on the way policy measures distances on the model. Asymptotic covariance is shown to be such a directional derivative with a particular choice of geometric primitives.
To put our work in perspective, let us disassemble empirical analysis in economics into a stack of layers and interfaces. At the top level, there is a model of economic quantities that are defined independently of data. This can be a structural model or a descriptive relationship between control and response variables, say, quantity $\vartheta$ is of interest to the researcher. For example, a price elasticity, a rate of return, a parameter of utility function, a location or scale parameter. At the bottom of the empirical analysis stack, there are data from unknown distribution $P$ on sample space $(\mathcal{X},\mathcal{A})$ that can be described with a statistical model $\matheuler{P}$. For example, a random sample from a parametric, nonparametric or semiparametric model. At the interface between the application and the data layers, high-level object $\vartheta$ is identified with a particular feature of the statistical model $\psi(P)$. Thus, the middle layer between data and application is a specification $\Psi:a\mapsto \psi_a$ that assigns a statistical parameter $\psi_a$ to the economic quantity $\vartheta$ under modeling assumptions $a$ of the researcher, say, index set $\matheuler{A}$ describes all specifications entertained by the researcher.
Three logically independent types of variation in empirical inference about $\vartheta$ can be distinguished based on the application, specification and data layer anatomy. Application model sensitivity analysis examines dependences within the mathematical relationships of the application layer, sobol1993sensitivity, 1404.2405. Specification sensitivity arises from variations at the interface layer in mapping $\Psi$. For example, $\vartheta$ can be identified with a coefficient in a linear regression model or an iv equation, both ols and iv can be set up with different sets of covariates or instruments. Omitted variable bias is the quintessential example of variation in specification. Both the statistical model $\matheuler{P}$ and the unknown distribution $P$ of sampled data remain fixed across different specifications, only the choice of statistical functional $\psi(P)$ that is used for inference about $\vartheta$ changes. Exploring specification variation for a fixed $P$ is analytically straightforward -- estimates of all interesting choices $\Set{\psi_a(P)}{a\in \matheuler{A}}$ can be obtained, hopefully uniform, inferences can be reported, a parametrization $\Psi$ can be differentiated with techniques from calculus to find local effects of changing specification.
This paper studies sensitivity of a fixed statistical functional defined on a statistical model $$ \psi:\matheuler{P}\rightarrow\mathbb{R} $$ to local variations of the data distribution $P$ within model $\matheuler{P}$. We work strictly at the data layer, holding specification fixed, but suggest both data level and application level interpretations. In mathematical terms, we consider differential calculus of functionals on the model manifold under different Riemannian geometries. Statistical model sensitivity is a directional derivative of the statistical functional. Since a typical statistical model behind economic applications is an infinite-dimensional space, it is helpful to identify a direction on the model with a tractable statistical parameter, denoted $\nu(P)$. Sensitivity with respect to parameter $\nu(P)$ is the partial derivative along its gradient vector $\nabla\nu$, denoted by $\partial_\nu$ operator: $$ \partial_\nu \psi \coloneqq \lim_{h\rightarrow 0} h^{-1} \big[ \psi(P + h \cdot \nabla\nu) - \psi(P) \big]. $$
Practical utility of sensitivity analysis comes from the fact that it is closely related to asymptotic approximations for a large class of estimators. The main observation is that influence functions are gradients according to the information geometry of the model. Gradients in any other geometry on the model are linear transformations of influence functions. By varying geometric primitives in the definition of sensitivity $\partial_\nu\psi$, researcher obtains different local counterfactual values of $\psi$, corresponding to different perturbations on $\matheuler{P}$ that change $\nu$ in a controlled way. One of such counterfactual is given by the asymptotic covariance of two regular estimators.
For a pair of estimators on statistical model $\matheuler{P}$ with standard asymptotic behavior
The $\Lambda,\Delta$ measures were introduced by Gentzkow and Shapiro (2015) Gentzkow2015 for the purpose of comparing a nontransparent estimator $\wh\psi_n$ to a tractable statistic $\wh\nu_n$. Andrews, Gentzkow and Shapiro (2017) andrews2017measuring interpreted $\Lambda$ as a measure of local specification sensitivity of gmm functionals implicitely parametrized by the population value of moments $E g(X, \psi_a(P) ) = a$.\footnote{From the fact that Jacobian $G$ of moments does not depend on specification parameter $a$, it follows that dependence of moments on $a$ must be additive.}
We define sensitivity directly on the model using techniques of differential geometry, rather than in terms of the asymptotic distribution of estimators as in Gentzkow2015. We then relate our sensitivity of functionals to asymptotic distributions of estimators using results from semiparametric efficiency theory. This relationship is similar to Newey (1994) newey1994asymptotic, but in our definition we allow for an explicit choice of geometric primitives. We show that, in information geometry of $\matheuler{P}$: Estimator sensitivity $\Lambda(\wh\psi,\wh\nu)$ is (i) the directional derivative $\partial_\nu\psi$ of $\psi(P)$ in the direction of $\nu(P)$. Estimator sufficiency $\Delta(\wh\psi,\wh\nu)$ is (ii) the square of cosine of the angle made by linear approximations to $\psi$ and $\nu$ at $P$, (iii) the relative size of partial derivative of $\psi$ along $\nu$ to total derivative of $\psi$, (iv) the efficiency gain in estimating $\psi$ obtained by fixing population value of $\nu$. With other geometries on $\matheuler{P}$, measures (i-iii) are available but not reflected in the asymptotic distribution of estimators.
Our investigation is inspired by Gentzkow2015 but we proceed in a different direction from their line of inquiry. The main objective of this paper is to provide interpretation of $\Lambda,\Delta$ measures from semiparametric efficiency perspective. This leads us to information geometry and motivates our local counterfactual interpretation of sensitivity, which we generalize by allowing a policy metric instead of the intrinsic information metric of the geometry behind statistical efficiency. Apart from generalizations, our inquiry fundamentally diverges from Gentzkow2015 in that we make a clear distinction between varying specification $a\mapsto\psi_a$, holding $P$ fixed, and varying distribution $P$, holding specification $a\mapsto\psi_a$ fixed, and consider only the latter exercise. By contrast, andrews2017measuring,Gentzkow2015 are primarily concerned with variation in the specification of moment conditions in gmm functionals, which are not deviations on the statistical model. This paper and andrews2017measuring,Gentzkow2015 obtain complementary interpretations for quantities $\Lambda,\Delta$ which should only increase their value in practice.
We suggest two types of applications of statistical model sensitivity. A data level interpretation as a measure of local alignment of two functionals can be used to compare competing specifications or target specifications to tractable statistics. An application level interpretation as a derivative can be used for local counterfactual analysis and policy evaluation.
Measures (ii-iv) above quantify the quality of local approximation of $\psi$ by $\nu$ in a neighborhood of $P$. Linear approximation of $\psi$ determines first order asymptotic behavior of estimates $\wh\psi$. Sensitivity thus provides an analytic tool for exploring inferences based on the asymptotic distribution of estimates of $\psi$. Reporting sensitivity to tractable parameters $\nu$ helps explain how inferences about $\psi(P)$ are obtained from $P$. See andrews2017measuring,Gentzkow2015 and references therein for a discussion on transparency and empirical examples. In the case with multiple specifications for $\vartheta$, the natural course is to report all estimates $\Set{\wh\psi_a}{a\in \matheuler{A}}$. This provides a one-point comparison of different specifications at the sampling distribution $P$. Reporting $\psi_a(P)$ similar to $\psi_b(P)$, positive estimator sensitivity $\Lambda(\wh\psi_a,\wh\psi_b)$ and estimator sufficiency $\Delta(\wh\psi_a,\wh\psi_b)$ close to one, can be offered as formal evidence that results are not sensitive to specification in a neighborhood of $P$. We call these applications estimator or information sensitivity. \footnote{ Note that identification and consistency of estimates $\wh\psi$ are global properties of the functional and the model and thus are outside of the scope of local sensitivity analysis. }
Directional derivatives (i) provide a simple description of the local behavior of functional $\psi$ at distribution $P\in\matheuler{P}$. For streams of random samples generated by $P_h = P + h \, \wt\nu_P$, where $\wt\nu_P$ is the gradient of $\nu(P)$, the limits under $P_h$ of estimators $\wh\psi$ and $\wh\nu$ are: $$ \wh\psi \xrightarrow[]{P_h} \psi(P) + h \cdot \partial_\nu\psi + o(h) \quad\text{and}\quad \wh\nu \xrightarrow[]{P_h} \nu(P) + h \cdot \partial_\nu\nu + o(h) . $$ We see that sensitivity $S(\psi,\nu)\coloneqq \partial_\nu\psi /\partial_{\nu}\nu $ is the local effect on the value of $\psi(P)$ of a ceteris paribus change in the value of $\nu(P)$ accomplished by changing the underlying distribution from $P$ along $P_h$. This is the local version of the counterfactual analysis that typically takes $\psi(P)$ to be some location parameter of a response variable $Y$ and $\nu(P)$ to be the marginal distribution of a policy variable $X$ stock1989nonparametric,heckman2007econometric,chernozhukov2013inference. Finally, one can use the identification of statistical functionals $\psi,\nu$ with economic quantities $\vartheta,\eta$ of the application layer and interpret the local relationship $S(\psi,\nu)$ as the partial derivative of $\vartheta$ with respect to $\eta$. We call these applications policy sensitivity and argue that it should be based on a geometry of $\matheuler{P}$ with a policy metric motivated by the application, rather then the information metric dictated by technicalities of asymptotic approximations.
Policy metric is a local notion distance on the model $\matheuler{P}$. Asymptotic inference implicitly relies on the information metric that measures “statistical” distances on the model. Metric determines the direction $\wt\nu$ on the model along which policy shifts in the value of $\nu$ are achieved. Thus, the combination of control functional $\nu$ and policy metric determines the path of counterfactual distributions $P_h$ along which sensitivity of target functional $\psi$ is measured. We describe a simple procedure for specifying and interpreting policy metrics, and illustrate the analysis with a Monte Carlo experiment. Parametrizing directions on the model by a control functional and a policy metric is a tractable and flexible way to reason about local counterfactuals.
The scope and contribution of this paper is to provide geometric foundation for statistical model sensitivity analysis and to highlight the importance of the metric of the model. We provide new geometrically motivated methodology for counterfactual analysis. This appears to be a novel use of geometry in econometrics and statistics. More specifically, we introduce the notion of a policy metric on a statistical manifold, including semiparametric and nonparametric models. We then define sensitivity as a directional derivative with respect to policy gradient of a control statistical parameter. This geometric formulation enables us to interpret policy sensitivity, including the covariance of asymptotic distribution, as a local counterfactual. In order to compute and estimate policy sensitivities, we obtain a result that relates policy gradients to influence functions. We provide high level conditions for consistency of estimated sensitivity. Detailed econometric analysis of estimation and inference for real-valued and distributional local counterfactuals is left to future work.
This paper draws on and contributes to several seemingly unrelated literatures. Geometric foundations in statistical inference have been investigated by many authors: Hotelling (1930) hotelling1930 considers the spaces of statistical parameters as curved surfaces embedded in Euclidean space, one of which can be seen in (ref). Mahalanobis (1936) mahalanobis1936generalized defines general distances between statistical populations and notes parallels with special relativity. Rao (1949) rao1949appendix writes down the information metric of a population space (parametric model) in local coordinates and describes geodesics between two distributions. Amari (1985, 2000) amari1985,amari2000 provides geometric insight into asymptotic efficiency in parametric models. To this literature we contribute by applying differential geometry to infinite-dimensional models and by new methodology motivated by geometry. Specification sensitivity analysis based on $\Lambda,\Delta$ was introduced by Gentzkow and Shapiro (2015) and Andrews, Gentzkow and Shapiro (2017) Gentzkow2015,andrews2017measuring. Semiparametric efficiency theory shows that variance of asymptotic Gaussian distribution in large statistical models is the information norm of the differential e.g. Stein (1956) stein1956, Koshevnik and Levit (1976) Koshevnik76, Pfanzagl (1982) pfanzagl1982contributions, van der Vaart (1991) vaart1991differentiable, Bickel et al. (1993) bickel1993efficient but does not make explicit use of modern geometry. We contribute to the efficiency literature by modelling large models as manifolds.
We organize the paper as follows: In (ref) we define sensitivity using econometrics language of semiparametric efficiency and provide a Monte Carlo example to illustrate the methodology. To make geometric ideas of this paper accessible without requiring familiarity with Riemannian geometry and semiparametric efficiency, we consider in (ref) the special case of a two-dimensional statistical model embedded in $\mathbb{R}^3$. This allows a graphical illustration of methodology and explicit calculations. In (ref) we review required foundations from differential geometry, state the general definition of sensitivity measures, explain how they depends on geometric primitives of the model, and discuss analytic interpretation of these measures. In (ref) we apply results of semiparametric efficiency theory to obtain information sensitivity from regular efficient estimators, relate policy gradients to influence functions, and briefly consider consistency of estimated policy sensitivity. We work out some simple examples in (ref) and give a self-contained summary of efficiency theory results we cite in (ref).
This section provides an informal introduction to sensitivity, explains how it relates to geometry of the statistical model and shows how to compute sensitivity for tractable policy metrics. We provide an axiomatic development and technical details in (ref), and focus on the main ideas below, all calculations are deferred to (ref).
Let $\matheuler{P}$ be a statistical model. We are interested in estimating parameter $\psi:\matheuler{P}\rightarrow\mathbb{R}$ or, possibly, a set of alternative specifications $\Set{\psi_a:\matheuler{P}\rightarrow\mathbb{R}}{a\in \matheuler{A}}$ defined on the same model. Statistical functionals estimable at the parametric rate $\sqrt{n}$ are smooth. Therefore we can define sensitivity as a directional derivative of $\psi$ along a tangent vector $v$ to the model $\matheuler{P}$ at the sampling (true) distribution $P_0$. Tangent vector $v$ is the score of a one-dimensional parametric submodel $t\mapsto P_t$ defined in a neighborhood of $0\in [0,\epsilon)$:
For the purposes of interpreting sensitivity, score $v$ stands for any submodel that satisfies above derivative condition in quadratic mean. All such submodels admit the same local counterfactual interpretation of sensitivity. The collection of different scores $v$, obtained from all smooth submodels through $P_0$, is called the tangent set, denoted $T_{P_0}\matheuler{P}$. On a fully nonparametric model, the tangent set is the space $L^2_0(P_0)$ of $P_0$ square-integrable functions with zero mean. Parametric and semiparametric models restrict the tangent set in significant ways. Because we are not concerned with efficiency here, we can assume that the tangent set is unrestricted.
Sensitivity of $\psi$ along the tangent vector $v\in L^2_0(P_0)$ is the local effect of changing the distribution in the direction of score $v$:
Here the perturbation $P_0+tv$ is understood to be any one-dimensional submodel $P_t$ with score $v$. For example, $dP_t = (1 + tv)dP_0$ or $dP_t=c(t)\exp(tv)dP_0$. To compute sensitivity we can use the influence function of $\psi$:
As defined above, sensitivity is not very useful. The problem is that tangent space $T_{P}\matheuler{P}$ typically does not have an obvious parametrization that would enumerate all scores and put different sensitivities into context of the application layer. To make sensitivity analysis convenient for the practitioner, the direction $v$ should be associated with a tractable parameter of interest to the researcher. This can be a statistical functional motivated by the application layer, e.g. a related economic quantity or an alternative specification of the same quantity. Or this can be a data level parameter that provides a tractable summary of distribution $P_0$, e.g. a mean or a quantile. We call this parameter a control functional and denote it by $\nu:\matheuler{P}\rightarrow\mathbb{R}$.
The natural direction to associate with $\nu(P)$ is the gradient where functional increases most rapidly. This is analogous to the way Cartesian coordinates work, if we think of coordinates as functions of the point. However, it is not enough to pick a control functional to specify the direction of sensitivity. This should not be surprising, because $T_{P}\matheuler{P}$ is a large space, for which we have not introduced any structure.
Gradients depend on the notion of distance on the model $\matheuler{P}$. A metric at $P\in\matheuler{P}$ is an inner-product norm $\norm{\cdot}_P$ on tangent vectors $T_P\matheuler{P}$. The distance between $P_0$ and $P_\epsilon$ along submodel $P_t$ is the sum of lengths of tangent vectors along the curve:
Different metrics define different distances on $\matheuler{P}$ and generate the different geometries.
Influence function $\wt\nu$ is the gradient of $\nu$ according to the information geometry of $\matheuler{P}$ that has metric $\norm{v}^2_P = \int v^2 \, dP$. Information $\norm{v}_{L^2(P)}$ measures statistical discrepancy between $P$ and a perturbation $P+\epsilon v$ in the direction of score $v$. Influence function is the direction on the model along which change in the value of the functional is greatest per statistical deviation away from $P$. This direction is least favorable on the model for estimating $\nu$ from random samples of $P$.
Calculation of influence functions is a standard exercise in efficiency literature, we refer to Ichimura and Newey (2015) ichimura2015influence for a modern treatment and use their formula as a convenient definition:
Let us fix a simple example. Let the target functional be a generic moment of data $\psi_\rho(P) = \int\rho(x)dP$, and let the control functional be a quantile of data $\nu_\tau(P) = F^{-1}_{X(j)}(\tau)$. The mean and the $\tau$-quantile have influence functions $$ \wt\psi_\rho(x) = \rho(x) - \psi_\rho(P) \quad \text{and} \quad \wt \nu_\tau (x)=\dfrac{\tau - 1_{[x_i,\infty)}(\nu_\tau (P))}{f_{X(i)}(\nu_\tau(P))} . $$ The information sensitivity of the mean to the quantile
is the asymptotic covariance of regular efficient estimators $\wh\psi,\wh\nu$.
To interpret this, rescale $\Lambda(\psi,\nu)\coloneqq \partial_\nu\psi / \norm{\wt\nu}^2_P$ and recall the original definition of information (in infinite-dimensional models) form Koshevnik and Levit (1976) Koshevnik76: $\Lambda$ is the effect on the mean $\psi(P_0)$ of a perturbation to $P_0$ along a one-dimensional submodel $P_h$ that satisfies two requirements:
Information sensitivity $\Lambda$ measures the effect of this perturbation on the counterfactual value of the mean: $$ \psi_\rho(P_h)= \psi_\rho(P_0) + h \Lambda(\psi_\rho,\nu_\tau) + o(h). $$
Sensitivity to perturbations along the least favorable submodel is interesting for comparing statistical properties of estimators. For example, if $\psi,\nu$ are two alternative specifications for the same economic quantity, then information sufficiency $\Delta(\psi,\nu)\coloneqq \abs{\partial_\nu\psi}^2/\norm{\wt\psi}^2_P\norm{\wt\nu}^2_P$ is a natural measure of local similarity of the two estimates. But the choice of least favorable submodel as the counterfactual distribution when measuring the response in $\psi$ to changes in $\nu$ has no structural or causal foundation. Our point is to make this choice explicit.
A general sensitivity of parameter $\psi$ can thus be specified by a combination of:
To contrast general sensitivity with information sensitivity, we will call the metric used to determine gradients a policy metric, the direction along which the sensitivity is measured a policy gradient, and the directional derivative itself a policy sensitivity. Control functionals and a policy metric provide a partial parametrization of the tangent space $T_P\matheuler{P}$ that enables local counterfactual analysis motivated by the application.
A tractable way to specify a policy metric is to postulate a policy distribution $Q_P$ whose density function $dQ_P(x)$ reflects the cost of displacing a unit of mass at location $x$ in the sample space. The choice $Q_P$ should be motivated by the application. The resulting policy metric is $\norm{v}_{L^2(Q_P)}^2 = \int \abs{v}^2dQ_P$. Policy sensitivity with this metric is
where the scaled gradient $v = \nabla\nu / \norm{\nabla\nu}^2_{L^2(Q)}$ is the score $\tfrac{d}{dh}_{|h=0} \log dP_h$ of a one-dimensional submodel $P_h$ that solves the following program for a sufficiently small $\epsilon$:
Under some regularity conditions, policy gradient of functional $\nu:\matheuler{P}\rightarrow\mathbb{R}$ with respect to policy metric $\norm{\cdot}_{L^2(Q)}$ is
The effect of changing the metric from information to policy is very intuitive: the influence function is rescaled by the likelihood ratio of information to policy and recentered. Policy sensitivity measures the effect on the counterfactual value of target functional $\psi$ from the perturbation to the value of control functional $\nu$ along any submodel with policy gradient $\nabla\nu$: $$ \psi(P_h)= \psi(P_0) + h S(\psi,\nu) + o(h). $$
Let $X,Y$ be continuously distributed according to joint distribution $P$ on the interval $[0,1]^2$, and suppose that $Y$ is a measure of income, $X$ is a measure of education. Application layer postulates that $Y$ is a response variable, whereas $X$ is a control variable of intereset. Let the target and control functionals $$ \psi(P) = \int_{[0,1]^2} y dP \qquad \text{and} \qquad \nu(P) = F_X^{-1} \big(\tfrac{1}{2} \big) $$ be the mean of response variable $Y$ and the median of control variable $X$. In our simulation, we take $$ Y|X \sim \mrm{Beta}(\alpha,\beta) \quad \text{with} \quad \alpha=2, \;\beta=5-5X \quad \text{so that} \quad E[Y|X] = 2/(7-5X) $$ the conditional mean of income given education is positively correlated with education. Marginal distribution of $X$ is shown on (ref), marked {\footnotesizesamplingPDF(x)}.
We are interested in the predictive effect on income, via target functional $\psi(P)$, of a policy that perturbs the marginal distribution of education. It is assumed that the perturbation does not change the conditional distribution $P_{Y|X}$. Policy is designed to increase the median level of education $\nu(P)$ by some prescribed amount ($0.1$ in the simulation). Three implementations of policy are proposed.
$P_X:$ The perturbation along the least favorable submodel in the direction of the influence function of the median $\nu$ has the effect $\Lambda(\psi,\nu) = 0.3041$ on the mean $\psi$. The influence function is marked {\footnotesizeinfluenceFunction(x)} in (ref). The density function of the counterfactual distribution $dP_h = (1+h \wt\nu)dP$ that produces $\nu(P_h)\approx 0.6$ is marked {\footnotesizeinfoCfPDF(x)} in (ref). The information counterfactual value of the mean is $\psi(P_h)\approx \psi(P) + 0.3041 \times 0.1$.
$Q_1:$ The first policy proposal minimizes the taxpayers' cost of policy. It is argued that increasing the proportion of highly educated workers and reducing the proportion of workers with most basic education is progressively more costly as one approaches the extremes of the distribution. This may be due to higher investment requirements of displacing workers at the extremes. This proposal is summarized with policy cost density function $dQ_1$, marked {\footnotesizepolicyPDF1(x)} in (ref). Distribution $Q_1$ defines policy metric $\norm{\cdot}_{L^2(Q_1)}$ on deviations from sampling distribution of education $P_X$ and produces a policy gradient function $\nabla_{Q_1}\nu$, marked {\footnotesizepolicyGrad1(x)} in (ref). The resulting counterfactual distribution, marked {\footnotesizepolicyCfPDF1(x)} in (ref), is closer to the original sampling distribution $P_X$ below the first and above the third quartiles, and further away at the interquartile range, compared to the information counterfactual. The counterfactual value of the median is $\nu(P+h\nabla_{Q_1}\nu) \approx 0.6$, and the policy sensitivity is $S_{Q_1}(\psi,\nu) = 0.2513$, so the counterfactual value of the mean is $\psi(P+h\nabla_{Q_1}\nu)\approx \psi(P) + 0.2513\times 0.1$.
$Q_2:$ The second policy proposal minimizes economic inequality by designing the perturbation to have the strongest effect at the lowest levels of education and tapering off toward the highest levels of education. This is achieved with policy distribution $dQ_2$, marked {\footnotesizepolicyPDF2(x)} in (ref), and confirmed by the counterfactual distribution marked {\footnotesizepolicyCfPDF2(x)} in (ref). The sensitivity $S_{Q_2}(\psi,\nu) = 0.2835$ fits in between the information sensitivity $\Lambda$ and the $Q_1$ policy sensitivity $S_{Q_1}$. This is explained by noting that the mean of $Y$ is positively related to the mean of $X$, and that deviations with more mass at the tails effect the mean stronger than deviations that displace more mass around the median of the distribution.
$Q_3:$ The third policy proposal minimizes the macroeconomic shock by assigning equal cost to deviations across all levels of education. Perturbation profile, the gradient $\nabla_{Q_3}\nu$, under policy metric $\norm{\cdot}_{L^2(Q_3)}$ is most similar to the influence function $\wt\nu$ of the information metric $\norm{\cdot}_{L^2(P_X)}$. This is because both the sampling distribution and the policy measure $Q_3$ are relatively flat. The similarity is reflected in the counterfactual distributions and sensitivities as well.
Counterfactual value of $\nu$ in each case is approximately $0.6$. We compute the sensitivity of $\psi$ to changes in $\nu$ under each of the four counterfactual distributions and report results in (ref).
In this section we illustrate sensitivity analysis with gmm and descriptive statistics. We consider gmm functionals on the nonparametric model $\matheuler{P}$ that is constrained only by regularity (smoothness, integrability) conditions. Application layer provides a parameter space $\Theta\subset\mathbb{R}^p$ for the economic quantity of interest $\vartheta$ and a vector of moment criterion functions $ g:\mathcal{X}\times\Theta \rightarrow \mathbb{R}^r, $ assumed to be sufficiently smooth in parameter $\theta$ and sufficiently integrable over the sample space $\mathcal{X} \subset \mathbb{R}^d$. Integrals with respect to distribution $P$ of data are written as $Pg(\theta) = \int_\mathcal{X} g(x,\theta)\,dP(x)$. It is assumed that the economic quantity $\vartheta$ is “over-identified”, meaning that $r>p$, and that $G\coloneqq P\partial_\theta g(\theta)$ and $\Omega\coloneqq P g(\theta)g(\theta)^T$ are full rank at $\theta=\psi(P)$.
Researcher specifies that the value of $\vartheta\in\Theta$ is given by a function $\psi:\matheuler{P}\rightarrow\Theta$ of the statistical model. gmm estimation is set up from the application layer assumptions that
These assumptions are typically optimality conditions of the interactions described by the application layer model or postulated by the researcher orthogonality conditions. Often these models are highly stylized and are not expected to describe real-world data precisely. Our view is that (ref) assumptions should not be taken literally to data, that the role of specification is nontrivial and deserves attention (but not our focus here). gmm functionals are defined by $$ \psi_W(P) \coloneqq \operatorname*{arg\,min}_{\theta\in\Theta} P g(\theta)^T W Pg(\theta). $$ In the over-identified case, weighting determines the functional and should be chosen based on application layer considerations. We consider only deterministic positive definite weighting matrices for now. We compute a set of information sensitivities to compare a given gmm functional $\psi_W$ to tractable summaries of the data and to alternative specifications $\psi_{A}$ obtained by using a different weighting. As directions we use descriptive statistics such as quantiles $q_\tau \coloneqq F_{X(i)}^{-1}(\tau)$ and generic moments $\nu_\rho(P)\coloneqq P\rho(X)$ of the data. Here the moment function $\rho:\mathcal{X} \rightarrow \mathbb{R}$ can be, for example, a component $\rho(x)=x_i$ of the data or a component of the moment criterion vector $\rho(x) = g_i(x,\psi_W)$ .
Information sensitivity is simple to compute and offers greater insight into inferences based on asymptotic approximations. Consider the gmm functional. The economic model that leads to formulation of functional $\psi_W$ may be complicated, but the asymptotic distribution of estimates, and inferences derived from it, are completely determined by the local behavior of the functional at $P$. Information sensitivities and the complementary sufficiency measures provide tractable one-dimensional summaries of this local variation: $$ \partial_\nu \psi_W = P \wt\psi_W \wt\nu \quad \text{and} \quad R(\psi_W, \nu) = (P \wt\psi_W \wt\nu)^2 / P \wt\psi_W^2 P \wt\nu^2 . $$ Information sufficiency is an $R^2$ statistic that indicates how well the control functional $\nu$ approximates local variation of the target functional $\psi_W$. Specifically, $R$ is the square of cosine of the angle made by tangent hyperplanes to $\psi$ and $\nu$. If $R(\psi_W,\nu)$ is close to one, then inferences based on asymptotic approximations around $\psi_W(P)$ are obtained from $P$ in the same way as inferences about $\nu(P)$. By making local comparisons of complicated structural functionals $\psi_W$ to simple features of the data $q_\tau, \nu_\rho$, the statistical part of the empirical analysis can be made transparent Gentzkow2015,andrews2017measuring. Another application is to compare two competing specifications $\psi_W$ and $\psi_A$ locally in the neighborhood of $P$. Reporting $R(\psi_W,\psi_A)$ close to one can be offered as formal evidence that the choice of weighting does not change results in a neighborhood of $P$. Conversely, observing $R(\psi_W,\psi_A)$ close to zero warrants careful examination of specification.
Asymptotic distribution of gmm estimators on misspecified models has been investigated by Imbens (1997) imbens1997one, Hall and Inoue (2003) hall2003large, we derive the influence function and policy gradients of the functional in order to provide sensitivity analysis. The influence function of the gmm functional on a fully nonparametric model where moment conditions (ref) are possibly violated is
The sign of sensitivity $\partial\psi_{W,i} /\partial\psi_{A,i} = P \wt\psi_{W,i} \wt\psi_{A,i}$ shows if the two specifications for $\vartheta_i$ move in the same direction at $P$, and sufficiency $R(\psi_{W,i},\psi_{A,i})$ quantifies the alignment of two specifications locally at $P$. Furthermore, sufficiency $R(\psi_{W,i}, \nu_{g(j)})$ measures the amount of local variation in the estimate of $\vartheta_i$ contributed by the local variability of $j$th moment function at $P$.
Policy sensitivity $$ S(\psi_W, \nu) = \int \wt\psi_W \Big[ \wt\nu - P \wt\nu \tfrac{dP}{dQ} / P \tfrac{dP}{dQ} \Big] \tfrac{dP}{dQ} \; dP / \int \wt\nu \Big[ \wt\nu - P \wt\nu \tfrac{dP}{dQ} / P \tfrac{dP}{dQ} \Big] \tfrac{dP}{dQ} \; dP $$ gives the local counterfactual value $\psi_W(P) + h\cdot S(\psi_W,\nu) + o(h)$ of the economic quantity $\vartheta$ identified with $\psi_W$ to the perturbation of size $h$ in the value of statistical parameter $\nu(P)$ according to policy metric $L^2(Q)$. Measure $Q$ can be a policy relevant reference distribution on the sample space. Taking the empirical measure $P$ as policy measure is a convenient choice in terms of estimation.
Let $\matheuler{P}$ be a collection of probability measures on a sample space $(\mathcal{X},\mathcal{A})$. The starting point for our investigation is to realize a statistical model as an object with intrinsic geometry -- a space with notions of smoothness, length and angle. In this section we consider a special case of a two-dimensional statistical model and employ graphical aid to provide a nontechnical exposition. The idea is to map a two-dimensional statistical model onto a surface in $\mathbb{R}^3$ while preserving the intrinsic metric properties of the model. We can then forget about the set of probability measures and work with the surface in $\mathbb{R}^3$. For details on geometry of surfaces we refer to carmo1976differential.
The natural space to host statistical models is the set of square-integrable functions $L^2(\mu)$, with some dominating measure $\mu$ for elements of the model $\matheuler{P}$. In this ambient space, probability distributions are identified with square-roots of their densities $dP^{1/2} \coloneqq \sqrt{\tfrac{dP}{d\mu}}$, the model $\matheuler{P}$ is a subset of the unit ball of $L^2(\mu)$, and the tangent set $T_P\matheuler{P}$ is a subset of a hyperplane in $L^2(\mu)$. This simple setup provides a lot of structure to the model $\matheuler{P}$, in particular, the information distance between two distributions $P_0$ and $P_1$ is the length of the shortest curve on the model joining them. The length of a curve $\alpha:[0,1]\ni t\mapsto P_t \in \matheuler{P}$ is obtained by adding magnitudes of velocity vectors along the curve:
The curve in $L^2(\mu)$ is $t\mapsto dP^{1/2}_t$. Its velocity at time $t$ is the tangent vector $v_t(x) = \tfrac{d}{dh}_{|h=t} dP_h^{1/2}(x)$, whose length, doubled for purely technical reasons, $\norm{v_t}^2 = \int_\mathcal{X} [2v(x)]^2 d\mu$ is the information metric norm. Finally, the sum of velocities along the trajectory of the curve $\int_{[0,1]} \norm{v_t} dt$ is, by definition, the length of the curve.
The problem of embedding $\matheuler{P}$ into $\mathbb{R}^3$ is to find a surface $S\subset\mathbb{R}^3$ such that length of the image of any curve $\alpha$ on $S$, computed according to the Euclidean geometry of $\mathbb{R}^3$, coincides with the value in (ref). Isometric embedding is an active area of research. Conditions for preserving the metric are formulated with a system of partial differential equations whose solvability requires enough degrees of freedom provided by the dimensionality of ambient space. A general 2-manifold can be embedded into $\mathbb{R}^{10}$ by Nash's theorem and its extensions han2006isometric. The metric ultimately determines the shape of the surface required for the embedding.
We consider three examples of statistical models with constant Gauss curvature:
From a statistical model $\matheuler{P}$ and its information metric we obtain a surface $S\subset\mathbb{R}^3$ and from a statistical functional $\psi(P)$ we obtain a function $f:S\rightarrow\mathbb{R}$ defined on the points of the surface. We use the surface to show that local behavior of $f$ at $P\in S$, summarized by its derivative, determines the asymptotic behavior of estimates of $\psi(P)$. Calculations near point $P$ on $S$ are carried out by means of a parametrization by an open subset $U\subset\mathbb{R}^2$. There are many choices of a parametrization $$ \mathbf{x}: \mathbb{R}^2\supset U\ni(u,v) \mapsto \big(x(u,v), y(u,v), z(u,v)\big)\in S\subset\mathbb{R}^3 $$ around a point $P$, the only requirements are that $\mbf{x}$ be differentiable with derivative $d\mbf{x}_q:\mathbb{R}^2\rightarrow\mathbb{R}^3$ that is full rank for all $q\in U$. For example, $\matheuler{P}_{\mrm{sph}}$ can be parametrized by $x,y$ or $y,z$ or $z,x$ coordinates of its points, or by latitude and longitude, or by points of the inscribed simplex. Parametrization deforms a flat two-dimensional neighborhood $U$ by stretching, shrinking and bending onto a neighborhood $V$ of the surface. Because of the deformation, distances and angles in $U$ are different from those in $V$. Calculations in each parametrization appear to be different but the values on the surface $S$ are invariant similarly to how mle is parametrization invariant.
Differential calculus works on tangent vectors that are the infinitesimals. At every point $P\in S$ there is a unique tangent plane $T_PS\subset\mathbb{R}^3$ to the surface. The derivative $d\mbf{x}_q$ of the parametrization maps vectors in $U$ anchored at $q$ into tangent vectors in $T_{\mbf{x}(q)} S$. Tangent vectors $\mbf{x}_u = d\mbf{x} e_1$ and $\mbf{x}_v =d\mbf{x} e_2$ span $T_{\mbf{x}(q)}S$ and are known as scores\footnote{I would appreciate a reference to the etymology of this terminology.}. Due to deformation by $\mbf{x}$, orthonormal vectors $e_1,e_2$ in $U$ have images $\mbf{x}_u, \mbf{x}_v$ that are not orthogonal and not unit length in $\mathbb{R}^3$. This is because the model $\matheuler{P}$ is not flat at $P$ in its metric. Consequently sensitivity $\partial_v u$ of parameters $u,v$ on $S$ is not zero. A function $f:S\rightarrow\mathbb{R}$ is differentiable if its expression in local coordinates $f\circ\mbf{x}$ is differentiable. The derivative $df_P:T_PS\rightarrow\mathbb{R}$ maps tangent vectors to $S$ at $P$ into vectors in $\mathbb{R}$ anchored at $f(P)$.
Recall that we took care to preserve distances and angles while mapping model $\matheuler{P}$ into surface $S$. The $\mathbb{R}^3$ inner product $\brk{\,\cdot}{\cdot\,}_P$ induced on vectors of the tangent plane $T_PS$ is in agreement with intrinsic metric structure of the statistical model $\matheuler{P}$. This intrinsic statistical metric determines the sensitivity $\partial_\nu\psi$ of statistical functional $\psi(P)$ to another parameter $\nu(P)$ as follows. By a basic fact of linear algebra, the linear map $df_P$ has a simple representation by the gradient vector $\nabla f_P$ of function $f$. The gradient $\nabla f_P \in T_PS$ is the unique tangent vector that satisfies
Gradient $\nabla f_P$ points in the direction on the surface along which values of $f$ increase most rapidly and has magnitude $\norm{\nabla f_P}_{\mathbb{R}^3}$ equal to the rate of the increase at $P$ on model $\matheuler{P}$. Since gradient $\nabla \nu_P$ determines the linearization $w\mapsto\brk{\nabla \nu_P}{w}_P$ of functional $\nu$ at $P$, it is natural to take it to be the “$\nu$-direction" of the model at $P$. This is in perfect analogy with the direction of $u$-axis in $U$ where the $u$ coordinate is the linear function $w\mapsto\brk{e_1}{w}_{\mathbb{R}^2}$ on $U$ with gradient $\nabla u = e_1$. This motivates our measure of local statistical dependence:
Next we use parametrization to compute the derivative $\partial_\nu\psi$ and establish that parameter sensitivity of (ref) and estimator sensitivity of (ref) agree for many estimators, specifically that $\sigma_{\nu\nu}\Lambda(\wh\psi,\wh\nu) = \sigma_{\psi\nu}=\partial_\nu\psi$. Let $ E(u_0,v_0) = \brk{ \mbf{x}_u}{ \mbf{x}_u }_P$, $F(u_0,v_0) = \brk{ \mbf{x}_u}{ \mbf{x}_v }_P$ and $G(u_0,v_0) = \brk{ \mbf{x}_v}{ \mbf{x}_v }_P$ denote the expression of the $\mathbb{R}^3$ inner-product on $T_PS$ in local coordinates. And let $$ I_{u,v} =
$$ denote the Fisher information matrix for this parametrization. Information matrix appears in the expression for sensitivity because it reconciles distorted distances in $U$ with statistical distances on $S$. In local coordinates, $f\circ\mbf{x}$ can be differentiated as usual to obtain partial derivatives $f_u,f_v$; these are the directional derivatives of $f$ on $S$ along scores $\mbf{x}_u,\mbf{x}_v$. From relationships $\brk{\nabla f}{ \mbf{x}_u } = f_u$ and $ \brk{\nabla f}{ \mbf{x}_v } = f_v$ we solve for the expression of $\nabla f$ in $\set{\mbf{x}_u,\mbf{x}_v}$ basis:
In this section we define general sensitivity measures of two statistical parameters. Sensitivity is defined through differential calculus of a statistical functional. Functionals are real-valued maps of a set of possible distributions of each observation. Sensitivity quantifies the local relationship between a functional of interest and any set of regular functionals. We can relate this to regression and designate the functional of interest as response or target and the set of regular functionals as controls. We only consider sensitivity to a single control, but the extension to a set of controls is straightforward and the partialling out reasoning of regression applies. Sensitivity measures the deviation in the value of response functional under the perturbation of the value of control functional. Unlike regression coefficients, sensitivity is a bona fide directional derivative. The direction depends on local properties of the control statistical functional and the notion of distance between two distributions. Sensitivity can be used to make local counterfactual inferences about economic quantities of interest in empirical work and to gain greater insight into asymptotic distributions.
The set of possible sampling distributions is generally not linear, but can be modeled, in a neighborhood of every point, as a smoothly transformed open subset of some linear space. The idea takes some effort to develop methematically but the result provides great intuition. We introduce necessary elements of Riemannian geometry for completeness, and refer to do Carmo (1976, 1992) carmo1976differential, carmo1992riemannian and Lang (1999) lang1999fundamentals for more details. Most elements of differential geometry that we need to define sensitivity are also employed in the semiparametric efficiency literature. However, efficiency theory makes use of the ambient Hilbert space $H_2$ of square roots of measures Koshevnik76. From $H_2$ the model inherits the differential structure (pathwise differentiability) and the information metric (Hellinger distance), similarly to our use of $\mathbb{R}^3$ in (ref). By contrust, we define sensitivity based on a development of differential calculus on the model without an ambient space and make dependence of sensitivity on the metric explicit. Our development is similar to the setup in van der Vaart (1991) vaart1991differentiable.
The point here is to allow local counterfactual “policy” analysis at the population level to be independent of the asymptotic approximations and statistical efficiency analyses. We allow a general Riemannian metric on the model manifold, which we call a policy metric, to be used for sensitivity measures at the population level. We describe how these policy sensitivities can be obtained from asymptotic distributions in (ref).
Let $M$ be a collection of distributions $P$ on sample space $(\mathcal{X},\mathcal{A})$. We introduce a differentiable structure on $M$; this enables us to consider smooth functions $\psi,\nu:M\rightarrow\mathbb{R}$ which can be approximated on $M$ at a given point $P$ along directions $v\in T_PM$ of the tangent space; differential $d\psi:T_PM\rightarrow\mathbb{R}$ provides linear approximation of $\psi$ along any direction $v$; metric $g$ is an inner-product on tangent spaces $T_PM$ that provides a Riesz representations $\nabla\nu_P$ of the differentials $d\nu_P$ of $\nu$; finally, the sensitivity $\partial_\nu\psi$ is the directional derivative $d\psi(\nabla\nu)$ of $\psi$ along the gradient direction of $\nu$.
We consider only the simplest case of an open manifold. Extensions that allow for manifolds with boundaries, corners, etc., common with statistical models, are possible but are not considered here. Tangent sets for the purposes of this paper are always complete linear spaces. Our approach to start with an arbitrary manifold structure and consider inclusion into the space of square roots of measures $H_2$ can be used to restrict the tangent space in an explicit way and allows us to consider any metric in definition of sensitivity.
Statistical models can have many parametrizations. For example, the $N(\mu, I_3)$ family usually parametrized by the vector of means $\mu\in\mathbb{R}^3$, can alternatively be specified using spherical coordinates $(\norm{\mu}, \tan^{-1}(\mu_2/\mu_1), \cos^{-1}(\mu_3/\norm{\mu})$; a (regression) function can be parametrized by the coefficients of different Fourier bases. Parametrizations are necessary for computation, but as long as we consider only compatible parametrizations, calculations we do and quantities we define will be invariant of the chosen parametrization. A differential structure is an equivalence class of compatible parametrizations. A manifold is a set with a differentiable structure.
An atlas on $M$ is a collection of local parametrizations (charts) $(U_i,\varphi_i)$ satisfying the following conditions:
Let $M, N$ be manifolds. A map $f:M\rightarrow N$ is differentiable if, given $P\in M$, there are charts $(U,\varphi)$ at $P$ and a chart $(V,\psi)$ at $f(P)$ such that $f(U)\subset V$ and the composition $\psi \circ f \circ \varphi^{-1}:\varphi U\rightarrow\psi V$ is differentiable as a map between normed linear spaces. The composition is called expression of $f$ in local coordinates. Similarly, we define directional and compact differentiation 0036-0279-22-6-R06,0036-0279-23-4-R02,penot2016analysis by applying the definition to the expression of $f$ in local coordinates.
Let $E,F$ be Banach spaces. A tangent vector in $E$ is a direction $v\in E$ with a position $P\in E$. Given a smooth curve $ [0,\epsilon)\ni t \mapsto P_t \in E $ with position $P$ and direction $v=\frac{d}{dt} P_{t} \in E$ at time $t=0$, and a differentiable map $f:E\rightarrow F$, we can associate the tangent vector $v$ with the directional derivative operator $$ \tfrac{d}{dt} f(P_t)_{\big |t=0} = Df_{P_t} \, (\tfrac{d}{dt} P_t )_{\big| t=0} = D_v f (P). $$ Let $M$ be a manifold modelled on a Banach space $E$. A curve on $M$ is a differentiable map $\alpha:[0,\epsilon)\rightarrow M$. A tangent vector at $\alpha(0)=P\in M$, corresponding to the direction of $\alpha$, is the directional derivative operator $\alpha'(0)$ on differentiable maps $f:M\rightarrow\mathbb{R}$ $$ \alpha'(0) f = \tfrac{d}{dt} (f \circ a)_{\big| t=0}. $$ The set $T_PM$ of all tangent vectors to $M$ at point $P$, obtained from all curves passing through $P$, is called the tangent space. A tangent vector $\alpha'(0)$ corresponds to the direction $v\in E$ of the expression $\varphi\circ\alpha_t$ of the curve in local coordinates. Tangent space $T_PM$ is in bijective correspondence with $E$ and has the same structure of a topological vector space.
Geometric primitives discussed above are closely related to the ideas employed in semiparametric efficiency. The next geometric primitive is implicit and fixed in the efficiency bounds theory but has an active role in our local counterfactual analysis of functionals. The idea is to give the statistical model $M$ a notion of distance by giving each tangent space an inner product. In semiparametric efficiency theory this object is called information and it measures the “statistical” (Hellinger) distances between distributions. However, the empirical researcher identifies statistical functionals $\psi,\nu:M\rightarrow\mathbb{R}$ with economic quantities $\vartheta,\eta$ and wants to understand the local relationship between $\vartheta$ and $\eta$ at the data generating point $P$ on the model $M$. There is no reason to assume that “economic” distances on $M$ coincide with “statistical” distances. Therefore we consider a completely general metric for policy analysis purposes.
A Riemannian metric on a statistical model $M$ is a correspondence $g$ that assigns to every point $P\in M$ a continuous bilinear symmetric positive-definite form $g(\cdot,\cdot)_P$ on the tangent space $T_P M$, and varies smoothly over $M$. For direction $v\in T_P M$ we can think of the norm $\abs{v}_g\coloneqq\sqrt{g(v,v)_P}$ as the economic cost of a deviation from $P$ on $M$ at rate $v$. We will call $g$ a policy metric to contrast it with the statistical metric given by information $\sqrt{ \int v^2 dP}$.
The metric determines gradient directions of functions of $M$ as follows.
From definition it is clear that gradient of $\psi$ depends on the metric $g$. The choice of metric determines the problem of approximating $\psi$ with a single tangent vector. By Cauchy-Schwarz, $$ \sup_{\abs{v}\le 1} d\psi_P(v) \le \abs{\nabla^g\psi_P}. $$ According to the metric $g$, gradient is the direction of most rapid increase in the value of the function. The norm $\abs{\nabla \psi_P}_g$ is the slope of the tangent to the restriction of $\psi$ along any curve through $P$ with unit speed and direction $\nabla \psi_P$.
Clearly numbers $\partial_\nu\psi, S_\nu\psi,R_\nu\psi \in \mathbb{R}$ and the linear map $\Pi_\nu\psi\in T_PM^*$ depend on the choice of metric $g$ through gradients of $\psi,\nu$. Directional derivatives $\partial_\nu\psi$ and $S_\nu\psi$ measure response in the value of $\psi(P)$ to a perturbation in the value of $\nu(P)$ that is achieved by a deviation from $P$ on $M$ in the direction of most rapid change in $\nu$. This is analogous to partial derivatives in linear spaces with respect to functionals of a coordinate system. Projection vector $\Pi_\nu\psi$ gives the local approximation of $\psi$ by its partial derivative along $\nu$ in all directions on $M$; this is the regression of $\psi$ onto $\nu$ locally at $P$. An interesting fact is that the coefficient of this local regression, the sensitivity, is a genuine derivative in this case. Sufficiency is the coefficient of determination in this regression and measures the alignment of $\psi(P)$ and $\nu(P)$ in a neighborhood of $P$ on M. Specifically, $R_\nu\psi$ is the square of cosine of the angle between $\nabla\psi$ and $\nabla\nu$. A value of $R(\psi,\nu)$ close to $1$ reflects high degree of similarity in the local behavior of $\psi,\nu$ at the data generating distribution; a value close to $0$ reflects that $\psi,\nu$ move in orthogonal directions of the model $M$. When $R(\psi,\nu)$ is close to $1$ any perturbation that moves $\nu$ will have a proportional effect on the value of $\psi$, where as with $R(\psi,\nu)$ close to $0$ any perturbation that significantly moves $\nu$ will have negligible effect on the value of $\psi$.
Extension to a set of control functionals is straightforward by analogy with regression. Here sensitivities are coefficients of the projection of $\nabla\psi$ onto the linear span of $\nabla\nu_1,\ldots,\nabla\nu_p$. The interpretation of sensitivity coefficient of $\nu_1$ is as above but for the local variation in $\psi$ and $\nu_1$ that is orthogonal to the linear span of $\nabla\nu_2\ldots\nabla\nu_p$ frisch1933,lovell1963.
Let $M$ be a semiparametric model described in (ref), let $\psi,\nu:M\rightarrow\mathbb{R}$ be Hadamard differentiable functionals on $M$. In this section we consider estimation based on random samples from $P\in M$ and relate the asymptotic distribution of estimators $ \mathchoice {\accentset{\displaystyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} {\accentset{\textstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} {\accentset{\scriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} {\accentset{\scriptscriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} , \mathchoice {\accentset{\displaystyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} {\accentset{\textstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} {\accentset{\scriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} {\accentset{\scriptscriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} $ to the sensitivity measures $\partial_\nu\psi, S_\nu\psi$ defined in (ref).
Efficiency bounds on the asymptotic distribution of regular estimators $( \mathchoice {\accentset{\displaystyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} {\accentset{\textstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} {\accentset{\scriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} {\accentset{\scriptscriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\psi}} , \mathchoice {\accentset{\displaystyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} {\accentset{\textstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} {\accentset{\scriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} {\accentset{\scriptscriptstyle \text{\smash{\raisebox{-1.3ex}{ $\widehatsym$}}}}{\nu}} )$ depend on local properties of functionals $\psi,\nu$ on the image $\matheuler{P}$ of the inclusion $$ i:M\rightarrow{} H_2 $$ of the model manifold into the Hilbert space of square roots of measures, see Koshevnik and Levit (1976) Koshevnik76 for the role of this embedding and Neveu (1965) neveu1965mathematical for the definition of the space $H_2$.\footnote{Koshevnik76 cite neveu1965mathematical but the English translation of Koshevnik76 references pages in the Russian translation of neveu1965mathematical.} We collect details of semiparametric efficiency theory in (ref). Our setup with inclusion of $M$ into $H_2$ is similar to van der Vaart (1991) vaart1991differentiable, but we emphasise the intrinsic geometry of the model where as vaart1991differentiable is concerned with pathwise differentiability. The following is a standard
Manifold $M$ determines the set of pathwise differentiable one-dimensional submodels $\matheuler{P}(P)$ and the tangent space $T_P\matheuler{P} = A[T_PM] = R(A)\subset L^2_0(P)$, which are important elements of the efficiency theory. Note that differential $A$ need not be isomorphic and need not be isometric. If range of $A$ is not closed in $L^2(P)$, then $A^{-1}$ is not bounded, and bilinear functional $g$ is not continuous on the tangent space $T_P\matheuler{P}$. For example, $T_PM=H^k$, the Sobolev space of $L^2(P)$ functions with $k$ derivatives.
We make the following stronger assumption that simplifies our functional analysis. Roughly speaking, we consider models that behave either like finite dimensional smoothly parametrized families or like fully nonparametric models.
It follows that metric $g$ is continuous on the embedded tangent space $(T_P\matheuler{P},\brk{\cdot\,}{\cdot\,}_P)$, and the $L^2(P)$ inner-product $\brk{\,\cdot}{\cdot\,}_P$ is continuous on the manifold tangent spce $(T_PM,g)$, and that inclusion differential $A$ has an adjoint $A^*:T_P\matheuler{P}\rightarrow T_PM$ such that $$ \brk{Au}{v}_P = g(u, A^*v) \quad \text{for every } u\in T_PM, v\in T_P\matheuler{P}. $$ Since $A$ is the derivative of the inclusion map, we can treat it as the identity operator on $T_PM$. Furthermore, from functional analysis identity $(\operatorname*{Ker} A)^\perp = \@ifnextchar^{{ \setbox0\hbox{${\mathaccent"0362{\operatorname*{Ran} A^*}}^H$} \setbox2\hbox{${\mathaccent"0362{\kern0pt\operatorname*{Ran} A^*}}^H$} \ifdim\ht0=\ht2 \begingroup \def\mathaccent#\operatorname*{Ran} A^*#0{ \if22 \let\macc@nucleus \fi \setbox\z@\hbox{$\macc@style{\macc@nucleus}_$} \setbox\tw@\hbox{$\macc@style{\macc@nucleus}_$} \dimen@\wd\tw@ \advance\dimen@-\wd\z@ \divide\dimen@ 3 \@tempdima\wd\tw@ \advance\@tempdima-\scriptspace \divide\@tempdima 10 \advance\dimen@-\@tempdima \ifdim\dimen@>\z@ \dimen@0pt\fi \kern0.6\dimexpr\macc@kerna\kern-\dimen@ \if21 \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\macc@nucleus\kern0.4\dimexpr\macc@kerna\kern\dimen@} \advance\[email removed]\dimexpr\macc@kerna \let\final@kern0 \ifdim\dimen@<\z@ \let\final@kern1\fi \if\final@kern1 \kern-\dimen@\fi \else \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\operatorname*{Ran} A^*} \fi } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \if21 \macc@nested@a\relax111{\operatorname*{Ran} A^*} \else \def\gobble@till@marker#\operatorname*{Ran} A^*\endmarker{} \futurelet\gobble@till@marker\operatorname*{Ran} A^*\endmarker \ifcat\noexpand A\else \def{} \fi \macc@nested@a\relax111{} \fi \endgroup \else \begingroup \def\mathaccent#\operatorname*{Ran} A^*#0{ \if12 \let\macc@nucleus \fi \setbox\z@\hbox{$\macc@style{\macc@nucleus}_$} \setbox\tw@\hbox{$\macc@style{\macc@nucleus}_$} \dimen@\wd\tw@ \advance\dimen@-\wd\z@ \divide\dimen@ 3 \@tempdima\wd\tw@ \advance\@tempdima-\scriptspace \divide\@tempdima 10 \advance\dimen@-\@tempdima \ifdim\dimen@>\z@ \dimen@0pt\fi \kern0.6\dimexpr\macc@kerna\kern-\dimen@ \if11 \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\macc@nucleus\kern0.4\dimexpr\macc@kerna\kern\dimen@} \advance\[email removed]\dimexpr\macc@kerna \let\final@kern0 \ifdim\dimen@<\z@ \let\final@kern1\fi \if\final@kern1 \kern-\dimen@\fi \else \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\operatorname*{Ran} A^*} \fi } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \if11 \macc@nested@a\relax111{\operatorname*{Ran} A^*} \else \def\gobble@till@marker#\operatorname*{Ran} A^*\endmarker{} \futurelet\gobble@till@marker\operatorname*{Ran} A^*\endmarker \ifcat\noexpand A\else \def{} \fi \macc@nested@a\relax111{} \fi \endgroup \fi }}{ \setbox0\hbox{${\mathaccent"0362{\operatorname*{Ran} A^*}}^H$} \setbox2\hbox{${\mathaccent"0362{\kern0pt\operatorname*{Ran} A^*}}^H$} \ifdim\ht0=\ht2 \begingroup \def\mathaccent#\operatorname*{Ran} A^*#1{ \if22 \let\macc@nucleus \fi \setbox\z@\hbox{$\macc@style{\macc@nucleus}_$} \setbox\tw@\hbox{$\macc@style{\macc@nucleus}_$} \dimen@\wd\tw@ \advance\dimen@-\wd\z@ \divide\dimen@ 3 \@tempdima\wd\tw@ \advance\@tempdima-\scriptspace \divide\@tempdima 10 \advance\dimen@-\@tempdima \ifdim\dimen@>\z@ \dimen@0pt\fi \kern0.6\dimexpr\macc@kerna\kern-\dimen@ \if21 \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\macc@nucleus\kern0.4\dimexpr\macc@kerna\kern\dimen@} \advance\[email removed]\dimexpr\macc@kerna \let\final@kern1 \ifdim\dimen@<\z@ \let\final@kern1\fi \if\final@kern1 \kern-\dimen@\fi \else \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\operatorname*{Ran} A^*} \fi } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \if21 \macc@nested@a\relax111{\operatorname*{Ran} A^*} \else \def\gobble@till@marker#\operatorname*{Ran} A^*\endmarker{} \futurelet\gobble@till@marker\operatorname*{Ran} A^*\endmarker \ifcat\noexpand A\else \def{} \fi \macc@nested@a\relax111{} \fi \endgroup \else \begingroup \def\mathaccent#\operatorname*{Ran} A^*#1{ \if12 \let\macc@nucleus \fi \setbox\z@\hbox{$\macc@style{\macc@nucleus}_$} \setbox\tw@\hbox{$\macc@style{\macc@nucleus}_$} \dimen@\wd\tw@ \advance\dimen@-\wd\z@ \divide\dimen@ 3 \@tempdima\wd\tw@ \advance\@tempdima-\scriptspace \divide\@tempdima 10 \advance\dimen@-\@tempdima \ifdim\dimen@>\z@ \dimen@0pt\fi \kern0.6\dimexpr\macc@kerna\kern-\dimen@ \if11 \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\macc@nucleus\kern0.4\dimexpr\macc@kerna\kern\dimen@} \advance\[email removed]\dimexpr\macc@kerna \let\final@kern1 \ifdim\dimen@<\z@ \let\final@kern1\fi \if\final@kern1 \kern-\dimen@\fi \else \overline{\kern-0.6\dimexpr\macc@kerna\kern\dimen@\operatorname*{Ran} A^*} \fi } \macc@depth\@ne \let\math@bgroup\@empty \let\math@egroup\macc@set@skewchar \mathsurround\z@ \frozen@everymath{\mathgroup\macc@group\relax} \macc@set@skewchar\relax \let\mathaccentV\macc@nested@a \if11 \macc@nested@a\relax111{\operatorname*{Ran} A^*} \else \def\gobble@till@marker#\operatorname*{Ran} A^*\endmarker{} \futurelet\gobble@till@marker\operatorname*{Ran} A^*\endmarker \ifcat\noexpand A\else \def{} \fi \macc@nested@a\relax111{} \fi \endgroup \fi }$ and continuity of $A^{-1}$, conclude that $A^*$ has a continuous inverse $(A^*)^{-1}:T_P\matheuler{P}\rightarrow T_PM$, and
Estimator sufficiency is informative of the local statistical relationship between $\psi$ and $\nu$, and specifically, to what extent regular estimates of parameter $\nu(P)$ determine inferences based on the asymptotic distribution of regular estimators of $\psi(P)$ in the sense of efficiency gain.
Here we consider a tractable example of policy and obtain explicit relationships between policy gradients and influence functions. Consider a nonparametric model $M$, fix a distribution $P\in M$, and let the tangent space $T_P M=T_P\matheuler{P}=L^2_0(P)$ be unrestricted. Suppose that policy metric $g$ on $T_PM$ is given by
where policy distribution $Q$ satisfies the following regularity condition
Probability measure $Q$ may be a social weighting on sample space $\mathcal{X}$ that is relevant for policy. Policy probability density $dQ(x)$ is the cost of displacing a unit of mass in $P$ at location $x$ of the sample space $\mathcal{X}$.
We want to find the policy relevant response $\partial_\nu \psi = g(\nabla\psi, \nabla\nu)$ of the change to economic quantity associated with statistical functional $\psi$ that would result from a perturbation to $\nu$. This can be computed from influence functions, obtained as part of the asymptotic distribution derivation for estimators of $\psi,\nu$ or from an efficiency bound calculation. We assume that policy regularity condition (ref) and find the gradient operator $A^*$, which we can then verify to be isomorphic.
From definition (ref) we have the following relationships
It follows that for every $v\in L^2_0(P)$
so that $$
\mathchoice {\accentset{ \smash{\raisebox{-1.3ex}{ $\widetildesym$}}}{\psi}} {\accentset{\textstyle \smash{\raisebox{-1.3ex}{ $\widetildesym$}}}{\psi}} {\accentset{ \smash{\raisebox{-1.3ex}{ $\widetildesym$}}}{\psi}} {\accentset{\scriptscriptstyle \smash{\raisebox{-1.3ex}{ $\widetildesym$}}}{\psi}} = \nabla\psi \tfrac{dQ}{dP} - P [\nabla\psi \tfrac{dQ}{dP}] \quad and \quad \nabla \psi = \Big[\wt\psi + P[\nabla\psi \tfrac{dQ}{dP}] \Big]\tfrac{dP}{dQ} \qquad a.e. P,Q. $$ To solve for the centering constant $P[\nabla\psi \tfrac{dQ}{dP}]$, use the fact that $P\nabla\psi = 0$, to find that $P[\nabla\psi \tfrac{dQ}{dP}] = - P\wt\psi\tfrac{dP}{dQ} / P \tfrac{dP}{dQ}$. Conclude:
Thus, we have expressed the policy gradients $\nabla\psi,\nabla\nu$ in terms of the influence functions and can compute the policy sensitivity as follows.
Reporting sensitivity in empirical work requires estimating it along with the asymptotic variance. We consider two distinct scenarios. If the policy metric $g_P$ has a fixed relationship with the distribution $P$ of data, then estimating sensitivity is straightforward and consistency follows (roughly) from consistency of the asymptotic approximation. If the policy metric $g_P$ depends on the distribution $P$ in a general way, then gradient operator $A^*_P$ needs to be estimated and consistency requires additional justification.
We consider the typical situation where one estimates a vector of parameters $\theta\in\Theta$ and obtains an estimate of the asymptotic variance by plugging in the estimate $\wh\theta$ to obtain influence functions $\wt\psi_{\wh\theta},\wt\nu_{\wh\theta}$. Here $\psi,\nu$ could be some functions of $\theta$. We first assume that $g_P$ has a fixed relationship to $P$ so that the gradient operator $A^*$ is known. We use notation $\mathbb{P}_n = n^{-1}\sum_{i=1}^n\delta_{X_i}$ for the empirical measure.
With the additional assumption that functions $\Set{\wt\nu_\theta\cdot A^*\wt\nu_\theta}{\theta\in\Theta}$ are Glivenko-Cantelli, one can form a consistent estimator of sensitivity coefficient $S_\nu\psi$. More primitive conditions can be based on e.g. bracketing entropy. If an estimate of bracketing numbers is available for functions $\wt\nu_\theta$ and $A^*$ preserves point-wise order at each $x\in\mathcal{X}$ like the multiplication operator (ref), then one can estimate bracketing numbers for $A^*\wt\nu_\theta$.
Estimator of sensitivity derivative when gradient operator $A^*_P$ depends on $P$ can be based on the plugin estimate with empirical distribution or mollified empirical distribution
E.g. the multiplication operator of (ref) is of this form because the likelihood ratio $\frac{dP}{dQ}$ of the data generating $P$ to policy cost distribution $Q$ depends on unknown $P$. We leave consistency of the general form (ref) to future work and consider consistency of policy sensitivity of (ref) formulated with a policy cost distribution $Q$.
Possible variation on above strategy is to assume that likelihood estimates are bounded and apply H\"{o}lder's inequality instead of Cauchy-Schwarz. An alternative strategy is to investigate uniform convergence of the product of influence functions and likelihood ratio approximations.
Here we continue with our example setup of a nonparametric model $M$ with full tangent space $T_PM=T_P\matheuler{P}=L^2_0(P)$ on sample space $\mathcal{X}=\mathbb{R}$ with Borel $\sigma$-algebra. Ichimura and Newey (2015) ichimura2015influence describe how influence functions can be computed. Their idea is to use Lebesgue differentiation to recover the influence function $\wt\psi\in L^2(P)$ from its integral in (ref). It is enough to consider a sequence of curves $P_{z,t}^j = (1-t)P + tG_z^j$ and compute
for an approximation to identity $G_z^j\rightarrow \delta_z$. In models with tangent sets that are a proper subspaces of $L^2_0(P)$, the efficient influence function is the projection onto the subspace.
functional $\psi_1(P) = \int_\mathbb{R} x \, dP(x)$ has information gradient $\wt\psi_1(x) = x - \psi(P) \in L^2_0(P)$.
functional $\psi_2(P)= P ( x-\psi_1(P))^2$ has influence function $\wt\psi_2(x) = (x-\psi_1(P))^2 - \psi_2(P)$.
The policy sensitivity derivative of the mean with respect to the variance according to policy metric $g$ as in (ref) is
of a continuous strictly increasing distribution is $\psi_3(P)=F_P^{-1}(p)$. IN formula allows to use paths through distributions with these properties. Influence function can be derived from the following algebraic identity $ F_tF_t^{-1} (p) = p $ or
Differentiating both sides with respect to $t$ and evaluating at $t=0$, obtain
Solving for the $\frac{d}{dt}\psi_3(P_t)$, simplifying and taking limit on $j$, obtain $\wt\psi_3(x)=\dfrac{p- 1_{[x,\infty)}(\psi_3(P))}{f(\psi_3(P))}$.
The policy derivative of the mean with respect to the $p$-quantile according to metric $g$ of (ref) is
We study gmm functionals on the nonparametric model $\matheuler{P}$ that is constrained only by regularity (smoothness, integrability) conditions. Application layer provides a parameter space $\Theta\subset\mathbb{R}^p$ and a vector of moment criterion functions $$ g:\mathcal{X}\times\Theta \rightarrow \mathbb{R}^q. $$
Specification layer maps the economic quantity $\vartheta\in\Theta$ to a function $\psi:\matheuler{P}\rightarrow\Theta$ of the statistical model. gmm estimation is setup from the application layer assumptions that
This assumption is usually an optimality condition of the interactions described by the application layer model. Often these models are highly stylized and are not expected to describe real-world data precisely. Our view is that this assumption should not be taken literally to data, and that the role of specification layer is important and deserves attention (but is beyond the scope of this paper). We derive the sensitivity measures to provide a local characterization of a given gmm functional. Specifically we describe the local identification of gmm functionals $\psi_W$ on the nonparametric model $\matheuler{P}$ by measuring the local dependence of the estimated parameter on the values of individual moments $$ \nu_i(P) \coloneqq P g_i(\theta)_{|\theta=\psi_W}. $$ The direction and absolute magnitude of the dependence is measured by the derivative $\partial_{\nu(i)}\psi_W$. The relative magnitude of dependence on $\nu_i$ to total local variation in $\psi_W$ is measured by local sufficiency $R(\psi_W,\nu_i)$. The latter also measures the extent to which (statistical) uncertainty about the value of $\nu_i(P)$ in the model $\matheuler{P}$ determines inference about $\vartheta=\psi_W$ in the application layer.
Asymptotic distribution of misspecified gmm estimators was first considered tangentially in imbens1997one and derived explicitly in hall2003large. We derive the influence function (information gradient) of the functional and use it to compute sensitivities. Our derivation provides a characterization of the tangent set to the classical gmm model $\matheuler{P}_0$ that is restricted by assumptions (ref) in the over-identified case $q>p$. As a bonus, this also shows directly the semiparametric efficiency of `optimally weighted' estimator $\wh\psi_{\Omega^{-1}}$ on $\matheuler{P}_0$ and of all gmm estimators $\psi_W$ of different functionals on the nonparametric model $\matheuler{P}$. Although the values of functionals $\psi_W$ coincide on $\matheuler{P}_0$ their sensitivities to directions ruled out by (ref) are different. Chamberlain (1987) chamberlain1987asymptotic first showed efficiency of over-identified \textsc{gmm} estimators via discrete approximations.
We consider only deterministic weighting matrices $W$. In the over-identified case weighting determines the functional and should be chosen based on application layer considerations (we call this specification). gmm functionals are defined by $$ \psi_W(P) = \operatorname*{arg\,min}_{\theta\in\Theta} P g(\theta)^T W Pg(\theta), $$ or locally by the first order condition
To establish differentiability (relative to $H_2$ embedding) of the functional and to find the influence function we assume it along with necessary regularity conditions to proceed with a formal calculation that yields a candidate for the information gradient. Once the gradient is found, Riesz representation implies differentiability\footnote{this method has the name a priori estimate in PDEs}. Let $t\mapsto P_t$ be a smooth curve in $\matheuler{P}$ with score vector $\xi\in L^2_0(P)$ at $t=0$. The functional $\theta_t=\psi_W(P_t)$ satisfies the (ref) along the curve $P_t$:
We use denominator layout for derivatives of vectors (so that $\partial g / \partial\theta$ is $q$ by $p$); our reference for matrix calculus is dhrymes1978mathematics. Differentiating with $\frac{d}{dt}$ in (ref) obtain
Derivative $\frac{d}{dt}$ in terms $I,II$ has two components: perturbing distribution $P$ in the direction $\xi$ changes the integrals and also the value of the functional $\psi_W$ which enters the moment criterion functions. Consider the first element of the vectorized term $I$ above
Define the $qp\times p$ matrix $H_1$ and a $qp\times 1$ vector $H_2$ by stacking the underlined terms in last screen
then $$ I = P[ H_1] \cdot \dot\theta + P[H_2 \cdot \xi]. $$ Similarly
Above manipulation implicitly assumes that $\theta=\psi_W(P)$ is differentiable relative to the embedding of statistical model into $H_2$. Recall that the differential $d\psi_W$ has a Riesz representation $$ \partial_\xi\theta = d\psi_W[\xi] = \brk{\wt\psi_W}{\xi}_{H_2} = \int \wt\psi_W \xi \; dP = \dot\theta. $$ The last equality is termed pathwise differentiability in bounds literature. The point of our work in (ref) was to argue that the notion of differentiability used in bounds literature is precisely the same as the one used with linear spaces and that directional derivatives can be naturally interpreted.
By differentiating with $\frac{d}{dt}$ in (ref) we obtained the following expression that relates the pathwise (directional) derivative $\partial_\xi\theta$ and an integral involving the tangent vector $\xi$:
From above expression we can solve for the
Above derivation relies on smoothness and integrability conditions of moment functions $g$ and its (parameter) derivatives. Since the tangent space $T_P\matheuler{P}$ is unrestricted, we conclude that (ref) is the information gradient of $\psi_W$ and that functional is smooth under these conditions. Note that at $P\in\matheuler{P}_0$ where moment assumptions (ref) hold, we have $M=0$, which reduces the gradient to the familiar expression. Although at $P\in\matheuler{P}_0$ all the functionals $\psi_W$ obtained from different choices of weighting $W$ coincide, their gradients are different along the directions $\xi$ that point outside the model $\matheuler{P}_0$.
Assumptions (ref) imply restrictions for tangent set $T_P\matheuler{P}_0$. We characterize these restrictions next. Differentiating along a path similarly to above in (ref), obtain $$ P g \, \xi = - G \cdot \dot\theta, \quad\text{where } g\coloneqq g(\psi_w), \quad G\coloneqq P\partial_\theta g(\theta)_{\big|\theta=\psi_W} . $$ This condition states that the change in the integral of criterion functions due to perturbing the measure must be offset by the change in the value of the parameter. Since moments $Pg$ can move in $q$ independent directions, where as parameter deviations $\dot\theta$ can span only $p=\text{rank}(G)$ of them, the condition is restrictive. Define continuous linear operator $$ A: L^2_0(P) \rightarrow \mathbb{R}^q \text{ by } A\xi \coloneqq Pg\xi, \quad \text{ so that } T_P\matheuler{P}_0 = \Set{\xi\in L^2_0(P)}{A\xi \in R(G) }. $$ We will derive projections $\Pi_0$ onto $T_P\matheuler{P}_0 \subset L^2_0(P)$ and $\Pi_0^{\perp}$ onto the orthocomplement $T_P\matheuler{P}_0(P)^\perp$. First we reduce the problem to finite dimensional spaces by splitting $$ L^2_0(P) = H_g \;\oplus\; H_g^\perp, \quad \text{ where } H_g \coloneqq \text{span} \set{g}, $$ and noting that any vector $\xi\in L^2_0(P)$ that is orthogonal to $H_g$ does not change the integral of the moment functions and therefore does not change the value of $\psi_W$, as evident from (ref). Hence, $H_g^{\perp}\subset T_P\matheuler{P}_0$.
It is then enough to consider operator $A:H_g\rightarrow \mathbb{R}^q$ which is an isomorphism. If $A\xi \in R(G)$ then $A\xi = G\theta$ or $\xi = A^{-1}G\theta$, therefore $$ H_g = H_G \oplus H_G^\perp \quad \text{ where } \quad H_G \coloneqq R(A^{-1}G), \quad H_G^\perp \coloneqq N((A^{-1}G)^*) $$ is the orthogonal decomposition of $H_g$ onto directions that are in $T_P\matheuler{P}_0$ and those that point outside the classical gmm model. We have the refined decomposition of nonparametric tangent space: $$ L^2_0(P) = \underbracket[0.2pt][2pt]{ H_g^\perp \oplus H_G }_{T_P\matheuler{P}_0} \quad\oplus\quad \underbracket[0.2pt][2pt]{ \;\; H_G^\perp \;\; }_{T_P\matheuler{P}_0^\perp} . $$
To compute the projection $\Pi_G$ onto the range $R((A^{-1}G)^*)$ we fix the orthonormal basis $\Omega^{-1/2}g$, where $\Omega\coloneqq Pgg^T$, then obtain the matrix of $A$ to be $A_{[\;]}=\Omega^{1/2}$ and apply the regression formula for projection onto the range of $\Omega^{-1/2}G$ $$ \Pi_G = (\Omega^{-1/2}G) \big[ (\Omega^{-1/2}G)^T(\Omega^{-1/2}G) \big]^{-1} (\Omega^{-1/2}G)^T, \quad \text{ then } \Pi_G^\perp = I_q - \Pi_G. $$ Finally the projection $\Pi_0 \xi$ onto $T_P\matheuler{P}_0$ of tangent vector $\xi\in L^2_0(P)$ is obtained by removing the $H_G^\perp$ component that can be computed by passing to coordinates and applying above projection matrix $$ \Pi_0\xi = \xi - P[ \xi g^T\Omega^{-1/2} ] \; \Big[ I_q - \Omega^{-1/2}G \big[ G^T\Omega^{-1}G \big]^{-1} G^T\Omega^{-1/2} \Big] \; \Omega^{-1/2} g. $$ The classical gmm model restricts $q-p$ dimensions off of nonparametric tangent space. Specifically, vectors of the form
are restricted, whose span is of dimension $\text{rank}(\Pi_G^\perp)$.
The efficient influence function for gmm on $\matheuler{P}_0$ is obtained by projecting any $\wt\psi_W$ in (ref) onto the (mildly) restricted $T_P\matheuler{P}_0$:
Consequently, the sensitivity $\partial_\zeta \psi_{\Omega^{-1}}$ of the “efficient” gmm functional to any direction $\zeta$ that points out of the model $\matheuler{P}_0$ is zero, where as sensitivities of $\psi_W$ are nonzero. Estimators $\wh\psi_W$ suffer larger asymptotic variance because they estimate the (local) values of the functional outside of $\matheuler{P}_0$ and have nonzero sensitivities to local deviations in those directions.