Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,326 characters · 15 sections · 91 citation commands
Minimum Sliced Distance Estimation in a Class of Nonregular Econometric Models
Classical likelihood-based estimation and inference may face challenges in structural econometrics models such as auction models (see e.g., Paarsch1992 and Donald2002) and equilibrium job-search models (see e.g.,Bowlus2001). In such models, the density or conditional density function of the observable variable of interest may have a jump at a parameter-dependent boundary such as the one-sided and two-sided models studied in Chernozhukov2004. In the one-sided model in Chernozhukov2004 and Hirano2003, the conditional density function of the dependent variable is assumed to be strictly bounded away from zero at the parameter-dependent boundary; In the two-sided models in Chernozhukov2004, the conditional density function is assumed to have a jump at the parameter-dependent boundary, see Example (ref) in the next section. Chernozhukov2004 develop non-normal asymptotic theory for MLE and Bayes estimation (BE) for both one-sided and two-sided models. Hirano2003 study efficiency considerations according to the local asymptotic minimax criterion for conventional loss functions in one-sided models.\footnote{Chernozhukov2004 shows that Bayesian credible intervals based on posterior quantiles, which are computationally attractive, are valid in large samples and perform well in small samples. }
Besides the complexity of the non-normal asymptotic distribution of MLE established in Chernozhukov2004, inference is further complicated by the dichotomy of the asymptotic theory: MLE is asymptotically normal when the conditional density function of the dependent variable has no jump and non-normal otherwise. To simplify inference in such models, Li2010 proposes a two-step estimator based on indirect inference. In the first step, an auxiliary linear regression model is estimated from both real data and simulated data drawn from the structural model; in the second step, a minimum distance between estimators of the parameters in the auxiliary model based on real data and those based on synthetic data is adopted to estimate the structural parameter of interest. Assuming that the auxiliary model ensures identification of the structural parameter, Li2010 shows that under standard regularity conditions, the indirect inference estimator is asymptotically normally distributed and inference is simple and standard.
The motivation for this paper is the same as that of Li2010. Unlike Li2010, we develop a class of (one-step) sliced minimum distance estimators of the structural parameter of interest without relying on an auxiliary regression model. Each estimator in this class makes use of a sliced $L_{2}$-distance between some \textquotedblleft empirical measure\textquotedblright\ of the observed data and a parametric or semiparametric measure induced by the structural model. We refer to the resulting estimator as a minimum sliced distance (MSD) estimator. Two prominent examples of sliced $L_{2}$-distances are sliced $2$-Wasserstein distance and sliced Cr\'amer distance\footnote{The sliced Cr\'amer distance is used in Zhu1997 for goodness-of-fit testing.} based on which we construct minimum sliced Wasserstein distance (MSWD) estimator and minimum sliced Cr\'amer distance (MSCD) estimator respectively. Wasserstein distances have recently been used in the statistics and machine learning literatures for testing and estimation of generative models. For example, Bernton2019 proposes minimum WD estimator based on the $1$ -Wasserstein distance and establishes non-normal asymptotic distribution for univariate unconditional parametric distributions; Nadjahi2020AsymSWD proposes a sliced version of the estimator in Bernton2019 and extends the non-normal asymptotic distribution in Bernton2019.\footnote{The technical analysis of Bernton2019 (and Nadjahi2020AsymSWD) relies critically on the equivalence between the $1$-Wasserstein distance and the $1$-Cr\'amer distance for univariate distributions. Such an equivalence no longer holds for the $2$ -Wasserstein distance adopted in the current paper and the analysis in Bernton2019 breaks down.}
This paper makes several contributions. First, we establish consistency and asymptotic normality of the MSD estimator under a set of high-level assumptions applicable to a wide range of models and sliced distances. In contrast to likelihood-based inference, inference using our MSD estimator is standard. Second, we verify the high-level assumptions for the MSCD estimator under primitive conditions for conditional models including the one-sided and two-sided models in Chernozhukov2004 and Hirano2003. For the latter, our primitive conditions imply that the MSCD estimator is asymptotically normally distributed regardless of the presence/absence of a jump in the conditional density function. Third, we confirm the accuracy of the asymptotic normal distribution of MSCD estimator in a simulation experiment based on the auction model in Paarsch1992 and Donald2002 and compare it with the indirect inference estimator of Li2010.
The rest of the paper is organized as follows. Section (ref) introduces our general framework and the motivating example, proposes the class of MSD estimators including MSWD and MSCD estimators, and presents an illustrative example comparing the different distributions of MLE, MSWD, and MSCD estimators for a one-sided and a two-sided uniform models. Section (ref) establishes the consistency and asymptotic normality of the MSD estimator under high level assumptions. Section (ref) verifies the high level assumptions for the MSCD estimator in the conditional model including one-sided parameter-dependent support under primitive conditions. Section (ref) presents numerical results using synthetic data generated from the auction model in Paarsch1992, Donald2002, and Li2010. Section (ref) concludes. A series of appendices contains verification of assumptions for both MSCD and MSWD estimators in one-sided and two-sided uniform models, verification of high level assumptions for MSCD in two-sided parameter-dependent support models, and technical proofs of the main results in the paper.
Let $\left\{ Z_{t}\right\}_{t=1}^T $ denote a random sample satisfying Assumption (ref) below.
We note that in the conditional model, \ the distribution of $X_{t}$\ is unspecified. In both models, we are interested in the estimation and inference for $\psi _{0}$.
When the density function of either $F\left( \cdot ;\psi \right) $ or $ F\left( \cdot |x,\psi \right) $ exists, MLE or the conditional MLE is a popular approach to estimating the unknown parameter $\psi _{0}$. However, the asymptotic distribution of MLE or conditional MLE and the associated inference depend critically on model assumptions. The classical Wald, QLR, and score tests rely on smoothness assumptions on the density function and the assumption that the true parameter $\psi _{0}$ is in the interior of $\Psi$. Many important structural models in economics violate one or more assumptions underlying the classical likelihood theory which motivates the development of alternative methods of estimation and inference such as those discussed in Section 1.
Below we present two examples of the conditional model stated in Assumption 2.1. Example (ref) includes the one-sided models in Chernozhukov2004 and Hirano2003 and two-sided models in Chernozhukov2004 for which likelihood-based estimation and inference are difficult to implement. We refer to both as parameter-dependent support models and use them to illustrate our assumptions/results in subsequent sections. Example (ref) is the independent private value procurement auction model formulated in Paarsch1992 and Donald2002 for which the winning bid follows the one-sided model in Example (ref).
Many structural econometrics models lead to one-sided or two-sided models satisfying the assumptions in Hirano2003 and Chernozhukov2004 such as ((ref)) in two-sided models so that the non-normal asymptotic theory they develop is applicable. On the other hand, if ((ref)) does not hold, then the asymptotic distribution of MLE is normal.
Let $\mu_0$ and $\mu(\psi)$ denote respectively the true probability measure of $Z_t$ and the probability measure induced by the parametric model with parameter $\psi\in\Psi$. A minimum sliced distance estimator of $\psi_0$ is based on an estimator of a sliced distance between $\mu_0$ and $\mu(\psi)$ denoted as $ \widehat{\mathcal{S}}(\psi)$. To introduce $ \widehat{\mathcal{S}}(\psi)$, we first present a brief review of two popular sliced distances: the sliced $2$-Wasserstein distance or simply the sliced $2$-Wasserstein distance and sliced Cr\'amer distance.
Let $\mathcal{P}_{2}(\mathcal{Z})$ denote the space of probability measures with support $\mathcal{Z}\subset \mathbb{R}^{d}$ and finite second moments. Further let $\mathbb{S}^{d-1}=\{u\in \mathbb{R}^{d}:\lVert u\rVert _{2}=1\}$ be the unit-sphere in $\mathbb{R}^{d} $. For two probability measures $\mu $ and $\nu $ from $\mathcal{P}_{2}( \mathcal{Z})$, we denote by $\mathcal{W}_{2}\left( \mu ,\nu \right) $ their $ 2$-Wasserstein distance or simply the Wassserstein distance. It is a finite metric on $\mathcal{P}_{2}(\mathcal{Z})$ defined by the optimal transport problem:
where $\Gamma (\mu ,\nu )$ is the set of probability measures on $\mathbb{R} ^{d}\times \mathbb{R}^{d}$ with marginals $\mu $ and $\nu $.
When $d=1$, the Wasserstein distance is easy to compute. Proposition 2.17 in Santambrogio2015 or Theorem 6.0.2 in Ambrosio2008 implies that
where $F_{\mu }(\cdot )$ and $F_{\nu}(\cdot)$ are the distribution functions associated with the measures $\mu$ and $\nu$, respectively, and $F_{\mu }^{-1}$ and $F_{\nu}^{-1}$ are the quantile functions.
When $d>1$, the Wasserstein distance $\mathcal{W}_{2}$ is difficult to compute. The sliced Wasserstein distance is introduced to ease the computational burden associated with the Wasserstein distance $ \mathcal{W}_{2}$ (c.f. Bonneel2015). For $u\in \mathbb{S}^{d-1}$ and $z\in \mathbb{R}^{d}$, let $u^{\ast }(z)=u^{\top }z$ be the 1D (or scalar) projection of $z$ to $u$. For a probability measure $\mu $, we denote by $u_{\sharp }^{\ast }\mu $ the push-forward measure of $\mu $ by $u^{\ast }$. The sliced Wasserstein distance $\mathcal{SW}(\mu ,\nu )$ is defined as follows:
where $\varsigma (u)$ is the uniform distribution on $\mathbb{S}^{d-1}$. It is well known that $\mathcal{SW}$ is a well-defined metric, see Nadjahi2020SDProperty. For each $u\in \mathbb{S}^{d-1}$, let
Define $G_{\nu }(s;u)$ similarly. Since
we obtain that
For generality, we introduce a weighted version of $\mathcal{SW}(\mu ,\nu )$.
Similarly, we introduce a weighted sliced Cram\'er distance between two measures $\mu$ and $\nu$ with support $\mathcal{Z}\subset \mathbb{R}^{d}$. In the univariate case ($d=1$), the Cram\'er distance (see e.g. Cramer_1928 and Szekely_2017) is defined as
For $\psi \in \Psi \subset \mathbb{R}^{d_{\psi }}$, the sliced distance $\widehat{\mathcal{S}}(\psi)$ is defined as
and the MSD estimator denoted by $\hat{\psi}_{T}$ is defined via
where $Q_{T}(s;u)$ is an \textquotedblleft empirical measure\textquotedblright\ of the projection of observed data $\{u^{\top}Z_t\}_{t=1}^{T}$ and $\widehat{Q}_{T}(\cdot ;u,\psi )$ is the corresponding parametric or semiparametric measure induced by the structural model.
Two examples of $\widehat{\mathcal{S}}(\psi)$ we focus on are estimators of $\mathcal{SW}^2_{w}(\mu_0 ,\mu(\psi) )$ and $\mathcal{SC}^2_{w}(\mu_0 ,\mu(\psi) )$, where
The form of $\widehat{Q}_{T}(\cdot ;u,\psi )$ differs for unconditional and conditional models.
To motivate $\widehat{G}_{T}(s;u,\psi )$ for conditional models, let $Z=\left( Y,X^{\top }\right) ^{\top }$. We note that the cdf of $Z$ and $u^{\top }Z$ are given by
Since the population cdf of $u^{\top }Z$, i.e., $G(s;u,\psi_0)$, depends on the unknown distribution of $X$, we make use of its estimator $\widehat{G}_{T}(\cdot ;u,\psi )$ to define $\widehat{Q}_T$.
Before establishing the formal asymptotic theory for $\hat{ \psi}_{T}$ in the next section, we illustrate the different asymptotic distributions of MSCD estimator, MSWD estimator, and MLE in a one-sided uniform model and a two-sided uniform model in this section.
The one-sided uniform model is $Y\sim U[0,\psi _{0}]$, where $\psi _{0}>0$ is the true parameter. The two-sided uniform model is characterized by the density function
This model is an example of the two-sided parameter-dependent support model for all $\psi_0\in (0,1)$ except when $\psi_0=1/4$ so the MLE of $\psi_0$ follows asymptotically a non-normal distribution when $\psi_0\neq 1/4$. When $\psi_0=1/4$, there is no jump in the density function $f(y;\psi _{0})$. With straightforward but tedious algebra, we show that the conditions for asymptotic normality of MLE in Theorem 5.39 of Vaart1998 are satisfied when $\psi_0=1/4$.
In addition to MSCD estimator, MSWD estimator, and MLE, we also included an oracle Generative Adversarial Network (GAN) estimator\footnote{Here, we consider the oracle GAN estimator when the size of the simulated sample is infinity to minimize variation caused by the simulated sample.} in the comparison. The oracle GAN estimator minimizes the Jensen–Shannon divergence (see Proposition 1 of Goodfellow2014 and Example 3 in Kaji2020):
where $f(y_t, \psi)$ is the density function of the model-induced parametric distribution.
We present three figures below. In each figure, the left panel plots the population objective functions (divergences/distances) of the four estimators and the right panel presents the corresponding QQ plots based on 3000 normalized values of the estimator, where each value is computed from a random sample of size 1000. We note that in this section and Section (ref), the normalized value of an estimator $\hat{\psi}$ is defined as $(\hat{\psi} - \psi_0) / s$, where $s$ is the Monte-Carlo standard deviation of the estimator.
Several conclusions can be drawn from Figures (ref)-(ref). First, the population objective functions for our MSWD and MSCD estimators are smooth in all cases. For the one-sided uniform model, the KL and JS divergences are not differentiable at the true parameter value: KL divergence is not defined when $\psi < \psi_0$. For the two-sided uniform model, the KL and JS divergences are bounded but may or may not be first-order differentiable at the true parameter value. For example, both are not differentiable when $\psi_0 \ne 1/4$. Second, for the one-sided uniform model, both MLE and oracle GAN estimators are non-normally distributed; both MSWD and MSCD estimators are approximately normally distributed. Third, for the two-sided uniform model, when $\psi _{0}=1/4,$ all four estimators are close to being normally distributed; when $\psi_0 $ is far from $1/4$, MLE and oracle GAN estimators are non-normally distributed but MSWD and MSCD estimators are again close to being normally distributed. In Section (ref) and Appendix (ref), we will verify conditions for asymptotic normality in Theorem (ref) for both MSCD and MSWD estimators in the one-sided and two-sided uniform models, providing theoretical justification for the results in Figures (ref)-(ref).
In this section, we establish consistency and asymptotic normality of the MSD estimator $\hat{\psi}_{T}$ defined in ((ref)) under a set of high-level assumptions on $Q_{T}(s;u)$ and $\widehat{Q}_{T}(s;u,\psi )$. The high-level assumptions allow for the results to be applicable to a wide range of models and estimators. In the next section, we will verify them under sufficient primitive conditions for the MSCD estimator in conditional models.
Assumption $\ref{assumption:convergence}$ requires that $Q_{T}(\cdot ,\cdot \dot{)}$ and $\widehat{Q}_{T}(\cdot ,\cdot ,\psi )$ converge to $Q(\cdot ,\cdot )$ and $Q(\cdot ,\cdot ,\psi )$ in weighted $L_{2}$-norm, respectively, and the latter is uniform in $\psi \in \Psi $. Under Assumption (ref), $Q(\cdot ,\cdot )=Q(\cdot ,\cdot ,\psi_0 )$.
Assumption (ref) below states that $\psi _{0}$ is well-separated.
To establish asymptotic normality of $\widehat{\psi}_T$, we follow Andrews1999 and Pollard1980 by verifying the following quadratic approximation of the sliced distance $\widehat{\mathcal{S}}(\psi)$ defined in ((ref)),
where
Let
where $\widehat{D}_T(\cdot; \dot, \psi_0)$ is an $L_2(\mathcal{S} \times \mathbb{S}^{d-1}, w(s)\mathrm{d}s \mathrm{d}\varsigma(u))$-measurable function. We adopt the following assumption to show the quadratic approximation in ((ref)).
First, we note that Assumption (ref) is implied by the norm-differentiability of $\widehat{Q}_{T}(\cdot ;u,\psi )$: for any $\tau _{T}\rightarrow 0$,
Moreover, norm-differentiability of $\widehat{Q}_{T}(\cdot ;u,\psi )$ implies twice continuous differentiability of the sliced distance $\widehat{\mathcal{S}}(\psi)$. Second, Assumption (ref) allows the model-induced function $\widehat{Q}_T$ to have a finite number of kink points with respect to $\psi$ such as in the parameter-dependent support models in Example (ref). In these models, one can take
Assumption (ref) (i) below strengthens Assumption (ref) (i) by specifying the rate of convergence.
The next assumption implies asymptotic normality of the gradient vector of the sample objective function evaluated at the true parameter value. It can be verified by appropriate CLTs.
We are ready to state our main theorem.
In this section, we impose smoothness conditions to verify the high level assumptions for consistency and asymptotic normality of the MSCD estimator in conditional models, where $Q(\cdot ;u,\psi )=G(\cdot ;u,\psi )$ defined in ((ref)) and $\widehat{Q}_{T}(\cdot ;u,\psi )=\widehat{G}_{T}(\cdot ;u,\psi )$ defined in ((ref)).
Throughout this section, we assume that the parameter space $\Psi $ is compact and the weight function $w(\cdot)$ is integrable, i.e., $\int_{\mathbb{R}} w(s) \mathrm{d}s < \infty$.
We impose Conditions (ref) and (ref) below on the conditional model and verify them for the one-sided parameter dependent support model.\footnote{Verification for the two-sided parameter-dependent support model is similar and postponed to Appendix (ref).}
Condition (ref) is implied by uniform Lipschitz continuity of $F(y|x,\psi )$ with respect to $\psi$, i.e., there exists a finite constant $C$ such that for all $y\in\mathcal{Y}$ and $x\in\mathcal{X}$,
It is easy to see that Condition (ref) is satisfied in the one-sided and two-sided uniform models.
We now show that it is satisfied in the one-sided parameter-dependent support models in Example 2.1 under very mild conditions.
For one-sided parameter dependent support models, uniform Lipschitz continuity of $F(y|x,\psi )$ with respect to $\psi$ holds if $F(y|x,\psi )$ is absolutely continuous with respect to $\psi $ for each $y$ and $x$, and $\sup_{y\neq g(x,\psi )}|\partial F(y|x,\psi )/\partial \psi |$ is uniformly bounded by a finite absolute constant. It follows from equation ((ref)) that
Here, $\psi = (\theta^{\top}, \gamma^{\top})^{\top}$ and $\partial g(x, \psi)/\partial \psi := [\partial g(x, \theta)/\partial \theta^{\top}, 0]^{\top}$ because $g(x, \theta)$ does not contain $\gamma$.
The first two assumptions (i) and (ii) in Lemma (ref) ensure validity of interchanging the differentiation operation and integration in equation ((ref)). The third and fourth assumptions (iii) and (iv) are included in Conditions C2 and C3 in Chernozhukov2004. The fifth condition (v) is stronger than the condition $\sup_{\psi} \int \int \left\Vert \frac{\partial f_{\epsilon }(y-g(x, \psi) |x,\psi )}{ \partial \psi } \right\Vert \mathrm{d}y \mathrm{d}F_X(x) < \infty$ in C2 in Chernozhukov2004. However, the assumption: $f_{\epsilon }(0|x,\psi )>\eta >0$ in Chernozhukov2004 (see also ((ref)) in Example 2.1) is not required in Lemma (ref). Consequently, the asymptotic distribution of our MSCD estimator is normal regardless of the presence/absence of a jump in $f_{\epsilon}(0|x,\psi )$.
Let
where $D(\cdot;\cdot,\psi _{0})$ is a $L_{2}(\mathbb{R} \times \mathbb{S} ^{d-1} ,w(s)\mathrm{d}s \mathrm{d}\varsigma)$-measurable function. We impose the analog of Assumption (ref) on the population function $Q$ and use it to verify Assumption (ref) for the random $\widehat{Q}_T$.
We now verify Condition (ref) in the one-sided parameter-dependent support model.
In the one-sided model, when $y > g(x, \psi)$,
Assumption (i) in Lemma (ref) ensures validity of interchanging the differentiation operation and integration. The first three conditions in Assumption (ii) are included in Conditions C2 and C3 in Chernozhukov2004. It is important to note that like in Lemma (ref), the assumption: $f_{\epsilon }(0|x,\psi )>\eta >0$ in Chernozhukov2004 (see also ((ref)) in Example 2.1) is not required in Lemma (ref).
We are now ready to state the main result in this section.
By Theorems 3.1 and 3.2, it is sufficient to verify Assumption (ref) (ii), Assumption (ref), Assumption (ref) (ii) and (iii), and Assumption (ref), given that $\{Z_t\}$ satisfies Assumption (ref). Assumption (ref) (ii) is straightforward to verify, Assumption (ref) relies on CLT, and the other assumptions make heavy use of tools for degenerate U-statistics. We discuss details in the rest of this section.
\paragraph{Verification of Assumption (ref) (ii)} Noting that $$\widehat{G}_{T}(s;u,\psi ) = \frac{1}{T}\sum_{t=1}^{T} \mathbb{E}[u^{\top} Z_t \le s | X_t, \psi],$$ we have
Hence, $\int_{u \in \mathbb{S}^{d-1}} \int_{-\infty}^{\infty} (\widehat{G} _T(s; u. \psi) - G(s; u, \psi))^2 w(s) \mathrm{d}s \mathrm{d}\varsigma(u)$ can be represented as a degenerate $V$-statistic of order 2, i.e.,
where $k(\cdot,\cdot;\psi)$ is a degenerate symmetric kernel function indexed by $\psi$:
Lemma (ref) in Appendix (ref) shows that Lipschitz continuity of $k$ is implied by that of $F$ in Condition (ref). Under the Lipschitz continuity of $k(x, x^{\prime},\psi )$ with respect to $ \psi$ for every $x$ and $x^{\prime}$, Corollary 4.1 of Newey1991 or Lemma 4 in the appendix of Briol2019 can be used to verify Assumption (ref) (ii).
To sum up, the following lemma holds.
\paragraph{Verification of Assumption (ref) (ii)} Note that under i.i.d assumption, we have
So Assumption (ref) (ii) holds.
\paragraph{Verification of Assumptions (ref) and (ref) (iii)}
Under Condition (ref), it is sufficient to show that
Note that the integral in the numerator on the left hand side of the above equation can be represented as a degenerate V-statistic of order 2:
where $ k_2(X_t, X_j, \psi, \psi_0)$ is a degenerate symmetric kernel given by
When $k_{2}(x, x^{\prime},\psi ,\psi _{0})$ is Lipschitz continuous with respect to $\psi $ for each $x$, $x^{\prime}$, we can use Corollary 8 in Sherman1994 and the proof of Lemma 4 in the Appendix of Briol2019 to verify Assumptions (ref) and (ref) (iii). Furthermore, Lipchitz continuity of $k_2$ is implied by that of $F(\cdot|x, \psi)$ in Condition (ref), see Lemma (ref) in Appendix (ref).
Summing up, we obtain the following result.
\paragraph{Verification of Assumption (ref)} Note that
By applying CLT, we obtain the following result.
In this section, we report simulation results on the accuracy of the asymptotic normal distribution of MSCD estimator in the auction model in Example (ref).
In the simulation, we follow Li2010, where $h(x, \psi) = \exp(\psi_1 + \psi_2 x)$ and $X$ follows the square of $U[0, 2]$. We set $m = 6$ and $\psi_0 = (1, 0.5)$, $(1, 3)$. The number of Monte-Carlo simulations is $2000$.
In addition to our MSCD estimator, we also apply Li2010's indirect inference estimator. Both estimators are asymptotically normally distributed leading to simple Wald-type inference. For the MSCD estimator, we choose $w(s) = \mathds{1}(s \in [-50000, 50000])$ and 100 projection vectors. The algorithm is similar to the Simulation Algorithm in Deshpande2018 except that we only draw projection vectors during optimization. A description of our algorithm is given in Algorithm (ref).
For Li2010's estimator, we draw 100 synthetic samples and run the auxiliary linear regression with regressor $(1, x)$. Also, we use the optimal weighting matrix for Li2010's estimator.
For both estimators, optimization is done via Adam optimizer, where the learning rate is 0.1 and tuning parameter $(\beta_1, \beta_2)$ is $(0.9, 0.999)$ in options. We run 1000 epochs for all optimization. The code is implemented in Pytorch.
We consider and summarize results in Case 1 and Case 2 below.
\paragraph{Case 1: $\psi_0 = (1, 0.5)$ and $T = 100$}
Figures (ref) and (ref) present respectively the qqplots and histograms of the normalized values of our MSCD estimator and Li2010's estimator. They imply that when $\psi_0 = (1, 0.5)$, the normal distribution approximates the finite sample distributions of both our estimator and Li2010's estimator well even when $T=100$.
\paragraph{Case 2: $\psi_0 = (1, 3)$ and $T = 100, 200$}
Figures (ref) and (ref) show respectively the qqplots and histograms of normalized values of our MSCD estimator and Li2010's estimator when $T=100$. The normal approximation for both estimators is less accurate than in Case 1. Furthermore, we find that Li2010's estimator could be sensitive to the initial value when $T=100$. We then increased the sample size to $T=200$ and experimented with different initial values for Li2010's estimator. Figures (ref) and (ref) show respectively the qqplots and histograms of normalized values of our MSCD estimator and Li2010's estimator when $T=200$. As expected, the accuracy of the normal approximation improves for both estimators as $T$ increases. Moreover, Li2010's estimator with initial value close to the true value performs comparably with the MSCD estimator.
Motivated by simple inference in nonregular structural models in economics, we have proposed the method of minimum sliced distance estimation based on empirical and model-induced measures of the true distribution. We have developed a unified asymptotic theory under high level assumptions and verified them for the parameter-dependent support model. The numerical illustration on an auction model confirms the efficacy of the proposed methodology.