Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
52,029 characters · 8 sections · 30 citation commands
Inference on Partially Identified Parameters with Separable Nuisance Parameters: a Two-Stage Method
In a moment inequalities system, researchers often encounter parameters of interest that are not point identified. The literature has developed a variety of methods to construct confidence intervals for these partially identified parameters. For instance, chernozhukov2007parameter proposes a criterion function and characterizes the identified set using the set of minimizers. rosen2008confidence constructs a quadratic-form test statistic, deriving its asymptotic behavior, which follows a chi-square distribution. After determining the critical value for the parameters of interest, the confidence set can be obtained through test inversion. andrews2010inference introduce their test and confidence set, building upon the generalized moment selection method. Further contributions to the literature include the work of canay2010inference, which focuses on empirical likelihood inference for partially identified models and establishes the validity of the bootstrap for EL-based test statistics. Additionally, romano2010inference develops a general framework for conducting inference on the identified set in partially identified econometric models that is uniformly valid across a broad class of models, including those with moment inequalities. These studies significantly expand the toolbox of techniques available for constructing confidence sets in partially identified moment inequality models, providing researchers with a diverse array of approaches to tackle complex estimation problems.
Researchers may sometimes only be interested in specific parameters within the moment inequalities system, focusing on "subvector inference." In these cases, handling nuisance parameters and eliminating unnecessary parameters can be a challenging task. Traditional methods, such as subvector projection and subsampling inference, often exhibit conservatism and limited asymptotic power. Fortunately, recent literature has introduced several progressive improvements in this area. bugni2017inference introduces a bootstrap-based inference method specifically designed for subvector inference, which addresses some of the shortcomings of traditional approaches. Similarly, chen2018monte proposes an algorithm that constructs confidence sets for subvectors based on Monte Carlo simulations, providing a more accurate estimation of the parameter space. Additionally, kaido2019confidence presents a bootstrap-based calibrated projection procedure for constructing confidence sets, tailored for a single component or a smooth function of the parameter vector. These methods generally control asymptotic size uniformly across a broad class of data generating processes, offering more reliable and effective techniques for subvector inference in moment inequalities systems compared to their traditional counterparts.
In the literature, there are studies that focus on inference for partially identified parameters in specific structures, which are frequently encountered in empirical applications. While these studies consider a limited class of data generating processes, they offer improved efficiency and better asymptotic convergence rates compared to more general methods. andrews2019inference examines moment inequalities that are linear with respect to the nuisance parameters. The authors develop a least favorable test that addresses the nuisance parameter and reduces conservativeness. cox2023simple considers a linear moment inequalities system and eliminates the nuisance parameter using a linear algebra method developed by kohler1967projections, combined with vertex enumeration. Subsequently, they develop the refined chi-squared (RCC) test to estimate the parameters of interest. By concentrating on specific structures, these studies offer valuable insights into the inference of partially identified parameters in commonly encountered situations. These tailored approaches can lead to computational savings, making them particularly useful for empirical applications.
In this paper, we address a specific structure in which the nuisance parameters are separable from the parameters of interest. In this structure, nuisance parameters are contained within a single function that is unrelated to the parameters of interest, and the moment inequalities system is linear with respect to this function. This formulation goes beyond the linearity restriction of nuisance parameters found in some previous literature, allowing for a wider range of application scenarios.
Under certain assumptions and circumstances, separable nuisance parameters can be eliminated through direct elimination, as demonstrated by cox2023simple. However, this approach requires considerable computational effort, particularly as the complexity of the vertex enumeration process grows rapidly with the dimensions of moment inequalities. Furthermore, this method introduces conservativeness into the final results.
To overcome these limitations, we develop a two-stage method for constructing a confidence set for the parameters of interest. The first stage involves estimating the nuisance parameters. This stage is compatible with a wide range of methods, such as Generalized Method of Moments (GMM) or Maximum Likelihood, provided that the first-stage estimator satisfies a specific assumption detailed later in the paper.
In the second stage, we base our method on the refined chi-squared test developed by cox2023simple, incorporating necessary adjustments. We construct a test statistic and derive its asymptotic properties, enabling the calculation of a confidence region for the parameters of interest through test inversion. Although our method requires an additional stage of estimation for the nuisance parameters, it produces a confidence set with the correct asymptotic rate of coverage and avoids the conservativeness issue associated with the direct elimination of nuisance parameters. This approach, therefore, offers a reliable alternative for estimating partially identified parameters in moment inequalities systems with separable nuisance parameters, at the cost of assuming the nuisance parameters can be estimated separately.
The rest of the paper is organized as follows: Section (ref) sets up the moment inequalities and describes the method of estimation, Section (ref) discusses the asymptotic and finite-sample properties of the approach, Section (ref) presents an empirical application with U.S. vehicle market, and Section (ref) concludes this paper.
In this section, we describe the method of inference. Let $W \in \mathbb{R}^{d_W}$ be an observable random vector with distribution $F$. We have $n$ observations of $W$, denoted as $\{ W_i \}_{i=1}^n$. Let $\theta \in \Theta \subseteq \mathbb{R}^{d_\theta}$ and $\delta \in \Delta$ be unknown parameter vectors, where $\Delta$ is a subset of a vector space which can be finite or infinite dimensional. Suppose $\theta$ is our parameter of interest, and $\delta$ is a nuisance parameter.
Let $M : \mathbb{R}^{d_W} \times \Theta \to \mathbb{R}^{d_M}$ and $N : \mathbb{R}^{d_W} \times \Delta \to \mathbb{R}^{d_N}$ be measurable functions. Let $B : \Theta \to \mathbb{R}^{k \times d_M}$, $C : \Theta \to \mathbb{R}^{k \times d_N}$, and $\rho : \Theta \to \mathbb{R}^{k \times 1}$ be matrix-valued measurable functions, respectively. Consider the following moment inequality restrictions:
Note that Equation (ref) can represent any general form of moment inequalities, expressed as \[ E_F[m(W, \theta, \delta)] \geq 0, \] by setting the following components: \( B(\theta) = 0 \), \( M(W, \theta) = 0 \), \( C(\theta) = 1 \), \( N(W, \theta, \delta) = m(W, \theta, \delta) \), and \( \rho(\theta) = 0 \). We adopt the specific structure in Equation (ref) to achieve a clear separation between the nuisance parameter and the moment inequalities. In this form, the nuisance parameter \( \delta \) only appears through \( N(W, \theta, \delta) \), ensuring that \( N \) remains non-divisible. This separation is aimed at simplifying the computational procedures in subsequent sections. Therefore, we still use "separable" to describe our moment inequalities even though the separation between parameters of interests and nuisance parameters is in fact not required. We bring two economic examples matching with this moment inequalities structure.
\newcounter{example1} \setcounter{example1}{\value{example}}
cox2023simple consider a similar structure, where $N(W, \delta)$ does not depend on $W$, i.e., $N(W, \delta) = \delta$. Let $d_{\delta}$ be the dimension of $\delta$. They propose a procedure to eliminate the nuisance $\delta$ by a matrix transformation following Theorem 4.2 of kohler1967projections. We briefly review their setup and key steps in their approach. For convenience, the notation we use in this section is following their paper and not to be confused with other part of this paper.
Their moment inequality model has a form
where $B$, $C$, and $d$ are $k\times d_M$, $k\times d_{\delta}$, and $k\times 1$ matrices, $\delta$ is an unknown nuisance parameter, $\theta$ is the unknown parameter of interest, and $\bar M_n(\theta)=\frac{1}{n}M(W_i,\theta)$. They use matrix reformation to eliminate the nuisance parameter $\delta$.
Following Lemma (ref), the moment inequality system is equivalent to
with $A=A(B,C)$ and $b=b(C,d)$. Let $\hat{\Sigma}_n(\theta)$ be an estimator of the covariance matrix of the moments ${\Sigma}_n(\theta)$, where
Then they build a test statistics
Let $\hat{\mu}$ denote the solution to the minimization question. Let $a_j'$ denote the $j$th row of $A$ and $b_j$ denote the $j$th element of $b$ for $j=1, 2, \dots, d_A$. Let $$ \hat{J}=\{ j\in\{ 1,2,...,d_A \}: a_j' \hat{\mu}= b_j \} $$ which is the indices set for active inequalities. Write $A_J$ as the submatrix of $A$ formed by the rows of $A$ corresponding to the elements in $J$, and let $\text{rk}(A_J)$ denote the rank of $A_J$. Define $\hat{r} = \text{rk}(A_{\hat{J}})$.
If $\hat{r}=1$, without loss of generality, suppose the first inequality is active and satisfies $a_1\neq 0$. Then for each $j=2,...,d_A$, let
where $\| a \|_{\Sigma} = (a' \Sigma a)^{1/2}$. Let
Then define
Then, the RCC test for $H_o:\theta=\theta_0$ is given by
The following result is a lemma from cox2023simple's argument.
This assumption is actually weak, as cox2023simple argues, since it can be implied by four ordinary conditions: i.i.d. dataset, positive variance of moments, significant covariance matrix determinant, and bounded standardized moments. It elicits the following result:
This lemma has illustrated the validity of cox2023simple's method. However, their method cannot be directly applied to our structure since when $N(W, \delta)$ is no longer a single parameter.
To address this issue, we propose a two-step estimation procedure. The first step is an estimation of the nuisance parameter, where the nuisance parameter is assumed to be point identified. Suppose the first stage estimation of the nuisance parameter $\delta$ is $\hat{\delta}$. We will make further assumptions on the properties of $\hat{\delta}$ later to ensure the validity of our method. The method in the first stage estimation can vary across realistic conditions.
In the second step, we can follow the approach of cox2023simple but make some adjustments. First, denote the sample average of $X$ as $\bar{X}_n = n^{-1}\sum_{i=1}^n X_i$, where $X$ can be any function of random variables. Write \[ p(W, \theta, \delta) =
\] Denote the sample mean by $\bar M(\theta) = \frac{1}{n}\sum_{i=1}^n M(W_i,\theta)$, $\bar N(\theta,\hat\delta) = \frac{1}{n}\sum_{i=1}^n N(W_i,\theta,\hat\delta)$, and $\bar{p}_n(\theta,\hat\delta) = \frac{1}{n}\sum_{i=1}^n p(W_i,\theta,\hat\delta)$. Let $\hat{\Sigma}_n(\theta, \hat{\delta})$ be an invertible estimator of the conditional variance $\text{Var}(\sqrt{n}\bar{p}_n(\theta, \hat{\delta}) | \hat{\delta})$. Construct the following test statistic:
Write $\kappa =[t,\eta]'$, then the constraint $B(\theta)t - C(\theta)\eta \leq \rho(\theta)$ can be written as
Let $A(\theta) = [B(\theta), -C(\theta)]$, the test statistic then becomes
Let $\hat{\kappa}(\theta, \hat{\delta})$ be the solution to the minimization problem mentioned previously. Let $a_l'(\theta)$ denote the $l$th row of $A(\theta)$ and $d_l(\theta)$ denote the $l$th element of $\rho(\theta)$ for $l=1, 2, \dots, d_A$, where $d_A$ is the dimension of rows of $A(\theta)$. Define $$ \hat{J}(\theta,\hat{\delta})=\{ l\in\{ 1,2,...,d_A \}: a_l'(\theta) \hat{\kappa}(\theta,\hat{\delta})= d_l(\theta) \} $$ to collect the indices associated with the restrictions that hold with equality. Write $A_J$ as the submatrix of $A$ formed by the rows of $A$ corresponding to the elements in $J$, and let $\text{rk}(A_J)$ denote the rank of $A_J$. Define $\hat{r}(\theta, \hat{\delta}) = \text{rk}(A_{\hat{J}(\theta, \hat{\delta})}(\theta))$.
Conditional on the set of active inequalities, we follow the refined chi-squared test as proposed by cox2023simple. When $\hat{r} \neq 1$, the critical value of the RCC test statistic is set to be the $\chi^2$-distribution with a degree of freedom $\hat{r}$ at the quantile $1-\alpha$, that is, $\chi^2_{\hat{r}, 1-\alpha}$, where $\alpha$ is the pre-supposed significant level. When $\hat{r} = 1$, this critical value can be too conservative, as the test statistic tends to mass at zero. To address this issue, the RCC test employs a quantile of $1-\beta_n(\theta, \hat{\delta})$ instead of $1-\alpha$, which depends on how far from being active these inequalities are and would restore the size of the test.
Without loss of generality, suppose the first inequality is active. To measure the distance from being active for other inequalities, define $z_j(\theta, \hat{\delta})$, $j=2, 3, \dots, d_M+d_N$, as
where $\| a \|_{\Sigma} = (a' \Sigma a)^{1/2}$ in definition. $z_j$ is equal to zero when the $j$th inequality is active, and positive when it is inactive. Then, let
Define
where $\Phi$ is the standard normal cumulative distribution function. Finally, the critical value of the RCC test is the $\chi^2$-distribution with a degree of freedom $\hat{r}$ at the quantile $1-\beta_n(\theta, \hat{\delta})$. The test for $H_0: \theta_0 = \theta$ is represented as
Therefore, a confidence interval for $\theta$ under the significant level $\alpha$ can be built as
after plugging in the first stage estimator $\hat{\delta}$.
In large samples, to ensure the size of the test, we have to make some high-level assumptions on the first-stage estimator $\hat{\delta}$. Since we assume point identification of $\delta$, for a given data generating process $F\in \mathcal{F}$ there is only one $\delta$ satisfying $\delta=\Delta_0(F)$. Write the first-stage estimator as $\hat{\delta} = \hat{\Delta}_0(F)$, we post another assumption to motivate the theorem of the asymptotic properties.
This assumption is a variant of Assumption (ref). In large samples, we assume the estimator $\hat{\delta}$ can ensure the asymptotic normality of $\bar{p}$ and the convergence of $\hat{\Sigma}$. If both $\theta$ and $\delta$ considered as parameters of interest and the first stage estimation is skipped, the assumption is directly implied by Assumption (ref). Nevertheless, this assumption is more than the consistency of the first-stage estimator $\hat{\delta}$, as it involves the large sample performance of $\theta_n$, $\delta_n$, and $\hat{\delta}_n$ simultaneously. The first-stage estimator $\hat\delta$ could have an impact on computing the variance $\hat{\Sigma}_n(\theta,\hat\delta)$. As exemplified later in Section (ref), we use an influence function to propagate the error into the second stage. We can only disregard this impact in the case where $E[p(x,\theta,\delta)]$ is orthogonal to the first-stage estimator $\hat\delta$, which means the Jacobian matrix derived later in Equation (ref) is zero evaluated at $\delta=\hat\delta$, and thus the influence function does not contribute to the sample variance.
We have the following theorem illustrating that the test has correct asymptotic size in large samples.
Therefore, when the number of the observations $n$ approximates infinity, the confidence set constructed from test inversion will asymptotically cover the parameters of interest at the presumed significant level.
In finite samples, this method is valid under certain conditions. The following theorem states that when sample mean is assumed to be normal and the estimation of covariance matrix is almost surely accurate, the test exhibits proper size under finite sample.
Theorem (ref) demonstrates the validity of the proposed test under finite sample conditions. Specifically, it ensures that the test maintains the correct size, assuming normality of the sample mean and accurate estimation of the covariance matrix.
This section continues Example (ref) by applying the methodology to the empirical context of the U.S. commercial vehicle market using the dataset from wollmann2018trucks. The objective is to provide empirical evidence on the relative magnitude of entry and exit costs across different vehicle attributes in the U.S. commercial truck market.
The dataset records transaction data and product characteristics for each brand and model within a year. The data spans from 1992 to 2012, covering sales information for every type of vehicle product within all mainstream brands.
The characteristics in the dataset include Gross Vehicle Weight Rating, Cab-Over-Engine Indicator, Compact-Front-End Indicator, and Long-Cap Indicator. Gross Vehicle Weight Rating measures the legal maximum load a vehicle can carry and is a key specification used by both regulators and manufacturers to classify vehicles. It also reflects the intended usage of the vehicle, with higher ratings typically indicating commercial applications. The Cab-Over-Engine Indicator captures whether the vehicle has a cab-over-engine design, where the driver sits directly above the engine and front axle, with improving maneuverability and visibility but reducing comfort. The Compact-Front-End Indicator identifies vehicles with a shortened front hood that balances the benefits of maneuverability and space while accommodating limited engine capacity. Lastly, the Long-Cap Indicator denotes whether a vehicle is equipped with an extended cab, which offers additional interior space and comfort, especially suited for long-distance or demanding operations.
In our analysis, we treat each brand as a firm and define two vehicles as being within the same type if their characteristics are identical. Consequently, all vehicles are classified into 30 distinct types.
In the first stage of our estimation, we recover the unobserved mean utilities and estimate the demand-side parameters using a procedure inspired by berry1995automobile and nevo2001measuring. We begin with the utility specification for consumer $i$ choosing product $j$:
where $x_j$, $p_j$, $\xi_j$, and $\varepsilon_{ij}$ are defined in Example (ref). $\delta=(\beta,\alpha)$ are random coefficients that capture consumer heterogeneity, and are assumed to follow a joint distribution $F_{\beta,\alpha}(\cdot)$. We denote it by $U_{ij}$ for simplicity.
Given the assumption on $\varepsilon_{ij}$, the predicted market share of product $j$ takes the multinomial logit form:
where the integration is with respect to the joint distribution $F_{\beta,\alpha}$. Because the integral in Equation (ref) does not yield a closed-form expression, we approximate it using Monte Carlo simulation. That is, for $R$ independent draws $(\beta_r,\alpha_r)$, the simulated market share is computed as
Write the mean utility for product $j$ as
To match the predicted market shares with the observed ones, we recover the vector $\zeta = (\zeta_1, \ldots, \zeta_J)$ using a fixed-point iteration. Let $\hat{s}_j$ denote the observed market share for product $j$. We start with an initial guess $ \zeta_j^{(0)} = \ln(\hat{s}_j) $, and update iteratively as follows:
until the change in $\zeta$ is below a specified tolerance.
A major econometric challenge in estimating demand is the potential endogeneity of prices $p_j$, which may be correlated with unobserved product quality $\xi_j$. To address this, we use an instrumental variable, $z_j$, which is assumed to satisfy \[ E[z_j \xi_j] = 0. \] Based on our model, the GMM objective function is then
where $Z$ is the instrument matrix, $\xi$ is the vector of $\xi_j$ after plugging the estimated mean utility $\zeta_j$ to Equation (ref), and $W$ is a chosen positive-definite weighting matrix. \footnote{In our application, the parameter vector $\theta$ is estimated by minimizing $Q(\delta)$ using the BFGS algorithm.}
Note that the simulation draws in Equation (ref) are not assumed to represent the true distribution of the random coefficients; rather, they are used as a computational device to approximate the integral in the market share model, which provides sufficient variation for the simulation while not overwhelming the signal in the data.
After the simulation step, the fixed-point iteration recovers the unobserved mean utilities $\zeta$ based on these simulated market shares. Although the initial draws for $\beta$ and $\alpha$ may be crude, they serve only as starting values in the simulation of market shares. In the subsequent GMM estimation step, these parameters are updated by minimizing the GMM objective function based on Equation (ref). Thus, even if the initial simulation of $\beta$ and $\alpha$ uses a simple normal distribution, the iterative GMM procedure refines these parameters to achieve consistency with the observed market shares and the imposed moment conditions.
Table (ref) presents the estimation results of the coefficients of product characteristics $\beta$ and the price sensitivity $\alpha$.
The estimated coefficients in Table (ref) are large in magnitude, which is driven by the scaling of the explanatory variables and the nature of multinomial logit model. The signs of the parameters are consistent with economic theory. The positive estimates for the $\beta$ coefficients indicate that enhancements in product characteristics increase the mean utility, whereas the positive value of $\alpha$ implies that higher prices detract from consumer utility. Among all the features of vehicles, an extended cap is most valued by drivers. A cab over the engine is also cherished, while a compact front end type is less valued by drivers. This first-stage estimation thus produces demand-side parameters that are later used to construct the structural model for product entry and exit decisions in the second stage.
Recall that $p(W, \theta, \delta)$ is constructed as $p(W,\theta, \delta) = [M(W,\theta)', N(W,\theta, \delta)']'$. In our framework, the variance of the sample average $\bar{p}_n(\theta,\hat\delta)$ must account not only for the intrinsic sampling variation in $p(W,\theta,\delta)$, but also for the uncertainty arising from the estimation error in the nuisance parameter $\hat\delta$ from the first stage. To estimate the variance of $\bar{p}_n(\theta,\hat\delta)$, we apply the influence function method introduced by newey1994large to capture how errors in the estimation of $\delta$ propagate into the second-stage estimation.
Let $P_{\delta}$ denote the Jacobian matrix:
and denote the influence function by $\psi(W,\delta)$ to correct for the error in estimating $\delta$, defined as
with
The corrected moment for each observation is then given by $p(W,\theta,\delta) + P_{\delta}\psi(W,\delta)$, and consequently, the variance of $\sqrt{n}\bar{p}_n(\theta,\hat\delta)$ can be estimated as
where
The influence function correction compensates for the additional variability due to the estimation error in $\delta$, as well as for any potential endogeneity in prices via the instrumental variable $z_j$. The resulting variance estimator in Equation (ref) is then used to conduct robust inference on the structural parameters in our model.
Having obtained the first-stage estimates $\hat{\delta}$ and the sample variance of the moment function $\hat{\Sigma}_n(\theta,\hat{\delta})$, we proceed to the second-stage estimation of the sunk-cost parameters $\theta = (\lambda, \eta)$ using the RCC test procedure described in Section (ref). Since the model imposes moment inequality restrictions, the parameters are only partially identified. As a result, unlike point-identified settings where a single estimate summarizes the inference, our procedure yields a high-dimensional confidence region that characterizes the identified set of $\theta$.
In practice, we evaluate the test statistic $T_n(\theta)$ over a finite grid in the $(\lambda,\eta)$ space and retain those values of $\theta$ for which the RCC test does not reject at level $\alpha$, that is, for which $\phi_n^{RCC}(\theta,\hat{\delta},\alpha) = 0$. These values constitute our estimate of the identified set.
To visualize this high-dimensional confidence region, we plot two-dimensional slices in the $(\lambda, \eta_1)$ plane while holding the remaining components $(\eta_2, \eta_3, \eta_4)$ fixed at representative values. Specifically, Figure (ref) presents four panels, each corresponding to a different choice of $(\eta_2, \eta_3, \eta_4)$. The shaded areas in these plots represent combinations of $(\lambda, \eta_1)$ that are included in the confidence region under the specified values of the other components.
We observe that the confidence region is most concentrated around $\lambda \approx 0.3$ and $\eta_1 \approx 0$, suggesting that the sunk costs exhibit low recoverability, and the gross vehicle weight does not have a significant influence on the sunk cost of establishing a production line. Moreover, changes in the remaining components $(\eta_2, \eta_3, \eta_4)$ affect the shape of the confidence region in the $(\lambda, \eta_1)$ plane. Starting from the baseline configuration $(\eta_2, \eta_3, \eta_4) = (0.10,\,0.10,\,0.05)$, increasing $\eta_2$, the sunk cost sensitivity to the cab-over-engine feature, or decreasing $\eta_3$, the sunk cost sensitivity to compact-front-end design, will slightly narrow the confidence region. This suggests that $(\eta_2, \eta_3, \eta_4) = (0.20,\,0.10,\,0.05)$ or $(0.10,\,0.05,\,0.05)$ is still consistent with the data and lie within the identified set. In contrast, increasing $\eta_4$, which captures the effect of long-cap design on sunk cost, results in a substantially reduced confidence region. These patterns indicate that the model imposes tighter restrictions on the influence of long-cap design characteristics on sunk costs, thus narrowing the plausible values of $(\lambda, \eta_1)$ that remain consistent with observed entry and exit decisions.
In this paper, we have proposed an econometric method for estimating partially identified parameters in moment inequalities and separable nuisance parameters. Our method is applicable to a wide range of models, and we have demonstrated its effectiveness by applying it to two examples: the US vehicle market model from wollmann2018trucks and the hospital referral model from ho2014hospital. The former example focuses on the structural estimation of vehicle markets, while the latter examines the complex decision process of patient referrals to hospitals.
Our method ensures correct asymptotic size in large samples. To achieve this, we have made high-level assumptions on the first-stage estimator $\hat{\delta}$. Under these assumptions, the confidence set constructed from test inversion will asymptotically cover the parameters of interest at the presumed significant level when the number of observations $n$ approaches infinity.
In conclusion, the proposed method offers a robust and versatile approach for estimating and testing models with moment inequalities and separable nuisance parameters. It is applicable to various types of models, as demonstrated by our examples, and exhibits valid finite-sample and asymptotic properties. Our findings contribute to the econometric literature and provide researchers with a powerful tool for analyzing complex economic models.
Future research directions could include exploring methods to relax the separability assumption, which may extend the applicability of the proposed estimation method to a broader range of models. Additionally, investigating alternative assumptions on the first-stage estimator could provide insights into the robustness and performance of the method under different estimation scenarios. This would further enhance the understanding of the practical implications of the method in various econometric applications.