Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,480 characters · 11 sections · 65 citation commands
Dyadic Double/Debiased Machine Learning for Analyzing Determinants of Free Trade Agreements
Empirical research investigating network formation and outcomes spans the literature covering labor economics, industrial organization, macroeconomics, development, and international trade, among others. In a large fraction of these studies, the process through which economic entities are determined -- e.g., free/preferential trade agreements, friendships, and financial relationships -- concern dyadic/network link formation models. While recent methodological advances have increased the popularity of empirical network formation models (see \citet*{graham2019dyadic} and graham2019econometric for surveys), the large majority of existing studies focus on parsimonious (low-dimensional) model specifications. The increasing availability of `big data' should, in principle, allow researchers to consider increasingly rich specifications and uncover new insights into the nature of network formation. Rather, researchers are hamstrung by two competing limitations. On one hand, there is no existing econometric methodology that allows for both dyadic robust inference and high-dimensional link formation specifications. { In the presence of significant data heterogeneity in the policy variable of interest, one needs to control a large number of covariates. However, simply estimating a rich specification across a wide set of covariates is not yet possible. } Moreover, we document that common off-the-shelf approaches, such as the conventional double/debiased machine learning for i.i.d. sampling, are likely to result in both biased estimates and misleading standard errors under dyadic dependence. On the other hand, the most common empirical approach, reducing the dimensionality of the network formation model by assumption, risks model mis-specification, is a likely source of estimation bias and/or misleading inference, and reduces the set of questions researchers can investigate. In light of the increasing availability of big data today, we advance state-of-the-art approaches to estimating network models by developing novel methods for root-$n$ consistent estimation and inference for high-dimensional dyadic regressions and high-dimensional dyadic/network link formation models.
Our work builds on the literature which suggests using (near) Neyman orthogonal scores to accommodate possibly slow convergence rates of general machine learners \citep*[e.g.,][]{BCK15,BCCW18,CCDDHNR18,CEINR18} in high-dimensional regression estimation (and more generally high-dimensional Z-estimation). In particular, our approach extends the branch of research which employs cross-fitting, in conjunction with orthogonal scores, to mitigate over-fitting biases. This combined method is referred to as the double/debiased machine learning \citep*[DML,][]{CCDDHNR18}. Although standard cross fitting requires independent sampling CCDDHNR18, \citet*{CKMS2019} develop a multiway cross fitting algorithm to extend the DML to multiway cluster dependent data with linear scores (their results, however, does not cover estimators defined by nonlinear scores, such as Logit). This recent contribution has paved the way for machine learning to be applied to cross-sectional dependent data, but, unfortunately, does not provide a result which allows for the estimation of network models in the presence of dyadic dependence. Dyadic dependence is different from (multiway) cluster dependence,\footnote{Specifically, while multiway clustered data are represented by separately exchangeable arrays, dyadic data are represented by jointly exchangeable arrays -- see Section 7 in \citet*{Kallenberg2006}. The mathematical characterizations of these two types of exchangeable arrays are different, and the separate exchangeability does not imply the joint exchangeability.} and hence there is no guarantee that these existing cross fitting algorithms will work for dyadic regressions and dyadic/network link formation models. In this light, we propose a novel cross fitting algorithm that is effective under dyadic dependence. Our dyadic machine learning approach and asymptotic theories guarantee that estimation and inference based on this method are robust against arbitrary dyadic dependence.
In this paper, we demonstrate the importance of addressing dyadic dependence and model specification by reexamining the determinants of free trade networks, a classic network setting. Accounting for both dyadic dependence and high-dimensional controls, we reconfirm the two important theoretical implications suggested by the international trade literature: namely, (A) a greater distance between economies makes an FTA less beneficial krugman1991move,frankel1993continental,frankel1995trading,frankel1996regional,baier2004economic and thus makes an FTA less likely to be formed; and (B) larger sizes of economies make an FTA more beneficial krugman1998comment,baier2004economic and thus make an FTA more likely to be formed. In contrast to the workhorse empirical model, we find that after accounting for dyadic dependence and high-dimensional controls suggest that differences in country-pair factor endowments are not an important determinant of free trade agreements. We demonstrate that traditional approaches lead to fragile estimates: researchers may be misled into concluding that factor endowments differences either encourage or discourage free trade agreements depending on which common estimation approach is employed. Further, estimates produced by our method turned out to differ from those of the simple logistic regression or the conventional double/debiased machine learning not accounting for dyadic sampling for all key determinants of free trade agreements. Indeed, our analysis suggests that the landmark empirical work in this area overstates the importance of bilateral country size to FTA network formation. Allowing for arbitrary dyadic dependence, the proposed method implies less statistical significance. That being said, our results still support the aforementioned theoretical predictions about the determinants of free trade networks despite the inflated standard errors for robustly accounting for the dyadic dependence as well as high-dimensionality.
{\bf 1.1. Relation to the Econometrics Literature:} The history of statistics on dyadic data dates back at least to 1970s, when \citet*{holland1976local} derive moments of dyadic sums and \citet*{silverman1976limit} derives asymptotic normality for the sum of jointly exchangeable dissociated arrays. \citet*{eagleson1978limit} further demonstrated the asymptotic normal mixture for the sum of a weakly exchangeable array.
Explicit accounts for dyadic dependence in econometrics arose in the context of gravity models, and the use of fixed effects was recommended to control for such dependence \citep*[see][]{matyas1997proper,matyas1998gravity}. \citet*{cameron2005estimation} point out the importance of controlling for dyadic clustering, and develop a FGLS estimation method accounting for the dyadic cluster dependence. \citet*{fafchamps2007formation} propose dyadic cluster robust variance estimators for the OLS and logit. \citet*{cameron2014robust} generalize the dyadic cluster robust variance estimator for GMM and M-estimation frameworks as well as others cases. \citet*{aronow2015cluster} also discuss consistent estimation of the dyadic cluster robust variance estimators. Along with most, if not all, other papers that develop inference theories for dyadic data, we take advantage of the forms of the dyadic cluster robust variance formulas developed in these pioneering papers.
The more recent econometrics literature includes \citet*{tabord2019inference} who studies the asymptotic behavior of the t-statistic based on the dyadic cluster robust variance estimator of \citet*{fafchamps2007formation}; \citet*{davezies2019empirical} who study the asymptotic behavior of empirical processes and their bootstrap counterparts for dyadic data; \citet*{graham2019kernel} who propose nonparametric density estimation for dyadic data, show that the convergence rate of the stochastic part of the estimator is the square root of the number of nodes, and derive the asymptotic distribution of the estimator (see also graham2021minimax for extension to local regression estimation); and chiang2020inference who develop methods of inference for high-dimensional parameters. \citet*{graham2019dyadic,graham2020network,graham2020sparse} provide reviews on the asymptotic distribution and variance estimation in dyadic data in the context of parametric models, along with other closely related topics.
Also related to, but different from, the literature on dyadic data is the set of studies which investigate multiway clustered data \citep*[e.g.,][]{cameron2011robust,thompson2011simple,cameron2015practitioner,menzel2018bootstrap,davezies2018asymptotic,mackinnon2019wild,CKMS2019}. The robust variance formulas exposited in this literature are related to those relevant for dyadic cases, but are sufficiently different that they cannot be applied in settings with dyadic dependence. In particular, the structure of cluster dependent data (also known as separately exchangeable arrays) is related to, but mathematically different from, the structure of dyadic data (also known as jointly exchangeable arrays) -- see \citet*{Kallenberg2006}. Consequently, methods and theories developed in this literature will not apply to the dyadic setting.
Turning to the machine learning literature, as we have already emphasized, we use (near) Neyman orthogonal scores to accommodate possibly slow convergence rates of various machine learners following \citet*{BCH14review,BCK15,farrell2015robust,BCCW18,CCDDHNR18,CEINR18,farrell2021deep}. In developing the dyadic link formation models, we also extend the framework of \citet*{BelloniChernozhukovWei2016} and \citet*{BCCW18} to deal with dyadic dependent data and obtain convergence rates for high-dimensional lasso logit under dyadic dependence, which we in turn use as sufficient conditions to apply our general dyadic machine learning theory in order to establish asymptotic theories for high-dimensional dyadic link formation models.
The proposed dyadic cross-fitting algorithm sheds light on handling dyadic networks in cross-validation literature. There are several network-oriented cross-fitting and cross-validation procedures in econometrics and statistics. In the study of stochastic block models, chen2018network propose network cross-validation algorithms for detecting the number of communities. li2020network propose cross-validation algorithms for general tool for model selection and parameter tuning that rely on using a subset of node pairs and apply a low-rank matrix completion algorithm to obtain a predicted adjacency matrix. It is thus fundamentally different from ours. viviano2019policy proposes a network cross-fitting algorithm for estimating treatment allocation rules under network interference. It is different from our algorithm as its unit of observations is a node while ours is a dyad. In addition, our algorithm allows for fully connected networks that are important for international trade applications.
{\bf 1.2. Relation to the Trade Literature:} Understanding the formation of free trade networks across countries dates back to at least \citet*{Viner1950} and its history is fraught with debate. A rich theoretical literature emerged over time elucidating complex economics and political trade-offs associated with free trade agreement formation. \citet*{Levy1997} and \citet*{Krishna1998} pioneered a set of research aimed at understanding the nature of preferential trade in a multilateral trade network.\footnote{\citet*{GrossmanHelpman1995} analyze the political-economy determinants of FTAs but do not consider the implications of FTAs for the multilateral trade system.} Subsequent research investigates network formation and stability \citep*{FK05,GJ06,FurusawaKonishi2007} and characterizes the differential impact that FTAs have on countries excluded and included from FTAs \citep*{KennanRiezman1990,BagwellStaiger1999,BondRiezmanSyropoulos2004,AghionAntrasHelpman2007}.
Country asymmetries, either economic \citep*{FurusawaKonishi2007,SaggiYildiz2010,SaggiWoodlandYildiz2013,Lake2017} or political \citep*{Ornelas2005,StoyanovYildiz2015}, are understood as central to the willingness of any pair of countries to form a stable free trade agreement and their optimal responses policy change within and outside of their region. Likewise, dynamic considerations, such as the likelihood of future FTAs, are expected to encourage or deter firms from joining FTAs in the present \citep*{McLaren2002,MissiosSaggiYildiz2016,Lake2017,LakeRoy2017,LakeNkenYildiz2018,Lake2019}. In other words, the determinants of FTAs systematically vary across the trading network with country characteristics and the evolution of the trading system.
By comparison there is a dearth of empirical results. baier2004economic's (baier2004economic) seminal article takes up krugman1993regionalism's (krugman1993regionalism) call to empirically investigate the determinants of free trade regions.\footnote{As \citet*{{krugman1993regionalism}} writes: “To make any headway, one must either get into detailed empirical work, or make strategic simplifications and stylizations that one hopes do not lead one too far astray. Obviously detailed empirical work is the right direction...”} They identify a parsimonious set of key economic determinants for the formation of free trade agreements: trade costs, the market size of the free trade zone, and the similarity of trading partners in terms of economic development and/or factor-endowments.
Subsequent empirical research, such as that by \citet*{ChenJoshi2010} and \citet*{BaierBergstrandMariutto2014}, \citet*{SaggiStoyanovYildiz2018} among others, build upon the original baier2004economic specification to capture novel features of FTA formation eminating from the theoretical literature: endogenous responses in trade policy, excluded country characteristics, common external trade partners and/or changes in the nature of FTAs over time. Econometrically, in each case, researchers enrich the benchmark specification with additional co-variates intended to capture mechanisms which were beyond the scope of the benchmark baier2004economic specification. {Our application reexamines the Baier and Bergstrand (2004) empirical structure as it is the most highly cited FTA network specification and the basis for all subsequent empirical work in this literature.}
Our work addresses this literature in two important dimensions. First, we demonstrate that ignoring dyadic dependence can lead to biased estimates of model parameters and standard errors. Second, we show that incorporating a wider set of co-variates understates true standard errors among classic determinants and can potentially lead to misleading conclusions. Specifically, our estimates also suggest that the importance market size in FTA formation may be biased and overstated when employing conventional methods, including conventional DML, and a rich specification. Further, we document that standard estimation approaches, such as logit or conventional DML, may lead researchers to conclude that larger differences in relative factor endowment may encourage or discourage FTA formation depending on their preferred estimation approach. Accounting for dyadic dependence, in contrast, increases standard errors sufficiently that we can no longer conclude that relative factor endowments are an important determinant of FTA formation. In this sense, we further confirm that the use of econometric methods which do not allow for high-dimensional controls and/or dyadic dependence can lead to misleading conclusions, either in terms of point estimates or standard errors in the context of FTA formation models.
These empirical features are common to a wide set of network formation models across fields (see \citet*{graham2019dyadic} and graham2019econometric for surveys), but are particularly striking in our context. Although trade agreements are increasingly common, they continue to be relatively rare events. Despite the rise of trade agreements and the availability of rich network data, incorporating a wide host of determinants, allowing for country asymmetry, and flexibly controlling for dyadic dependence creates a particularly demanding empirical setting beyond the reach of standard tools. As we document below, our dyadic DML approach leverages benefits to dyadic robust inference and high-dimensional link formation to identify and robustly quantify the key determinants of FTA formation. In this sense, our work yields insights which bridges a long series of theoretical research in international trade and allows for a rich of set of new empirical investigations into the nature of international trade agreement networks.
In this section, we introduce the model and present our proposed method of dyadic machine learning. A theoretical guarantee that this method works will be presented in Section (ref). We start by fixing notations to be used to describe dyadic data. Let $\overline{\mathbb{N}^{+2}} = \{(i,j) \in \mathbb{N}^{+} \times \mathbb{N}^{+}: i \neq j\}$ denote the set of two-tuples of $\mathbb{N}^+$ without repetition, where $A^+:=A\cap (0,+\infty)$ for any $A\subset \mathbb{R}$. A dyadic observation is written as $W_{ij}$ for $(i,j)\in \overline{\mathbb{N}^{+2}}$, and we assume that this random vector $W_{ij}$ is Borel measurable throughout. Assume that a dyadic sample contains $N$ nodes with no self link. Let $\{\mathcal P_N\}_N$ be a sequence of sets of probability laws of $\{W_{ij}\}_{ij}$, let $P=P_{N}\in \mathcal P_N$ denote the law with the sample size $N$, and let ${\rm E}_{P}$ denote the expectation with respect to $P$. We write the dyadic sample expectation operator by $\mathbb{E}_N[\cdot]=\frac{1}{N(N-1)}\sum\limits_{(i,j) \in \overline{[N]^2}} [\cdot]$, where $[r]=\{1,...,r\}$ for any $r \in \mathbb{N}$. For any finite set $I$ with $I\subset [N]$, we let $|I|$ denote the cardinality of $I$, and let $I^c$ denote the complement of $I$, namely, $I^c=[N]\setminus I$. We write the dyadic subsample expectation operator by $\mathbb{E}_{I}[\cdot]:=\frac{1}{|I|(|I|-1)}\sum_{(i,j)\in \overline{I^2}}[\cdot]$.
The economic model is assumed to satisfy the moment restriction
for some score function $\psi$ that depends on a low-dimensional parameter vector $\theta \in \Theta \subset \mathbbm R^{d_\theta}$ and a nuisance parameter $\eta \in T$ for a convex set $T$. The nuisance parameter $\eta$ may be finite-, high-, or infinite-dimensional. The true values of $\theta$ and $\eta$ are denoted by $\theta_0 \in \Theta$ and $\eta_0 \in T$, respectively.
Let $\widetilde T=\{\eta - \eta_0 : \eta \in T\}$, and define the Gateaux derivative map $D_r: \widetilde T \rightarrow \mathbbm R^{d_\theta}$ by $ D_r[\eta-\eta_0]:=\partial_r \Big\{ {\rm E}_{P}[\psi(W_{ij};\theta_0,\eta_0+r(\eta-\eta_0))]\Big\} $ for all $r\in[0,1)$. Let its limit denoted by $ \partial_\eta{\rm E}_{P}\psi(W_{ij};\theta_0,\eta_0)[\eta - \eta_0]:=D_0[\eta-\eta_0]. $ We say that the Neyman orthogonality condition holds at $(\theta_0,\eta_0)$ with respect to a nuisance realization set $\mathcal T_N \subset T$ if the score $\psi$ satisfies ((ref)), the pathwise derivative $D_r[\eta-\eta_0]$ exists for all $r\in[0,1)$ and $\eta\in \mathcal T_N$, and the orthogonality equation
holds for all $\eta\in \mathcal T_N$. Furthermore, we also say that the $\lambda_N$ Neyman near-orthogonality condition holds at $(\theta_0,\eta_0)$ with respect to a nuisance realization set $\mathcal T_N\subset T$ if the score $\psi$ satisfies ((ref)), the pathwise derivative $D_r[\eta-\eta_0]$ exists for all $r\in[0,1)$ and $\eta\in \mathcal T_N$, and the orthogonality equation
holds for all $\eta\in \mathcal T_N$ for some positive sequence $\{\lambda_N\}_N$ such that $\lambda_N=o(N^{-1/2})$. We refer readers to CCDDHNR18 for detailed discussions of the Neyman (near) orthogonality, including its intuitions and a general procedure to construct scores $\psi$ satisfying it.
Throughout, we will consider structural models satisfying the moment restriction ((ref)) and either form of the Neyman orthogonality conditions, ((ref)) or ((ref)). As a concrete motivating example of the above abstract formulation, Example (ref) below presents the logit dyadic link formation models, which we consider for our main application in this paper. In addition, we also present a couple of simpler examples with the linear regression models and the linear IV regression models in Appendix (ref).
For the class of models introduced in Section (ref), we now propose a novel dyadic cross fitting procedure for estimation of $\theta_0$. With a fixed positive integer $K$, randomly partition the set $[N] = \{1,...,N\}$ of indices of $N$ nodes into $K$ parts $\{I_1,...,I_K\}$. For each $k \in [K] = \{1,...,K\}$, obtain an estimate $$\widehat \eta_{k}=\widehat \eta\left((W_{ij})_{(i,j)\in \overline{([N]\setminus I_k )^2}}\right)$$ of the nuisance parameter $\eta$ by a machine learning method (e.g., lasso, post-lasso, elastic nets, ridge, deep neural networks, and boosted trees) using only the subsample of those observations with dyadic indices $(i,j)$ in $\overline{([N]\setminus I_k )^2}$. In turn, we define the dyadic machine learning estimator $\widetilde \theta$ for $\theta_0$ as the solution to
where we recall that $\mathbb{E}_{I_k} [f(W)] = \frac{1}{|I_k|(|I_k|-1)}\sum_{(i,j)\in \overline{I_k^2}} f(W_{ij})$ denotes the subsample empirical expectation using only the those observations with dyadic indices $(i,j)$ in $\overline{I_k ^2}$. If an achievement of the exact 0 is not possible as in (ref), then one may define the estimator $\widetilde{\theta}$ as an approximate $\epsilon_N$-solution:
where $\epsilon_N=o(\delta_N N^{-1/2})$ and restrictions on the sequence $\{\delta_N\}_N$ will be formally discussed in Section (ref). Under the assumptions to be introduced in Section (ref), this dyadic machine learning estimator $\widetilde\theta$ enjoys the root-$N$ asymptotic normality $ \sqrt{N}\sigma^{-1}(\widetilde \theta - \theta_0) \leadsto N(0,I_{d_\theta}), $ where a concrete expression for the asymptotic variance $\sigma^2$ will be presented in Theorem (ref) ahead.
Note that, for each $k\in [K]$, the nuisance parameter estimator $\widehat\eta_{k}$ is computed using the subsample of those observations with dyadic indices $(i,j) \in\overline{ ([N]\setminus I_k )^2} $, and in turn the average score $\mathbb{E}_{I_k}[\psi(W; \cdot,\widehat\eta_{k})]$ is computed using the subsample of those observations with dyadic indices $(i,j) \in\overline{ I_k ^2}$. This two-step computation is repeated $K$ times for every $k \in [K]$. We call this cross fitting procedure the $K$-fold dyadic cross fitting. It differs from and complements the cross fitting procedure of CCDDHNR18 for i.i.d. data and the cross fitting procedure of CKMS2019 for multiway clustered data -- neither of these existing cross fitting methods will work under dyadic dependence. If the dyadic sample $\{W_{ij}\}_{ij}$ would reduce to the monadic sample $\{W_i\}_i$, then the $K$-fold dyadic cross fitting would accordingly boil down to the $K$-fold cross fitting procedure of CCDDHNR18.
Figure (ref) illustrates the dyadic cross fitting for the case of $K=2$. We let $N=8$ for simplicity of illustration, although actual sample sizes should be much larger. Suppose that a random partition of $[8] = \{1,...,8\}$ entails two folds with one consisting of $I_1 = \{1,...,4\}$ and the other consisting of $I_2 = \{5,...,8\}$. In the left panel, we compute $\widehat\eta_1$ with a machine learner applied to the subsample $\overline{(I_1^c)^2}$ marked by “Nuisance” in the gray shades at the bottom right quarter of the grid, and then evaluate the subsample mean $\mathbb{E}_{I_1}[\psi(W;\widetilde \theta,\widehat \eta_{1})]$ of the score using the subsample $\overline{I_1^2}$ marked by “Score” in the gray shades at the top left quarter of the grid. In the right panel, in turn, we compute $\widehat\eta_2$ with a machine learner applied to the subsample $\overline{(I_2^c)^2}$ marked by “Nuisance” in the gray shades at the top left quarter of the grid, and then evaluate the subsample mean $\mathbb{E}_{I_2}[\psi(W;\widetilde \theta,\widehat \eta_{2})]$ of the score using the subsample $\overline{I_2^2}$ marked by “Score” in the gray shades at the bottom right quarter of the grid. Although the figure may apparently suggest as if we were discarding the information in the off-diagonal half of the dyad, our dyadic machine learning method uses the full information of $N$ nodes for estimation of $\theta$. Specifically, this means that Equation (ref) evaluates the score with all of the $N$ nodes -- see Equation (ref) ahead for a precise mathematical expression in support of this argument. The white off-diagonal blocks reflect that the $(K-1)/K$ fraction of data are used to estimate the nuisance parameter $\eta_k$ for each $k \in [K]$, but this will not affect the asymptotic behavior of the estimator $\widetilde\theta$ of interest due to the Neyman orthogonality, as is the case with the conventional DML of CCDDHNR18. As a matter of fact, the next section shows that our estimator $\widetilde\theta$ achieves the root-$N$ asymptotic normality with the asymptotic variance being the same as the one in the case where an oracle would let us know the true nuisance parameter $\eta$.
In this section, we present a formal theory that guarantees that the dyadic machine learning estimator introduced in Section (ref) works. For convenience, we fix additional notations. Let $\{a_N\}_{N\geq 1}$, $\{v_N\}_{N\geq 1}$, $\{D_N\}_{N\geq 1}$, $\{B_{1N}\}_{N\geq 1}$, $\{B_{2N}\}_{N\geq 1}$ be some sequences of positive constants, possibly growing to infinity, where $a_N\geq N\vee D_N$ and $v_N\geq 1$, $B_{1N}\geq 1$, $B_{2N}\geq 1$ for all $N\geq 1$. Let $\{\delta_N\}_{N\ge 1}$, $\{\Delta_N\}_{N\ge 1}$, and $\{\tau_N\}_{N\geq 1}$ be sequences of positive constants that converge to zero such that $\delta_N \ge N^{-1/2}$. We use $a\lesssim b$ to mean $a\leq cb$ for some $c>0$ that does not depend on $n$, use $a\gtrsim b$ to mean $a\geq cb$ for some $c>0$ that does not depend on $n$, and use the notations $a\vee b=\max\{a,b\}$ and $a\wedge b=\min\{a,b\}$. For any vector $\delta$, we define the $l_1$-norm by $\|\delta\|_1$, $l_2$-norm by $\|\delta\|$, $l_{\infty}$-norm by $\|\delta\|_{\infty}$, and $l_0$-seminorm (the number of non-zero components of $\delta$) by $\|\delta\|_0$. For any function $f\in L^q(P)$, we define $\|f\|_{P,q}=(\int |f(w)|^q dP(w))^{1/q}$. We use $\|x_{ij}'\delta\|_{2,N}$ to denote the prediction norm of $\delta$, namely, $\|x_{ij}'\delta\|_{2,N}=\sqrt{\mathbb{E}_N[(x_{ij}'\delta)^2]}$.
The following assumption formally states the dyadic sampling under our consideration.
Part (i) requires the identical distribution of $W_{ij}$ while we do allow for dyadic dependence. Part (ii) requires that two groups of links in the dyad are independent whenever they do not share a common node. On the other hand, we do allow for arbitrary statistical dependence between any pair of links whenever they share a common node. Assumption (ref) replaces the i.i.d. sampling assumption of CCDDHNR18, and is a dyadic counterpart of Assumption 1 in CKMS2019 for multiway clustered data. Assumption (ref) (i) and (ii) together imply the Aldous-Hoover-Kallenberg representation (e.g. Kallenberg2006), which states that there exists an unknown (to the researcher) Borel measurable function $\tau_n$ such that
where $\{U_i, U_{\{i,j\}}: i,j\in [N], i\ne j\}$ are some i.i.d. latent shocks that can be taken to be $\text{Unif}[0,1]$ without loss of generality -- see aldous1981representations. This is a mathematically different structural representation (Aldous-Hoover-Kallenberg representation; see e.g. Kallenberg2006) from that in either of these preceding papers and thus none of the existing DML methods apply here -- see chiang2020inference for a detailed comparison and discussion of the differences. Finally, we remark that monadic variables (i.e., coordinates of $W_{ij}$ that do not vary with $i$ or $j$) are not ruled out from our sampling framework of Assumption (ref).
We further state the following two conditions.
The last two assumptions are counterparts of Assumptions 2.1 and 2.2 in BCCW18, except that the data in this paper are allowed to exhibit dyadic dependence. { One observation is that although there are multiple $N$-dependent sequences of constants appearing in these assumptions that might be, at first glimpse, interacting in complicated ways, their rates are pinned down by $\delta_N$ in Assumption (ref) (vii), analogously to the high level conditions in BCCW18 and CCDDHNR18. } Although these conditions are high-level statements, we provide explicit lower-level primitive conditions that imply these assumptions in the context of the logit dyadic link formation model (Example (ref)) in Section (ref).
The following theorem establishes the root-$N$ asymptotic normality of the dyadic DML estimator $\widehat\theta$ along with its consistent asymptotic variance estimation.
The asymptotic distribution is different from both that of the prototypical DMLCCDDHNR18 for i.i.d. data and that of the multi-way cluster-robust DML CKMS2019. In addition to the differences in dependence structures, a yet major departure from Theorem 1 in CKMS2019 is that this result allows a large class of nonlinear score, whereas only linear scores are permitted in Chiang et al. To handle the complications resulted from the nonlinearity, it is necessary to handle empirical processes that consist of dyadic random elements. Coping with these technical issues requires delicate works and hence Theorem (ref) is not a straightforward extension of results in CKMS2019.
In this section, we demonstrate an application of the general theory of dyadic machine learning from Section (ref) to high-dimensional dyadic/network link formation models. Consider a class of logit dyadic link formation models where the binary outcome $Y_{ij}$, which indicates a link formed between nodes $i$ and $j$, is chosen based on an explanatory variable $D_{ij}$ and $p$-dimensional control variables $X_{ij}$. Specifically, we write the model
where $\Lambda(t)=\exp(t)/(1+\exp (t))$ for all $t \in \mathbbm R$. Suppose that we are interested in the parameter $\theta$, while $x \mapsto x'\beta_0$ is treated as a nuisance function $\eta$.
Under i.i.d. sampling environments, BelloniChernozhukovWei2016 propose a method of inference for $\theta_0$ in high-dimensional logit models. We demonstrate that a modification of their suggested procedure with a flavor of our proposed dyadic cross fitting enables root-$N$ consistent estimation and asymptotically valid inference for $\theta_0$ under dyadic sampling environments by virtue of Theorem (ref). In Section (ref), we will supplement this procedure with a formal discussion of lower-level primitive conditions for the high-level statements in Assumptions (ref) and (ref) that are invoked in Theorem (ref).
As in BelloniChernozhukovWei2016, our goal is to construct a generated random variable $Z_{ij}=Z(D_{ij},X_{ij})$ such that the following three conditions are satisfied: \begingroup \allowdisplaybreaks
\endgroup Equations ((ref)) and ((ref)) are used for consistent estimation of $\theta_0$, while Equation ((ref)) obeys the Neyman orthogonality condition -- see Section (ref).
To this end, following BelloniChernozhukovWei2016, consider the weighted regression of $D_{ij}$ on $X_{ij}$:
where $f_{ij}:=w_{ij} / \sigma_{ij}$, $\sigma_{ij}^{2}:={\rm V}\left(Y_{ij} | D_{ij}, X_{ij}\right)$, $w_{ij}:=\Lambda^{(1)}\left(D_{ij} \theta_{0}+X_{ij}^{\prime} \beta_{0}\right)$, and $\Lambda^{(1)}(t)=\frac{\partial}{\partial t}\Lambda(t)$. Under the logit link $\Lambda$, these objects are specifically given by $f_{ij}^2=w_{ij}$ and $\sigma_{ij}^{2}=w_{ij}=\Lambda(D_{ij} \theta_{0}+X_{ij}^{\prime} \beta_{0})\{1-\Lambda(D_{ij} \theta_{0}+X_{ij}^{\prime} \beta_{0})\}$. We let $Z_{ij}=D_{ij}-X_{ij}'\gamma_0$.
From the log-likelihood function of the logit model $ L(\theta, \beta)=\mathbb{E}_{N}\left[L(W_{ij};\theta, \beta)\right], $ where $L(W_{ij};\theta, \beta)=\log \{1+\exp(D_{ij} \theta+X_{ij}' \beta)\}-Y_{ij}(D_{ij} \theta+X_{ij}' \beta)$, one can derive the following Neyman orthogonal score
where $\eta=(\beta',\gamma')'$ denotes the nuisance parameters.
With this procedure, applying Theorem (ref) yields the asymptotic normality result
where the components of the variance $ \sigma^2:=J_0^{-1}\Gamma(J_0^{-1})' $ take the forms of \begingroup \allowdisplaybreaks
\endgroup Section (ref) will present formal theories to guarantee that this application of Theorem (ref) is valid, based on lower-level sufficient conditions tailored to the current specific model.
Summarizing the above procedure, we provide a step-by-step algorithm below for dyadic machine learning estimation and inference about $\theta_0$. For any set $I$, we let $I^c$ denote its complement, and let $|I|$ denote the cardinality of $I$. We remind readers of the notation for the dyadic subsample expectation operator $\mathbb{E}_{I}[\cdot]:=\frac{1}{|I|(|I|-1)}\sum_{(i,j)\in \overline{I^2}}[\cdot]$.
In this section, we provide lower-level sufficient conditions that guarantee the root-$N$ asymptotic normality (ref) in the application to the high-dimensional logit dyadic link formation models. For convenience of stating such conditions, we first introduce additional notations. We continue to use $Z_{ij}=D_{ij}-X_{ij}'\gamma_0$ from the previous subsection. Let $a_N=p\vee N$. Let $q$, $c_1$, $C_1$ be strictly positive and finite constants with $q>4$, $M_N$ be a positive sequence of constants such that $M_N\geq \{{\rm E}_{P}[(D_{12}\vee \|X_{12}\|_{\infty})^{2q}]\}^{1/2q}$. For any $T\subset [p+1]$, $\delta=(\delta_1,...,\delta_{p+1})'\in \mathbbm R^{p+1}$, denote $\delta_T=(\delta_{T,1},...,\delta_{T,p+1})'$ with $\delta_{T,j}=\delta_j$ if $j\in T$ and $\delta_{T,j}=0$ if $j\not \in T$. Define the minimum and maximum sparse eigenvalues by
The following assumptions constitute the lower level sufficient conditions. Let $s_N\ge 1$ be a sequence of integers.
Assumption (ref) is a sparse eigenvalue condition similar to the one in bickel2009simultaneous under i.i.d. setting. Sufficient conditions for the sparse eigenvalue conditions under the related multiway clustering is provided by Proposition 3 in chiang2020inference. Assumption (ref) imposes a restriction on the number of non-zero components in the nuisance parameter vectors by $s_N$, where a constraint on the growth rate of $s_N$ in turn will be introduced in Assumption (ref) (iv). Assumption (ref) is standard and requires the parameter space $\Theta$ to be bounded as well as the true nuisance parameter vectors $\beta_0$ and $\gamma_0$ to have bounded $\ell_2$-norms. It permits $\|\beta_0\|_1$ and $\|\gamma_0\|_1$ to be growing as $N$ increases. Assumption (ref) (i) and (ii) impose lower bounds on eigenvalues and variances of some population conditional moments. Note that the $U_j$'s here come from the Aldous-Hoover-Kallenberg representation introduced in Equation ((ref)). Assumption (ref) (iii) requires the fourth moments of the covariates to be bounded in a uniform manner. Finally, Assumption (ref) (iv) constraints the rates at which the sparsity index $s_N$, the dimensionality $p$, as well as the bound for the $2q$-th moments of the variable of interest and the maximal components of the covariates vector can grow. It also implicitly requires these $2q$-th moments to exist (albeit the bounds can be diverging with $N$). As an illustration, suppose that $\|(D_{ij},X_{ij}')\|_\infty=1$. Then $q$ can be taken arbitrarily large and $M_N=1$. In this case, it essentially requires $N^{-1}s_N^2\log^2 a_N =o(1)$. These rate requirements are comparable to those assumed in state-of-the-art results in the literature under i.i.d. sampling (e.g. BCCW18).
These assumptions are sufficient for the high-level conditions invoked in Theorem (ref), as formally stated in the following lemma.
As a consequence of this lemma and the general result in Theorem (ref), we obtain the following result of the root-$N$ asymptotic normality of the dyadic machine learning estimator $\check\theta$ in the high-dimesnional logit dyadic link formation models.
Focusing on the application to the high-dimensional dyadic link formation model introduced in Section (ref), we study and present finite sample performance of the proposed method of dyadic machine learning in this section through Monte Carlo simulations.
We design the data generating process for our Monte Carlo simulation studies as follows. For each $(i,j) \in \overline{[N]^2}$, generate the random vector $(D_{ij},X_{ij}',\varepsilon_{ij})'$ according to \begingroup \allowdisplaybreaks
\endgroup where $\{(\widetilde D_{i}, \widetilde X_{i}')'\}_{i \in [N]}$, $\{(\widetilde D_{ij}, \widetilde X_{ij}')'\}_{(i,j) \in \overline{[N]^2}}$, $\{\widetilde \varepsilon_{i}\}_{i \in [N]}$ and $\{\widetilde \varepsilon_{ij}\}_{(i,j) \in \overline{[N]^2}}$ are independently drawn with the laws
and $\Sigma$ is the variance-covariance matrix whose $(r,c)$-th entry given by $\Sigma_{rc} = 5^{-|r-c|}$. While $D_{ij}$ is designed as a scalar random variable, we will vary the dimension of the random vector $X_{ij}$ across sets of simulations.
From these primitive variables, construct the dyadic binary decision $Y_{ij}$ in turn by the logistic threshold-crossing model
Note that this data generating model implies the logistic regression
We thus obtain a dyadic sample $\{(Y_{ij},D_{ij},X_{ij}')'\}_{(i,j) \in \overline{[N]^2}}$, which we assume is observed by a researcher. For the true parameter values, set $\theta_0 = 1$, and let the $c$-th coordinate of $\beta_0$ be $2(-2)^{-c}$ for $c \leq \lfloor \sqrt{N} \rfloor$ and 0 otherwise. We will experiment with various sample sizes $N$ as well as various dimensions of $X_{ij}$ across sets of simulations. The number of Monte Carlo iterations is set to 2,500 for each set of simulations.
Table (ref) summarizes Monte Carlo simulation results. The top panel of the table shows results for the conventional double/debiased machine learning without the dyadic cross fitting or the dyadic robust standard errors. The bottom panel of the table shows results for our proposed dyadic double/debiased machine learning with the dyadic cross fitting and the dyadic robust standard errors. Displayed are simulation results that vary with the sample size $N$, the dimension $\text{dim}(X)$ of $X_{ij}$, and the number $K$ of folds for the dyadic cross fitting. Displayed statistics include the simulation mean of $\check\theta$, the simulation bias of $\check\theta$, the simulation standard deviation of $\check\theta$, the simulation root mean squared error of $\check\theta$, the simulation inter-quartile range of $\check\theta$, and the simulation coverage frequencies for nominal sizes of 90% and 95%.
First, compare the top panel and bottom panel of the table in terms of simulation statistics of the estimates. Notice that the conventional DML (without accounting for dyadic data) produces biased estimates of $\check{\theta}$, while the dyadic DML produces less biased estimates. Recall that the conventional cross fitting works under independent sampling, while dyadic data are not independent. As such, for dyadic data, conventional cross fitting may not be able to remove over-fitting biases. On the other hand, the simulation results support the claim that dyadic cross fitting accounts for dyadic dependence and is successful in removing over-fitting biases.
Second, compare the top panel and bottom panel of the table now in terms of simulation coverage frequencies. Notice that the conventional DML (without accounting for dyadic data) suffers from severe under-coverage especially for larger sample sizes, while the dyadic DML maintains correct coverage. This difference may be imputed to two sources. One source of this difference is that estimates by the conventional DML are biased while those by the dyadic DML is less biased as discussed in the previous paragraph. As the sample size increases, it becomes less likely that shorter confidence intervals centered around the biased estimates $\check{\theta}$ contain the true parameter value $\theta_0$. Another source of the difference is that the conventional standard error without accounting for dyadic sampling is simply too small compared to the dyadic robust standard error.
In summary, the dyadic DML produces less biased estimates and asymptotically accurate coverage. Failure to account for dyadic dependence in cross fitting results in biased estimates, while failure to account for dyadic dependence in cross fitting and/or standard errors results in severe under-coverage.
In this section, we reconsider the classic network setting: the formation of free trade agreements across countries. We compare our proposed method of dyadic DML to traditional methods for analyzing determinants of free trade agreements (FTA) using both parsimonious and rich empirical models. In this sense, we contribute towards addressing the “inherent complexity and ambiguity” of the “pure economic theory of trading blocs,” krugman1993regionalism and add to the literature confirming the assertion “detailed empirical work is the right direction” needed to make headway. baier2004economic pioneered empirical studies in this field, and a large number of subsequent studies have followed since that time based on their foundational specification. We aim to strengthen our understanding of the conclusions produced in this literature by conducting a new empirical analysis with the following econometric details in mind: 1. we control for high-dimensional covariates in order to mitigate unobserved endogeneity or unobserved confoundedness; and 2. we account for the dyadic dependence inherent in trade data and, for the first time, report standard errors robust to this source of misleading inference.
Consider the empirical model
where $Y_{ij}$ is the dummy variable if there is a bilateral FTA between $i$ and $j$ as of year 2000. The parsimonious specification includes four classic explanatory variables $(D_{ij},X_{ij}')'$: (A) the logarithm of the population-weighted bilateral distance between $i$ and $j$ in kilometers, (B) the sum of the logarithms of the per-capita GDPs of $i$ and $j$ in the baseline year, (C) the absolute difference of the logarithms of the per-capita GDPs between $i$ and $j$ in the baseline year, and (D) the absolute difference of the logarithms of the capital-labor ratios between $i$ and $j$ in the baseline year. Following baier2004economic we complete our most parsimonious specification by including in the square of the difference of the logarithms of the capital-labor ratio.
Bilateral distance is a standard proxy for trade costs in this literature and helps capture the fact smaller bilateral trade costs raise the benefits from a trade agreement, ceteris paribus. As highlighted by baier2004economic, a standard prediction of economic theory is that “[t]he net gain from an FTA between two countries increases as the distance between them decreases.” Also see krugman1991move and frankel1993continental,frankel1995trading,frankel1996regional. This hypothesis motivates (A).
Market size is measured by the sum of log country-level GDPs, conditional on trade costs, while similarity in economic development is captured by the absolute difference of log GDPs. The former reflects the common finding that “[t]he net welfare gain from an FTA between a pair of countries increases the larger are their economic sizes (i.e. average real GDPs),” while the latter investigates the notion that “[t]he net welfare gain from an FTA between a pair of countries increases the more similar are their economic sizes (i.e. real GDPs).” Also see krugman1998comment. These hypotheses motivate (B) and (C).
Last, we also investigate the degree to which differences in factor endowments increase the likelihood of an FTA through the inclusion of the difference of the log capital-labor ratio across potential trade partners. baier2004economic state that “[t]he net welfare gain from an FTA between a pair of countries increases with wider relative factor endowments, but might eventually decline due to increased specialization.” Accordingly, we include the absolute difference of the log capital-labor ratio to capture the notion that countries with a greater difference in initial endowments are more likely to have greater differences in comparative advantage across industries and, as such, greater potential gains from trade with each other. This hypothesis motivates (D). As in baier2004economic we also include the square of the difference of the log capital-labor ratio as a control for a diminishing impact of differences in factor endowments.
Of course, there are wide host of additional bilateral characteristics which affect realized trade flows across countries and, as such, the likelihood of successfully forming an FTA. For instance, research confirms that time differences, colonial ties, language or ethic overlap, common legal origins, political ties, among many other country-pair characteristics, affect trade flows. Indeed, to the extent that these co-variates reflect differential market access, trade costs or the potential future benefits from trade agreements, there exclusion from existing work may present a source of omitted variable bias.
To investigate the impact of restricting attention to a parsimonious specification, we add a wide set of additional controls to our benchmark empirical model including the time difference between $i$ and $j$ in hours, a religious proximity index and a wide set of dummy variables capturing whether $i$ and $j$ have ever been in a colonial relationship, whether they have been in a colonial relationship after 1945, whether they had a common colonizer after 1945, whether they share a common official or primary language, whether there is a common language spoken by at least 9% of the population in both countries, whether one country is current or former hegemon of the other, whether $j$ is current or former hegemon of $i$, whether they have ever been two colonies of the same empire, whether they shared a common legal origin before or after transition to independent statehood. To saturate our set of potential confounders we also include all of the powers and interactions for each of these variables -- see Appendix (ref) for details of how we construct each co-variate. In total, there are more than 140 explanatory variables in the rich specifications that we consider, and these dimensions are high relative to the effective sample size $N$ of the dyad which consists of $N=229$ economies. Following baier2004economic, we set year 1960 as the baseline year for variable measurement, which is roughly the year when FTAs started to be signed across across global trading partners.
The dependent variable reflects whether any two countries have an existing FTA and comes from the data used by head2010erosion and baier2009estimating. Distance and gross domestic product data are retrieved from the CEPII Distance Dataset and World Development Indicators (World Bank), respectively. The information about colonial relationships, languages, wars, hegemon relationships, and sibling relationships come from the data used by head2010erosion, while the nature of bilateral legal relationships and religious proximity are respectively reported in la2008economic and disdier2007je. See the data descriptions in Appendix (ref) for further details of this data set.
Table (ref) summarizes estimation and inference results based on 50 iterations of resampled cross fitting. In this table, we report estimates and standard errors for the coefficients of log population-weighted distance, the sum of the log GDPs, the absolute difference of the log GDPs, and the difference of the log capital-labor ratios. These four explanatory variables correspond to hypotheses (A), (B), (C) and (D) suggested in the previous paragraph. The first column (I) reports results based on the simple, parsimonious logistic regression not including high-dimensional regressors and not accounting for dyadic sampling. The second column (II) reports results based on the simple logistic regression, while including high-dimensional regressors. It does not account for dyadic sampling. Columns (III) and (IV) report results based on conventional machine learning with $K=5$ and 10, respectively. The last two columns (V) and (VI) report results based on our proposed dyadic machine learning where we again set $K=5$ and 10, respectively.
We observe the following sharp results. First, the simple logit replicates the sign on each of the included co-variates in baier2004economic. Likewise, each co-variate is also statistically significant at conventional levels, save one.\footnote{This extends to the square of the log difference in capital labor ratios. The estimated coefficient for this co-variate is -0.169 with a standard error of 0.021 in column (I), and -0.117 with as standard error of 0.036 in column (II). We do not report the estimated coefficient for the square of the log difference in capital labor ratios in columns (III)-(VI) since it is treated as a control variable in $X_{ij}$ rather than part of $D_{ij}$ itself.} While baier2004economic find that greater economic similarity, measured as the difference in log GDPs, is a significant determinant of FTAs, our simple logit suggests that this coefficient is small and insignificantly different from zero. This likely reflects differences in the underlying data. Our model is based on FTAs signed by the year 2000, four years after the analysis conducted in baier2004economic and a period when a significant number of FTAs were signed. Further, improved data availability allows us to extend their analysis of FTA formation among 54 countries to a setting with 167 countries, including many additional developing countries.\footnote{baier2004economic also use a probit model rather than a logit model. We investigated whether difference in methodology resulted in significantly different results in column (I). Rather, probit results were very similar to the logit specification. We focus on the logit since it is most directly comparable to the conventional ML and dyadic ML approaches in columns (III)-(VI).}
Second, rows (A), (B) and (D) each highlight a separate, but equally important, feature of dyadic machine learning in this classic setting. For instance, in row (A) the dyadic ML indicates that the coefficient of (A) the logarithm of the population-weighted distance is negative and significantly different from zero. In this sense, the estimated coefficient is similar to that estimated by the parsimonious logit, the full logit, and conventional ML. However, the estimated standard errors in columns (V) and (VI) are 48-53 percent larger than those in column (II), the rich logit specification, and 37-46 percent larger than those in columns (III)-(IV), where conventional machine learning is employed. Nonetheless, our estimates support the hypothesis that as distance, and trade costs, between two countries decrease the net gain from an FTA between them rises krugman1991move,frankel1993continental,frankel1995trading,frankel1996regional,baier2004economic. The dyadic ML estimates again confirm that this common empirical conclusion is robust to the inclusion of high-dimensional covariates and controlling for dyadic dependence among FTA formation data.
In row (B) we observe that the dyadic ML also indicates that the coefficient of (B) economic size, measured as the sum of log GDP, is positive and significantly different from zero. Again, this finding is similar to that indicated in the parsimonious logit, the full logit, and the conventional ML and supports the hypothesis that greater economic size between a pair of countries increases the net welfare gain from an FTA krugman1998comment,baier2004economic. However, our results also suggest that ignoring dyadic dependence is likely to overstate the importance of market size in the formation in FTAs.
Specifically, the magnitude of the estimated coefficient on the sum of the log GDPs changes significantly across columns and is consistent with claim that conventional methods may yield biased estimates in rich specifications. For instance, in columns (II), (III) and (IV), the estimation routines all return estimates which are substantially larger than those in columns (V) and (VI) where we control dyadic dependence. These empirical results are consistent with the findings in our Monte Carlo simulations: conventional ML was likely to produce biased estimates in rich specifications. In Table 2, we find that FTA market size is likely to receive outsized importance; the estimated coefficients on the size variable in columns (V) and (VI) are 27-32 percent smaller than those columns (III) and (IV).
Row (D) illustrates a different source of bias. In particular, the simple logit indicates that the coefficient on (D) relative factor endowments, measured as the difference of the log capital-labor ratio, is positive and significant in columns (I) and (II). This result is consistent with existing empirical results that confirm that differences in factor endowments are important determinants of FTAs. However, once we implement machine learning variable selection in the high-dimensional controls (columns (III)-(IV)), just the opposite pattern emerges: greater similarity in capital-labor ratios are found to encourage FTA formation when we employ conventional ML. In this sense, benchmark analyses which do not directly account for the inclusion and confoundedness associated with high dimensional co-variates are likely to deliver misleading results. For instance, baier2004economic do control for a number of the above co-variates in the robustness checks of their seminal analysis and, like our results in column (II), find that the inclusion of additional co-variates in the simple logit has little impact on their benchmark findings.
At the same time allowing for high dimensional co-variates and dyadic dependence in columns (V) and (VI) also has a large impact on the recovered standard errors. Indeed, in this case, it changes preceding conclusions: while we maintain the positive estimated coefficient on the relative factor endowments variable after controlling for dyadic dependence, the standard error more than doubles. We can no longer conclude that relative factor endowments are an important determinant of FTA formation, overturning this classic result. More generally, accounting for dyadic dependence significantly increases the recovered standard errors for all of our explanatory variables by at least 37 percent relative to conventional ML, which was already relatively conservative.
Finally, all the methods indicate that the coefficient of (C) economic similarity, measured as the absolute difference of log GDP, is consistently negative, but statistically insignificant. In this sense, we do not find strong evidence, regardless of methodology, to support the hypothesis that greater economic similarity between a pair or countries increases the likelihood of an FTA being formed between them.
In sum, we confirm that trade costs and market size are key determinants of FTA formation, even after controlling for high-dimensional covariates and robustly accounting for the dyadic dependence in data. Allowing for high-dimensional covariates and robustly accounting for the dyadic dependence are both generally found to substantially increase standard errors and reduce $t$-statistics. Our estimates also suggest that the importance of market size in FTA formation may be biased and overstated when employing conventional methods, including conventional ML, and a rich specification. Indeed, we document that standard estimation approaches, such as logit or conventional ML, may lead researchers to conclude that larger differences in relative factor endowment may encourage or discourage FTA formation. Accounting for dyadic dependence, in contrast, increases standard errors sufficiently that we can no longer conclude that relative factor endowments are an important determinant of FTA formation. In this sense, we further confirm that the use of econometric methods which do not allow for high-dimensional controls and/or dyadic dependence can lead to misleading conclusions, either in terms of point estimates or standard errors in the context of FTA formation models. Last, in contrast to the existing literature, we do not find evidence that differences in economic similarity or factor endowments are particularly important determinants of FTA formation regardless of which estimation approach we employ.
This paper presents new methods and theories for the use of machine learning when data are dyadic. A novel dyadic cross fitting algorithm is proposed to remove over-fitting biases under arbitrary dyadic dependence. Together with the use of Neyman orthogonal scores, this dyadic cross fitting method enables double/debiased machine learning for dyadic data. Applying the general results (Theorem (ref)), we demonstrate how to estimate and conduct inference for high-dimensional logit dyadic link formation models.
With this method applied to empirical data of international economic networks, we reexamine FTA determinants as links formed in the dyad composed of world economies. We reconfirm two important theoretical implications suggested by the international trade literature, even after robustly accounting for both dyadic dependence and high-dimensional controls. Namely, (A) a greater distance between economies makes an FTA less beneficial and thus makes an FTA less likely to be formed; and (B) larger sizes of economies make an FTA more beneficial and thus make an FTA more likely to be formed. In the context of a rich specification, we further document that conventional methods return biased FTA market size coefficients. Similarly, while standard approaches suggest larger differences in relative factor endowment are an important determinant of FTA formation, once we account for dyadic dependence we can no longer support this standard conclusion.
Motivated by the application to the dyadic link formation model for the analysis of FTA determinants, we focus our econometric methodology on unconditional moment restrictions in this paper. In other applications, such as the instrumental variable regression in dyadic data, conditional moment restrictions are more relevant AiChen2003,AiChen2007,ChenLintonKeilegom2003,ChenPouzo2015. We conjecture that our proposed method with dyadic cross fitting will extend to estimation and inference for models based on conditional moment restrictions. Formal theoretical investigation in this important direction is left for future research, given the length of the current paper.