EconBase
← Back to paper

State Dependence and Unobserved Heterogeneity in the Extensive Margin of Trade

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

98,450 characters · 11 sections · 91 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

State Dependence and Unobserved Heterogeneity in the Extensive Margin of Trade

{0em} {0em} {1em} {1em}

\thispagestyle{empty}

abstractWe study the role and drivers of persistence in the extensive margin of bilateral trade. Motivated by a stylized heterogeneous firms model of international trade with market entry costs, we consider dynamic three-way fixed effects binary choice models and study the corresponding incidental parameter problem. The standard maximum likelihood estimator is consistent under asymptotics where all panel dimensions grow at a constant rate, but it has an asymptotic bias in its limiting distribution, invalidating inference even in situations where the bias appears to be small. Thus, we propose two different bias-corrected estimators. Monte Carlo simulations confirm their desirable statistical properties. We apply these estimators in a reassessment of the most commonly studied determinants of the extensive margin of trade. Both true state dependence and unobserved heterogeneity contribute considerably to trade persistence and taking this persistence into account matters significantly in identifying the effects of trade policies on the extensive margin.

JEL Classification Codes: C13, C23, C55, F14, F15 \\ Key Words: Dynamic binary choice, extensive margin, high-dimensional fixed effects, incidental parameter bias correction, trade policy

\setcounter{page}{1}

Introduction

What induces country pairs to trade? In 2019, still more than one third of potential bilateral trade relations reported zero trade flows.\footnote{According to data from IMF DOTS.} Comparing these zero trade flows with trade relations in 2018, these zeros turn out to be extremely persistent: 88.6 percent of country pairs that did not trade in 2018 did not trade in 2019 either, as can be seen in the transition matrix depicted in Table (ref). And similarly, 92.3 percent of pairs that did trade in 2018 continued to do so in the year after.\footnote{ Note that throughout the paper, “country pair” refers to a directed pair of countries, i.e.\ Germany-France and France-Germany are two distinct country pairs.}

table[table omitted — 422 chars of source]
figure[figure omitted — 435 chars of source]

Both intuitively and based on existing theoretical and empirical insights, we would expect geographically close and economically large country pairs to have the greater bilateral trade potential and thus be more likely to engage in international trade. As distance is time-invariant and economic size does not change abruptly from one year to another, these gravity-like characteristics may explain (part of) the observed persistence. Figure (ref) breaks down the share of non-zero trade flows in 2019 along the percentiles of four different ad hoc indicators of trade potential: bilateral distance; product of GDPs; “naive” gravity, i.e.\ the product of GDPs divided by the countries' bilateral distance; and the latter when excluding country pairs in FTAs, with common currencies or common colonial history. The x-axis indicates the potential trade volume, i.e.\ the joint economic size and/or proximity of any two countries. All four plots paint a common picture: the black dots, covering all country pairs, show a strong general relationship between trade potential and actual non-zero trade. The blue and red dots split the country pairs according to whether the two did or did not engage in trade in the previous year. The clearly separated pattern for the two groups highlights the remarkable persistence of trade relations, even after controlling for differences in trade potential in terms of distance, size, and bilateral trade policy. More than 50 percent of those country pairs in the lowest percentiles of trade potential trade again in 2019, provided they already did so in 2018. On the other hand, even comparably large and close pairs are likely not to trade in 2019 if they did not trade in 2018 either.\footnote{A very similar pattern emerges for other points in time (see Figure (ref) in Appendix (ref) where the same graph is reproduced for the years 2000--2001). If longer time intervals are considered, a similar picture remains, but the relationship becomes considerably weaker (see Figure (ref) in Appendix (ref) for the years 2000--2019).}\\

Two potential features of the extensive margin of trade that can generate the pattern documented by Table (ref) and Figure (ref) are what Heckman1981 termed “true state dependence” and “spurious state dependence”. The former characterizes countries to actually be more likely to trade because they did so in the previous period, whereas the latter describes persistence due to unobservable factors continuously driving bilateral trade potential.\\

In this paper we introduce estimators for the determinants of the extensive margin of international trade that explicitly take its persistence due to observable characteristics, true state dependence, and unobserved heterogeneity into account. We introduce features from the firm dynamics literature into a heterogeneous firms model of international trade with bounded productivity to derive expressions for an exporting country's participation in a specific destination market in a given period. These expressions depend on partly unobserved (i) exporter-time, (ii) destination-time, and (iii) exporter-destination specific components, as well as on (iv) whether the exporter has already served the market in the previous period, and on (v) exporter-destination-time specific gravity-type trade cost determinants. We estimate the model making use of recent computational advances in the estimation of binary choice estimators with high-dimensional fixed effects to address (i)-(iii). The inclusion of fixed effects in a binary choice setting induces an incidental parameter problem that is potentially aggravated by the dynamics introduced by (iv). To mitigate this bias, we propose new analytically and jackknife bias-corrected estimators for coefficients and average partial effects in three-way fixed effects specifications. Additionally, we provide an expression for long-run partial effects. Extensive simulation experiments demonstrate the desirable statistical properties of our proposed bias-corrected estimators.\\

The empirical application provides evidence that both unobserved bilateral factors and true state dependence due to entry dynamics contribute strongly to the high persistence. Taking this persistence into account changes the estimated effects of the most commonly studied potential determinants considerably: The impact of a common currency is reduced from almost 10 percentage points to less than 4 percentage points, the effect of a common regional trade agreement drops from 5.4 to 1.2 percentage points, and joint membership in the WTO/GATT increases the trading probability by 1.3 rather than 3.4 percentage points. For all three variables, only about two thirds of these effects are realized immediately, while the remaining third is added gradually over time. Furthermore, specifications with a lagged dependent variable and/or bilateral fixed effects yield better predictions for which country pairs will trade than specifications that fail to account for state dependence appropriately.\\

Our paper builds on recent insights from three flourishing strands of literature. First, our paper is related to the literature on the extensive margin of international trade. A number of theoretical frameworks have sought to propose mechanisms behind the decisions of firms to export, and their aggregate implications of zero or non-zero trade flows at the country pair level. Analogous to the intensive margin counterpart, these theories have established gravity-like determinants, such as two countries' bilateral distance, a free trade agreement, a common currency and joint membership in the WTO/GATT. Egger2011a and \citet*{Egger2011b} append an extensive margin to an Anderson2003-type model by assuming export participation to be determined by (homogeneous) firms weighing operating profits and bilateral fixed costs of exporting. Helpman2008 build a model of international trade with heterogeneous firms and bounded productivity in which a country only exports to a given destination if the most productive firm can afford to overcome the fixed costs of exporting. Eaton2013 move away from the arguably simplifying notion of a continuum of firms and develop a model of a finite set of heterogeneous firms. Here, no firm may export to a given market because of their individual efficiency draws. Our model proposed in this paper directly builds on Helpman2008 and extends it by features from the literature on firm dynamics. In this firm-level literature, Das2007 develop a dynamic discrete-choice model in which current export participation depends on previous exporting, and hence sunk costs, and observable characteristics of profits from exporting Roberts1997,Bernard2004. Alessandria2007 embed the distinction between sunk costs and “period-by-period” fixed costs into general equilibrium.\footnote{A number of recent contributions also stress the dynamic character of firms' exporting behaviour and additionally provide alternative rationales for dynamic feedbacks beyond sunk costs of entry, such as “demand learning” or consumer accumulation \citep*[see e.g.][]{Bernard2017,Ruhl2017,Berman2019,Piveteau2019}.} We aim at reconciling the estimation of the aggregate extensive margin with the insight from the firm-level literature that dynamics feature prominently in the determination of the exporting decision by deriving an econometric specification that explicitly incorporates previous export experience at the country pair level. \\

Second, our paper builds on advances in the literature on the gravity equation and the intensive margin of international trade. With the advent of what has now been coined structural gravity Head2014, the gravity framework has gained rich microfoundations. Anderson2003 and Eaton2002 each formulate an underlying structure for exporting and importing countries that in estimations can easily be captured by appropriate two-way country(-time) fixed effects, as first noted by Feenstra2004 and Redding2004. Since Baier2007, it has furthermore become standard to include country pair fixed effects to tackle unobservable bilateral trade cost determinants. Additionally taking into account the multiplicative structure of the gravity equation following SantosSilva2006, nonlinear estimation with exporter-time, importer-time, and country pair fixed effects has become the gold standard for the intensive margin.\footnote{In a linear intensive margin setting, this most general fixed effects structure was already proposed by \citet*{Baltagi2003}.} Estimating the model introduced in this paper similarly calls for three sets of fixed effects, specific to exporters and importers in a given year, as well as to a given country pair over time. The binary nature of the decision whether to export to a destination market at all, also clearly asks for a nonlinear estimator. Therefore, in this paper, we put the estimation of the extensive margin on a par with the intensive margin gold standard by introducing a respective three-way fixed effects binary choice specification. \\

Third, the paper builds on and contributes to the literature on estimating nonlinear fixed effects models. As it is known since Neyman1948, the inclusion of fixed effects potentially introduces an incidental parameter problem (IPP). Although the maximum likelihood estimator (MLE) is consistent if all dimensions of the panel grow large, it has an asymptotic bias in its limiting distribution leading to invalid inference Fernandez-Val2018. Recently, there have been a number of advances to deal with the IPP Fernandez-Val2018. In the context of the aggregate extensive margin, only approaches for cross-sectional bilateral data with importer ($j$) and exporter ($i$) fixed effects have been suggested. Cruz-Gonzalez2017 apply the bias correction of Fernandez-Val2016a\footnote{The bias corrections of Fernandez-Val2016a were originally developed for classical panel data models with individual and time fixed effects and cover a wide range of nonlinear models.}, and Charbonneau2017 proposes a conditional logit estimator. Many trade data sets however consist of bilateral cross-sections over time, i.e.\ network panel data. The theory-consistent estimation of our model includes fixed effects for exporter-time ($it$) and importer-time ($jt$). On closer inspection, one finds that this two-way case can basically be covered by the bias corrections of Fernandez-Val2016a for individual and time fixed effects.\footnote{Similarly, it is possible to adapt the estimator of Charbonneau2017 to the setting with exporter-time ($it$) and importer-time ($jt$) fixed effects. However, her approach has some limitations: 1. it is limited to logit models, 2. it precludes the possibility to estimate average partial effects, 3. it is computationally infeasible in cases where the number of levels per fixed effects becomes large.} However, the literature lacks a suitable method to estimate our preferred specification, which additionally includes a third, bilateral ($ij$), set of fixed effects. Our contribution is to develop suitable three-way fixed effects binary choice estimators for network panel data with potentially weakly exogeneous regressors under asymptotics where $I, J$ and $T$ are large. Therefore, our article complements the work of Weidner2020 on estimating the intensive margin of trade, who examine the IPP for the three-way fixed effects Poisson pseudo maximum likelihood (PPML) estimator under fixed $T$ asymptotics and suggest appropriate bias corrections. \\

The remainder of the paper is structured as follows. In Section (ref) we build a dynamic model of the extensive margin of international trade. The model yields aggregate predictions that can be structurally estimated using a probit model with high-dimensional fixed effects. In Section (ref) we describe the new bias-corrected three-way fixed effects estimator. We demonstrate its performance in Monte Carlo simulations in Section (ref), before finally showing the estimator in action by estimating the model in Section (ref). Section (ref) concludes.

An Empirical Model of the Extensive Margin of Trade

We start by setting up a model of the extensive margin of trade that will later guide our econometric specification. We consider a stylized dynamic Melitz2003-type heterogeneous firms model of international trade. Following Helpman2008 we assume a bounded productivity distribution, like a truncated Pareto in HMR's case. We deviate from HMR by explicitly stating a time dimension and, unlike in the standard Melitz setting, separate fixed exporting costs into costs of entering a new market and costs of selling in a given market Alessandria2007,Das2007.\\

There are $N$ countries, indexed by $i$ and $j$, each of which consumes and produces a continuum of products. The representative consumer in $j$ receives utility according to a CES utility function:

align[align omitted — 216 chars of source]

where $q_{jt} (\omega)$ is $j$'s consumption of product $\omega$ in period $t$, $\Omega_{jt}$ is the set of products available in $j$, $\sigma$ is the elasticity of substitution across products, and $\xi_{ijt}$ is a log-normally distributed idiosyncratic demand shock (with $\mu_\xi=0$ and $\sigma_\xi=1$) for goods from country $i$ in country $j$ and period $t$ Eaton2011. Demand in country $j$ for good $\omega$ depends on this demand shock, $j$'s overall expenditure $E_{jt}$, and the good price $p_{jt}(\omega)$ relative to the overall price level as captured by the price index $P_{jt}$:

align*[align* omitted — 263 chars of source]

Each country has a fixed continuum of potentially active firms that have different productivities drawn from the distribution $G_{it}(\varphi)$, where $\varphi\in(0,\varphi_{it}^{*}]$. The productivity distribution evolves over time and firms' ranks within the productivity distribution can also change from period to period, though firms that in the last period did not export to a market already served by a domestic competitor are assumed not to directly jump to being the country's most productive firm in the next period.\footnote{Note that we could in principle also allow for new firm entry into the pool of potential producers without changing our final expression for the extensive margin as long as the new entrants cannot become the country's most productive firm right away.} Each period, a firm can decide to pay a fixed cost $f_{it}^{prod}$ and start production of a differentiated variety using labour $l$ as its only input, such that $l_{t}(\omega) = f_{it}^{prod} + q_{t}(\omega) / \varphi_{t}(\omega)$. A firm's marginal cost of providing one unit of its good to market $j$ consists of iceberg trade costs $\tau_{ijt}$ and labour costs $w_{it}/\varphi_{t}(\omega)$. Firms compete with each other in monopolistic competition and charge a constant markup over marginal costs. Therefore, the price of a good $\omega$ produced in $i$ and sold in $j$ is:

align*[align* omitted — 102 chars of source]

A firm's operating profits in market $j$ are hence given by:

align*[align* omitted — 184 chars of source]

If a firm wants to export to a market $j$ in period $t$, it has to pay a fixed exporting cost $f_{ijt}^{exp}$. The exporting fixed cost is higher by a market entry cost factor $f^{entry}\geq 1$ if the firm has not been active in the respective market in the previous period. For tractability, the entry cost factor is assumed to be constant across countries and time. Capturing the export decision by a binary variable $y_{ijt}(\omega)$, i.e.\ equal to one if the firm decides to serve market $j$ in period $t$, we can formalize a firm's realized profits in market $j$ as follows:

align*[align* omitted — 143 chars of source]

In the absence of entry costs, a firm would simply compare its operating profits to the fixed exporting cost and decide to serve a market if the former are greater than the latter. With market entry costs, a firm might be willing to incur a loss in the current period if expected future profits from that same market outweigh the initial loss. Firms discount future profits at a rate $\delta$ per period. To keep things tractable and allow us to derive a theory-consistent estimation expression below, we assume that firms expect their future operating profits from and fixed costs of serving a given market to be equal to today's values, i.e.\ $\operatorname{\mathbb{E}}_{t}[\tilde{\pi}_{ij(t+s)}]=\tilde{\pi}_{ijt}$ and $\operatorname{\mathbb{E}}_{t}[f^{exp}_{ij(t+s)}]=f^{exp}_{ijt}$ $\forall s\in\mathbb{N}$.\footnote{Note that our final expression for the extensive margin also holds if firms instead expect their operating profits from serving an export market to grow at a constant rate $\bar{g}<\delta$.} The current value of today's and all future operating profits from market $j$ is then given by $\sum_{s=0}^{\infty}(1-\delta)^{s}\tilde{\pi}_{ijt}=\frac{\tilde{\pi}_{ijt}}{\delta}$. A firm will decide to serve a destination market if these discounted expected profits exceed the sum of today's and discounted future fixed costs of entry and exporting, given by

align*[align* omitted — 217 chars of source]

Given this model setup, the question whether a country exports to another country at all can be considered by looking at the most productive firm (with $\varphi_t^{*}$) only. Denoting that firm's product by $\omega^{*}$, we can capture the aggregate extensive margin by the binary variable $y_{ijt}$ as follows:

align[align omitted — 401 chars of source]

Country $i$ is hence more likely to export to country $j$ in period $t$ if (i) bilateral variable trade costs are lower; (ii) wages in $i$, and hence production costs, are lower; (iii) the productivity of the most productive firm is higher, again reducing production costs; (iv) competitive pressure, inversely captured by the price index, in $j$ is lower, corresponding to the idea of inward multilateral resistance coined by Anderson2003 in the intensive margin context; (v) the market in $j$ is larger; (vi) bilateral fixed costs of exporting are smaller; or (vii) $i$'s most productive firm already served market $j$ in the previous period and therefore does not have to pay the market entry cost. Note that (i) to (iv) all act via higher operating profits and depend on the elasticity of substitution between goods. The higher this elasticity, the stronger the reaction of profits to changes in any of these factors. At the same time, a higher elasticity reduces the mark-up firms can charge and hence makes it generally harder to earn enough profits to mitigate the fixed costs of exporting. Further note that the importance of the entry costs depends on the discount factor. Intuitively, if agents are more patient, the one-time entry costs matter less compared to the repeatedly earned profits. Empirically, (vii) induces true state dependence. As previous exporters do not have to incur entry costs, they are more likely to stay active in the destination market and the extensive margin becomes more persistent than would be implied merely by the persistence of productivity, market potential, and trade costs.\\

In order to turn equation (ref) into the empirical expression that we will bring to the data, we take the natural logarithm and group all exporter-time and importer-time specific components and capture them with corresponding sets of fixed effects. Further, we need to specify the fixed and variable trade costs. In keeping with the existing literature, we model them as a linear combination of different observable bilateral variables, such as geographic distance, whether $i$ and $j$ are both WTO/GATT members, and whether $i$ and $j$ share a common currency. In our most general specification, we additionally include country pair fixed effects. Following Baier2007, this is common practice in the estimation of the determinants of the intensive margin of trade in order to avoid endogeneity due to unobserved heterogeneity. Further, these bilateral fixed effects may capture (part of) the strong persistence documented above.\footnote{If the trade costs further include any exporter(-time) or importer(-time) specific components, these are captured by the aforementioned corresponding sets of fixed effects.} Note, however, that the nature of the persistence captured by these fixed effects is different from the one that is due to the entry dynamics. This additional state dependence is “spurious” in the sense that countries are not actually more likely to export to a destination because of the prior experience, but because they keep incorporating the same unobserved factors over time. With the three sets of fixed effects and our parametrization for time-varying trade cost determinants, we arrive at the following econometric model:

align[align omitted — 258 chars of source]

where $\kappa=-\sigma\log(\sigma)-(1-\sigma)\log(\sigma-1)-\log(1+\delta (f^{entry}-1))$, $\lambda_{it}=(1-\sigma)(\log(w_{it})-\log(\varphi_{it}^{*}))$, $\psi_{jt}=(\sigma-1)\log(P_{jt})+\log(E_{jt})$, $\beta_{y}=\log(1+\delta (f^{entry}-1))$, $\mathbf{z}_{ijt}'\mbox{\boldmath${\beta}$}_{z}+\mu_{ij}=(1-\sigma)\log(\tau_{ijt})-\log(f_{ijt}^{exp})$, and $\zeta_{ijt}=-\log(\xi_{ijt})\sim \mathcal{N}(0,1)$. The error term distribution implies that a probit estimator is the appropriate choice to estimate our model. Alternatively, we could deviate from Eaton2011 and assume a log-logistic distribution for the idiosyncratic demand shocks, which would lead to a logit specification. As mentioned above, we capture the three sets of unobserved components by introducing according sets of fixed effects. A supposed alternative using random effects is actually not possible, at least for the $it$ and $jt$ effects, as they are implied by the theoretical model. We therefore cannot make the distributional assumptions required in a random effects setting. The theory is silent about the exact form of the bilateral heterogeneity. We decide for a third set of fixed effects as the most general option in order to avoid assumptions on its distribution or its correlation to observed factors. \\

Our theoretical framework implies a flexible empirical specification that can reconcile the extensive margin estimation with the stylized fact presented in Section (ref). Note that we chose to make a number of simplifying assumptions in order to achieve the clear theory-consistent interpretation of specification (ref). An alternative interpretation of equation (ref) as a reduced-from representation of a more elaborate and realistic model Roberts1997 is equally justifiable. At the same time, while our model is written along the lines of Helpman2008, which remains the benchmark for the empirical assessment of the (aggregate) extensive margin of trade, it is not decisive for our empirical specification that zero trade flows result from a truncated productivity distribution instead of a discrete number of firms Eaton2013 or from fixed exporting costs in a Krugman1980-type homogeneous firms setting Egger2011a,Egger2011b.

Binary Choice Estimators with Three-Way Fixed Effects

Having set up the empirical framework, we now turn to the estimation. As equation (ref) demands three-way fixed effects to capture unobservable characteristics, we describe how to implement suitable binary choice estimators. In a first step, we review the standard maximum likelihood estimator (MLE) for probit and logit models with fixed effects. In a second step, we explain the consequences of the incidental parameter problem on the MLE and characterize new bias corrections to address the induced incidental parameter problem.

Model and Maximum Likelihood Estimation

Our empirical model (ref) can be be written in a general way by the following three-way fixed effects binary choice model:

align[align omitted — 270 chars of source]

where $i = 1, \dots, I$, $j = 1, \dots, J$, and $t = 1, \dots, T$ are the indexes denoting exporters, importers, and time periods, respectively, $\mathbf{x}_{ijt}$ is a $p$-dimensional vector of regressors, $\mbox{\boldmath${\beta}$}$ are the corresponding structural parameters, $\boldsymbol{\psi} = (\psi_{11}, \dots, \psi_{JT})$, $\boldsymbol{\lambda} = (\lambda_{11}, \dots, \lambda_{IT})$, and $\boldsymbol{\mu} = (\mu_{11}, \dots, \mu_{IJ})$ are the incidental parameters, $\mathbf{x}_{ij}^t = (\mathbf{x}_{ijt}, \mathbf{x}_{ij(t-1)}, \dots, \mathbf{x}_{ij1})$, and $F_{\zeta}$ is either the logistic or standard normal cumulative distribution function.\footnote{To keep notation simple, we abstract from the cases of no-self flows, $i \neq j$, and unbalanced panels. The no-self flows are unproblematic for our estimators. For unbalanced panels, we have to additionally assume, that the attrition process is random conditional on $\mathbf{x}_{ij}^t$ and the fixed effects.} Note that we allow $\mathbf{x}_{ijt}$ to contain weakly exogenous regressors. This is important, because our empirical model (ref) contains a lagged dependent variable, which violates the strict exogeneity condition, i.e. $\zeta_{ijt} \mid \mathbf{x}_{ij}^T, \boldsymbol{\psi}, \boldsymbol{\lambda}, \boldsymbol{\mu} \sim F_{\zeta}$ with $\mathbf{x}_{ij}^T = (\mathbf{x}_{ijT}, \mathbf{x}_{ij(T-1)}, \dots, \mathbf{x}_{ij1})$. Even in the absence of a lagged dependent variable, the strict exogeneity condition is often too restrictive, since it rules out any dynamic feedback between regressors and the dependent variable.\footnote{For instance, one can imagine that country pairs that do not trade with each other in period $t$, might sign a trade agreement in the future, because they did not trade in $t$.}\\

A standard estimator for the parameters of interest $\mbox{\boldmath${\beta}$}$ and the incidental parameters $\boldsymbol{\alpha} = (\boldsymbol{\lambda}, \boldsymbol{\mu}, \boldsymbol{\psi})$ is the following maximum likelihood estimator (MLE)

equation[equation omitted — 360 chars of source]

where $\ell_{ijt}(\boldsymbol{\beta}, \psi_{jt}, \lambda_{it}, \mu_{ij} ) = y_{ijt} \log(F_{ijt}) + (1-y_{ijt}) \log(1-F_{ijt})$ is the log-likelihood contribution of the $ijt$-th observation, and $F_{ijt}$ denotes the cumulative distribution function chosen for $\zeta_{ijt}$ evaluated at the linear index $\eta_{ijt} = \mathbf{x}_{ijt}'\mbox{\boldmath${\beta}$}+ \psi_{jt}+\lambda_{it}+\mu_{ij}$.\footnote{Note that the incidental parameters $\mu_{ij}, \lambda_{it}, \psi_{jt}$ are not uniquely identified due to collinearity issues and we need to impose some restrictions. This can be achieved by adding a penalty on the log-likelihood which imposes a normalization on the incidental parameters.}\\

The brute-force estimation of equation (ref) quickly becomes computationally demanding, if not impossible due to the large number of parameters that need to be estimated. For example, in a balanced data set ($I=J=N$) without self flows, the number of parameters to be estimated is $\approx N (N-1) + 2 N T$. In a trade panel data set with 200 countries and 50 years, the number of fixed effects in this case amounts to 59,800 parameters. However, recent developments in computational econometrics reduce this computational burden by employing a straightforward strategy called pseudo-demeaning, which mimics the well-known within transformation for linear regression models Stammann2018. We give a brief summary of the pseudo-demeaning algorithm in Appendix (ref). \\

The parameters of the variables of interest can only be directly interpreted in terms of their signs and relative magnitudes, but are not informative in themselves about the absolute strength of the effects they represent. Another quantity of interest are therefore the average partial effects:

align[align omitted — 132 chars of source]

where the partial effect of the $k$-th regressor $\Delta_{ijt}^k$ is either $\Delta_{ijt}^k = \partial F_{ijt} / \partial x_{ijtk}$ in the case of a non-binary regressor or $\Delta_{ijt}^k = F_{ijt}|_{x_{ijtk = 1}} - F_{ijt}|_{x_{ijtk = 0}}$ in the case of binary regressors. Here $F_{ijt}|_{x_{ijtk = z}}$ indicates that all values of the $k$-th regressor are replaced by $z$. In dynamic models, the simple average partial effect $\delta_k$ does not provide the full picture of how the export probability is affected by a change in a regressor. Rather, there are additional feedback effects.\footnote{In our context, the introduction of a permanent trade policy that increases the probability to export to a destination implies that in the next period, entry costs are more likely to have already been paid, and hence the impact becomes higher with increasing duration of the policy.} To derive expressions for long-run effects, we make use of the long-run probability of $y_{ijt}=1$ for a given set of regressors and fixed effects, also mentioned in Carro2003 and Browning2010:

equation[equation omitted — 90 chars of source]

where $\Delta_{ijt}^y = F_{ijt}|_{y_{ij(t-1) = 1}} - F_{ijt}|_{y_{ij(t-1) = 0}}$. Long-run average partial effects are then given by

equation[equation omitted — 327 chars of source]

in the case of non-binary or binary regressors, respectively. Estimators of ((ref)) and ((ref)) can be formed by plugging in the MLE defined in ((ref)).

Incidental Parameter Bias Correction

The MLE of many nonlinear fixed effects models --- including binary choice models --- suffers from the well-known incidental parameter problem (IPP) first identified by Neyman1948. The problem stems from the necessity to estimate many nuisance parameters, which contaminate the estimator of the structural parameters and average partial effects. It can be further amplified by the inclusion of predetermined regressors like a lagged dependent variable. This amplification can be interpreted as a Nickell-type bias Nickell1981. Although the MLE is consistent (under appropriate asymptotics), it has an asymptotic bias in its limiting distribution leading to invalid inference Fernandez-Val2018. The literature suggests different types of bias corrections to reduce these incidental parameter and Nickell-type biases Fernandez-Val2018. Jackknife corrections, like the leave-one-out jackknife proposed by Hahn2004, or the split-panel jackknife (SPJ) introduced by Dhaene2013, are the simplest approaches to obtain a bias correction, at the expense of being computationally costly. In contrast to analytical corrections, their application only requires knowledge of the order of the bias components to form appropriate subpanels that are used to reestimate the model and to form an estimator of the bias terms. For analytical bias correction (ABC), it is necessary to derive the asymptotic distribution of the MLE, in order to obtain an explicit expression of the asymptotic bias. This is then used to form a suitable estimator for the bias terms. \\

The IPP affects all quantities of interest mentioned in the previous subsection, i.e. $\hat{\boldsymbol{\beta}}$, $\hat{\delta}_k$, and $\hat{\delta}_k^{LR}$.\footnote{Note that we do not make a notational distinction between direct and long-run APEs in the following. In all expressions for APEs, one can move between the direct and long-run case by substituting $\Delta_{ijt}^{LR}$ for $\Delta_{ijt}$.} Its consequences become clear when looking at the asymptotic distribution of the estimators. Under asymptotic sequences, where all panel dimensions grow at a constant rate as $I, J, T \rightarrow \infty$, the MLE is consistent but asymptotically biased because its asymptotic distribution is not centered around the true parameter value $\boldsymbol{\beta}^0$:

equation*[equation* omitted — 215 chars of source]

where $\overline{\mathbf{b}}_{ \infty}^{\beta}$ denotes the asymptotic bias and $\overline{\mathbf{V}}_{\infty}^{\beta}$ is the covariance matrix Fernandez-Val2018. Fernandez-Val2018 derive a simple heuristic to determine the order of the bias induced by the incidental parameters: $bias \sim p / n$, where $p$ is the number of the incidental parameters and $n$ is the sample size. Based on this heuristic Fernandez-Val2018 conjecture that the bias of the three-way fixed effects estimator we consider in this paper is of order $(IT+JT+IJ)/(IJT)$ and of the form $B_1 / I + B_2 / J + B_3 / T$. In line with the bias structure in two-way error component models, the inclusion of importer-time and exporter-time fixed effects entails two bias terms of order $1 / I$ and $1 / J$, respectively. Intuitively, the inclusion of dyadic fixed effects induces another bias of order $1/T$ because there are only $T$ informative observations per additionally included parameter. Although the bias decreases with increasing $I,J,T$, its order is larger than the order of the standard deviation of the MLE, $1/\sqrt{IJT}$.\footnote{As mentioned by Fernandez-Val2018, even if $I = J = T$, the order of the bias is $1/I$ and the order of the standard deviation is $\sqrt{1/I^3}$.} As a consequence, confidence intervals do not have the desired nominal confidence level, i.e. they undercover, leading to invalid inference Fernandez-Val2018. The asymptotic distribution of the APEs is affected similarly by the IPP:

equation*[equation* omitted — 183 chars of source]

where $\overline{b}_{k, \infty}^{\delta}$ denotes the asymptotic bias and $\overline{V}_{k, \infty}^{\delta}$ is the variance. The conjecture of Fernandez-Val2018 does not incorporate explicit expressions for bias-corrected estimators. Thus, based on their conjecture, we propose novel analytical and jackknife bias corrections for three-way fixed effects models, which deal with the IPP (including the Nickell-type bias) and the associated inference problem.\footnote{In Appendix (ref), we formulate the asymptotic distributions of $\hat{\boldsymbol{\beta}}$ and $\hat{\delta}_k$. In Appendix (ref), we illustrate the statistical problem and the working of bias corrections with a version of the prominent Neyman1948 variance example.} \\

In the following we adapt and extend the analytical and split-panel jackknife bias corrections proposed by Fernandez-Val2016a in the context of nonlinear models with individual and time fixed effects to our three-way error component structure.\footnote{In Appendix (ref), we also derive the bias corrections for a two-way fixed effects model in our $ijt$ network panel structure. Previous two-way bias corrections considered either classical $it$ panel structures or $ij$ pseudo-panels.}$^{,}$\footnote{We do not elaborate on the leave-one-out jackknife bias correction because it requires all variables to be independent over time and thus rules out predetermined and serially-correlated regressors Fernandez-Val2018.}

For the split-panel jackknife bias correction, the aforementioned three-part bias structure implies that we need to split our panel across three dimensions, leading to the following estimator for the structural parameters:

align[align omitted — 983 chars of source]

where $\lfloor \cdot \rfloor$ and $\lceil \cdot \rceil$ denote the floor and ceiling functions. To clarify the notation, the subscript ${\{i:i\leq \lceil I/2 \rceil \},J,T}$ denotes that the estimator is based on a subsample, which contains all importers and time periods, but only the first half of all exporters. Note that similar to Fernandez-Val2016a, the split-panel jackknife bias correction requires a homogeneity assumption of the distribution of $y_{ijt}$ and $\mathbf{x}_{ijt}$ across the dimensions $I$, $J$, and $T$. For instance, this assumption is violated in the case of time trends or structural breaks, where the subsample estimates of splitting dimension $T$, i.e. $\widehat{\mbox{\boldmath${\beta}$}}_{\{I,J,t:t\leq \lfloor T/2 \rfloor \}}$ and $\widehat{\mbox{\boldmath${\beta}$}}_{\{I,J,t:t \geq \lceil T/2 + 1\rceil \}}$, are systematically different Fernandez-Val2016a.\footnote{The homogeneity assumption can be tested with a Wald test Dhaene2013, Fernandez-Val2016a.} This additional homogeneity assumption is not required for the analytical bias correction. \\

table[table omitted — 1,023 chars of source]

To characterize the analytical bias correction, we need to introduce some additional notation. Let $\partial_{\iota^r}g(\cdot)$ denote the $r$-th order partial derivative of an arbitrary function $g(\cdot)$ with respect to some parameter $\iota$. A collection of further required expressions are reported in Table (ref). Let $\mathbf{D}$ denote the dummy matrix corresponding to the fixed effects and $\mathbf{X}$ denote the matrix with the regressors of interest. Define the residual projection $\widehat{\operatorname{\mathbb{M}}} = \mathbf{I}_{IJT} - \widehat{\operatorname{\mathbb{P}}} = \mathbf{I}_{IJT} - \mathbf{D}(\mathbf{D}^{\prime} \widehat{\boldsymbol{\Omega}} \mathbf{D})^{-1} \mathbf{D}^{\prime} \widehat{\boldsymbol{\Omega}}$, where $\mathbf{I}_{IJT}$ is an ${IJT} \times {IJT}$ identity matrix, and $\widehat{\boldsymbol{\Omega}}$ is a diagonal weighting matrix with elements $\hat{\omega}_{ijt}=\partial_{\eta}\hat{F}_{ijt}$.\\

Combining insights from the classical panel structure in Fernandez-Val2016a, the pseudo-panel setting in Cruz-Gonzalez2017, and the three-way fixed effects conjecture by Fernandez-Val2018, we propose the following analytical bias correction:

align[align omitted — 1,845 chars of source]

The first two correction terms in equation (ref) are generalizations of the corresponding components in the $ij$ pseudo-panel structure of Cruz-Gonzalez2017 to our $ijt$ structure. The inclusion of a third set of ($ij$) fixed effects additionally leads to the third correction term that mimics the correction for individual fixed effects in an $it$-panel setting. In contrast to $\widehat{\mathbf{B}}_1$ and $\widehat{\mathbf{B}}_2$, the expression $\widehat{\mathbf{B}}_3$ includes a part to correct the Nickell-type bias that arises by imposing the weak exogeneity condition instead of the strict exogeneity condition. The parameter $L$ is a bandwidth used for the estimation of truncated spectral densities Hahn2007. In a model in which all regressors are strictly exogenous, $L$ is set to zero, such that the second part in the numerator of $\widehat{\mathbf{B}}_3$ vanishes. Note that the strict exogeneity condition is often too strong in the context of panel data, because it rules out any dynamic feedback between regressors and the dependent variable. In case of weakly exogeneous regressors, for instance if one of the regressors is the lagged dependent variable, Fernandez-Val2016a suggest conducting a sensitivity analysis with $L \in \{1,2,3,4\}$.\\

Moving to the APEs, the split-panel jackknife estimator is formed by replacing the estimators for the structural parameters with estimators for the APEs in formula (ref). The analytically bias-corrected estimator is given by

align[align omitted — 1,642 chars of source]

where $\widehat{\boldsymbol{\Psi}}_{ijt} = \partial_{\eta} \widehat{\boldsymbol{\Delta}}_{ijt}/\hat{\omega}_{ijt}$. The last part in the numerator of $\widehat{\mathbf{B}}_3^{\delta}$ is again dropped if all regressors are assumed to be strictly exogenous. Note that all quantities are evaluated at bias-corrected structural parameters and the corresponding estimates of the fixed effects.\footnote{For this purpose, we use a computationally efficient offset algorithm as in Czarnowske2019.} An appropriate covariance estimator for the APEs of the three-way fixed effects model is

align[align omitted — 733 chars of source]

where $\widehat{\bar{\boldsymbol{\Delta}}}_{ijt} = \widehat{\boldsymbol{\Delta}}_{ijt} - \hat{\boldsymbol{\delta}}$, $\widehat{\boldsymbol{\Delta}}_{ijt} = [\widehat{\Delta}_{ijt}^1, \dots, \widehat{\Delta}_{ijt}^m]^{\prime}$, $\hat{\boldsymbol{\delta}} = [\hat{\delta}_1, \dots, \hat{\delta}_m]^{\prime}$, and

align*[align* omitted — 567 chars of source]

Note that the term $v_2$ refers to the variation induced by $\hat{\boldsymbol{\theta}}$. The terms $v_1$ and $v_3$ are in the spirit of Fernandez-Val2016a to improve the finite sample properties. These are, on the one hand, the variation induced by replacing population by sample means ($v_1$). On the other hand, if we are concerned about the strict exogeneity assumption (as we are in the case of dynamic three-way error structure models), the covariance between $v_1$ and $v_2$ should be incorporated ($v_3$).\\

Just as the calculation of the MLE, the application of the bias corrections is computationally demanding due to the high number of parameters. Therefore, we suggest to use the algorithms developed by Stammann2018 and Czarnowske2019.

Monte Carlo Simulations

In this section, we conduct extensive simulation experiments to investigate the properties of different estimators for both the structural parameters and the APEs. The estimators we study are MLE, ABC and SPJ.\footnote{We do not include OLS as an alternative estimator for APEs in our discussion since --- apart from a Nickell-bias in a dynamic three-way fixed effects model --- it suffers from a kind of misspecification bias. In a standard gaussian data generating process, where the dependent variable is continuous instead of binary, the bias corrected ordinary least squares estimator proposed by Hahn2002 works as desired, however it fails in our probit data generating process. The reason is that a large share of the predicted probabilities obtained by OLS exceed the boundaries of the unit interval, such that estimates of APEs become heavily biased. This kind of misspecification bias already becomes evident, if we consider a dynamic two-way fixed effects data generating process, which does not suffer from a Nickell bias. We investigated this issue in additional extensive simulation experiments and provide these results on request.} Our main focus are the biases and inference accuracies. To this end, we compute the relative bias and standard deviation (SD) in percent, the ratio between bias and standard error, the ratio between standard error and standard deviation (SE/SD), and the coverage probabilities (CPs) at a nominal level of 95 percent.\\

For the simulation experiments we adapt the design for a dynamic probit model of Fernandez-Val2016a to our $ijt$-panel structure with three-way fixed effects.\footnote{Further simulation experiments including dynamic panel models with two-way fixed effects and static panel models with three-way fixed effects are presented in Appendices (ref) and (ref). In an earlier version of this article we additionally report simulations results for static two-way fixed effects models.} In line with our theoretical model, the simulations include unobserved components captured by fixed effects in the $it$, $jt$, and $ij$ dimensions, as well as the lagged dependent variable. Specifically, we generate data according to

align*[align* omitted — 313 chars of source]

where $i=1,\ldots, N$, $j = 1, \dots, N$, $t = 1, \ldots, T$, $\beta_{y} = 0.5$, $\beta_{z} = 1$, $\psi_{jt} \sim \mbox{$\mathrm{iid}.\ $} \mbox{$\mathcal{N}$}(0, 1 / 24)$, $\lambda_{it} \sim \mbox{$\mathrm{iid}.\ $} \mbox{$\mathcal{N}$}(0, 1 / 24)$, $\mu_{ij} \sim \mbox{$\mathrm{iid}.\ $} \mbox{$\mathcal{N}$}(0, 1 / 24)$, and $\epsilon_{ijt} \sim \mbox{$\mathrm{iid}.\ $} \mbox{$\mathcal{N}$}(0, 1)$.\footnote{We again follow Fernandez-Val2016a and incorporate the information that $\{\lambda_{it}\}_{IT}$, $\{\psi_{jt}\}_{JT}$, and $\{\mu_{ij}\}_{IJ}$ are independent sequences, and $\lambda_{it}$, $\psi_{jt}$, and $\mu_{ij}$ are independent for all $it$, $jt$, $ij$ in the covariance estimator for the APEs. The explicit expression is provided in Appendix (ref).} The exogenous regressor is modeled as an AR-1 process, $z_{ijt} = 0.5z_{ijt-1} + \psi_{jt} + \lambda_{it} + \mu_{ij} + \nu_{ijt}$, where $\nu_{ijt} \sim \mbox{$\mathrm{iid}.\ $} \mbox{$\mathcal{N}$}(0, 0.5)$ and $z_{ij0} \sim \mbox{$\mathrm{iid}.\ $} \mbox{$\mathcal{N}$}(0, 1)$. We consider different sample sizes, specifically $N \in \{50, 100, 150\}$ and $T \in \{10, 20, 30, 40 , 50\}$ and generate 1,000 data sets for each.\\

Tables (ref) -- (ref) in Appendix (ref) summarize the extensive simulation results. For ABC we report two different choices of the bandwidth parameter, $L=1$ and $L=2$, which are indicated by values in parentheses. In the following, we discuss the biases and coverage probabilities for $N \in \{50, 150\}$ which are shown in Figures (ref) -- (ref) for the sake of clarity. We focus on these two statistics in particular because they give us a good idea about the quality of the inference of an estimator. \\

figure[figure omitted — 344 chars of source]
figure[figure omitted — 196 chars of source]

We start by considering the different estimators for the structural parameters summarized in Figures (ref) and (ref). For both kinds of regressors, MLE exhibits a severe bias that decreases with increasing $T$. However, even with $N=150$ and $T=50$, the estimator shows a distortion of 11 percent in the case of the predetermined regressor and 5 percent in the case of the exogenous regressor. We also find that the inference is not valid, since the CPs are zero or close to zero. The bias corrections bring a substantial improvement. First, they reduce the bias considerably. For example, the MLE estimator of the predetermined regressor shows a distortion of 64 percent for $T=10$ and $N=150$. ABC reduces the bias to 8 percent and SPJ to 20 percent. In the case of the exogenous regressor, MLE exhibits a bias of 23 percent, whereas ABC has a bias of 1 percent and SPJ of 7 percent. Irrespective of the type of the regressor, both bias-corrected estimators also converge quickly to the true parameter value with growing $T$. Second, the bias corrections improve the CPs. For the exogenous regressor the CPs of ABC are close to the desired level of 95 percent for all $T$, whereas SPJ remains far away from 95 percent even at $T = 50$. In the case of the predetermined regressor, the CPs of both corrections approach the nominal level when $T$ rises. This happens faster for ABC.\\

We proceed with the APEs, where we distinguish between direct and long-run APEs. We first look at the direct APEs in Figures (ref) and (ref). Overall, we obtain similar findings as for the structural parameters. MLE is distorted over all settings, but the bias decreases as $T$ increases. The distortion is especially severe in the case of the the predetermined regressor, where it ranges from roughly $70$ percent for the settings with $T=10$ to $15$ percent for $T=50$. In contrast, the bias of the MLE for the exogenous regressor ranges from roughly $4$ percent ($T=10$) to $1$ percent ($T=50$). The bias corrections bring a substantial reduction in both cases. Whereas ABC shows only a small distortion of less than $1$ percent in the case of the exogenous regressor at $T=10$, SPJ is even more heavily distorted than MLE. However, with increasing $T$, both SPJ and ABC quickly converge to the true APE. Looking at inference, unlike ABC, SPJ needs a sufficiently large number of time periods to get its CPs close to 95 percent. For the predetermined regressor, these convergence processes last longer than for the exogenous regressor. In contrast, the coverage probabilities of the MLE are always far below the desired nominal level. This is especially noteworthy for the exogenous regressor with $T \geq 40$, where the bias of MLE appears to be negligibly small for practical work. However, the bias of the MLE is still large compared to the standard error. This becomes evident, when looking at the Bias/SE ratio in Tables (ref) -- (ref), which is much higher for MLE than for the bias corrected estimators. Turning to the results of the long-run APEs reported in Figure (ref), we note that the results are qualitatively similar to the results of the direct APE of the exogenous regressor. As $T$ increases, both bias corrections quickly bring the bias towards zero and the coverage probabilities quickly reach the nominal level. Again, the convergence of ABC is faster and inference of MLE remains invalid despite small biases in settings with larger $T$. \\

Overall, our three-way fixed effects simulation results confirm the conjecture of Fernandez-Val2018 about the general form of the bias and lend support to our bias corrections. First, we find that the bias corrections indeed substantially mitigate the bias. Second, as already found in other studies, analytical bias corrections outperform split-panel jackknife bias corrections (see among others Fernandez-Val2016a, and Czarnowske2019). For samples with shorter time horizons, ABC is often less distorted and its dispersion is generally lower. This is also reflected by better CPs. Generally, in the three-way fixed effects setting, a sufficiently large number of time periods appears to be crucial to obtain reliable results for the bias-corrected estimators. As a main takeaway, our simulation results suggest that estimates based on MLE should be treated with great caution, because even in situations where we expect small biases, the inference may be invalid.

Determinants of the Extensive Margin of Trade

Having described the estimation and bias correction procedures, we now turn to the estimation of the determinants of the extensive margin of international trade outlined in Section (ref).\\

Recall equation (ref) that relates the incidence of non-zero aggregate trade flows to exporter-time and importer-time specific characteristics, trade in the previous period, time-invariant unobservable trade barriers and bilateral trade policy variables:

align*[align* omitted — 250 chars of source]

This yields the following dynamic three-way fixed effects probit model:

align[align omitted — 272 chars of source]

$y_{ij(t-1)}$ is the lagged dependent variable, $\mathbf{z}$ is a vector of observable bilateral variables, and $\beta_{y}$ and $\beta_{z}$ are the corresponding parameters. We largely follow Helpman2008 and the wider literature on the determinants of the intensive margin of trade Head2014 in the choice of these variables: distance, a common land border, the same origin of the legal system, common language, previous colonial ties, a joint currency, an existing free trade agreement, or joint membership in the WTO/GATT. The effect of all time-invariant variables will only be identified in specifications in which we omit the bilateral fixed effects.\\

For data, we turn to the updated Gravity dataset provided by CEPII Conte2021, which encompasses annual information on bilateral trade flows and these variables of interest for 198 countries from 1948 -- 2019. As only non-zero flows are reported in the data, we construct zeros analogous to Head2010: Whenever an exporter reports at least one non-zero flow and an importer reports at least one non-zero flow, their bilateral flow --- if not non-zero already --- is coded as zero. This procedure results in an unbalanced data set with 1,652,296 observations.\footnote{The updated Gravity dataset by Conte2021 supersedes the gravity dataset provided alongside Head2010, which provided similar data from 1948 -- 2006. In the latter, some non-zero flows are coded as zero, in order to correct for implausible values reported in the original IMF DOTS data Head2010. } \\

Out of the 38,912 country pairs included in our analysis, 9,984 never switch their trading status, 4,609 switch a single time, and the average (median) number of switches is 3.9 (3), ranging up to 29 for trade from Benin to Sweden and 32 from Barbados to Bolivia.\footnote{See Figure (ref) in Appendix (ref) for the full distribution.}

Main Results

Before turning to the regression results, we repeat the descriptive analysis about the persistence of the bilateral trade flows from Section 1, now considering the transition probabilities into export from period $t-1$ to $t$ for the time horizon from 1948 -- 2019. Table (ref) confirms the high level of persistence: 89.7 percent of the country pairs which did not trade in the previous year did not trade in the following year and 91.5 percent of the pairs which did trade in the previous year continued to trade in the following year. Thus, the probability to export in period $t$ is 81.2 percentage points higher for those country pairs that already engaged in trade in period $t-1$ compared to those that did not.\footnote{The number is computed as the difference between the probability of exporting in period $t$ conditional on exporting and not exporting in period $t-1$.} However, Table (ref) does not reveal any information about the kind of persistence. In the following analysis we will investigate the importance of using dynamic model specifications which allow us to disentangle the observed persistence into two sources: (i) true state dependence and (ii) observed and unobserved heterogeneity.\\

table[table omitted — 317 chars of source]

In the following analysis we only consider analytically bias-corrected estimates, mainly for two reasons: (1) the simulation results demonstrate that ABC performs much better than SPJ; (2) SPJ requires the additional homogeneity assumption, which is likely to be violated in our application. Progressing economic integration over time --- be it through an increasing number of WTO/GATT memberships or through bilateral and multilateral integration in the form of currency unions and free trade agreements --- is at odds with the required stationary pattern.\footnote{See the corresponding homogeneity tests for model specifications (2) -- (5) of table (ref) in table (ref) in Appendix (ref). We find evidence that the homogeneity assumption is violated across all splitting dimensions.}\\

Table (ref) reports average partial effects of several static and dynamic fixed effects probit specifications.\footnote{Coefficient estimates are reported in Table (ref) in the Appendix.} Bias-corrected estimates and their corresponding standard errors are printed in bold. For comparison, the uncorrected estimates are also shown. In column (1) we first mimic the static specification estimated by Helpman2008.\footnote{Helpman2008 use a dataset that ranges from 1970 to 1997. They also include dummy variables for whether both countries are landlocked or islands, or follow the same religion. Hence our estimates deviate somewhat from theirs, while remaining qualitatively similar.} Their specification includes exporter, importer, and time fixed effects.\footnote{Note that following Fernandez-Val2018 the incidental bias problem is small enough to ignore in this setting with $i$, $j$ and $t$ fixed effects, since the order of the bias is $1/IT + 1/JT + 1/IJ$, which in our case becomes negligible small since $I$, $J$ and $T$ are large.} All average partial effects have the expected sign, indicating a negative impact of distance on the probability to trade, while having a common border, the same origin of the legal system, a shared language, or a joint colonial history are all estimated to have a positive impact. Also note the strong and highly significant impact of a common currency, free trade agreement or joint membership of the WTO/GATT. Ceteris paribus, each is estimated to increase the probability of non-zero flows by between 5.2 and 7.1 percentage points. \\

table[table omitted — 4,388 chars of source]

Column (2) introduces a stricter set of fixed effects, namely at the exporter-time and importer-time level. This specification can be considered a theory-consistent estimation of the model by HMR, and of our model if entry costs are zero and other bilateral trade cost determinants are fully observable. The average partial effects are qualitatively the same and quantitatively similar for most variables to those in column (1). However, e.g.\ the estimated effects of contiguity and joint WTO/GATT membership are reduced by more than a third, while the estimated effect of sharing a common currency increases by 2.6 percentage points or 37 percent. \\

Specification (3) keeps the same fixed effects, but adds a lagged dependent variable and thus controls for one type of persistence. Assuming no unobservable bilateral heterogeneity, this specification correctly estimates the model set up in Section (ref). As the partial effects of static and dynamic models are not directly comparable due to the feedbacks involved in the latter, we show two types of average partial effects for our dynamic specifications: the usual direct effects and the long-run effects described in Section (ref). The first result to note for the third specification is the highly significant average partial effect for the lagged dependent variable, which reflects the strong impact of previous non-zero trade flows on current ones. Ceteris paribus, the average partial effect shows a 39.4 percentage points higher probability of non-zero trade, given the two countries were also engaged in trade in the previous year. This implies that 48.5 percent of the observed persistence are attributed to true state dependence in this first dynamic specification.\footnote{This value is calculated as the ratio of the estimated average partial effect and the unconditional effect of trade in the previous period on the exporting probability ($39.4 / 81.2$).} In terms of our model, this suggests a vast effect of market entry costs on the aggregate extensive margin. The second observation is that direct APEs are about 50 percent smaller than in column (2) across the board. However, once dynamic adjustments are taken into account, the average partial effects resulting from specifications (2) and (3) become very similar, suggesting that accounting for the market entry dynamics mainly matters for getting the timing of trade policy effects right, rather than for the overall magnitude of the effects. \%Note, however, that the APEs of the dynamic specification represent only the immediate effects, while the long-run effects should be larger due to the positive state dependence.

Specification (4) takes one step back and one forward. While not including the lagged dependent variable in the estimation, it introduces a bilateral fixed effect that controls for a second type of persistence --- bilateral unobserved heterogeneity.\footnote{Note that we do not want to make the assumption of strict exogeneity in this specification, but instead we allow the regressors to be weakly exogenous. For instance, strict exogeneity is too restrictive if we want to allow that country pairs sign a trade agreement in the future, because they did not trade with each other in period $t$.} This also follows the important insight by Baier2007, who show that controlling for unobserved bilateral heterogeneity produces a considerably different estimated impact of free trade agreements, among other variables, on the intensive margin of trade. While now an identification of many of the variables of interest is no longer possible because of their time invariance, this specification reveals a much reduced estimated impact of the time-varying variables. The impact of a common currency on the probability of exporting is reduced to 3.4 percentage points, while those of a common free trade agreement and WTO/GATT are decreased to 1.2 and 1.3 percentage points, respectively. These results highlight the importance of controlling for unobserved country pair heterogeneity to avoid endogeneity problems associated with trade policy variables. \\

Finally, in the last two columns we present the results from our preferred specification (5), which implements equation (ref). The estimation again includes the “full set” of fixed effects, i.e.\ exporter-time, importer-time and bilateral fixed effects, now combined with the lagged dependent variable, and therefore controls for both kinds of persistence simultaneously. Again, the average partial effect on the lagged dependent variable is highly significant. It now entails a partial effect of about 18.7 percentage points, i.e.\ roughly 23 percent of the observed persistence can be attributed to true state dependence and 77 percent to other observed factors and unobserved heterogeneity. Failure to account for unobserved heterogeneity in specification (3) hence overestimated the importance of the lagged dependent variable (corresponding to entry costs in our model) roughly by a factor of two and therefore mislabelled a substantial part of spurious as true state dependence. Considering the effects of the time-varying trade policy variables, small but statistically significant direct average partial effects in column (5) are estimated for a common currency at 2.2 percentage points and for both joint FTA and WTO memberships at 0.8 percentage points. Just as in the comparison between specifications (2) and (3), the average partial effects again become very similar when the static effects are compared to the long-run effects in the dynamic specification. \\

When comparing bias-corrected and uncorrected average partial effects throughout Table (ref), it is noticeable that both differ only slightly for the exogenous regressors within the different specifications. The most significant impact is observed on the average partial effect for the predetermined variable, which in specification (5) differs by more than 26 percent. These results are in line with the theoretical properties of the estimators and the findings of our simulation study.\footnote{For the two-way models in columns (2) - (3) the order of the bias is $1/I + 1/J$, i.e.\ the bias only depends on the number of exporters/importers. For the three-way models in columns (4) - (5) the order of the bias is $1/I + 1/J + 1/T$, i.e.\ the bias additionally depends on the number of time periods. In our application the number of exporters/importers is relatively large but the number of time periods is substantially lower.} Despite the small biases for the exogenous variables, a bias correction is still necessary for our three-way fixed effects specifications because the biases are not negligible relative to the standard errors and thus inference of the uncorrected estimator is invalid. Note that for applications with a shorter time horizon, the biases will be more evident.

Predictive Analysis

After we have seen that the presented innovations matter significantly for the estimation of extensive margin determinants, we now consider the predictive performance of the different specifications discussed above. As an additional benchmark, we also look at a naive --- purely descriptive --- approach that predicts export decisions in period $t$ solely based on the export decision in period $t-1$. For the resulting altogether six options, we evaluate the predictive power using the following measures: the total share of correctly predicted export decisions (accuracy), the share of correctly predicted decisions conditional on not exporting (true negative rate), and the share of correctly predicted decisions conditional on exporting (true positive rate). As becomes clear in Table (ref), the naive approach already works very well for predictive purposes. Close to 90 percent of the exporting decisions in period $t$ are correctly predicted by simply reproducing exporting decisions from $t-1$, irrespective of the measure we are considering. As it turns out, both the static model with $i$, $j$, and $t$ fixed effects (specification (1)) and the static model with $it$ and $jt$ fixed effects (specification (2)) produce poorer predictions than the naive approach --- particularly clearly so by about six percentage points in the former case. However, this changes once persistence is explicitly taken into account in specifications (3) to (5). All of them at least slightly improve the predictions of the naive approach. Whether persistence is incorporated only via the lagged dependent variable or via pair fixed effects turns out to yield very similar predictive quality. Combining both in our preferred specification --- the dynamic three-way fixed effects models --- yields the highest predictive power. We are able to correctly predict 92.6 percent of the export decisions, 91.4 percent of the decisions conditional on not exporting, and 93.5 percent of the decisions conditional on exporting, implying an improvement compared to the naive approach by around 2 percentage points. All in all, the predictive analysis underlines the importance of controlling for true state dependence and unobserved bilateral factors simultaneously.

table[table omitted — 660 chars of source]

Robustness Checks

table[table omitted — 4,381 chars of source]

We next consider two robustness checks for our main results from Section (ref). First, while our preferred estimation of the extensive margin of trade with a dynamic three-way fixed effects binary choice estimator follows from the stylized facts and our theoretical model, the decision on which binary choice estimator to use hinges on the distributional assumption for the error term (and hence for the demand shock in our model). We followed Eaton2011 and assumed log-normal shocks, leading us to using a probit for our main estimations. We now consider log-logistic shocks and the resulting logit estimator instead. Table (ref) displays the results for the same specifications as in Table (ref), but estimated using a logit. Reassuringly, the average partial effects are very similar to the probit case for all variables in all specifications. Introducing the different innovations step-by-step changes the estimated effects in the same way as for the probit. The insights discussed above therefore do not hinge on the specific distributional assumption made in the parametrization of the probability of success.\\

Second, in the estimation of our preferred specification (5) we have discretionary power in one dimension for the exact form of bias correction used. Specifically, the bandwidth parameter $L$ used to estimate the spectral densities has to be chosen. We therefore follow the recommendation by Fernandez-Val2016a and investigate the sensitivity of the results with respect to $L$. Table (ref) depicts the direct and long-run average partial effects obtained with the bias-corrected dynamic three-way fixed effects probit estimator for $L \in \{1,2,3,4\}$. Again, our results turn out to be very robust. There only appears to be a slight upward trend in the estimated true state dependence for larger $L$.\footnote{In Table (ref) in Appendix (ref), we show the same sensitivity check for the dynamic three-way logit estimates and find the same robust pattern.}

table[table omitted — 1,698 chars of source]

Conclusion

In this paper we reexamine the determinants of the extensive margin of international trade. We set up a model that exhibits a dynamic component and allows for time-invariant unobserved bilateral trade cost factors, generating persistence --- a feature in the data that has so far been given little attention. For the estimation of our model, we propose new bias-corrected fixed effects binary choice estimators to account for the incidental parameters problem and to deal with weakly exogenous regressors. Finally, we show that our estimates of the determinants of the extensive margin of trade differ significantly from previous ones. This highlights the importance of true state dependence and unobserved heterogeneity and therefore strongly supports the use of our bias-corrected dynamic fixed effects estimator.\\

The extensive margin of trade obviously extends beyond the aggregate level, warranting further research at lower levels of aggregation, in particular in the context of firms. While our model's prediction and its empirical specification rely on some abstractions, it provides a very tractable and flexible framework that can be estimated with recently established estimation procedures, when combined with the bias correction technique we introduce.\\

A further natural extension for future research would be to define the extensive margin as the number of exporting sectors leading to a fractional instead of a binary outcome SantosSilva2014,Anderson2020. Following the insights for such conditional moment models by Fernandez-Val2016a, our split-panel jackknife estimator can be directly applied and the analytical bias correction has to only be slightly adjusted for the fractional outcome setting.

{ \printbibliography }

\setcounter{table}{0} \setcounter{figure}{0} \setcounter{equation}{0}