Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
46,261 characters · 12 sections · 55 citation commands
Uniform Inference in Linear Error-in-Variables Models: Divide-and-Conquer
To account for measurement error in independent variables, higher-order moment estimators have been used to analyse the relation between R&D expenditures and patent applications (lewbel1997constructing), to test the $q$-theory of investment in finance (erickson2000measurement), and to investigate firm saving behaviour (riddick2009corporate). Underlying these estimators are two crucial assumptions: first, there needs to be sufficient skewness in the latent regressor. Second, the coefficient $\beta_{0}$ that relates the latent regressor to the observed outcome cannot equal zero. In this paper we focus on the situation where this second assumption potentially fails. In particular, we consider the standard third-order moment estimator that can be attributed to geary1941inherent. When $\beta_{0}=0$, this estimator converges to a ratio of correlated (mean zero) normal random variables.
As a solution to this problem, we propose an estimator based on a divide-and-conquer strategy: the data is split into equal sized blocks. In each block, we calculate the denominator and numerator of the estimator (by geary1941inherent) on non-overlapping subsets of the data. This breaks the dependency between the numerator and denominator, and ensures that the estimator is median unbiased regardless of the value of $\beta_{0}$. Subsequently taking the median over the estimators from the different blocks yields a consistent and asymptotically normal estimator irrespective of the value of $\beta_{0}$. However, the rate of convergence of the proposed estimator depends on the number of blocks when $\beta_{0}=0$ and increases to the standard rate when $\beta_{0} \neq 0$.
To test the proposed estimator in practice, we revisit an empirical test of corporate finance's $q$-theory. This theory states that investment fluctuations are driven by the marginal $q$: the market value of capital relative to the shadow value of capital.\footnote{An overview of the history of $q$-theory can be found in erickson2000measurement} Empirically, $q$-theory appeared discredited with for example blundell1992investment finding a significant role for internals funds after controlling for $q$. However, erickson2000measurement show that these findings can be explained by the substantial measurement error in Tobin's $q$ measure of marginal $q$. This measurement error drives down the estimated coefficient on $q$, while increasing the coefficient of controls such as internal funds. Accounting for measurement error via the use of higher-order moment estimators showed that $q$-theory is not at odds with the data, see also erickson2012treating and andrei2019did.
We first consider the simulation set-up of erickson2014minimum, which is geared to the environment in which $q$-theory can be tested. We find that the divide-and-conquer estimator performs well regardless of the value of $\beta_{0}$, unlike the third-order moment estimator that is ill-defined when $\beta_{0}=0$.
We then analyse the data on firm investment, Tobin's $q$ and cash flow from erickson2014minimum. When using the full sample of available data, ranging from 1970 to 2011, the divide-and-conquer estimates are in line with those in erickson2014minimum. In particular, we do not find evidence for the effect of cash flow on firm investment. Motivated by the notion of erickson2014minimum that there exists variation over time in the estimates, we then re-estimate the coefficients using an expanding window of data starting with the data in 1970-1980 and ending with the full sample 1970-2011. We indeed find substantial changes in the estimated relation between investment, Tobin's $q$ and cash flow over time. In particular, we find rather extreme point estimates and confidence intervals in time periods before the mid-1980s. These occur precisely in time periods in which the divide-and-conquer estimator is not significantly different from zero. The divide-and-conquer estimates for the effect of Tobin's $q$ is found to be rather stable over time.
The paper relates to three different strands of the literature. First, the error-in-variables model has been thoroughly studied in econometrics and statistics. Estimation based on third-order moments as the one we employ can be subscribed geary1941inherent. Identification in the error-in-variables model is considered by (among others) reiersol1950identifiability,kapteyn1983identification,bekker1986comment. pal1980consistent extends this estimator for the single regressor case based on higher-order cumulants, while dagenais1997higher, cragg1997using and lewbel1997constructing consider other functions of mismeasured regressors. An overview of the literature can be found in wansbeek2000measurement and more recently schennach2016recent.
A general framework for higher-order moment estimators in models with multiple mismeasured and perfectly measured regressors is proposed in erickson2002two. Instead of the moments, it is somewhat simpler to rely on higher-order cumulants as in erickson2014minimum as the estimators are available in closed-form. Nonparametric identification and semiparametric estimation are considered in schennach2013nonparametric.
The second strand of literature concerns divide-and-conquer estimators. These estimators are developed for settings where a massive data set cannot be loaded into memory, see for instance shi2018massive. The idea is to construct a sequence of estimators on independent subsets of the data, and then aggregate the estimators into a single estimator. This is sometimes also referred to as distributed inference. The common aggregation method is to take the mean. Applied to monotone regression, banerjee2019divide document a superefficiency property of this method. Distributed quantile regression is studied by chen2019quantile and volgushev2019distributed.
Finally, rather than taking the mean of the subsample estimators, we rely on the median. This is shared with median-of-means (MoM) estimators that are used to robustify machine learning algorithms in the presence of heavy-tailed data. MoM estimators were originally proposed by nemirovskij1983problem, jerrum1986random, and alon1999space. In recent years, various authors show that these estimators attain optimal rates of convergence under weak assumptions on the data (hsu2014heavy, lugosi2019risk, and lecue2020robust). However, theoretical results are only available under the assumption that the random variables have finite second moment. The estimator we propose is closer to a median-of-ratios-of-means estimators, which in the worst case scenario does not have any finite moments.
The paper proceeds as follows. In (ref) we describe the model, the proposed estimator and its implementation. In (ref) we describe asymptotic results for the divide-and-conquer estimator. The simulation study and empirical application are presented in (ref). (ref) concludes. Proofs are deferred to (ref).
As a basis of our analysis we consider the following simplified error-in-variables model for one variable
for $i=1,\ldots, n$. Here the observed variables are $(y_{i},x_{i})$, while the remaining variables are latent. The main parameter of interest is $\beta_{0}$, the causal effect of marginal change in $\xi_{i}$ on $y_{i}$. As $\xi_{i}$ is not observed, one could naively consider estimating $\beta$ by running the OLS regression of $y_{i}$ on $x_{i}$:
This estimator, however, generally is inconsistent as in the limit $n\to\infty$:
provided that all stochastic quantities are mutually independent. Here we denote $\mathbb{E}[\xi^2_{i}] = \sigma_{\xi}^2$, $\mathbb{E}[\xi_{i}^3] = \sigma_{\xi}^{(3)}$, $\mathbb{E}[\xi_{i}^4] = \sigma_{\xi}^{(4)}$ and similarly for $\varepsilon_{i}$ and $u_{i}$ we have $\mathbb{E}[\varepsilon_{i}^2] = \sigma_{\varepsilon}^2$, $\mathbb{E}[\varepsilon_{i}^3] = \sigma_{\varepsilon}^{(3)}$, $\mathbb{E}[\varepsilon_{i}^4] = \sigma_{\varepsilon}^{(4)}$, $\mathbb{E}[u_{i}^2] = \sigma_{u}^{(2)}$.
As the OLS estimator is generally inconsistent, in what follows we limit our attention to a class of method-of-moments estimators that use the information from the higher order moments of the data. The parameter $\beta_{0}$ can then be estimated with the estimator attributed to geary1941inherent.
In what follows, we assume that all random variables are mutually independent.
A direct implication of (ref) is that in population
This implies the following moment condition,
for $\beta=\beta_{0}$. Associated with the above population orthogonality conditions is the following estimator, attributed to geary1941inherent,
As it is evident from (ref), the identifying power of this estimator is sensitive to the exact values of $\beta_{0}$ and/or $\sigma_{\xi}^{(3)}$. In particular, if $\beta_{0}=0$ and/or $\sigma_{\xi}^{(3)}=0$, the probability limit of the estimator is ill-defined.
This is also evident from the expression for the asymptotic variance given by
Clearly, when $\beta_{0}=0$ or $\sigma_{\xi}^{(3)}=0$, the variance is ill-defined.
In what follows, we describe the main intuition behind the procedure put forward in this paper. Notice that when $\beta_{0}=0$, the estimator (ref) is the ratio of two correlated sums that both have expectation zero. Borrowing on the idea of the Split Sample IV estimator of angrist1995split, we can easily break the dependence between the limiting value of the numerator and the denominator by evaluating the denominator and numerator on independent samples. As a result, the resulting estimator will be centered at the true value $\beta_{0}=0$. However, the estimator will still remain non-normal in the limit for $\beta_{0}=0$, and normal for $\beta_{0}\neq 0$, invalidating any standard inference approaches.
The above problem, however, can be solved. Define two subsets $R_{j,1}\subset \{1,\ldots,n\}$ and $R_{j,2}\subset\{1,\ldots,n\}$ such that $R_{j,1}\cap R_{j,2}=\emptyset$, denote $|R_{j,1}|=|R_{j,2}|=b/2$. The index $j$ will be used to index different choices of the subsets. Consider now the estimator,
When $\beta_{0}=0$, (ref) is the ratio of two independent sums that both have expectation zero. Intuitively, we would expect that as $b\rightarrow\infty$, $\hat{\beta}_{j}$ converges weakly to a Cauchy-type random variable with the CDF $F_{c}(x)$. Let $F_{j,b}(x)$ be the corresponding CDF of $\hat{\beta}_{j,b}$, then in (ref), we show that indeed
for some constant $M>0$ when $\beta_{0}=0$. Note that as Cauchy-type random variables are symmetric, their median is $0$, which happens to be exactly the value $\beta_{0}=0$ we are after.
In cases when the underlying data is a pure cross-section, this facilitates the following strategy to recover the median (thus also $\beta_{0})$ of the limiting random variable with the CDF $F_{c}(x)$. Randomly partition the $n$ observations into $B$ blocks, so that the number of observations within each block equals $b=n/B$. Within each block, split the sample to form (ref), assuming that $b/2$ is an integer. This yields a sequence of independent estimators $\{\hat{\beta}_{j,b}^{DC}\}_{j=1}^{B}$. Our estimator $\hat{\beta}_{b}^{DC}$ for $\beta$ is the median of this sequence.
The algorithm for the Divide-and-conquer estimator is as follows. \setcounter{bean}{0}
We prove below that for any fixed $\beta_{0}$ (including $\beta_{0}=0)$, $\hat{\beta}^{DC}_{b}$ is asymptotically normal. The convergence rate, however, depends on whether $\beta_{0}=0$ or $\beta_{0}\neq 0$. In the first case, the convergence rate is $B^{1/2}$, so it is determined only by the number of blocks $B$. In the latter case, we end up with the standard (pooled) convergence rate of $\sqrt{B\cdot b} = \sqrt{n}$. Hence, the fact that we use the median to estimate $\beta$ is free of any asymptotic costs, as long as the convergence rate of the estimator is concerned.
The choice of the number of subsample estimators $B$ is important. The theoretical results indicate the $B/b$ should go to zero for consistency of the estimator. If $B$ is too large relative to $b$, finite sample bias in the subsample estimators does not wash out over the different blocks. Our simulation results show that we should choose the value of $B$ relatively small. There, we have a total of 60,000 observations and set $B=20$, so that $B/b = 0.0066$ and $B=40$ so that $B/b = 0.0267$. For the latter, we find an increase in bias in the coefficients and confidence intervals that undercover for regressions that include fixed effects. When we increase the cross-section dimension to $n=6,000$ in (ref), so that for $B=40$ we get $B/b = 0.0133$ and coverage approaches the nominal rate. Optimal selection of the value of $B$ is an important topic for further research. We suggest that empirical researchers report their estimates for a range of different values of $B$ as a robustness check of their results.
As the asymptotic variance of the estimator depends on the unknown parameter $\beta_{0}$, we construct confidence intervals from a nonparametric bootstrap procedure. First, we draw $B$ samples $\tilde{e}_{j,b}$ with replacement from $\{\pm e_{1,b},\ldots,\pm e_{B,b}\}$ where $e_{j,b} = \hat{\beta}_{j,b}^{DC}-\hat{\beta}_{b}^{DC}$. Note that in view of the asymptotic distribution, we enforce symmetry by including each $e_{j,b}$ both with a plus and a minus sign. We then calculate $\tilde{\beta}_{i}^{DC}=\text{median}(\tilde{\beta}_{1,b}^{DC},\ldots, \tilde{\beta}_{B,b}^{DC})$, where $\tilde{\beta}_{j,b}^{DC} = \hat{\beta}_{b}^{DC}+\tilde{e}_{j,b}$ and $i=1,\ldots,B_{n}$. The $\alpha/2$ and $1-\alpha/2$ quantiles $(q_{\alpha/2}^{DC},q_{1-\alpha/2}^{DC})$ of the empirical distribution of $\tilde{\beta}_{i}^{DC}-\hat{\beta}_{b}^{DC}$ are then used to construct a confidence interval as $(\hat{\beta}_{b}^{DC}+q_{\alpha/2}^{DC},\hat{\beta}_{b}^{DC}+q_{1-\alpha/2}^{DC})$.
Two extensions to the above methodology are needed to implement the divide-and-conquer estimator in our empirical setting. The first is the addition of perfectly measured control variables to the model. The second is to consider a panel data setting.
First, the model (ref) can be expanded by including a set of perfectly measured controls, so that the model is of the form
In this case, define
and similarly for $\dot{y}_{i}$. Denote by $\ddot{x}_{i}$ and $\ddot{y}_{i}$ the corresponding estimators on the set $R_{j,2}$. We then adapt (ref) in the divide-and-conquer algorithm to
This retains the independence between the subsample estimators. Note that for inference on $\bs\gamma$ we can adapt the procedure from (ref) and construct confidence intervals from the quantiles of $\tilde{\bs \gamma}^{DC}_{j,b} = (\sum_{i=1}^{n}\bs z_{i}\bs z_{i}')^{-1}\sum_{i=1}^{n}\bs z_{i}(y_{i}-x_{i}\tilde{\beta}^{DC}_{j,b})$.
As a second extension, we consider a linear panel data model that will be used in our empirical application,
In this case, we construct subsample estimators on cross-sections of the data. To account for individual fixed effects, we perform a within transformation on the data. Moreover, to account for time fixed effects in the model, we demean the data within blocks. Perfectly measured controls are handled in the same way as above.
The algorithm for the Panel Divide-and-conquer estimator is as follows. \setcounter{bean}{0}
We start by making the following assumption on the moments of $(\xi_{i},u_{i},\varepsilon_{i})$ that facilitate a Berry-Esseen inequality in the proof of the main theorem. We limit our attention to the basic estimator we suggest in Section (ref). In the following, denote by $\Phi(x)$ and $\phi (x)$ the CDF/PDF of a standard normal r.v.\ and by $F_{c}(x)$/$f_{c}(x)$ the CDF/PDF of a Cauchy r.v.
We impose the following assumption on the number of blocks $B$, which is chosen as a function of the sample size within each block $b$.
Note that when $\beta_{0}=0$, the estimator can be decomposed as
where $|R_{j,1}|=|R_{j,2}| =b/2$ and $j=1,\ldots,B$. Denote $\sigma_{v}^2 = \mathbb{E}[v_{j,b}^2]$ and $\sigma_{w}^2=\mathbb{E}[w_{j,b}^2]$. When $\beta_{0}\neq 0$, define $w_{i} = x_{i}\varepsilon_{i}^2 + 2\beta_{0} (\xi_{i}^2 + \xi_{i}u_{i})\varepsilon_{i}$ and $\sigma_{\beta}^2 = (\beta_{0}\sigma_{\xi}^{(3)})^{-2}\mathbb{E}\left[w_{i}^2\right]$.
This brings us to our main result.
We see that regardless of the value of $\beta_{0}$, the estimator is asymptotically normal. The asymptotic distribution, on the other hand, is generally discontinuous at $\beta_{0}$. In particular, for $\beta_{0}=0$, the convergence rate is $\sqrt{B}$. Since (ref) requires that $B/b = B^2/n\rightarrow 0$, the maximal convergence rate is below $n^{1/4}$. When $\beta_{0}\neq 0$, we recover the standard $\sqrt{n}$ convergence rate.
We also see that the asymptotic variance depends on the, evidently unknown, value of $\beta_{0}$. Inference therefore proceeds via the non-parametric bootstrap discussed in the previous section that accommodates cases $(a)$ and $(b)$ simultaneously.
In this section we consider the simulation set-up from erickson2014minimum. Their setup is calibrated in order to reflect several characteristics of the data used in their analysis. In particular, we consider a simple linear panel data model of the form,
Here $i=1,\ldots,n$ with $n=3,000$ and $t=1,\ldots,T$ with $T=20$. The errors $u_{it}$ and $e_{it}$ are i.i.d.\ Gamma random variables standardized using the population mean and standard deviation. The scale parameter is set to 1 and the shape parameter is set to $0.32$ for $u_{it}$ and $0.09$ for $e_{it}$.
The perfectly measured controls and the mismeasured regressor are generated using an autoregressive processes of order 1 - AR(1), as follows
Here we set $\phi_{\xi} = 0.78$ and $\phi_{z}=0.48$. The intercepts are set to match the first moment of Tobin's $q$ and cash flow in the data as $\delta_{\xi} = 0.570$ and $\delta_{z} = 0.094$. The errors $(v_{it}^{\xi},v_{it}^{z})$ are again independent Gamma random variables. The scale parameter is set to 1, the shape parameter for $v_{it}^{\xi}$ equal to 0.007 and for $v_{it}^{z}$ equal to 2.08. We follow the standard practise in the literature and initialize the processes for $\xi_{it}$ and $z_{it}$ in the recent past by setting $\xi_{i,-10}=z_{i,-10}=0$ for all $i$. We then generate $T+10$ time series observations for both processes and drop the first 10 periods.
After having generated $z_{it}$ and $\xi_{it}$, we transform these to ensure that the covariance matrix calculated using all available data is equal to the covariance matrix of $x_{it}$ and $z_{it}$ in the data, with the variance of Tobin's $q$ multiplied by $\tau^2 = 0.45$ and the covariances multiplied by $\sqrt{\tau^2}$. Numerically, the covariance matrix is
Finally, we generate the observed variables as follows
Here $\mu_{y}$ and $\sigma_{y}^2$ are set to match the population mean and variance of the investment variable in the data. We consider two choices for $\beta$. In the first $\beta=0.025$ as in erickson2014minimum, while in the second $\beta=0$. The coefficient for control variable $z_{it}$ is fixed to $\gamma=0.05$.
We estimate the parameters in (ref) using OLS (OLS), the third-order moment estimator (3M) and the divide-and-conquer estimator (DC). For the divide-and-conquer estimator, we vary the number of blocks per time period from 1 to 2, so we take the median over $B\cdot T = 20$ and $B\cdot T = 40$ subsample estimators. In each time period, the partition of the firms is random.
Bootstrap confidence intervals are constructed as outlined in (ref) with $B_{n}=399$. For the third-order moment estimator, confidence intervals are calculated as in erickson2002two.
(ref) shows the mean and standard deviations of the estimates for the $q$-coefficient $\beta$ and the cash flow coefficient $\alpha$. When $\beta_{0}=0.025$, we see that the OLS estimates for $\beta$ and $\alpha$ are badly biased due to the measurement error in the process. The third-order moment estimators as well as the divide-and-conquer methods show no to very little bias. The standard errors of the divide-and-conquer estimators are larger relative to the third-order moment estimator.
When $\beta_{0}=0$, we see that the third-order moment estimator is badly defined, with the standard error going up substantially. The divide-and-conquer estimators on the other hand perform well, with standard deviations nearly identical to those when $\beta_{0}=0.025$.
(ref) shows the coverage rates for the various methods. As expected, the bias in the OLS estimator when $\beta_{0}\neq 0$ results in confidence intervals that do not include the true parameter. The confidence intervals corresponding to the 3M estimator are somewhat conservative, especially when we include fixed effects. For the DC estimator, we find overall close to nominal coverage except when $\beta_{0}=0.025$ and we include fixed effects. (ref) shows that increasing the sample size to $n=6,000$ largely resolves this issue.
We consider an application from the corporate finance literature that seeks to determine the relation between actual investment and unobserved investment opportunities in capital as in fazzari1988financing. The investment opportunities are proxied by Tobin's $q$, which contains considerable measurement error. The third-order moment estimator (ref) and generalizations thereof have been applied to account for this measurement error in erickson2000measurement, erickson2012treating and erickson2014minimum. While identification problems due to lack of skewness were investigated in erickson2012treating, problems that arise because $\beta_{0}$ is (close to) zero, have not been addressed.
We use the data of erickson2014minimum to estimate the linear panel data model
where $i=1,\ldots,n$ indexes the firm, and $t=1,\ldots,T$ the year. The data ranges from 1970 to 2011. After removing firms for which only one year is available, the sample consists of 121,733 firm-year observations. The outcome variable $y_{it}$ is investment as constructed by erickson2014minimum. The perfectly measured control variables $\bs z_{it}$ include an intercept and a measure of cash flow, $x_{it}$ denotes Tobin's $q$.
We estimate the effect of Tobin's $q$ on investment in (ref) by (1) OLS, (2) the third-order moment estimator (ref) and (3) the divide-and-conquer algorithm. We calculate confidence intervals for (ref) as described in erickson2002two. For the divide-and-conquer method, we apply the bootstrap procedure from (ref).
For the divide-and-conquer estimator, for every year we divide the data into 2 blocks, which for the full sample corresponds to 80 blocks. We also report results for 1 block per year. Since for each block we split the data into two halves to calculate the numerator and denominator of the subsample estimators, in each year we discard a randomly selected subset of firms so that the number of firms is a multiple of 2 (for DC with 1 block per year) or a multiple of 4 (for DC with 2 blocks per year).
We first report the full sample pooled estimators and standard errors in (ref). For the OLS estimators, we see the significant effect of cash flow on investment. This questions $q$-theory, which states that in a regression of investment on $q$, control variables must be insignificant. Indeed, we see that this theory can be aligned with empirical evidences by the third-order moment estimators. These conclusions are upheld by the divide-and-conquer methods: Tobin's $q$ has a significant positive effect on investment, while cash flow does not. The exception is formed by the results from divide-and-conquer with both firm and year effects. In this case, the confidence intervals for the Tobin's $q$ include zero.
Secondly, erickson2014minimum document that the coefficient estimates may be varying over time. We therefore apply the same estimation method on subsamples of the data. The first subsample is 1970-1980, which we then expand by one year at a time until we reach the full-sample estimates in 2011. For this exercise, we limit our attention to the model with firm fixed effects. We report results for the divide-and-conquer estimator with $2$ blocks per year. Results for $1$ block per year are reported in (ref).
(ref) shows the estimates for Tobin's $q$ alongside the 95% confidence intervals. Interestingly, in the first half of the time period, we find that both the third-order moment estimator and the DC estimator include zero in the confidence interval. This is paired with the observation that the third-order moment estimator takes a value for Tobin's $q$ which is an order of magnitude larger than the full-sample estimator.
(ref) shows the estimated cash flow coefficient. While the confidence intervals for the third-order moment estimator include zero for all time periods, the divide-and-conquer estimator shows some mild evidence for a cash flow effect, especially in the early years of the sample. This aligns with the findings by andrei2019did who find that $q$-theory has become increasingly efficient in explaining investment patterns.
Estimators based on higher-order moments of the data can eliminate the effect of measurement error. However, if the coefficient $\beta_{0}$ of the latent regressor is zero, the estimators converge weakly to a ratio of correlated mean zero normal random variable. As a result, standard inference procedures are not applicable in this case.
As a solution to this problem we suggest the divide-and-conquer estimator that estimates the parameter of interest using the median of subsample estimators. We show this estimator is consistent and suggest a bootstrap based inference procedure that is applicable irrespective of the coefficient $\beta_{0}$. The estimator and inference procedure are intuitive and easy to implement by practitioners. Monte Carlo results indicate that this bootstrap procedure well approximates the finite sample distribution of the estimator.
Finally, we analyse the data on firm investment, Tobin's $q$ and cash flow from erickson2014minimum and do not find significant evidence for the effect of cash flow in firm investment. Interestingly enough, we find substantial changes in the estimated relation between investment, Tobin's $q$ and cash flow over time. In particular, we find rather extreme point estimates and confidence intervals in time periods before the mid-1980s for the third-moment estimator of erickson2014minimum. These occur precisely in time periods in which the divide-and-conquer estimator is not significantly different from zero. On the other hand, the divide-and-conquer estimates for the effect of Tobin's $q$ is found to be rather stable over time. These observations confirm the empirical relevance of our proposed statistical procedure.
\paragraph{Acknowledgements} We thank Gerard van den Berg, Simon Broda, Noud van Giersbergen, Toru Kitagawa, Frank Kleibergen, Ruud Koning, Tom Wansbeek, and seminar participants at the Tinbergen Institute Amsterdam and the University of Groningen, for helpful comments. We thank Toni Whited for sharing the data. Financial support from the Netherlands Organization for Scientific Research (NWO) under research grant number $201\text{E}.011$ (TB) and $451-17-002$ (AJ) is gratefully acknowledged.
\paragraph{Disclosure statement} The authors report there are no competing interests to declare.
\paragraph{Data availability statement} The authors confirm that the code and data supporting the findings of this study are available within the supplementary materials. {0pt plus 0.2ex}