Ricardo Masini, Marcelo Medeiros
arXiv 19 Feb 2025 · Statistics — Methodology
arXiv:2502.13438 · PDF · DOI · OpenAlex · Extracted main text
Traditional parametric econometric models often rely on rigid functional forms, while nonparametric techniques, despite their flexibility, frequently lack interpretability. This paper proposes a parsimonious alternative by modeling the outcome $Y$ as a linear function of a vector of variables of interest $\boldsymbol{X}$, conditional on additional covariates $\boldsymbol{Z}$. Specifically, the conditional expectation is expressed as $\mathbb{E}[Y|\boldsymbol{X},\boldsymbol{Z}]=\boldsymbol{X}^{T}\boldsymbol{\beta}(\boldsymbol{Z})$, where $\boldsymbol{\beta}(\cdot)$ is an unknown Lipschitz-continuous function. We introduce an adaptation of the Random Forest (RF) algorithm to estimate this model, balancing the flexibility of machine learning methods with the interpretability of traditional linear models. This approach addresses a key challenge in applied econometrics by accommodating heterogeneity in the relationship between covariates and outcomes. Furthermore, the heterogeneous partial effects of $\boldsymbol{X}$ on $Y$ are represented by $\boldsymbol{\beta}(\cdot)$ and can be directly estimated using our proposed method. Our framework effectively unifies established parametric and nonparametric models, including varying-coefficient, switching regression, and additive models. We provide theoretical guarantees, such as pointwise and $L^p$-norm rates of convergence for the estimator, and establish a pointwise central limit theorem through subsampling, aiding inference on the function $\boldsymbol\beta(\cdot)$. We present Monte Carlo simulation results to assess the finite-sample performance of the method.
appendix boundary found by appendix_command · 47% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Friedberg, R., Tibshirani, J., Athey, S., and Wager, S (2021) Local linear forests | 1.000 | 5 | 3 | 100% |
| 2 | Athey, S., Tibishirani, J., and Wager, S (2019) Generalized random forests | 0.585 | 3 | 1 | 100% |
| 3 | Dagenais, M. G (1969) A threshold regression model | 0.511 | 2 | 1 | 100% |
| 4 | Goldfeld, S. M. and Quandt, R (1972) Nonlinear Methods in Econometrics | 0.511 | 2 | 1 | 100% |
| 5 | Hastie, T. and Tibishirani, R (1993) Varying-coefficient models | 0.511 | 2 | 1 | 100% |
| 6 | Hansen, B (2000) Sample splitting and threshold estimation | 0.405 | 1 | 1 | 100% |
| 7 | Breiman, L (2001) Random forests | 0.405 | 1 | 1 | 100% |
| 8 | Tong, H. and Lim, K (1980) Threshold autoregression, limit cycles and cyclical data (with discussion) | 0.405 | 1 | 1 | 100% |
| 9 | Fan, J., Zhang, C., and Zhang, J (2001) Generalized likelihood ratio statistics and wilks phenomenon | 0.405 | 1 | 1 | 100% |
| 10 | Fan, J. and Zhang, W (1999) Statistical estimation in varying coefficient models | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 34 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | 0.5cm dpd LGB+: A Macroeconomic Forecasting Road Test . 0.25cm | 0.405 | 1 | 1 |
| 2 | Global Testing in Multivariate Regression Discontinuity Designs | 0.000 | 1 | 1 |