Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,790 characters · 47 sections · 33 citation commands
Reconstructing Subnational Labor Indicators in Colombia: An Integrated Machine and Deep Learning Approach
\noindentKeywords: Labor statistics; temporal disaggregation; machine learning; deep learning; Colombia
Accurate, timely, and spatially disaggregated labor statistics are crucial for diagnosing employment dynamics, informing policy decisions, and advancing regional equity within decentralized governance frameworks ILOCEPAL2022,OECD2021. However, in Colombia, structural constraints have resulted in significant gaps in subnational labor monitoring. As of 2025, the national household survey (GEIH, Gran Encuesta Integrada de Hogares) provides consistent indicators for only 23 of the 33 departments. Since December 2021, city-level series for 32 major urban areas have been released; however, these data do not directly correlate with departmental labor conditions in Amazonas, Arauca, Caquetá, Chocó, Guainía, Guaviare, Putumayo, San Andrés, Vaupés, or Vichada \footnote{Despite the fact that Gross Domestic Product (GDP) estimates are produced for all 33 departments,these capture production rather than labor outcomes.}. This gap underscores the necessity for methodological efforts to align city and departmental labor statistics, which is the central focus of this study.
The limited availability of historical, department-level labor data constrains the study of long-term trends, cyclical fluctuations, and regional disparities. These limitations are particularly significant for analyzing informality, which constitutes a substantial portion of the Colombian workforce DANE2025. To address this gap, we introduce a unified estimation framework that reconstructs consistent monthly labor indicators for all 33 Colombian departments over the period 1993-2025. The framework encompasses standard labor aggregates and distinguishes between formal and informal employment \footnote{Subemployment and other highly disaggregated labor indicators are omitted due to insufficient and inconsistent information across time and departments.} It integrates temporal disaggregation, demographic anchoring, statistical learning techniques, and the enforcement of labor accounting identities, with all estimates calibrated to national benchmarks.
The outcome is a demographically consistent and statistically coherent panel that aligns fully with official aggregates. For the first time, the resulting dataset provides comprehensive monthly coverage of seven core variables: employment, unemployment, inactivity, labor force (PEA, Población Económicamente Activa), working-age population (PET, Población en Edad de Trabajar), total population, and the informality rate. This enables systematic comparisons across various time periods and regions, including those prior to 2001, when no official high-frequency statistics were available.
Validation against official GEIH departmental annual data (2007-2024) indicates that the in-sample Mean Absolute Percentage Errors (MAPEs) are below 2.3% for all core indicators. Informality estimates align with city-level monthly benchmarks (2007-2025) within acceptable margins (MAPE$<2.1\%$). The reconstructed series effectively captures significant episodes in Colombia’s labor history: consistent gains in participation during the late 1990s, a sharp contraction during the 1999 financial crisis, partial recovery in the mid-2000s, renewed stress during the 2008 global recession, and disruptions caused by the COVID-19 pandemic.
Beyond its technical contribution, the reconstructed panel provides a robust empirical foundation to trace long-run regional labor trajectories, assess the heterogeneous effects of macroeconomic crises, and analyze structural asymmetries across departments. A composite Employment Quality Index (EQI) is incorporated to characterize territorial labor markets, explicitly penalizing configurations where high employment coexists with high informality. This multidimensional approach facilitates a more nuanced evaluation of convergence, divergence, and resilience within Colombia’s labor landscape.
To our knowledge, no existing dataset provides consistent monthly labor indicators, including informality, for all Colombian departments over such an extended period. Current official statistics remain fragmented, offering only partial subnational coverage. By addressing this gap, our dataset facilitates, for the first time, a systematic analysis of both the quantity and quality of employment across the national territory over more than three decades. This supports historical inquiry and evidence-based policy design.
The broader significance of this work extends to other countries facing similar statistical limitations. It illustrates how statistical learning methods can complement household surveys to expand the scope of labor monitoring in data-constrained contexts. Simultaneously, the results highlight a fundamental limitation: while model-based reconstruction can generate coherent and policy-relevant indicators, it cannot replace primary data collection in structurally underrepresented regions.
The remainder of this paper is structured as follows. Section (ref) reviews the relevant literature on labor estimation, temporal disaggregation, and supervised prediction. Section (ref) details the reconstruction pipeline. Section (ref) presents the main empirical findings and validation metrics. Section (ref) concludes and outlines future research directions.
This section reviews recent academic and technical contributions related to the estimation of labor market indicators under conditions of incomplete data coverage. The review is organized into five thematic areas: (i) applications of machine learning to labor statistics, (ii) temporal disaggregation of labor indicators, (iii) estimation under spatial gaps, (iv) regional segmentation and convergence, and (v) accounting consistency in predictive modeling. We prioritize studies conducted in Colombia while also considering significant contributions from Latin America that have methodological relevance. Table (ref) summarizes the main references discussed.
Supervised learning methods have increasingly been applied to estimate labor statistics in contexts with limited direct observation. These approaches utilize flexible algorithms and auxiliary covariates to infer labor outcomes in situations of incomplete coverage. A key advantage of these methods is their capacity to learn complex, nonlinear relationships from the data, enabling the models to generalize patterns that may not be evident using traditional econometric techniques.
vanDijk2022 developed high-resolution labor maps for Vietnam by downscaling district-level employment data using a super learner ensemble. Their model integrated satellite imagery, nightlight intensity, and household surveys to estimate spatial employment distributions across diverse areas. This study demonstrates the effectiveness of ensemble methods in interpolating labor variables in data-scarce contexts.
In Colombia, Vidal2024 proposed a composite labor index that integrates informality rates, labor expectations, and behavioral indicators derived from Google Trends. Their model employs tree-based learners and feature selection techniques, effectively capturing short-term dynamics in urban labor markets despite the absence of high-frequency official data. Although their analysis was limited to city-level series, it demonstrates the predictive potential of nontraditional data sources.
Recent studies in Colombia have utilized supervised learning techniques to predict unemployment, labor force participation, and informality by incorporating economic variables, alternative data sources (e.g., Google Trends), and official survey data. OrozcoCastaneda2024 A composite labor indicator will be constructed through machine learning models trained on city-level inputs, highlighting the potential of these methods to capture complex urban labor dynamics. PerezRosero2025 Explainability will be emphasized by integrating Uniform Manifold Approximation and Projection (UMAP) with Gaussian Processes to develop interpretable models of unemployment. These efforts underscore the increasing empirical relevance of machine learning techniques in Colombian labor research.
These studies confirm the feasibility of applying supervised learning to labor market modeling, provided that the models incorporate robust anchor variables and context-sensitive predictors.
Temporal disaggregation techniques facilitate the estimation of high-frequency time series from low-frequency aggregates. These techniques are particularly pertinent when official labor statistics are available solely on an annual basis at the subnational level.
In contexts where low-frequency data are present, temporal disaggregation is employed to construct high-frequency labor series. In Colombia, GonzalezHerrera2018 applied the Denton and Chow-Lin methods to interpolate employment data while maintaining consistency constraints. Although these methods are standard in national accounts, their application to social indicators is limited in practice. Sanchez2015 establishes the foundation for reconstructing intra-annual variation in labor aggregates when direct monthly data are unavailable.
Chow1971 introduced a regression-based approach that distributes annual data into monthly or quarterly values based on their covariation with an auxiliary indicator. Denton1971 proposed a quadratic minimization procedure that smooths the interpolated path while preserving the trend of the indicator. Fernandez1981 generalized these techniques to accommodate integrated processes, and Litterman1983 introduced a stochastic extension incorporating autoregressive innovations.
Recent developments, such as Mosley2022, incorporate sparse regularization (LASSO) for the selection among multiple noisy predictors during the disaggregation process. This approach enhances robustness in high-dimensional settings, where collinearity and overfitting are prevalent risks.
However, traditional disaggregation methods typically assume complete data at the aggregate level and are infrequently applied at the subnational scale. Their adaptation to contexts with fragmented spatial and temporal coverage is constrained.
A significant empirical challenge in Colombia is the lack of monthly labor data for departments not included in the GEIH. Existing studies have employed methods such as spatial extrapolation, synthetic estimation, and predictive modeling. For example, ILOCEPAL2022 provides technical guidelines for constructing consistent subnational labor indicators using auxiliary data and partial survey coverage.These methods represent viable strategies for estimating labor indicators in regions lacking direct observation, particularly when aligned with annual benchmarks. These methods are effective for estimating labor indicators in regions without direct observation, especially when correlated with annual benchmarks.
Understanding regional labor disparities necessitates modeling structural heterogeneity. GarciaPena2017 analyzes convergence in informal employment across Colombian departments using spatial econometrics. Their findings indicate persistent segmentation and path dependence in subnational labor markets. Similar studies in Latin America have utilized clustering, principal components, or typology frameworks to categorize territories based on labor structures. These methods support region-specific modeling strategies and facilitate the identification of structural gaps in labor conditions across territories.
Labor indicators are interdependent through accounting identities (e.g., labor force = employed + unemployed). However, few empirical models explicitly enforce these constraints. Recent work, such as OECD2021, explores multi-output machine learning models and structural regularization to ensure coherence among predicted indicators. Incorporating such restrictions enhances interpretability and consistency, particularly when the outputs are utilized for policy formulation or official statistics. In Latin America, this area remains open for methodological development, particularly regarding the integration with machine learning pipelines.
This study significantly contributes to the labor statistics literature by addressing longstanding territorial and temporal gaps in official measurements. First, it introduces a harmonized monthly panel of labor market indicators for all 33 Colombian departments, covering the period from 1993 to 2025. This representation effectively bridges a critical empirical gap in subnational labor monitoring. Second, it develops an integrated estimation pipeline that combines supervised machine learning, econometric disaggregation, deep learning, and accounting-based reconstruction. Third, it implements a structured calibration protocol, ensuring consistency with national aggregates while maintaining internal coherence across labor indicators.
Figure (ref) summarizes the multi-stage pipeline developed to reconstruct monthly labor indicators for all Colombian departments from 1993 to 2025. This process is hierarchical, progressing from national anchors to city-level dynamics and finally to departmental series. At each stage, consistency is ensured through the integration of official data sources, demographic controls, and statistical learning methods.
A consistent national baseline is established by compiling data from the World Bank, the International Labour Organization (ILO), the General Household Survey (GEIH, Gran Encuesta Integrada de Hogares), and official demographic projections. Concepts are harmonized, and totals are calibrated to official national aggregates. Monthly series are derived through temporal disaggregation, providing the reference against which all subnational reconstructions are reconciled. This process ensures that departmental estimates remain coherent with national labor accounts.
To extend coverage prior to the GEIH, we reconstruct Colombia’s national labor market series using international data from 1993 to 2000. Annual indicators from the World Bank’s World Development Indicators worldbank2025, which incorporate ILO estimates ILOSTAT2025, provide figures for employment, unemployment, and population. These indicators are temporally disaggregated and integrated with monthly GEIH data from 2001 onward, ensuring demographic consistency and continuity.
However, global estimation frameworks often fail to capture dynamics specific to Colombia. For example, the surge in unemployment following the 1998–1999 recession, the mild impact of the 2008-2009 financial crisis, and the sharp shock of COVID-19 in 2020 are only partially reflected Medina2013, Grosh2014, Alvarez2021, DANE2025, BanRep2025. To address these deficiencies, we structurally recalibrate the series by smoothing employment and unemployment to align with macroeconomic cycles, reconstructing participation based on its empirical relationship with employment, and deriving aggregates using official projections of the working-age population.
The resulting annual estimates encompass employment, unemployment, labor force (PEA), inactivity, and the working-age population (PET). All estimates are internally consistent, and their comparison with GEIH benchmarks (2001-2024) demonstrates that average absolute percentage errors (MAPE) are below 5%. These series function as macro anchors for monthly disaggregation, providing a harmonized baseline for the period 1993-2025.
The demographic baseline is based on official population projections published by DANE dane_proyecciones2023. These projections cover the period from 1993 to 2050 and are available in three vintages (1993-2004, 2005-2019, 2020-2050), which have been harmonized to ensure consistency. Departmental projections extend through 2035. These projections provide denominators for labor rates (e.g., participation and unemployment) and facilitate the conversion of rates into absolute levels that align with official totals.
In parallel, monthly labor indicators are compiled from the GEIH geih_dane, producing a harmonized panel extending from 2001 to 2025. Together, the DANE projections and GEIH indicators establish the national trajectory that underpins all departmental estimates.
To extend labor indicators beyond the GEIH window, annual World Bank estimates are disaggregated into a monthly frequency using macroeconomic covariates. Instead of interpolating absolute levels, we reconstruct ratios of employment, unemployment, participation (PEA), working-age population (PET), and inactivity relative to reference populations. This approach minimizes seasonal noise and inconsistencies among sources.
Monthly trajectories are derived using Chow–Lin models calibrated on macroeconomic predictors, including the Consumer Price Index (CPI, Índice de Precios al Consumidor), Producer Price Index (PPI, Índice de Precios al Productor), exchange rate (TRM, Tasa Representativa del Mercado), and real minimum wage. The selection of predictors is based on their empirical correlation with labor cycles stock1999inflation, gali1999techshock. The optimal predictor minimizes the Mean Absolute Percentage Error (MAPE) relative to GEIH benchmarks from 2001 to 2024, with all ratios resulting in errors below 1.5.
Ratios are converted into levels using DANE projections. PET is linearly smoothed to avoid discontinuities, while other aggregates are reconstructed utilizing accounting identities. Coherence is maintained consistently, ensuring that the sum of employed and unemployed individuals equals PEA, and the sum of PEA and inactive individuals equals PET. Figures (ref) and (ref) illustrate the reconstructed ratios and residual adjustments. The result is a demographically consistent, high-frequency national baseline for the years 1993-2025, which serves as the foundation for all departmental reconstructions. Comprehensive technical details, including equations, diagnostics, and sensitivity checks, are presented in (ref).
The reconstruction process of the subnational labor market begins with the 32 principal cities, for which monthly labor market indicators have been available since December 2021, alongside a synthetic series for Cundinamarca (excluding Bogotá). The signals from these cities are extended backward using the 23 GEIH urban domains as references. In instances where a city lacks a direct match, the backward extension employs the domain exhibiting the highest correlation. To mitigate noise, the extensions are initially calculated at an annual frequency and subsequently disaggregated into monthly series that align with national totals. This approach results in a coherent framework of 33 city-level signals, continuously extended from 1993 to 2025, which serve as predictors for the departmental reconstruction.
City-department mappings are established based on pairwise correlations of monthly indicators. Spearman’s rank correlation serves as the primary criterion, supplemented by Kendall’s tau as a robustness check. Pairs with fewer than ten overlapping observations are excluded from the analysis. For departments lacking direct observations, the donor with the strongest correlation is selected. In the case of Cundinamarca, departmental results are utilized directly to ensure comprehensive coverage. This approach results in a stable donor structure applicable across all extensions.
The 32 observed cities and the synthetic Cundinamarca series are aggregated to annual frequency (medians, 2021-2025) and linked to GEIH-compatible domains. Direct matches are paired directly, while unmatched cities utilize donor series. Proportional retropolarization extends these annual signals back to 2007, ensuring consistent levels and growth rates. Monthly disaggregation is then performed using the optimal Chow–Lin method (average conversion), anchored to national monthly indicators. The final panel for 1993-2025 is produced through proportional splicing against national totals, ensuring consistency in population, employment, unemployment, inactivity, PEA, and PET. Labor accounting identities are enforced, and national projections are aligned. Full implementation details are provided in (ref).
The second stage addresses the departments, which are the primary targets for estimation. For the 23 departments directly covered by the GEIH, official annual aggregates are disaggregated, labor accounting identities are applied, and demographic projections are integrated. For the 23 departments covered by the GEIH (2007-2024), annual aggregates are converted to a monthly frequency using the Chow-Lin method, which is anchored to national indicators. This approach preserves both annual totals and temporal continuity.
Historical extensions to 1993 are included only when sufficient evidence is available (threshold of 0.95), thus avoiding artificial interpolations. Unlike city-level predictors, which are systematically extended to provide inputs for modeling, departmental series are treated as final outputs and, therefore, necessitate higher accuracy. These reconstructed series serve as the dependent variables in the predictive stage, with city signals and other covariates functioning as explanatory inputs.
A unified set of monthly predictors is compiled to model departmental labor indicators through supervised learning. Covariates are selected based on three criteria: (i) availability since 1993, (ii) monthly frequency, and (iii) empirical relevance to labor market dynamics.
Macroeconomic predictors include the nominal exchange rate (TRM, Tasa Representativa del Mercado), consumer price index (CPI, Índice de Precios al Consumidor), producer price index (PPI, Índice de Precios al Productor), exports, imports, real value unit (UVR, Unidad de Valor Real), and the industrial production index (IPI, Índice de Producción Industrial). Institutional predictors consist of the legal minimum wage (Salario Mínimo Legal Vigente) and the transport subsidy (\textit{Subsidio de Transporte}). A real minimum wage proxy is derived by deflating the nominal wage using the CPI. GDP and departmental production are excluded due to their annual frequency and delayed update. Consequently, the final set of covariates primarily comprises monthly indicators, supplemented by a limited number of retrospectively expanded annual series. Comprehensive definitions and data sources are reported in (ref).
To enhance generalization, departments are grouped into clusters that capture structural similarities in demographic, geographic, and economic characteristics. Clustering allows unobserved departments to borrow information from observed departments, thereby incorporating variability that is not entirely explained by labor indicators . A strict constraint is imposed: every cluster must contain at least one department with GEIH coverage. This constraint prevents the formation of synthetic clusters composed exclusively of unobserved units, ensuring that all extrapolations are empirically grounded (Table (ref)).
Model outputs are initially derived as monthly labor shares (employment, unemployment, and inactivity). These shares are converted into absolute population counts using official demographic projections and accounting identities: employment = employment rate $\times$ PET, unemployment = unemployment rate $\times$ PEA, and inactivity = inactivity rate $\times$ PET. This approach ensures consistency between rates and levels over time.
For departments covered by the GEIH, splicing aligns the observed and reconstructed series prior to level conversion. For unobserved departments, the trained model generates monthly shares, which are transformed in a similar manner. Details regarding the splicing formulation and reconstruction logic are provided in (ref).
All elements are integrated into a department-month feature matrix suitable for supervised estimation. Two partitions are defined: (i) a training set consisting of 23 departments with observed labor outcomes and (ii) a prediction set comprising 33 departments, which includes both observed and unobserved cases. This setup facilitates out-of-sample predictions, with generalization validated through cross-department predictive checks. The variables were selected by prioritizing those with high frequency and timely updates, ensuring that the methodological framework allows for the continuous and up-to-date reconstruction of indicators.
To reconstruct monthly labor indicators for the 33 departments from 1993 to 2025, we employ a supervised learning framework for multi-output prediction. This framework captures non-linear dynamics, leverages shared structures across outputs, and integrates heterogeneous predictors. The five core targets, working-age population (PET, Población en Edad de Trabajar), economically active population (PEA, Población Económicamente Activa), employed individuals, unemployed individuals, and inactive individuals, are estimated jointly, ensuring that temporal and structural patterns are consistently modeled. The framework relies on the strength of city-level signals, which provide sufficiently rich information to extrapolate departmental trajectories without necessitating more complex architectures.
The modeling strategy employs gradient-boosted regression trees (XGBoost). Each indicator is formulated as a function of demographic and institutional covariates. Targets are modeled as population shares to enhance stability and comparability across departments. Following estimation, absolute levels are derived by rescaling predicted shares with the population feature, which is intentionally excluded from inputs to prevent data leakage. The boosting process iteratively fits residuals, enabling the model to capture complex interactions and non-linearities with high predictive accuracy and generalization capabilities.
Training follows systematic procedures, including feature scaling to stabilize optimization. Performance is evaluated using three standard metrics: Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and Root Mean Squared Error (RMSE). These metrics provide a comprehensive assessment of accuracy and enable comparability across indicators and departments.
Target transformation and scaling. Targets are normalized as shares to enhance numerical stability and comparability across departments and over time, while preventing data leakage from absolute levels. After making predictions, absolute counts are recovered by rescaling the estimated shares using official population projections for the period 1993-2035.
Input features. The design matrix incorporates demographic, macroeconomic, and labor-related signals:
Training and validation strategy. All features and targets are standardized prior to training. The training set encompasses the 23 GEIH departments. Model validation adheres to four schemes:
Performance is evaluated based on share-based targets. The stratified holdout method yields an average Mean Absolute Percentage Error (MAPE) of 1.38%. LOGO errors increase to 10–12%, highlighting the challenges associated with spatial extrapolation in the absence of direct anchors. These results are accessible on (ref).
Prediction and generalization. The final model is trained on all observed departments and generates monthly predictions for all 33 departments. For departments that lack GEIH data, results are primarily influenced by macroeconomic predictors, demographic projections, and cluster-level proxies.
Post-processing and alignment. Predictions are adjusted to ensure demographic and national consistency:
The resulting panel yields internally consistent departmental labor indicators that align with national benchmarks, rendering it suitable for longitudinal analysis and policy applications. All technical details can be found on (ref), which includes a description of the estimation procedure, an evaluation of model performance, relevant metrics, and a visual comparison of the primary labor market results against official annual data.
Informality presents a unique challenge due to the absence of official departmental series to serve as benchmarks. While employment, unemployment, and inactivity can be reconstructed by mapping city-level survey data to departments, informality is reported only for 23 metropolitan areas in the GEIH (Gran Encuesta Integrada de Hogares), with no official departmental totals. This lack of data necessitates the extrapolation of urban evidence to departments with significant rural sectors, where the dynamics of informality differ substantially. Consequently, this situation requires a more complex approach than a typical optimized machine learning model.
To address this gap, we implement a custom neural network framework specifically designed to capture the heterogeneous and non-linear drivers of informality. Standard methods, such as gradient boosting or baseline multilayer perceptrons, tend to regress toward the mean and perform poorly in regions of high informality (MAPE $>18\%$ on average). Our architecture introduces domain-informed inductive biases and robust training strategies, thereby enhancing extrapolation to departments lacking direct survey coverage.
A consistent national benchmark for the years 1993 to 2025 is established through four steps:
This process enforces coherence across different levels, prevents spurious breaks, and ensures additivity to the national employment base.
The departmental model is implemented as a custom multilayer perceptron (MLP) that incorporates residual connections, dropout, and optional monotonicity constraints. Key elements include:
Training employs early stopping, adaptive learning rates, and systematic input scaling. Predictions are limited to $[0,1]$, ensuring interpretability as rates.
Predicted departmental informality rates are converted into counts of informal workers by multiplying them with reconstructed employment levels. Monthly rescaling aligns departmental totals with the national benchmark, after which the rates are reinstated. This process ensures additivity and consistency across levels.
Validation is conducted using four protocols: (i) in-sample fit, (ii) temporal 80/20 split, (iii) Leave-One-Group-Out (LOGO), and (iv) Leave-$k$-Out (LKO). The custom neural network consistently outperforms simpler methods, particularly in high-informality regimes where baseline models exhibit failure. Results are summarized in Table (ref).
In conclusion, the neural network framework offers nationally consistent estimates of departmental informality. Comprehensive details, including the formal specification, MLP architecture, optimization strategies, calibration and validation protocols, and visual comparisons with official monthly data, are presented in (ref).
To assess departmental labor market performance, we construct a composite Employment Quality Index (EQI). The EQI consolidates multiple indicators into a standardized monthly metric that reflects both the level and structural quality of employment. A penalty adjustment is included to account for scenarios in which favorable outcomes, such as high employment rates, coexist with a significant degree of informality.
Each underlying indicator is smoothed to reduce seasonality and to emphasize long-term dynamics. It is then converted into percentile scores, with directionality defined by whether higher values indicate improvements (e.g., employment, participation) or deteriorations (e.g., unemployment, informality). Standardized indicators are aggregated using normalized weights, which are applied uniformly unless specified otherwise. Subsequently, departments are classified into ordered categories—Very Low, Low, Medium, High, and Very High—based on their average Employment Quality Index (EQI), resulting in a typology of labor market quality across regions.
To analyze dynamics, we compute periodic averages and departmental rankings, which are visualized through bump charts that illustrate improvement, stagnation, or decline over time. In addition to ranking, the EQI supports the segmentation of departments by clustering regions based on long-run average EQI, thereby creating comparable profiles of employment quality. This framework is used to contextualize trajectories and facilitate peer comparisons (see Table (ref)). Comprehensive methodological details, including indicator selection, weighting, smoothing procedures, and robustness checks, are provided in (ref).
The methodological framework is structured in seven interconnected stages: establishing a national baseline with international benchmarks and GEIH information; reconstructing city-level signals and extending them with historical domains; generating departmental series with demographic consistency and accounting identities; compiling a unified feature set of macroeconomic and institutional covariates; predicting departmental labor indicators with a supervised XGBoost model; estimating informality through a tailored neural network with residual connections and monotonicity constraints; and constructing the Employment Quality Index (EQI) to integrate level and quality dimensions of labor outcomes.
This dataset facilitates a systematic examination of national and subnational dynamics, encompassing patterns of convergence and divergence, as well as heterogeneous responses to significant shocks, such as the 1999 financial crisis, the 2008 global recession, and the COVID-19 pandemic. By employing methodologies such as trend decomposition, correlation analysis, Granger causality tests, composite index construction, and comparative visualization, the dataset elucidates the interplay among participation, employment, unemployment, inactivity, and informality over time and across regions. Beyond its descriptive capacity, the dataset establishes an empirical foundation for policy by highlighting persistent gaps in employment quality, identifying clusters of departments with shared trajectories, and supporting targeted interventions to reduce informality, enhance participation, and strengthen territorial resilience.
The reconstruction pipeline produces a harmonized and internally consistent set of labor market indicators that ensure comprehensive temporal and spatial coverage. The final outputs comprise:
These outputs provide a statistical foundation for labor analysis, territorial diagnostics, and the design of employment policies. Depending on the analytical horizon, we recommend three standard panels:
A design principle applicable to all products is demographic and accounting consistency. Official population projections serve as
This approach preserves additivity across both temporal and geographic dimensions while maintaining internal identities.
Unlike employment, unemployment, or inactivity, the informality rate lacks an official departmental counterpart. Consequently, departmental data are entirely extrapolated by the neural network. These series represent statistically consistent reconstructions rather than direct survey values.
Following the complete estimation and calibration pipeline, we evaluated accuracy by comparing the reconstructed annual departmental aggregates against the official GEIH results. This assessment captures the definitive reconstruction error, reflecting both the supervised predictions and the post-processing adjustments, which include the enforcement of accounting identities and national calibration.
Table (ref) presents the Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE) for each indicator. These metrics are expressed in absolute counts and averaged across all departments and years from 2007 to 2024.
The results indicate strong predictive accuracy. The overall Mean Absolute Percentage Error (MAPE) across all variables is 1.8%. Errors are lowest for the employed and economically active population (PEA), both below 1.3%. Slightly higher errors are observed in the inactive population (PET, 1.5%) and the total population (2.3%), reflecting the compounding effect of demographic inputs. Unemployment exhibits a MAPE of 2.2%, consistent with its smaller scale and greater volatility.
Table (ref) disaggregates MAPE by department and indicator. Performance is consistent across regions: in most departments, errors for core indicators remain below 3%, with employed, PEA, and PET typically ranging between 1.0% and 1.6%. The largest deviations are observed in Cundinamarca (up to 4.4%) and Quindío (2.9%), although these remain within acceptable thresholds. Overall, the pipeline demonstrates reliable spatial generalization and robust accuracy in subnational labor market estimation.
Performance metrics are computed solely for the 23 departments covered by the GEIH. For the remaining 10 departments, the lack of ground-truth values precludes direct error measurement. As a result, evaluation for these regions depends on structural plausibility, consistency with national totals, and the generalization properties of the model.
MAPE is emphasized as the primary metric because it provides scale-invariant comparisons across indicators and territories. Absolute deviations (e.g., $\pm$10,000 individuals) are less informative due to the heterogeneity in departmental population sizes, while percentage-based errors facilitate more interpretable assessments of relative accuracy.
The distribution of estimation errors is uniform across departments, highlighting the model's ability to generalize across heterogeneous regional labor markets. Most departments demonstrate Mean Absolute Percentage Errors (MAPEs) below 2.5% for the majority of indicators, with only a few instances exceeding 3%. This uniformity is significant considering the structural heterogeneity in demographic size, economic base, and labor informality across Colombian regions.
The estimation pipeline shows strong spatial adaptability. Predictive accuracy is preserved in both large and diverse departments, such as Antioquia, Santander, and Valle del Cauca, as well as in regions with more unstable dynamics, such as Chocó and Meta. This suggests that the model effectively captures both national and subnational structural patterns without overfitting to specific territories.
These results provide empirical validation of the modeling strategy. The low and stable MAPE values across departments and indicators suggest that the model captures structural relationships rather than merely reproducing noise or idiosyncratic fluctuations. The consistency of errors, even in departments with contrasting economic conditions, demonstrates the adaptability of the pipeline and its ability to generalize beyond the observed domains while utilizing all available data.
A detailed graphical comparison between annual observed and estimated values for each indicator is included in (ref) These plots confirm the alignment of the reconstructed series with GEIH aggregates across departments. The coherence of levels and trends supports the reliability of the estimation pipeline, following calibration and enforcement of identity. Each figure contrasts annual predictions with observed values, providing visual evidence of the performance.
To estimate monthly labor informality at the departmental level, we first reconstructed a continuous time series and benchmarked it against observed data from the 23 primary cities for the period 2007-2025. This benchmarking enables a direct evaluation of the model's ability to replicate observed labor market dynamics within the sample.
The in-sample performance across all cities indicates high accuracy. The RMSE of 0.0149 implies that deviations between predicted and observed informality rates average 1.49 percentage points on the original scale. The mean MAPE of 2% confirms that relative errors consistently remain low compared to the magnitude of the observed rates.
At the city level, MAPE values exhibit a narrow range around the 2% band. The smallest errors are observed in Santa Marta (1.40%), Florencia (1.66%), and Bucaramanga A.M., while the median across cities is 1.96%. Most results cluster near this median value, indicating a strong overall fit; however, a few cities, such as Bogotá D.C. and Barranquilla A.M., demonstrate slightly higher deviations that suggest potential for targeted refinements.
As in other labor market evaluation results, error metrics are computed solely for the 23 primary cities covered by the GEIH, which are the only domains with consistently high-frequency ground-truth data. For departments outside this coverage, direct error quantification is not feasible. In such cases, validation relies on indirect checks of structural plausibility, consistency with national and regional aggregates, and the expected generalization capacity inferred from the city-based sample.
The convergence of findings across these schemes supports the robustness of the estimation pipeline and its applicability in reconstructing monthly informality and labor market series beyond the directly observed domains. A visual comparison of observed and reconstructed informality rates across the 23 cities is presented in Figure (ref), thereby reinforcing the previously discussed quantitative results.
Table (ref) reports the validation metrics for the six core labor variables reconstructed through temporal disaggregation and multi-output regression. For these variables, the Mean Absolute Percentage Error (MAPE) remains below 2.3%, with the lowest relative errors observed for Participation in the Economically Active Population (1.2%) and employment (1.3%).The aggregate error across all variables results in a MAPE of 1.8%.
For the informality rate, estimated using the neural network approach, the in-sample evaluation over the same period yields an RMSE of 0.0149 (in proportion units), an MAE of 0.0111, and a MAPE of 2% (Table (ref)). These values are directly comparable to the relative error magnitudes obtained for the reconstructed variables, thereby reinforcing the internal consistency of the modeling framework.
The in-sample results and the complementary validation schemes provide strong and convergent evidence for the stability and generalization capacity of the proposed estimation framework. The consistency of error magnitudes across cities, variables, and test protocols indicates that predictive performance is not an artifact of a specific calibration subset but a consequence of structurally robust modeling choices.
This subsection presents monthly national estimates of key labor market aggregates and rates from 1993 to 2025. The series are obtained by aggregating the reconstructed departmental indicators and serve as the benchmark reference for the estimation framework. They integrate official population projections, GEIH microdata, and historical benchmarks from the World Bank and ILO, harmonized through supervised learning, temporal disaggregation, and accounting-consistent smoothing techniques.
Figure (ref) displays six national aggregates: working-age population (PET), total population, economically active population (PEA), employed individuals, unemployed individuals, and inactive individuals. PET and total population exhibit stable long-term growth, consistent with demographic projections. Employment and inactivity increase gradually, reflecting both the expansion of the labor market and evolving patterns of participation. Informal employment declines steadily until the late 2010s, while formal employment experiences growth.
The COVID-19 pandemic disrupted these trends. Both informal and formal employment fell sharply, followed by a quicker recovery of informal jobs, which is consistent with their greater flexibility and lower entry barriers. The PEA and employment series reflect these disruptions: a contraction in 2020, a surge in inactivity, and temporary exits from the labor force. Unemployment reached historically high levels before stabilizing around pre-pandemic values by 2022. Informality, which had been declining, rose temporarily during the recovery, highlighting the role of informal job creation in the economic rebound.
Figure (ref) presents the corresponding national rates. The participation rate increases from the mid-1990s to the early 2000s, stabilizing near 65% thereafter. This trend is consistent with higher education levels, increased female participation, and urbanization. The unemployment rate fluctuates between 9% and 15%, peaking during the 1999 financial crisis and the 2020 pandemic, followed by gradual recoveries. The informality rate declines from above 68% in the early 1990s to approximately 56% by 2024, exhibiting cyclical fluctuations and a rebound related to the pandemic. The inactivity rate mirrors participation and remains above 35%, highlighting persistent barriers to inclusion.
Overall, the national indicators demonstrate credible long-term dynamics, coherent responses to shocks, and consistency with demographic constraints and accounting identities. Informality has exhibited a sustained decline over three decades, interrupted only by significant crises, particularly the COVID-19 pandemic, after which the downward trend resumes. These patterns confirm the structural coherence of the reconstructed series and suggest persistent regional heterogeneity.
The table (ref) summarizes national monthly labor market statistics for the period 1993-2025. The employment rate ranges from 0.427 to 0.660, with a mean of 0.587 and a low coefficient of variation (CV) of 5.5%, indicating stable labor absorption over the long term. The unemployment rate spans from 0.071 to 0.216, averaging 0.120, but exhibits the highest variability (CV = 23.5%), confirming its sensitivity to cyclical shocks. The participation rate lies between 0.534 and 0.719, with a mean of 0.667 and minimal variation (CV = 4.2%). Inactivity ranges from 0.281 to 0.466, averaging 0.333 (CV = 8.3%), as expected due to its complementarity with participation. Informality fluctuates between 0.553 and 0.708, averaging 0.627, with low dispersion (CV = 6.9%). Overall, participation, employment, and informality remain structurally stable, while unemployment is the most volatile indicator.
Table (ref) compares the response of labor indicators to four major macroeconomic shocks occurring between 1993 and 2025, using the two years prior to each event as a baseline. This approach captures both the magnitude and direction of short-term disruptions, as well as structural shifts.
The 1999 financial crisis is notable for being the most significant contraction. In comparison to 1996-1997, participation and employment, accompanied by a sharp rise in unemployment and inactivity. In addition to the job losses associated with the banking collapse, voluntary withdrawals from the labor force exacerbated the downturn. Although informality increased only modestly, its already high baseline indicates a structural deterioration in job quality. The prolonged recovery further underscores the severity of the shock.
The 2008 global recession resulted in only mild adjustments. In comparison to the years 2006-2007, participation and employment increased slightly, unemployment remained stable, and inactivity decreased marginally. Informality also decreased, consistent with favorable external conditions and counter-cyclical policies. No evidence of long-term scarring has been observed.
The COVID-19 pandemic resulted in an abrupt and synchronized economic collapse. In comparison to the years 2018-2019, participation and employment declined sharply, unemployment surged, and inactivity increased. Informal employment rose after a decade of decline, highlighting the vulnerability of urban service-sector jobs to lockdowns and the limited capacity for telework. This episode combined immediate disruption with regressive structural effects.
The post-pandemic recovery (2021-2022 baseline) indicates partial normalization. Participation and employment rates have improved, unemployment has decreased and inactivity has declined. Informality has returned to pre-pandemic levels; however, persistently high rates in several metropolitan areas underscore the concentration of new jobs in lower-quality segments.
Taken together, these episodes reveal distinct dynamics: the 1999 crisis caused long-term scarring, the 2008 recession was absorbed with limited adjustment, the pandemic exposed systemic vulnerabilities, and the recovery underscored the gap between quantitative rebounds and qualitative improvements. These contrasts suggest the need for differentiated policy strategies: structural reconstruction in deep crises, countercyclical stabilization in moderate recessions, and hybrid approaches that combine emergency relief with long-term reform in response to systemic shocks.
The correlation matrix (Figure (ref)) illustrates the anticipated relationships among labor market indicators. The unemployment rate and employment rate are strongly negatively correlated (-0.63), while the employment rate and participation rate exhibit a high positive correlation (0.76). As expected, the inactivity rate and participation rate demonstrate a near-perfect negative correlation (-1.0) due to their complementary definitions. With respect to informality, positive correlation with unemployment (0.50) suggests that higher joblessness is associated with increased informal employment. Conversely, the negative association with inactivity (-0.54) may indicate lower informal participation in more inactive populations. The near-zero correlation with the employment rate (0.05) suggests that the overall employment level is not directly indicative of job formality.
Granger causality tests (Figure (ref)), utilizing pair-specific optimal lags selected based on the Bayesian Information Criterion (BIC) ($max~lag = 6$), indicate that unemployment and inactivity Granger-cause the informality rate ($p < 0.05$), whereas the employment rate does not Granger-cause it ($p = 0.430$). Conversely, the effect of informality on unemployment is not significant ($p = 0.190$). These results are consistent across pairs, suggesting that informality primarily responds to broader labor market conditions in the short run rather than exerting a causal influence on them.
This subsection presents the reconstructed annual trajectories of four key labor market indicators across Colombia's 33 departments: the unemployment rate, the labor force participation rate, the employment rate, and the inactivity rate. These indicators were estimated for the period 1993-2025 using a unified methodological framework that integrates supervised learning, temporal disaggregation, and demographic harmonization. The resulting panel adheres to labor accounting identities and facilitates detailed diagnostics of both cyclical dynamics and long-term structural asymmetries across regions.
Figure (ref) illustrates a diverse yet bounded distribution of unemployment rates over time. While national crises, such as the 1999 recession and the 2020 pandemic, generate noticeable transitory spikes in many departments, these events do not induce persistent unemployment divergence across the territory. Most departments exhibit unemployment rates fluctuating within a 7--14% range, with no evident outliers exceeding 20% in recent years. Departments previously highlighted, such as Guainía, Arauca, and Chocó, no longer systematically stand out; instead, transient peaks appear in various locations and dissipate over time. This observation suggests that the reconstructed series smooths local volatility while preserving the spatial-temporal imprint of systemic shocks.
The participation rate (Figure (ref)) displays moderate heterogeneity across departments but converges toward a central range between 55% and 70% throughout the examined period. Contrary to prior assumptions of extreme disengagement in Amazonian or frontier regions, the updated estimates reveal more homogeneous dynamics. Departments such as Cesar, Chocó, and Caquetá exhibit gradual long-term increases in participation, with noticeable boosts following 2020. These shifts may indicate improved demographic integration, post-pandemic recovery in institutional reporting, or structural transitions in the incorporation of informal labor. Furthermore, no department consistently exceeds 75% or falls below 50%, suggesting relatively contained behavior across regions.
As expected, the employment rate (Figure (ref)) largely aligns with the participation rate, given that unemployment remains stable across regions. The estimates confirm consistent labor absorption dynamics, with rates ranging from 45% to 65% in most departments. No regions exhibit structural employment collapses; instead, employment variation follows cyclical and symmetric patterns over time and space. Economically strong departments, such as Antioquia and Santander, display moderate rates, consistent with sectoral structures that are more stable but less intensive in informal or precarious labor.
Figure (ref) completes the labor identity through the lens of inactivity. The pattern aligns with participation and employment trajectories, with rates typically ranging between 30% and 45%. No department exhibits chronic detachment from the labor market. In contrast to prior interpretations that emphasized structural exclusion in regions such as Huila, Risaralda, or Sucre, the updated estimates suggest progressive reductions in inactivity and a more evenly distributed rebound following the COVID-19 shock.
Figure (ref) provides a long-term view of informality across Colombian departments. The heatmap highlights marked and persistent regional disparities. Departments such as Bogotá D.C., Valle del Cauca, Santander, and Cundinamarca consistently exhibit lower levels of informality, generally between 50% and 65%. In contrast, territories including Chocó, Vichada, La Guajira, and several Amazonian departments exhibit structurally higher informality rates, remaining above 70% throughout the period. This reflects persistent territorial gaps rather than uniform convergence. The COVID-19 shock in 2020 is evident across most departments, followed by partial recoveries at varying speeds. Some regions recovered quickly to pre-pandemic levels, while others continued to exhibit elevated informality rates, underscoring the asymmetric nature of labor market adjustments.
In summary, the departmental trajectories from 1993 to 2025 reveal a labor market that, while sensitive to macroeconomic shocks, maintains a high degree of structural cohesion across regions. Unemployment fluctuations are transitory and spatially diffuse; participation and employment rates converge toward relatively narrow ranges, and inactivity shows a gradual decline with only temporary reversals during crises. Informality, although still heterogeneous, follows a long-term downward trend, exhibiting convergence across most departments with a contained rebound after COVID-19. These patterns suggest that the reconstructed series capture both the cyclical synchronization and the slow-moving structural transformations of Colombia’s labor market, providing a robust empirical foundation for diagnosing territorial asymmetries, evaluating policy impacts, and anticipating differential regional responses to future shocks.
The trajectories of the EQI reveal persistent heterogeneity across Colombian departments (Figure (ref)). Departments such as Bogotá D.C., Santander, and Antioquia consistently achieve high scores, while Caquetá, Sucre, and Vaupés remain at the lower end of the distribution. This stratification reflects structural disparities that are not easily offset by short-term fluctuations. The penalty mechanism built into the EQI further emphasizes cases where high employment rates coexist with high informality, thereby preventing inflated assessments of fragile labor structures. Over time, the evidence indicates limited convergence: the gap between the best- and worst-performing departments has tended to persist or widen.
Relative rankings across three subperiods (1993-2003, 2004-2014, and 2015-2025) confirm these patterns (Figure (ref)). Bogotá D.C. and Santander retain their positions as leaders throughout, with Cundinamarca consolidating its place in the top group more recently. Conversely, Chocó, Caquetá, Arauca, and Sucre consistently occupy the lowest ranks. Some mobility is observed: Nariño and Risaralda made improvements in the most recent decade, whereas Norte de Santander and Huila declined. These movements suggest that structural disadvantages are resilient but not immutable, indicating that enhancements in labor market institutions and conditions can alter trajectories over time.
Table (ref) further illustrates these disparities. The Muy bajo group, composed of seven departments, exhibits the weakest labor market conditions, characterized by persistent informality and fragile employment structures. The Bajo group, with six departments, encounters significant but less acute deficits in formality and stability. The Medio group reflects intermediate outcomes, with some departments demonstrating potential for upward mobility. The Alto group, also consisting of six departments, achieves comparatively favorable results, marked by stronger participation rates and lower informality levels. Finally, the Muy alto group, encompassing seven departments, including the largest urban economies, represents the core of high-quality employment in Colombia.
Overall, the evidence reveals a polarized pattern: a set of departments consistently positioned at the top, a group persistently at the bottom, and a middle segment with limited capacity for mobility. This persistence underscores the structural and territorial nature of labor market inequalities in Colombia.
\paragraph{National level.} The monthly panel facilitates the monitoring of both structural and cyclical labor market dynamics over more than three decades, providing a consistent benchmark for assessing countercyclical responses to crises such as those in 1999, 2008, and during COVID-19. Results indicate that informality tends to rise alongside unemployment and inactivity, highlighting the importance of integrated strategies that promote job creation while incentivizing formalization. The demographic alignment of the series ensures that labor market projections remain consistent with population aging and demographic transitions, which is essential for pension sustainability, education planning, and social protection systems.
\paragraph{Regional level.} The Employment Quality Index (EQI) reveals persistent heterogeneity across departments. High-performing regions (e.g., Bogotá, Antioquia, Santander, Valle del Cauca, Cundinamarca) can serve as benchmarks for the diffusion of best practices, while structurally weaker territories (e.g., Chocó, Vaupés, Arauca, Sucre, Putumayo) require targeted measures. These measures may include fiscal incentives to promote formal employment, stronger enforcement of labor regulations, and support for productive sectors with greater potential to generate stable jobs. The evidence also highlights differentiated exposure to shocks: while some regions exhibit rapid recovery, others remain persistently vulnerable. This suggests the need for tailored stabilization mechanisms, such as localized unemployment insurance schemes or emergency employment programs.
\paragraph{Institutional level.} The methodological framework establishes the foundation for an early-warning system in labor markets. By integrating national benchmarks, survey microdata, and machine learning-based extrapolations, this approach generates consistent estimates even in contexts with limited survey coverage. Rather than replacing primary data collection, it complements it by reducing delays and ensuring the timely availability of policy-relevant information. The transparency and replicability of the pipeline also create opportunities for adoption by statistical agencies across Latin America that face similar data constraints.
In summary, the findings emphasize the need to move beyond national averages to tackle territorial disparities in labor market outcomes. The reconstructed panel provides policymakers with the ability to identify structural deficits, design region-specific interventions, and evaluate their effectiveness with a temporal and spatial resolution that was previously unavailable.
The reconstructed labor series provide a coherent and demographically aligned panel of monthly indicators for all 33 Colombian departments from 1993 to 2025. Accounting identities are enforced, estimates are calibrated to national and departmental benchmarks, and validation confirms robust predictive performance, particularly in departments with GEIH microdata. Nevertheless, several limitations remain:
This paper develops the first harmonized and territorially exhaustive panel of monthly labor indicators for Colombia's 33 departments, covering the period from 1993 to 2025. The reconstruction pipeline integrates temporal disaggregation, supervised learning, deep neural architectures, retro-polarization, and demographic calibration to produce consistent series of employment, unemployment, participation, inactivity, informality, and population aggregates. The result is a high-frequency dataset with national errors below 2.5% MAPE, while maintaining demographic and labor accounting identities.
Three central findings emerge. First, supervised and deep learning models, when carefully regularized and anchored to structural proxies, can extrapolate labor indicators into territories lacking direct survey coverage. This is accomplished by leveraging signals from metropolitan markets and departmental aggregates. The neural network model demonstrated particular effectiveness in capturing heterogeneous informality dynamics that conventional econometric approaches could not replicate. This underscores the potential of modern learning methods to enhance labor statistics in contexts with incomplete territorial data.
Second, enforcing demographic and labor identities ex-post ensures numerical coherence and flexibility; however, these constraints are not inherently learned by the models. While this choice enhances predictive accuracy, it diminishes structural interpretability. Future work should explore architectures that embed these identities directly, allowing them to emerge endogenously during training.
Third, although the overall reconstruction is accurate, performance varies across space and time. Peripheral or resource-dependent departments exhibit larger errors and slower recovery following crises, whereas diversified urban economies converge more rapidly to national averages. These differences underscore the necessity of enhancing primary data collection in marginalized areas and formulating differentiated territorial policies.
Beyond technical performance, the reconstructed panel offers empirical leverage for both research and policymaking. At the national level, it reflects Colombia's labor dynamics during significant shocks, including the 1999 financial crisis, the 2008-2009 global recession, and the COVID-19 pandemic, while maintaining internal balance. At the subnational level, it reveals persistent asymmetries: high informality and low participation in peripheral regions; low informality and high participation in diversified urban economies; and intermediate groups demonstrating mixed outcomes. Recovery trajectories also vary, with formalized regions recovering more rapidly from crises.
Methodologically, the modular pipeline provides a replicable framework for countries with urban-biased labor surveys or incomplete territorial coverage. This approach can be expanded to reconstruct additional indicators (e.g., underemployment, hours worked, or formality), incorporate sectoral and occupational decompositions using CIIU-linked auxiliary data, and adopt Bayesian or multivariate architectures to enhance uncertainty quantification. Integrating labor accounting identities directly into neural architectures represents a key frontier for improving both accuracy and interpretability.
In summary, the reconstructed panel offers a novel statistical framework for analyzing regional convergence, resilience to shocks, and informality gaps in Colombia. Although it does not serve as a replacement for direct statistical production, it complements official sources by addressing territorial and temporal gaps, facilitating retrospective diagnostics, and informing policy strategies for more inclusive labor markets.
This study relies exclusively on aggregated, anonymized, and publicly available data sourced from official statistical agencies, including the Departamento Administrativo Nacional de Estadística (DANE), the World Bank, and the International Labour Organization (ILO). At no stage were individual-level, confidential, or personally identifiable records accessed, stored, or processed. No ethical approval was required due to the non-sensitive nature of the data. The study adheres to the principles of research transparency and reproducibility by documenting all methodological steps and ensuring the open availability of derived results.
The author acknowledges the institutional efforts of DANE, ILOSTAT, and the World Bank in providing unrestricted access to labor market and demographic statistics. The views, interpretations, and conclusions expressed in this study are solely those of the author.
This research did not receive funding from public agencies, private institutions, or nonprofit organizations. The author declares no competing interests, whether financial, professional, or personal.
All reconstructed labor market indicators generated in this study, including monthly national and departmental series from 1993 to 2025, are publicly available at: \url{https://doi.org/10.5281/zenodo.16899686}
The repository contains machine-readable data files, methodological documentation, and version-controlled metadata. All input data used in the reconstruction process, including GEIH labor annexes, World Bank indicators, and ILOSTAT aggregates, are publicly accessible through their respective institutional repositories.
All source code utilized for data processing, reconstruction, and visualization is publicly available in the GitHub repository archived on Zenodo: \url{https://doi.org/10.5281/zenodo.16899686} (also accessible at \url{https://github.com/jaimevera1107/colombian-labor-market})
The codebase is version-controlled, thoroughly documented, and comprises reproducible scripts necessary for replicating all analyses and figures presented in this paper.