Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,721 characters · 4 sections · 46 citation commands
Enhancing the Accuracy of Regional Input--Output Table Estimation: A Deep Learning Approach
{ Keywords \ Deep learning, regional input--output table, machine learning, data augmentation}
In the context of a macroeconomic quantitative analysis of regional economies, input--output tables represent a fundamental dataset. The creation of input--output tables requires substantial data sources. Regions with limited data sources face considerable challenges in creating these tables. Therefore, much research has focused on developing and improving estimation methods for regional input--output tables.
An input--output table records the flow of products within a target region from two perspectives: producers (input) and consumers (demand). In general, input--output tables are compiled at the national level. Those compiled for domestic regions are referred to as regional input--output tables.
Regional input--output tables are used in various analyses focusing on regional economies. Recently, these tables have frequently been utilized to analyze the effects of environmental policies and disasters. Okuyama2007 reviewed past studies that analyzed the impact of disasters on regional economies. Some of these studies used regional input--output tables. Koks2016 evaluated the impact of disasters on regions using various methods based on regional input--output tables, and compared these approaches. Kjaer2015 calculated the environmental footprint using Denmark's environmental input--output table.
However, regional input--output tables are rarely compiled and published for small domestic regions. In Japan, for example, the national input--output table is compiled based on primary statistics. \footnote{Within the administrative divisions of Japan, the country is comprised of 47 prefectures, each of which contains municipalities, such as cities, towns, and villages.} Prefectural input--output tables are estimated by combining available partial primary data with data from original surveys. On the other hand, at the municipal level, the availability of primary statistics is even more limited. Consequently, very few municipalities publish their own tables.
In light of this, numerous methods have been developed for estimating regional input--output tables. Particular focus has been placed on a series of methods known as non-survey methods. They estimate input--output tables from published statistics without conducting original surveys in the target region.
Most non-survey methods are designed to estimate input coefficients. An input coefficient is defined as the ratio of intermediate input to the gross output of an input industry. Input coefficients play a crucial role in obtaining the entire input--output table and in calculating economic effects.
The primary non-survey methods employed are the location quotient (LQ) method and the RAS method. These methods require less data and therefore facilitate the estimation of input coefficients, even in small regions where primary statistics are limited.
The ratio of the industry's share of economic activity in the target economy to its share in a reference economy is denoted as the location quotient Isserman1977. The LQ method estimates input coefficients for a target region based on the input coefficients of a reference region and the location quotient. Derivatives include CILQ, RLQ, and FLQ Flegg2016. Moreover, new derivatives have recently been developed, such as FLQ+ and 2DLQPereiraLopez2020, Flegg2021. Studies such as Bonfiglio2008, Flegg2012, and Lamonica2018 have verified the accuracy of these derivatives.
The RAS method is regarded as the technique initially introduced in Stone's researchBacharach1970. It estimates the input coefficient matrix for a target region by adjusting the initial matrix using the intermediate demands, intermediate inputs, and gross outputs by industry in that region. Hewings1977 demonstrated that the RAS method can estimate input coefficients for the target region with a high degree of precision. Similar to the LQ method, numerous derivatives and improvements have been developed for the RAS method Junius2003,Lenzen2007,Lemelin2009,Lenzen2014,Temursho2021. Furthermore, studies such as Hiramatsu2016 and Holy2023 have sought to improve the RAS method focusing on regional input coefficients.
However, these major non-survey methods require strong assumptions about economic activity, and their validity has been the subject of prolonged debateRound1983. According to Riddington2006, the LQ method requires that productivity and consumption per employee must be identical across all regions and that there must be no cross-hauling of products within the same industry. For the RAS method, Miller2022 as noted, differences in production technology between the referenced and target regions must be decomposed into substitution and fabrication effects. \footnote{Regarding the RAS method, some view it merely as a calculation to adjust the matrix to satisfy constraintsMiller2022.}
For these non-survey methods, as less data is required, the information is also selective. Consequently, the accuracy may be lower than when using more information.
Furthermore, both the LQ and RAS methods necessitate additional data. In the LQ method, the referenced region and economic activity must be selected when calculating the location quotient. The RAS method requires selecting the initial input coefficient matrix and specifying intermediate demands, intermediate inputs, and gross outputs by industry in the target region. These data selection and specification influence the estimation accuracy. For example, Fukui2025 demonstrated that the estimation accuracy of input coefficients varies depending on the reference region in FLQ and the initial value setting of the input coefficient matrix in RAS.
In recent years, deep learning has been utilized in many fields. A neural network is a function constituted by aligning simple-shaped functions, known as activation functions, and connecting them in layers. Moreover, a neural network with a large number of layers and structures is called a deep neural network. Due to its structure, a deep neural network can precisely capture nonlinear relationships between variables, making it particularly useful for predicting target variables with high accuracy. Training deep neural networks with large datasets is referred to as deep learning. Deep learning serves as the foundational framework for current generative artificial intelligence (AI). It has also frequently been utilized to predict and evaluate asset pricesDing2020, Sriviney2022, Chen2023.
Some studies have sought to incorporate neural networks into the estimation of input coefficients. Papadas2002 predicted input coefficients of the United Kingdom using a neural network. Pakizeh2022 estimated input coefficients of nine regions within Japan using machine learning methods, including neural networks. Both studies attempted to ensure sufficient data size by applying a single neural network model to all input coefficients. Nevertheless, it can be inferred that deep neural networks were not applied due to the insufficient size of the data. Training deep neural networks with small data may result in overfitting, which reduces prediction accuracy. In such cases, it is imperative to design the neural network layers to be relatively shallow. Consequently, it appears that the above studies could not achieve a level of prediction accuracy that clearly surpasses that of conventional methods.
Fukui2025 developed a method for estimating input coefficients with deep learning. This method addressed overfitting through a data augmentation technique called mixup, as described in Zhang2018. Compared to conventional methods, the method proposed by Fukui2025 had the advantage of not necessitating assumptions about economic activity. The method achieved higher precision than conventional methods for estimating input coefficients in the input--output table for Japan. Moreover, as no additional estimation or selection of other data was required, the accuracy of the method was stable.
While a number of methods have been developed to estimate input coefficients (i.e., the intermediate inputs section of an input--output table), few studies have presented standardized methods to estimate other sections, such as final demand, gross value added, and gross output. Simple estimation methods for these sections involve using the ratio of macroeconomic variables between the target and referenced regions. For instance, when estimating gross outputs by industry in a region, one might calculate the ratio of workers in each industry in the region compared to the country. Then, these ratios are multiplied by the gross outputs by industry in the country to obtain the estimatesKronenberg2009. While this method is straightforward, it may not accurately reflect the region's or industry's actual conditions because it disregards information beyond the calculated ratios. Alternatively, there are methods that estimate the items in the sections of final demand, gross value added, and gross output based on data available for the target regionHosoe2014. These methods can better reflect regional characteristics than the simpler approach. However, the feasibility of these methods depends not only on data availability but also on the researcher's knowledge and experience.
Given the current state of regional input--output table estimation, this study proposes a novel, deep learning-based method. This method extends the approach in Fukui2025 to encompass items beyond input coefficients, including final demand, gross value added, and gross output. The generation process for each item in these sections is approximated by a deep neural network that takes various regional economic data as inputs. We expect that this approach will yield highly accurate and stable estimates for items in the input--output table.
Section 2 explains the estimation method for regional input--output tables using a deep learning. Section 3 estimates the 2015 input--output table for Japan and verifies the results. Section 4 discusses the properties and limitations of the proposed method.
This section presents the estimation method for the 2015 input--output tables, which cover Japan as a whole and the individual cities. \footnote{The input--output tables for regions in Japan are published approximately every five years with a delay of about five years. As of this writing, only a few regions had published their 2020 tables. Therefore, deriving a high-precision estimation model for the tables remained challenging, even with data augmentation. Consequently, this article focuses on the 2015 tables.} The data used to train the model were derived from the 45 prefectures and four cities of Japan, as shown in Table (ref). \footnote{Among the 47 prefectures of Japan, Tokyo includes the headquarters sector in its input--output tables, and Okinawa uses its own industry classification. The format of their tables differs from that of the other 45 prefectures. Therefore, these two regions were excluded from the subsequent analysis.}
The input--output table employed in this study was of the competitive import type, and the industries were reclassified as shown in Table (ref). Final demand and gross value added were classified as in Tables (ref) and (ref), respectively.
In our method, net exports were the target for estimation. Net exports are defined as the difference between exports and imports. As noted in the appendix of Fukui2025, in a competitive import type input--output table, the sum of exports from two regions does not equal the exports when those regions are combined into one. The same applies to imports. This property does not satisfy the criteria for the data augmentation in the method, a topic that will be elaborated on later. Thus, exports and imports are not suitable for estimation by our method. Conversely, the criteria holds for net exports, rendering them suitable for our model.
In the estimation of input--output tables for cities, the sum of the net exports and the net outflow to other regions within the country was estimated. Net outflow is defined as the difference between outflow and inflow of the city.
The data used for this study are displayed in Table (ref). In Table (ref), the subscripts $i$ and $j$ represent the order of the industries in Table (ref). The subscripts $g$ and $h$ denote the order of final demand in Table (ref) and gross value added in Table (ref), respectively. Furthermore, $k$ indicate regions, and $c$ is used to denote industry classifications in the Economic Census. For minor industry classifications, $c = 1, \ldots, 619$ and for major classifications, $c = 1, \ldots, 17$.
Data augmentation was performed based on the data in Table (ref) , which was collected from the regions in Table (ref), in order to generate data for model training. Using these data of the regions directly for deep learning could lead to overfitting and a subsequent decline in prediction accuracy due to the small data size. To address this issue of overfitting, Fukui2025 employed a data augmentation technique known as “mixup”, which was proposed by Zhang2018. In applying this technique, Fukui2025 posited the following assumptions: For each quantitative variable underlying the model variables, the value measured by combining multiple regions equals the sum of the values across those regions, and the value measured by scaling a region equals the scaled value of the region. These assumptions were used to establish the prior knowledge that “for a virtual region obtained by linear interpolation (consisting of the above combination and scaling), the feature vectors should lead to an associated targetFukui2025.” Given this prior knowledge, we generated data for virtual regions through linear combinations of data vectors from multiple regions, thereby increasing the amount of data.
In this study as well, this data augmentation was implemented to mitigate the risk of overfitting. During the augmentation process, two to five regions were randomly selected from those in Table (ref). The weights for the linear combination of these regions were generated using random numbers from a Dirichlet distribution with all elements of the parameter vector set to one. \footnote{It is important to note that regions in an inclusive relationship, such as Hokkaido and Sapporo City, could not be selected concurrently. This is because interpreting their combination is extremely difficult. }
Moreover, this study introduced a method described in Fukui2025 to address the decline in prediction accuracy caused by the difference in scale between the trained and predicted regions. Since most of the observations were prefectures, the data augmented by the above method notably reflected the scale of prefectures. Consequently, neural network models trained using the generated data showed reduced accuracy for our prediction targets, such as the entire country or municipalities. To address this issue, the quantitative data of the prefectures was converted into per capita data for the population aged 15 and over in 2015, and the data was augmented. This augmented quantitative data was then multiplied by such population of each target area to yield training data reflecting the scale of the prediction targets.
To obtain the training data for Japan as a whole, the augmented data was multiplied by the population aged 15 and over for Japan in 2015. To augment data on cities, we first generated a set of uniform random numbers within the range of the minimum and maximum values of the population aged 15 and over across all cities in Japan in 2015. Next, we multiplied these random numbers by each observation in the augmented data to obtain the city training data.
The values of the explanatory and target variables were calculated from the augmented data. As shown in Table (ref), the explanatory variables were derived from the quantitative variables presented in Table (ref). The indices of the explanatory variables are identical to those in Table (ref).
In order to eliminate the potential impact of regional economic scale, the ratio to the total of gross outputs for each item in the input--output table was used as the target variable. Given that economic scale positively correlates with each item in the input--output table and with multiple quantitative variables included in the explanatory variables, it can be inferred that a similar correlation exists between the items in the input--output tables and the explanatory variables. When this correlation is pronounced, the trained model reflects of the relationship notably. To mitigate the influence of this correlation, the following variables were used as the target variables:
No model was set for the target variables for which all values in the regions of Table (ref) were zero, and their predicted values were also set to zero.
Consequently, estimating total of gross outputs is required separately for the method of this study.
Prior to model training, 50,000 samples were generated using the aforementioned data augmentation. Of these, 80% (40,000 samples) were randomly selected as the training data, and the remaining 20% (10,000 samples) served as the test data. \footnote{Our model showed almost no difference in estimation accuracy compared to the model with data augmented to 100,000 samples. However, considering the computational time required, we used data augmented to 50,000 samples.} We estimated the model parameters using the training data. Subsequently, predictions were made on the test data, and the model was verified by calculating the prediction error.
Figure (ref) illustrates the model for each item in an input--output table. Prior to the calculation in the model, the explanatory variables in Table (ref) were standardized and converted to principal component scores. These scores were processed through a rectified linear unit (ReLU) layer comprising 512 nodes, followed by a stack of ten layers. Each layer in the stack comprised a ReLU layer (512 nodes), a skip connection, and a batch normalization layer. The value calculated from the stack was then used by the output layer to generate the target variable value. \footnote{ The following activation function is called rectified linear unit (ReLU): \[ y = \max (0, a + \boldsymbol{x}' \boldsymbol{b}), \] $\boldsymbol{x}$ is the input vector to ReLU, $y$ is the output from ReLU, $a$ is the bias, and $\boldsymbol{b}$ is the weight for $\boldsymbol{x}$. }
In a batch normalization layer, the inputs are normalized with each training data batch, followed by a linear transformationIoffe2015. In a skip connection, the intermediate layer explains the differences between the inputs and outputsHe2016. Both batch normalization and skip connections are said to improve the estimation accuracy of model parameters. Furthermore, dropout layers were positioned after the fourth and ninth layers in middle layer to prevent overfitting. A dropout layer is constituted by nodes that discard their inputs with a specified probabilitySrivastava2014.
Different output layers were configured depending on the target variable. Assuming the values of each $y_{i,k}$ and $a_{i,j,k}$ are non-negative, a sigmoid function was established for their output layer. \footnote{The sigmoid function is equivalent to the logistic function in logistic regression.} However, when the values of $y_{i,k}$ and $a_{i,j,k}$ were too small, optimization processes were susceptible to failure due to gradient vanishing in the output layer. To circumvent the problem, we adopted a target value transformation similar to that demonstrated in Fukui2025. The following transformed variable $y^*$ for each of the variables $y_{i,k}$ and $a_{i,j,k}$ was designated as the target variables:
In this equation, $y$ represents the each variable of the $y_{i,k}$ and $a_{i,j,k}$. Moreover, $\min(y)$ and $\max(y)$ denote the minimum and maximum values of $y$, respectively. On the other hand, given that each of $d_{i,g}$ and $v_{i,h}$ can be both positive and negative, an identity function was implemented as the output layer. The target variable in the output layer was set to the the standardized variable of each $d_{i,g}$ and $v_{i,h}$.
These transformations for the model variables were performed on the training data. The values employed for the transformation (means, standard deviations, minimum and maximum values, principal component loadings, and so on) were used in the prediction phase.
The training was carried out in the following settings:\footnote{For more information on these settings, refer to Geron2019_1.} The mean squared error was configured as the loss function. The loss function incorporated an $L_1$ regularization term with a parameter of $10^{-5}$. The optimization method employed was stochastic gradient descent with a mini-batch size of 32. Additionally, for parameter updates in the stochastic gradient descent, Nesterov's accelerated gradient method was applied with a parameter of $0.9$. The learning rate was updated exponentially and periodically according to the method proposed by Smith2017. The initial value and lower bound was $10^{-6}$, the upper bound was $0.01$, and the step size between the upper and lower bounds was 10. \footnote{The learning rate corresponds to the step size in the quasi-Newton method.} The updates to the learning rate and parameter values were performed simultaneously. The maximum number of epochs for the training was set to 200, with early stopping. First, 20% of the training data was randomly selected for validation during the model training. Then, the training process was terminated when the error on the validation data (validation error) increased for 10 consecutive epochs. Finally, the parameter values immediately preceding the increase in the validation error were adopted as the training result.
Subsequently, the trained neural network was used to predict each element of the regional input--output table. For the region under prediction, the explanatory variables were standardized, and their principal component scores were calculated. This calculation employed the mean, standard deviation, and principal component loadings recorded for each explanatory variable before training. The principal component scores were fed into the trained neural network. The output of this process was each transformed element of the regional table. As previously mentioned, given that the outputs were also transformed, an inverse transformation was performed using the means, standard deviations, and the values of $y_l$ and $y_u$ of the target variables used for the transformation prior to training. For each variable $y_{i,k}$ and $a_{i,j,k}$, the predicted value $y$ was obtained by performing the inverse transformation of Eq. ((ref)) to the output $y^*$ from the model as follows, using the values of $y_l$ and $y_u$ of the training data: \[ y = y_l + y^* (y_u - y_l). \] For each $d_{i,g}$ and $v_{i,h}$, the inverse transformation from the model output $y^*$ to the predicted value $y$ was as follows: \[ y = \mu_y + s_y y^*. \] Here, $\mu_y$ and $s_y$ are the mean and standard deviation of the training data, respectively.
In order to estimate the regional input--output table using our trained model, the outputs from the model must be multiplied by $\sum_i Y_{i,k}$.
The gross outputs ($Y_{i,k}$) were predicted for each industry individually. Consequently, the sum of these predicted outputs was not guaranteed to match the actual total of gross outputs. To ensure that, the following processing was applied to each $Y_{i,k}$ predicted by the model. The processed value $Y^*_{i,k}$ was then used as the predicted gross output. \[ Y^*_{i,k} = \frac{\tilde{Y}_{i,k}}{\sum_i \tilde{Y}_{i,k}} \sum_i Y_{i,k} \] where $\tilde{Y}_{i,k}$ is the gross output for industry $i$ from the model for the region $k$.
A matrix balancing procedure was subsequently implemented to ensure that the predicted values satisfied the input--output table constraints. The RAS method, a non-survey technique, is frequently employed for matrix balancing. However, it should be noted that the RAS method is not directly applicable to matrices that contain negative values. To address this issue, Junius2003 proposed the generalized RAS (GRAS) method. The GRAS method is a matrix balancing technique based on cross-entropy maximization. \footnote{McDougall1999 noted, in situations where the RAS method could be applied, it was equivalent to the cross-entropy maximization method.} In this study, cross-entropy maximization based on GRAS was performed for matrices containing negative values, subject to the following two conditions. First, the row sums and column sums of the input--output table were equal to the gross outputs by industry. Second, the total of consumption expenditures outside households in the gross value added and final demand sectors were equal to each other. \footnote{The appendix provides details regarding matrix balancing in this study.}
Under the configurations described in the previous section, we estimated the 2015 input--output table for Japan, and the accuracy of our method was verified. \footnote{Various computations related to deep learning were performed using Python and the PyTorch library. Data augmentation was implemented using the F\# programming language.} The input--output tables for Japan that are available to the public are derived using a survey method, and their high levels of accuracy make them suitable as benchmarks.
The total of gross outputs required in the estimation process used the actual values of the 2015 input--output tables. The model input consisted of the top 60 principal component scores with the highest contribution.
No large difference in prediction error was observed between the training and test data for the trained model. Therefore, each model was considered valid.
Table (ref) presents various indicators of the prediction error of the input--output table for Japan estimated with the method of this study. These indicators were selected based on Hosoe2014 and calculated according to the following definitions: \footnote{In the definition used in Hosoe2014, the denominator of the STPE was $\sum X_{i}$. This definition could lead to overestimation when $X_i$ assumes both positive and negative values. Accordingly, this study designated the denominator as $\sum \left| X \right|_{i}$.}
For any indicator, a smaller value indicates a smaller prediction error. In these definitions, $i$ represents index for each item in the input--output table, $X_{i}$ is the actual value for $i$, $\tilde{X}_{i}$ is the corresponding estimated value. $N_1$ denotes the number of items where $X_{i} \neq 0 \ \text{and} \ \tilde{X}_{i} \neq 0$, and $N_2$ is the total number of items whose actual values are neither zero nor empty in the actual input--output table. \footnote{As items with both actual and estimated values of zero contribute to an underestimation of the MAD and RMSE, these items were excluded from the calculation of these indicators. Furthermore, several items had zero actual values in the national input--output table but not in the prefectural tables. Our method estimated these items as well, and their estimates were almost never zero. Due to the inability to calculate the error rates for these items, they were excluded from the MAPE calculation.}
Figures (ref) and (ref) are heatmaps representing the levels of error and error rates, respectively, for each item in the input--output table for Japan. These figures follows the structure of the input--output table. The labels “I1”, $\cdots$, “I12” correspond to the order in Table (ref), “D1”, $\cdots$, “D6” correspond to the order in Table (ref), “V1”, $\cdots$, “V6” correspond to the order in Table (ref). Moreover, the label “Y” denotes gross outputs by industry.
Figure (ref) reveals a positive correlation between the magnitude of error and the gross output of the corresponding industry. Industries with smaller gross output, such as agriculture, forestry, fisheries, and mining, tend to have smaller errors. In Figure (ref), the error rates are generally high for items of final demand, particularly net exports.
Furthermore, Figure (ref) shows a bar plot of the error rates. The figure indicates that the prediction error rates for 159 items fall within the range of [-10%, 10%), which corresponds to approximately 55% of all items (including empty items, for a total of 288).
We compared the accuracy of the proposed method with that of conventional non-survey methods. As Morrison1974 stated, the RAS method is usually more accurate than the LQ method when the actual values for row and column totals are available. However, input--output tables generally contain negative values. This makes it inappropriate to use the RAS method directly.
Therefore, as a conventional method, we adopted matrix balancing via cross-entropy maximization utilized also in our method. In this matrix balancing, the actual values of gross outputs by industry of Japan in 2015 were used as the row and column totals. Moreover, the 2011 input--output table for Japan and the 2015 input--output tables for each prefecture were employed as the initial values for the matrix balancing.
Table (ref) shows the estimation errors associated with the conventional methods. Using the 2011 table for Japan as the initial value yielded smaller errors than using the 2015 tables for prefectures. Compared to these errors, our method generally produced even smaller errors, except that the STPE and MAD were slightly larger than the errors of the estimates based on the 2011 table for Japan.
Table (ref) shows the estimation errors for each sector using the conventional method. This initial value of the method was the 2011 input--output table for Japan. Compared to Table (ref), the errors for intermediate inputs, final demand (excluding MAPE), and gross value added were larger than the errors from our method. However, for net exports in the final demand sector, the errors were smaller than those produced by our method.
The overall errors were closer between the two methods compared to the errors by sector. This was due to the presence or absence of estimation of gross outputs by industry. In our method, estimation errors occurred for these gross outputs because they were estimation targets. In contrast, the conventional method used actual measured values for these gross outputs. Consequently, as shown in Table (ref), the errors were zero. It is plausible that the discrepancies in the estimation of gross outputs by industry affected the overall error. Excluding them, the STPE and MAD were $0.0743$ and $481,307.12$, respectively, using our method, and $0.0991$ and $640,813.13$, respectively, using the conventional method. Our method yielded smaller values for other indicators as well.
Using the method developed in this study, we estimated input--output tables for each of the three cities in Japan: Sapporo, Gujo, and Okayama. The estimation employed the published gross production values for each city, and the top 100 principal component scores with the highest contribution were selected as the inputs for the model. Subsequently, we compared the input--output tables published by these three cities with the estimates.
The estimates derived from our method differed relatively large from the published values of the three cities for certain items. The sub-tables in Table (ref) show the differences between the estimated and published values for the three cities, as measured by STPE, MAD, $\text{U}_2$, RMSE, and MAPE. The sub-figures in Figure (ref), (ref), and (ref) display heatmaps of the differences and difference rates for each item in the input--output tables for each city.
Examining the STPE, $\text{U}_2$, and MAPE in Table (ref), which are independent of the magnitude of the amounts, exhibit that their values for the three cities were generally higher than those for Japan. Particularly for the final demands of Okayama City, the MAPE was notably large. As illustrated in Figure (ref), (ref), and (ref), there were positive correlations between the gross outputs by industry and these levels of differences similar to when the input--output table for Japan was estimated. Moreover, these figures reveal a tendency for high difference rates in inputs for mining in all cities. Sapporo City exhibited high difference rates in inputs associated with agriculture, forestry, and fisheries, as well as manufacturing. Gujo City exhibited high difference rates in inputs related to electricity, gas, and water supply.
However, it is important to note that the above results do not necessarily represent estimation accuracy. Therefore, it was not possible to reach a definitive conclusion about the validity of the method developed in this study for estimating input--output tables for three cities in Japan. Estimation accuracy is determined by comparing estimates with true values or values close to them. The input--output tables published by these cities, however, were estimated using non-survey or hybrid methods due to the insufficient primary statistics required for their compilation. These methods have larger estimation errors compared to survey method and do not guarantee proximity to true values. Consequently, the estimation accuracy of the input--output tables of these cities could not be determined by comparing the estimates with the published values.
This study proposed a method for estimating regional input--output tables using deep learning. The properties of this method were verified by estimating input--output tables for regions in Japan.
The proposed method generally achieved higher accuracy in estimating the input--output table for Japan compared to conventional methods. When estimating the 2015 input--output table for Japan, the conventional matrix balancing using the 2011 table and the 2015 gross outputs by industry, both of which were obtained through survey methods, represents an estimation under particularly ideal assumptions. In comparison with the conventional method based on these assumptions, our method generally produced lower errors. This is a noteworthy observation.
The high accuracy of this method is primarily attributable to the implementation of deep learning. This study addressed overfitting in deep learning by generating virtual regions through data augmentation. When the model was trained on the augmented data, the prediction accuracy extrapolated to the input--output table for Japan was satisfactory, indicating that data augmentation effectively mitigated overfitting.
Compared to more precise conventional methods, this study may achieve equivalent accuracy. Hosoe2014 estimated the 2005 Japan input--output table by balancing the past tables revised with the available data at that time. The results showed that MAPE was approximately 40%, and 53% of the table's 486 items had an error rate within $[-10\%, 10\%)$. In contrast, the method proposed in this study yielded a MAPE of approximately 43% and an error rate within $[-10\%, 10\%)$ for approximately 55% of all items. Based solely on these metrics, the prediction accuracy of the two methods can be considered comparable.
The method proposed in this study circumvents the instability associated with selecting reference values, which was necessary in conventional estimation methods. Both cross-entropy maximization and the RAS method require an initial table for calculation. The LQ method requires selecting the target region to derive location quotients. These selections can affect estimation accuracy. As shown in Table (ref), the conventional cross-entropy maximization exhibited varying estimation accuracy depending on the initial table selected. In contrast, while our method requires the establishment of explanatory variables, it does not necessitate selecting a small number of reference values as conventional methods do. Therefore, it does not exhibit the accuracy fluctuations observed in conventional methods.
In comparison with cross-entropy maximization and the RAS method, our approach has the additional advantage of requiring fewer specified values in advance. For example, the cross-entropy method requires gross outputs by industry, which serve as the row and column sums during balancing. In contrast, our method merely requires the total of gross outputs. In general, the estimation of the total is less challenging than that of the gross outputs for each industry. For Japanese municipalities, estimating the total of gross outputs is relatively straightforward because various data sources, such as the Economic Census, are available.
Although our method demonstrated higher accuracy than conventional approaches in estimating the input--output table for Japan than conventional approaches, its efficacy in smaller regions, such as cities, remains to be elucidated. In three cities of Japan, relatively large differences were observed between our estimates and the published values for certain items. If these differences are attributable to city-specific characteristics, our model may not fully capture them. Furthermore, the regions used for training were predominantly prefectures. Consequently, the model may not accurately reflect the properties of cities. However, it is important to acknowledge that the published values for these cities are also estimates and may not accurately reflect the true values.
The accuracy of certain items of final demand, particularly net exports, could be improved. One potential improvement would be to introduce explanatory variables that correlate more closely with these items. However, this study could not identify suitable variables for small-area estimation.
The method of this study does not entirely supplant conventional methods. Because our method alone is incapable of satisfying the constraints of input--output tables, the conventional matrix balancing method must be used alongside it. Moreover, our method is limited in its ability to directly estimate exports and imports. Therefore, alternative methods are required to estimate them.
While the proposed method is amenable to improvement and has certain limitations, it can serve as a foundation for obtaining input--output tables in regions where compiling them using primary data is difficult. Such regions are not exclusive to Japan. The compilation of national-level tables may also pose challenges in certain countries. For these regions, our methodology could facilitate the creation of more precise and stable input--output tables compared to conventional methods.