Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,193 characters · 16 sections · 47 citation commands
Min(d)ing the President: A text analytic approach to measuring tax news
\onehalfspacing
In this paper we propose a novel approach for incorporating textual information into the empirical analysis of causal relations in economics. We complement the literature with three main contributions. First, we develop a two-step semi-supervised topic model that automatically extracts specific information about future policy changes from textual data. Second, we provide methodology for transforming such textual information into economically meaningful time series to be included in causal econometric models as variable of interest or instrument. Third, we study fiscal foresight by extracting information about future tax reforms from speeches of U.S. presidents. We find that economic agents react to these news signals.
News about future policy changes is likely to have effects today. When receiving new information, economic agents’ forward-looking behavior implies that economic decisions are also contemporaneously affected. There is a growing empirical literature supporting these findings.\footnote{For example jaimovich2009can,blanchard2013news,beaudry2014news,forni2017noisy investigate news shocks related to future productivity, Ramey11 considers fiscal news, monetary policy news shocks are analysed in NakamuraSteinsson18; see also Ramey16 for a comprehensive overview of the recent literature.} Taking foresight into account is particularly important when analyzing effects of fiscal policy: economic agents typically receive clear signals about future tax reforms long before a particular bill is enacted. yang2007chronology shows that, in the U.S., almost all implemented tax changes were preceded by legislative lags ranging between a quarter and three years. The Tax Reform Act of 1986 provides an illustrative example of how fiscal foresight impacts economic behavior. This Act stipulated an increase of the effective maximum tax rate on capital gains from 20% to 28%, to be implemented in the following year. As a result, tax revenues from capital gains jumped by 90% before the act came into force auerbach1997economic. Assessing this particular tax reform's effect is flawed if fiscal foresight is not accounted for, and one may easily confuse directions of causality.
As a general critique, leeper2013fiscal argue that standard empirical approaches such as structural vector autoregressions (SVARs) often fail to properly account for fiscal foresight, as the variables typically included do not span the full information set available to economic agents. Estimated tax change effects using such models may as a result be biased and misleading; tax multipliers might even be of the wrong sign. Several studies therefore use instead a narrative approach to trace the arrival of information on future tax policy changes from the legislative process. romer2009narrative,RomerRomer10 (henceforth RR) identify the key motivation behind all legislated post-war tax changes in the U.S. and determine their impact on government revenues. Using various sources of narrative records, such as presidential speeches and Congressional records, they identify those tax changes that are not systematically related to changes in output and classify them as exogenous.\footnote{The exogeneity of RR's tax narratives is a matter of debate brueckner2021fiscal. Nevertheless, it is widely used in the literature as measure of exogenous legislated tax changes in the U.S.} RR additionally define a series of tax news as the present value of tax changes discounted back to the date of enactment (in contrast to the exogenous tax changes which are assigned to the implementation date), with the aim of capturing anticipation effects. mertens2012empirical,mertens2014reconciliation further split the RR series of exogenous tax changes into two components and use these to identify anticipated and unanticipated tax shocks.
We find that neither RR’s tax news, nor derivatives of RR's series as for example mertens2012empirical,mertens2014reconciliation, fully get the timing of arrival of new information right. In addition, narrative approaches that identify tax shocks solely from tax changes that were eventually implemented do not take into account that a news signal might be noisy and that policy plans, and alongside them economic agents’ expectations, are subject to revision over time. Before a tax bill is signed into law, information on the exact design of the tax reform is imprecise. Moreover, some proposals of tax changes initially put forward by the administration never come to fruition at all, while nonetheless influencing expectations.
While similar to RR, our approach is conceptually different in two important aspects. First, while analyzing similar auxiliary sources of data, our goal is not to search for the motivation behind each tax change but rather to trace the arrival of information regarding future tax reforms. In the U.S., the president is the main driving force behind tax policy legislation yang2007chronology,RomerRomer10. Presidential speeches are therefore a highly informative source for signals about future tax reforms. We review the U.S. president’s speeches and communications and quantify how prominently tax reforms, as well as their direction (cuts vs. hikes), are featured on the political agenda at a given point in time. To implement this idea, we build on the following hypothesis: if a policy maker repeatedly emphasises the importance of future tax changes, the public should expect that these changes will likely be implemented in the future. As a result we obtain two measures of prevalence: one capturing the importance of tax cuts on the administration's policy agenda and one capturing the importance of tax hikes.
Second, to construct these prevalence measures from a considerable number of documents (we consider all of the President's public statements since 1949), we introduce a novel semi-supervised text-analytic approach. While a standard, unsupervised Latent Dirichlet Allocation (LDA) topic model would allow us to quantify the prevalence of specific political issues -- henceforth referred to as topics -- and distinguish speeches about tax reforms from speeches on other topics, further differentiating between tax hikes and tax cuts is not possible. Our proposed novel strategy is to feed additional information into the LDA topic model in order to construct more informative priors for the tax hike/cut topics. To this aim, we combine lexical knowledge of a priori selected terms related to the direction of tax changes with the results from an unsupervised model. It is important to stress that our dictionary of selected terms only `nudges' the model towards the terms we deem to be important: the topic estimation remains data-driven and robust to misspecification.
We find that our tax prevalence series predict future federal tax reforms regardless of how we measure tax changes: both cyclically adjusted revenue changes as well as all RR narrative measures are Granger-caused by our prevalence series. In contrast, we do not find evidence for predictability in the opposite direction. Interestingly, our series also Granger-cause implicit tax rates, which are often used to proxy tax news in the U.S.\footnote{To construct implicit tax rates, i.e. a measure of expected future tax rates, leeper2012quantitative use the yield spread between taxable treasury bonds and tax-exempt municipal bonds to identify economic agents' expectations about future personal tax changes.} Again, we do not find evidence for predictability in the opposite direction. These findings suggest that our tax prevalence series capture the timing of the arrival of tax news more accurately. Additionally, and importantly we do not find evidence that our tax prevalence series are driven by business cycle conditions in the (recent) past.
We use our constructed measures to analyse the effects of tax news on economic activity. Along the lines suggested in SW2018, we attach a scale, or monetary value, to the news by using the tax prevalence measures as instruments for future tax changes. We find that output expansion is driven by (post-) implementation effects of tax changes. However, allowing our measures to also instrument tax changes in the more distant future also reveals pre-implementation effects. More specifically, anticipation effects trigger initially a contraction in GDP if the anticipation horizon is long enough. We also find that the longer the anticipation horizon, the weaker the expansion due to implemented tax changes.
The remainder of this paper is organised as follows. Section (ref) illustrates the type of textual information we aim to quantify and gives an intuitive overview of our approach. In Section (ref) we outline the semi-supervised topic model in more detail and discuss how we estimate tax policy signals. The identified topics, including the tax prevalence measures, are presented in Section (ref). In Section (ref) we use our prevalence measures to analyse the impact of tax news on economic activity. Section (ref) concludes. Additional methodological details and further empirical results are collected in several Appendices.
In the U.S. the president is the main driving force behind tax policy legislation yang2007chronology,RomerRomer10, making presidential speeches a highly informative source for signals about future tax reforms. Below we illustrate the information flow and the various types of signals we aim to quantify. All four statements were made by Ronald Reagan in the months leading up to enactment of the Economic Recovery Act of 1981.
Throughout his term, the president outlines tax reforms he wants to pursue, albeit in general terms:
When the president believes the tax law should be changed he recommends that to the House of Representatives:
In that announcement, as well as in the following months, the president may offer details about particular measures included in the upcoming tax law changes:
Once the bill is passed through the Congress, the president provides final remarks as he signs it and ends the legislative process:
Quantifying individual statements in terms of monetary value is difficult as they often lack sufficient details about the proposed reform. While the Congressional Budget Office (CBO) might publish an estimate of its tax revenue impact, this is done only in the last stages of the legislative process, once the law is drafted. Moreover, the president’s statements are often subject to revisions: statements might be retracted, announced changes might not come to fruition or they might include different measures than initially planned. The goal of our approach is to capture those signals as well.
Our solution is to instead quantify how prominently each direction of tax reforms (cuts vs. hikes) is featured on the political agenda of the president. We do that by measuring the proportion of presidential speeches in a given quarter that is devoted to tax cuts and tax hikes respectively.
The idea of measuring how often a particular issue is mentioned in a body of texts is not new. Most prominently, baker2016measuring develop an Economic Policy Uncertainty (EPU) index by quantifying the proportion of news article referencing various types of uncertainty. The main challenge of such an approach is determining which texts, or which parts of a specific text, are relevant. The simplest approach is to check whether (or how many of) the words in a document belong to a pre-defined set of terms (a so-called lexicon). For example, baker2016measuring use a rule-based extension of this approach, which consists of checking if a combination of certain terms appears in a given text document. Lexicon-based approaches are particularly useful in sentiment analysis, since readily available sentiment lexicons can often be used in a variety of applications. shapiro2020measuring combine multiple lexicons to produce a robust analysis of economic news sentiment and its effects. In the case of our application however a lexicon-based approach would require that we exactly specify the words and phrases which the president uses to reference tax reforms. This introduces a risk of misspecification, especially if the relevant terms can appear in different contexts.
Topic modelling solves this problem by jointly considering the whole text instead of looking for particular terms. We can compare the terms used in a document with the term distributions used to discuss a certain topic to determine to what degree the text is about it. Crucially, in contrast to lexicon- or rule-based approaches, a topic model “learns” those distributions from the data in an unsupervised fashion. Given a collection of documents and a specified number of topics, we find topics which best explain the way the words are used in the texts. The method we chose is the Latent Dirichlet Allocation (LDA) model proposed by blei2003latent which is discussed further in the following section. It allows us to identify the way in which the president talks about tax policy changes, expressed as a probability distribution over the vocabulary, and determine which texts discuss them.
A similar approach has been used for example by larsen2019value. Their LDA-based analysis of news outlets shows the impact that various types of news have on financial markets. We are also not the first ones in the economic literature who use topic models to analyse the statements of policymakers. LDA has been widely used to analyse the economic impact of the communication by central banks. For instance hansen2016shocking,10.1093/qje/qjx045 analyse the minutes of the deliberations of the Federal Open Market Committee to identify topics of discussion and measure their impact on the economy. In a study similar to ours, dybowski2018economic also use a LDA model to identify the part of presidential speeches devoted to taxation in general. Using a lexicon-based approach they further measure how optimistic or pessimistic the identified tax communication is and investigate whether the effect of the communicated tax changes on economic activity depends on the tone. Crucially, the effect they capture is by design that of changes in perceived uncertainty. Since those are only very loosely connected with the communicated tax changes, their approach is unfortunately not well suited for analyzing the effects of tax news. The above examples are part of a fast growing body of literature using text-mining in economic research; for a recent survey see gentzkow2019text.
Our approach is meant to explicitly distinguish between signals about tax hikes and tax cuts. Because standard LDA estimates the topics in an unsupervised fashion, there is little control over their composition. Intuitively, the discussion about both tax increases and decreases relies on a relatively similar, tax-oriented subset of vocabulary. As a result, when using the standard approach, the two are “grouped” together into a single topic, preventing us from determining the direction of the discussed changes. By reading through (some of) the speeches that we know are tax related we can gather precise information about terms which differentiate the two types of signals. Including this additional (prior) information in the model requires however a departure from the conventional LDA approach. In our two-step approach we combine the so obtained information with the results from the unsupervised approach to construct informed priors for the topics. By using those priors in an LDA model we are able to differentiate between the content devoted to discussing tax hikes and tax cuts. We aggregate the per-document results for each quarter to obtain a measure of the relative prominence of each political issue on the presidents agenda which we refer to as prevalence. In Section (ref) we show that the prevalence measures of the two tax topics contain information about future tax changes.
In this section we outline our methodology to extract information about tax policy from presidential statements using topic modelling. Intuitively, our model assumes that when the president discusses different topics, he uses a different distribution over words (his vocabulary). Identifying these distributions, and their occurrence in each statement, therefore allows us to identify the topics the president talks about. We hypothesise that one or more of these topics can be linked to instances when the president talks about tax policy. Indeed, as we show in Section (ref), after estimating an unsupervised LDA model, one of the estimated topics can be labeled as the tax topic. However, as shown in Section (ref), such a tax topic contains information too imprecise to be useful for predictive or causal analysis. We therefore aim to explicitly differentiate the discussion about tax policy into tax increase and tax decrease topics. For this purpose we introduce a two-step topic modelling approach. Before we discuss the topic model in detail, we first describe the data and the steps we take in pre-processing.
Our analysis is based on the Public Papers of the Presidents, a compilation of all documents originating from the president. We obtain the raw texts from the American Presidency Project (APP).\footnote{\href{http://www.presidency.ucsb.edu}{www.presidency.ucsb.edu}, retrieved on 2019-03-25.} We analyse 59,214 texts spanning from 1949-01-20 to 2017-01-19.
The raw text of the speeches needs to be pre-processed to quantify relevant features of the text data, facilitating further statistical analysis. This process is described in detail in Appendix (ref).
First, the documents in the dataset range from short remarks consisting of several sentences to long speeches, such as the State of the Union Address, which cover a variety of otherwise unrelated issues. We therefore split the texts into a total of 1,119,200 individual paragraphs of roughly the same length. For the remainder of the analysis we treat those as separate documents.
Next, we prepare a matrix of word counts which is used as the input to our algorithm. Informally, for each text we count how often every word is used in it. To define the possible `words' we take two steps. First, we clean the text and exclude function words such as “a” or “and” and rare words. In this step we also transform words to their `root' form, e.g. “taxes” and “tax” are both counted as “tax” and “implements” and “implemented” both count as “implement”. In the second step we identify combinations of two subsequent words that occur frequently together (so-called bigrams), which are then also counted as a `word'. This is important for our analysis as combinations of words such as “tax cut” may contain very different information than either of the words would contain separately. To avoid confusion with the actual words used in the speeches, henceforth we refer to each of the `words' we count as terms, which then refers to either individual (pre-processed) words or the identified bigrams.
In this subsection we present the Latent Dirichlet Allocation (LDA) topic model. For this purpose we first define the following elements of the pre-processed data:
Our approach is based on the LDA topic model proposed by blei2003latent. LDA is a relatively straightforward, unsupervised approach that often leads to easily interpretable results, and thus has become one of the most popular choices for topic modelling jelodar2019latent. In LDA the process of creating a document is modelled as a series of independent draws from a particular distribution over the terms in the vocabulary. The key idea behind topic modelling is that the particular distribution changes depending on the topic, i.e., what the text is about. For example, we expect the president to use certain terms with different frequency when talking about education versus for example military build-up. As such, with $K$ topics,\footnote{In the model that follows, the number of topics $K$ is assumed to be known. Of course, in practice this has to be estimated. The particular choice of $K$ for our dataset is described later.} the distribution with which terms are used overall, is a mixture of the $K$ distributions per topic, which we label as the topic-term distributions. Furthermore, each document can consist of multiple topics. The proportions of each topic's occurrence in a single document is labelled the mixing proportion of that document, which is of key interest for our analysis, as informally it addresses how much each topic (such as tax) is discussed in the document. The mixing proportions therefore allow us to identify documents which relate predominantly to tax policy. It is crucial to allow documents to discuss multiple topics in varying proportions. For instance, the president will generally not discuss tax in isolation, but in conjunction with other topics such as the need to balance the budget, stimulating the economy or creating the funding for particular investments such as military expenditures in times of war.
Hence, the creation of a document -- and the tokens inside that document -- can be seen as a two-step process. First, the creator (in our case the president), decides on the proportions of each topic to be discussed in that document (the mixing proportion). Then, given that proportion, each token in the document is randomly assigned to one topic. Next, given this topic assignment, a term from the vocabulary is drawn from the distribution corresponding to the assigned topic. We now formalise this as follows.
We estimate the topic-term distributions and the mixing proportions in a Bayesian procedure by means of the posterior parameter distributions using Gibbs sampling. We denote them as $\hat{\bm{\Phi}}$ and $\hat{\bm{\Theta}}$ respectively. In order to do so, LDA puts Dirichlet priors on both the mixing proportions $\boldsymbol{\Theta}$ and the topics' term distributions $\boldsymbol{\Phi}$. Appendix (ref) provides an intuitive description of the properties of the Dirichlet distribution $\bm{\theta}_d \sim \text{Dir}(\bm{\alpha}), \quad d= 1,\ldots,D$ and the implications of its use on estimation. In particular, we assume that $\boldsymbol{\theta}_1, \dots, \boldsymbol{\theta}_D$ share the same Dirichlet distribution with $K$-dimensional parameter vector $\boldsymbol{\alpha} = \left( \alpha_{1},\dots,\alpha_{K} \right)^T$. Following the recommendations of wallach2009rethinking, we do not treat $\bm\alpha$ as a fixed hyper-parameter but place another level of (Gamma) priors on its elements, such that $\bm\alpha$ is estimated along with the lower-level parameters.
Dirichlet priors are also placed on the topic-term distributions. How we do that exactly depends on the first or second step of our model, and is the source of our methodological innovation in the second step. Generally, each topic-term distribution has a Dirichlet prior $\bm{\phi}_k \sim \text{Dir}(\bm{\eta}_k), \quad k= 1,\ldots,K$ with $V$-dimensional parameter vector $\boldsymbol{\eta}_k=\left( \eta_{k,1},\dots,\eta_{k,V} \right)^T$. In the first step we use the conventional LDA setup where no `expert knowledge' about topic composition is formulated. In particular, all topics share the same symmetric Dirichlet prior, such that $\eta_{k,i} = \eta$ for all $k=1,\ldots,K$ and $i=1,\ldots,V$ - the prior probability of every term in every topic is equal. This parameter $\eta$ is then also estimated by assuming a Gamma prior wallach2009rethinking. These uninformative priors imply that the topic-term distributions is fully driven by the data.
Crucially however, and as discussed before, a pure data-driven, unsupervised approach is unsuited to differentiate between tax increase and tax decrease topics, for which purpose we add a novel second step LDA estimation. Based on our first-step estimates $\hat{\bm{\Phi}}$, we first identify a general tax topic and use it to develop informed priors for the two topics of interest for the second step of LDA estimation. The exact process is described in the following section. The result is a set of parameter vectors $\bm{\eta}_k, \quad k=1,\ldots,K$ which specify separate, informed, prior for each topic, where we postulate that particular terms occur either more frequently, or less frequently, in specific topics. These priors are then used in another Gibbs sampling to estimate the posterior means of the parameters, which are denoted as $\doublehat{\bm{\Phi}}$ and $\doublehat{\bm{\Theta}}$. The full details of the estimation for both steps are provided in Appendix (ref).
The approach described above relies on the assumption that the number of topics $K$ is known, but in practice this number needs to be estimated. An incorrect specification of the number of topics decreases the overall quality of the model and the interpretability of the topics. While the choice of $K$ can be guided by metrics based on the likelihood of observing the corpus given the model, we base this choice on the interpretability of our results instead. Specifically, our approach requires that discussion about tax policy forms a distinct topic. Intuitively, specifying too few topics is likely to lead to a topic that includes also other issues, not necessarily related to tax policy. On the other hand, setting the number of topics too high introduces unnecessary computational complexity and can result in the tax topic being split, for example into personal and corporate tax policy discussion. As a result we choose the number of topics to be $K=25$ in the first step. In the second step, because we split the general tax policy topic into the tax increase and tax decrease topics, we use a total of $26$ topics. In section (ref) we substantiate our choice by inspecting the estimated topic-term distributions and mixing proportions.
By using different priors for the topic-term distributions we can include `expert knowledge' as a priori beliefs about topic composition to `steer' the topics into the direction we want them to go, in our case topics about tax increase and tax decrease. Here we describe how we achieve this in our two-step procedure. The main challenge is now to correctly translate a priori beliefs about occurrence of particular terms into meaningful prior parameters $\bm\eta_i$ of the Dirichlet distribution. To illustrate, note that for any topic-term distribution $\bm{\phi}_k \sim \text{Dir}(\bm{\eta}_k)$ the Dirichlet prior implies that $\mathbb{E}[\phi_{k,i}|\boldsymbol{\eta}_k] = \frac{\eta_{k,i}}{\sum_{j=1}^V \eta_{k,j}}$. Hence, in order to obtain sensible priors we need to formulate beliefs about the relative occurrence of all terms. For example, we may believe that terms belonging to a certain set $L$, a so-called lexicon, are likely to be mostly associated with a particular topic of interest. Such a belief could be expressed by choosing a topic $k$ and setting $\eta_{k,i}=a \mu_L$ if $v_i \in L$ and $\eta_{k,i}=a \mu_{\neg L}$ if $v_i \notin L$ where $\mu_L > \mu_{\neg L}$. However, the shape of the resulting a priori distribution is not in line with the way terms occur in text in general, as empirical term distributions tend to roughly follow a power law sato2010topic. Setting such a prior would therefore lead to a highly distorted topic distribution; instead we must make sure the prior parameters are in line with the empirical properties of the text.
To address this issue, we first learn about the general shape of the topic-term distributions from the data in the first-step unsupervised estimation and then modify these distributions using lexicons of terms relevant to the topics of interest. From the first-step unsupervised estimation we are able to identify a single topic that encompasses all the discussion about tax changes, which we refer to as the general tax topic,\footnote{This interpretation is based on the fact that this distribution assigns uniquely high probabilities to terms related to tax policy (e.g. “tax”), which is discussed in detail in Section (ref).} and denote the corresponding estimated topic-term distribution as $\hat{\bm{\phi}}_{k^*}$.
We then construct the priors for the tax increase and tax decrease topics based on the assumption that both topics of interest have a similar distribution over the vocabulary, except for some key differentiating terms. Based on reading of documents related to tax policy we manually identify the terms which are predominantly used when discussing changes in a particular direction.\footnote{For this purpose we use all off the speeches identified by romer2009narrative and yang2007chronology as announcements of tax changes.} When then group them into two lexicons - $L_\text{inc}$ for terms related to tax increases, and $L_\text{dec}$ for terms related to tax decreases. The composition of the lexicons and details concerning their creation are presented in Appendix (ref).
We then `guide' the algorithm towards the two tax topics by modifying the prior probabilities of terms which are contained in either of the lexicons. For the tax increase prior $\boldsymbol{\eta}_\text{inc}$, we modify $\hat{\bm{\phi}}_{k^*}$ by multiplying probabilities of terms in $L_\text{inc}$ by a constant $m_1>1$ to `up-weight' them, and simultaneously multiplying probabilities of terms in $L_\text{dec}$ by $m_2<1$ to down-weight those. For the prior of the tax decrease topic we do the reverse. For the priors of the other topics we use their respective distributions estimated in the first step, without modifying their shape. This allows us to `fix' the other topics while splitting up the general tax topic. The last step in the modification of those priors is to choose the strength of the effect of our prior, which is done by multiplying the vectors by a scalar $m_{3,k}$. Hence, we construct our prior parameters for all $i=1, \ldots,V$ as
We set $m_1=100$ and $m_2=100^{-1}$, while $m_{\text{inc}},m_{\text{dec}}, m_{k}$ are chosen such that the sum of the vectors is $\sum_i\eta_{\text{inc}, i} = \sum_i\eta_{\text{dec}, i} = \sum_i\eta_{k,i} =10,000$. This prior can be interpreted as postulating a prior belief equivalent to an additional observation of 10,000 tokens assigned to a given topic from the corresponding topic-term distribution; a relatively weak prior given the size of our dataset. In practice, the modified $\boldsymbol{\eta}_\text{inc}$ and $\boldsymbol{\eta}_\text{dec}$ act as `seeds', during each iteration of the estimation `nudging' the distributions of the relevant topics through the mechanism described in (ref).
While the priors help us discern between the two topics of interest, the distributions are still predominantly determined through learning from the data. In particular, we see that a tax increase (decrease) topic does not necessarily assign high probabilities to all the terms included in its lexicon. Moreover, each tax topic features terms that are not included in its lexicon, and can even feature terms from the other topic's lexicon. As a result, our approach is more robust to semantic misspecification than typical lexicon-based approaches as their impact is relatively small compared to that of the observed dataset.
The additional benefit of using an LDA model is that the topic assignment on any given token is based not only on the topic distributions, but also on the other tokens in the document. In Appendix (ref) we compare the two approaches, and show that prevalence measures derived from a two-step LDA approach have much higher predictive power for future tax changes.
The final step of our text-analytic approach is to transform the estimation results from the topic model into numerical measures that reflect the prominence of tax-related discussions by the president. That is, we aim to construct a quantitative measure of the prominence of signals about tax policy over time, which does not follow directly from our topic model. The estimated topic model allows us to estimate the mixing proportion $\doublehat{\bm{\theta}}_d$ for each document $d=1, \ldots,D$ as the posterior means of the second-step LDA. These estimates signify what proportion of its words is associated with a given topic: documents for which $\doublehat{\theta}_{d,k}$ is estimated to be high discuss topic $k$ for a significant portion relative to documents for which $\doublehat{\theta}_{d,k}$ is low. We now use this property to aggregate estimation results for individual documents to create a measure for the topics' prevalence, or popularity, and its evolution over documents registered over time. Because we are using our measure together with macroeconomic data in the empirical models discussed later, we opt for aggregating the documents to quarterly frequency.
Formally, let $T$ denote the total number of quarters, and let $T_d \in (0, T]$ denote the normalised date corresponding to the publication of document $d$. For the measure in quarter $t$, we then average over all documents published in the period $(t-1,t]$ to obtain the measure
To illustrate the proposed prevalence measure, consider two extreme cases:
Arguably, the myriad of choices made in pre-processing, topic model specification and measure creation may make our final measures seem arbitrary. Indeed, while we made those decisions generally in accordance with standards used in the literature, alternative choices appear equally plausible to justify. Therefore, rather than arguing that our choices are the optimal ones, we empirically investigate if our resulting estimates have the properties we attribute to a measure of tax policy signals. In the next section we assess the topics estimated through our two-step LDA approach and evaluate our model's ability to properly classify the tax content of the documents.
In this section we investigate in how far the fitted topic model captures our concepts of policy topics, with special attention to the tax (increase and decrease) topics. We first evaluate the two steps of constructing tax topics in detail. Next, we also briefly consider the other topics in order to understand how well the topic model captures the general essence of the speeches.
The general tax topic from the first-stage, unsupervised LDA, combined with our lexicons, is used in the second-stage guided LDA to obtain the tax increase and tax decrease topics presented in Figures (ref) and (ref), respectively. The terms in our lexicons driving the prior for tax increase (decrease) are highlighted in blue (orange). Although both distributions have many terms in common with the general tax topic estimated in the first step,\footnote{The wordcloud for the first-stage general tax topic is shown in Appendix (ref).} there are crucial differences extending beyond the terms whose prior probability was modified. Those include such terms as `social security' which is used predominantly when discussing tax increases, and `small business' which is referenced often when discussing tax decreases. Lastly, we can see that the topics still feature, to some extent, terms which we determined relate to tax changes in the other direction (e.g. `cut tax' still appears in the `tax increase' topic). This is an example of how the data overrides the priors. Apparently, even when discussing increasing taxes, the presidents sometimes make a reference to tax cuts.
In both steps we verify that the identified tax topics capture all of the content devoted to tax policy changes. We do that by inspecting the probabilities assigned to tax-related terms by the non-tax topics. For example, in the first step we find that the topic devoted to state and local policies has the second highest probability of using the term “tax”, which is still over 100 times smaller than the probability assigned to that term by the tax topics. Overall, in both steps the identified tax topics account for over 99% of the usage of the term “`tax”.
For our approach, the estimated mixing proportions of each document are more important than the topic-term distributions. For our method to make sense, we specifically need to verify that the estimated mixing proportions for the two tax topics do indeed reflect our interpretations of the texts. Given the size of our dataset with 1,119,200 individual paragraphs, such analysis is possible only for a subset of documents. We first consider paragraphs from speeches that RomerRomer10 and yang2007chronology list as announcements of tax policy changes. Figure (ref) shows that, as expected, documents coming from speeches announcing tax hikes tend to have much higher mixing proportion for the tax increase topic, and the same is true for those announcing tax cuts and the tax decrease topic. In case of some documents we see that the opposite mixing proportion is high. This can be at least partially explained by the fact that some legislated changes include measures in both directions, lowering some taxes while raising others. Finally, in both groups a lot of documents do not relate to tax changes at all, since some of the speeches considered (e.g. State of the Union Addresses) concern more issues than just taxation.
Out of those documents we randomly select 100 for which either $\doublehat{\theta}_{d,inc}>0.3$ or $\doublehat{\theta}_{d,dec}>0.3$ and compare the estimated mixing proportions with our interpretation of their content. Additionally, we consider the estimated maximum a posteriori (MAP) topic assignment labels for particular tokens.\footnote{For a token $w_{d,n}=v_i$ the maximum a posteriori (MAP) topic assignment estimate is $z_{d,n}^{*} = \argmax_{k} \doublehat{\theta}_{d,k} \doublehat{\phi}_{k,i}$} While those labels are not directly necessary for our analysis, they help visualise the clustering property of LDA. In Tables (ref) and (ref) we present an example for each direction of change, where tokens assigned to tax increase (tax decrease) topic are highlighted in blue (orange).\footnote{Words not in bold are removed during pre-processing.} We present details of this part of analysis and summarise our findings for the rest of the selected speeches in Appendix (ref). We find that overall the estimated mixing proportions correctly capture the direction of the implied changes.
Looking at the estimated topic assignments for individual tokens we can see the clustering property of LDA. In many cases terms that are not clearly related to tax changes are classified as referring to one of the tax topics. This is the result of the context in which they appear - because so much of the rest of the speeches relate to tax increase or tax decrease the likelihood that they relate to one of those topics is higher. Even terms that intuitively refer to changes in one direction can be classified as referring to changes in the opposite direction (e.g. `tax break' assigned to the tax increase topic in the 2012-11-05 speech). This property of the topic model is a crucial improvement over a simple lexicon-based approach, making it much less sensitive to misspecification.
It is important to note that despite the clustering property of LDA some paradoxical classifications can occur. This happens when the terms used in the text are not sufficient to capture the meaning contained in the syntax. For instance, speeches along the lines of “we are not going to increase taxes” would likely be classified as information about a future tax hike. This is inevitable given the bag-of-words assumption underlying the LDA topic model. Addressing this would require methods able to recover the true intention of statements and their interpretation in a broader context. However, in our context, this is difficult and often impossible even for well-informed political observers. Thus, it cannot be reasonably expected to be achieved perfectly by any algorithm, and we believe it should therefore also not be held against our approach based on the bag-of-words assumption. Moreover, such mistakes are likely to average out over all speeches within each quarter.
Our final diagnostic check in this section concerns the aggregated quarterly prevalence measures of tax increase and decrease, as presented in Figure (ref), along with periods of legislative lag of tax hikes (blue) and cuts (orange) as identified by yang2007chronology. Visual inspection shows that the topic prevalence of the relevant direction - but not the opposite one - generally increases leading up to an actual tax change in that same direction, with the peak around the enactment date. While by no means a formal analysis, this does seem to indicate that our identified tax increase and decrease topics have predictive power for actual tax changes. We investigate this more formally in the next section, but first we briefly investigate the other topics.
In addition to the two tax topics, we identify 24 other topics. While those topics are not the focus of our study, they can give an indication of the overall quality of the model. Their distributions remains relatively unchanged between the first and the second step of the topic model and the majority can clearly be attributed to specific policy issues such as public health, trade, and foreign policy. The distributions of all the topics are presented in Appendix (ref).
In several cases we can quite intuitively see how the changing importance of certain issues is reflected in our prevalence measures. Perhaps the best examples are the Cold War and War on Terror topics displayed in Figure (ref). The prevalence of the Cold War topic has oscillated for decades, reaching a sharp peak in the late 1980s after which it declined considerably. In case of the War on Terror topic we see a slight peak in the early 1990s - likely the discussion surrounding the First Gulf War, a steep rise after the 9/11 attacks and a considerable decline since Barack Obama entered office, reflecting a different approach of the new president. This also shows that our LDA approach is flexible enough to handle variation over time.
In this section we investigate how well the tax prevalence measures predict future tax changes. We initially estimate predictive regressions of various tax change series on our tax prevalence measures, assessing the strength of the correlation between tax changes and the lags of our measures by $F$-tests for the non-significance of the measures in the predictive regressions. We find that the general tax prevalence measure obtained from the first-stage unsupervised LDA is not a powerful predictor for any tax change measure considered. However, the second-step prevalence measures -- capturing tax cuts and hikes -- are strong predictors of future tax changes (regardless of how they are measured), with $F$-statistics well above the cut-off. The superior predictive power of the second-step measures is also reflected in the much higher partial $R^2$ statistics. Details of this analysis can be found in Appendix (ref).
To study (non-)predictive relationships in more detail, as well as to investigate the (conditional) exogeneity of our tax topics and to what extent they are driven by other factors such as macroeconomic conditions, we carry out Granger causality tests. We investigate contemporaneous correlations in Appendix A.1.
To test for Granger causality we consider a large VAR which includes a large number of possibly relevant macroeconomic and financial variables. We also include variables that incorporate information on future policies concerning spending Ramey11,Ramey18 and taxation leeper2012quantitative, as well as prevalence measures for the other topics. This allows us to determine in how far our tax prevalence series contain unique predictive power for the various tax change measures that is not found in other series. Similarly, by testing for Granger causality from a variety of macroeconomic variables (including several tax change measures) to the two tax prevalence series we can determine what the causes of potential endogeneity are. To handle the high-dimensionality of the VAR, we use the post-selection test of hecq2023granger which provides valid inference about Granger causality after selection of relevant covariates via the lasso.\footnote{All tests are based on a VAR with six lags and an intercept. Nonstationary macro variables are transformed to first differences of logs.}
The left column of Figure (ref) illustrates which measures of tax changes are Granger-caused by our prevalence measures. It also shows whether any tax measure is predicting our tax prevalence measures. The tax prevalence topics Granger-cause all tax change measures. While predictive power is higher for aggregate measures of tax changes (federal revenue, or narratives of legislated tax changes), the tax topics also Granger-cause corporate, income, and payroll taxes. Most notably, we find that even RR's exogenous news narrative, as well as \citepos{mertens2014reconciliation} unanticipated tax narrative, are Granger-caused by the tax prevalence measures. Those results support our hypothesis that tax policy changes are signalled by the president well ahead of time, even when the motivation is exogenous to recent economic conditions and the implementation lag is short. We further find evidence (p-value $=0.075$) that our prevalence measures Granger-cause implicit tax rates, which are considered as proxy for tax news in the literature.\footnote{We use the risk-adjusted implicit tax rate from leeper2012quantitative. We use rates with maturity of one year which are considered by the authors to best predict tax changes in the near future.} Since investors are forward looking, the yield spread between taxable treasury bonds and tax-exempt municipal bonds likely mirrors anticipated federal tax changes. The results above seem to suggest that investors' expectations are driven - at least to some extent - by tax-related communications of the president. That is, if the president signals that federal taxes (on treasury bond income) may increase, investors will, as a response, demand higher yields on treasury bonds. Thus, in that case, our prevalence measure - and not the yield spread - would get the timing of the arrival of news regarding future tax changes right. In contrast, we do not find that any of the considered tax (news) measures contain information that help predict the tax prevalence measures.
We next investigate whether the prevalence measures predict macroeconomic and financial variables and vice versa. Results displayed in the right column of Figure (ref) show that the tax prevalence topics Granger-cause the information typically included in a fiscal VAR. We find evidence for predictability in the other direction only for government debt. This suggests that the prevalence measures capture only information about tax changes that are not systematically correlated with short-run changes in economic activity, but are rather driven by other (long-run) policy goals, for example reducing the debt burden.
We also evaluate predictive relationships between the tax topics and other prevalence measures (results are shown in Figure (ref) in Appendix (ref)). We find that some topics related to policies that affect the federal budget Granger-cause the tax topics. In particular, the War on Terror topic Granger-causes the tax prevalence measures. We will use this insight when selecting control variables among the topics in the causal analysis in Section (ref) to eliminate possible endogeneity issues.
Our tax prevalence measures do not have a direct, economically meaningful (monetary) quantitative interpretation. In order to enable a causal analysis using these measures, we link our tax prevalence measures to data that are informative about the monetary value of expected tax changes. We cast the analysis in the LP-IV framework of SW2018, treating the prevalence series as imperfect measures of an unobservable tax news shock. That is, we assume that the prevalence measures (and their lags) are correlated with the unobserved tax news shock but contain measurement error. This setup requires an observable, endogenous variable driven by the tax news shock, such as some monetary measure of future taxes, in order to treat the prevalence measures (and their lags) as instruments for this endogenous variable and estimate relative impulse responses by LP-IV.
As argued above, we believe that expectations and decisions of economic agents are influenced not only by news about implemented future tax liability changes but also by `noise'; that is, new information about possible future policy changes that are deemed plausible ex-ante, even though such policy plans never come to fruition ex-post. For LP-IV to work, the unobserved tax news shock must lead on average to a change in the observable endogenous variable. Identifying the causal effect of a `noisy' tax news shock using our prevalence measures as instruments is possible as long as new ex-ante information on future policies does not systematically differ from (ex-post) implemented policy outcomes. Hence, we assume that economic agents' expectations regarding future tax changes may be noisy but are correct on average.\footnote{We do not identify noise and news separately. The information we hope to capture relates to \emph{both} news about future fundamentals as well as changes influencing agents' beliefs only. One might argue that not being able to identify either component limits a (deep) structural interpretation. This criticism holds, however, for a variety of empirical papers. The previous investigation on the informational content of the tax prevalence measures indicates that the news signal about future tax changes, although not noise free, is strong enough allowing for a meaningful interpretation. Generally, noise and news are tightly related, further distinguishing both components and identifying them separately -- if this is at all possible -- is beyond the scope of this paper. We refer to chahrour2018news for a detailed analysis of this relationship.}
Impulse responses are based on the following regressions
where $x_{1,t}$ is an endogenous measure of future tax liability changes and ${\beta}_{1,h}$ is the the coefficient of interest, capturing the response to a tax news shock in the outcome variable $y_t$, $h$ periods after the shock occurred. Further, ${\bm x}_{2,t}$ gathers relevant control variables and ${\bm \gamma}_{h}$ captures deterministic components. Estimation and inference via two-stage least squares is straightforward.
We also need to choose the appropriate endogenous variable that carries the information about future tax changes. The construction of the tax prevalence measures does not rely on an accurate definition of the timing of future policy changes. Our prevalence measures predict tax hikes or cuts implemented as early as the following quarter, and are even stronger predictors for policy changes three or four quarters ahead. Thus, our measures contain information about (expected) tax changes implemented throughout the following year(s) and hence do not allow us to pin down the precise timing of future tax liability changes. Indeed, economic agents are likely uncertain about the exact implementation timing of future tax changes, and this uncertainty is reflected in a predictive power over a longer time horizon. If we were to instrument (endogenous) tax changes in a given quarter only (vs. over a longer period of time), we would likely disregard meaningful information regarding (expected) tax changes at a different point in time and only partially capture expectations.
Bearing this in mind and to broadly capture `future tax changes' over a longer period of time, we propose to use the following rolling average as our endogenous measure of future tax changes:
where $\Delta T_t$ is the (endogenous) tax liability change in period $t$. This simple aggregation scheme is better suited for an analysis of expectations, considering the uncertainty of economic agents regarding the exact timing of expected future policy changes. In addition, this makes specifying the exact timing of when a tax change is expected no longer necessary.
The choice of window start- and endpoints $M$ and $N$ is of crucial importance for identifying possible anticipation effects. For example, setting $M=1$ and $N=4$ aggregates all tax changes throughout the coming year, starting the following quarter. If the tax prevalence measures were most strongly correlated with tax changes implemented one quarter ahead, short-term (exogenous) variation in tax liabilities would drive the response of output to a tax news shock and we would inadvertently purge out information contained in the prevalence measures about expected tax changes in the more distant future. In addition, the resulting short anticipation horizon would likely not allow for large anticipation effects, since economic agents would have little time to react, and a potential impact from anticipation would likely by dominated by implementation effects. effects.\footnote{mertens2012empirical also find that the shorter the anticipation horizon, the weaker the (negative) pre-implementation (or anticipation) effects on output.} In contrast, by increasing $M$ (and $N$) we would put less emphasis on the correlation between the prevalence measures and soon-to-be implemented tax changes and our prevalence measures would instead instrument tax changes that are expected in the more distant future only. This would allow for a longer anticipation horizon and potentially stronger anticipation effects.
Based on the local projection regressions in ((ref)) we estimate responses of log real output and log real government tax receipts to (noisy) news of future tax cuts. To construct our measure of future tax changes as in ((ref)), we use RR's rich narrative account of all legislated federal tax liability changes in the US from 1947 to 2006. RR link tax changes directly to the legislative process and, thus, their narrative record contains precise information about the timing of tax changes. We use our tax prevalence measures (including eight lags) as instruments. The deterministic components ${\bm \gamma}_{h}$ consist of an intercept, a linear, and a quadratic trend. 12 lags of the dependent variable are always included in the set of regressors. This is a rather conservative choice and possibly more than necessary to capture the temporal dependency in $y_t$. However, it is in line with the suggestions in MOPM20 regarding robust inference in the presence of persistent data.
Furthermore, ${\bm x}_{2,t}$ includes 12 lags of the following macroeconomic controls: log real GDP, log real government spending, log real government debt, and the 3-month Treasury Bill. Based on the analysis in section (ref), it may be necessary to include lags of other, non-tax topics as well. In particular, the two prevalence measures for War on Terror and Natural Resources, Energy & Technology Granger-cause the tax topics and are related to political decisions likely to affect the federal budget.
In addition, we take into account other policy changes that may coincide with tax changes. In Appendix (ref), we investigate the contemporaneous relationships between the tax topics and the other identified topics. This analysis reveals that the topic related to regulatory policies (Public Administration) as well as the one we associate with polices aiming at improving long-run economic conditions (Economic Development) correlate significantly with the two tax prevalence measures. Thus, we add contemporaneous values of these two series to the set of regressors in ${\bm x}_{2,t}$. Finally, the president may at times refer to past tax changes in his speeches. This does not add new information about future tax changes and can therefore not be considered as news. 12 lags of RR's series of all tax changes are therefore included as additional controls in ${\bm x}_{2,t}$.\footnote{Our baseline specification does not include contemporaneous values of most controls; in Section (ref) we investigate the sensitivity to this choice. In Appendix (ref) we further investigate whether results are sensitive to the choice of control variables. We find that using various subsets of the regressors discussed above, does not affect the shape and magnitude of responses very much.}
Before discussing the response of economic activity to news about future tax changes, we investigate the strength of the correlation between the tax prevalence measures and our endogenous measure of future tax changes $x_{1,t}$. Table (ref) shows $F$-statistics on the joint exclusion of our tax prevalence measures from the first-stage of the LP-IV regressions, as specified above, for various choices of $M$ and $N$. The tax prevalence measures are fairly strong predictors of future tax changes implemented throughout the coming two years. Predictability is weaker for soon-to-be implemented tax changes. Predictive power also rapidly declines for tax changes implemented in the more distant future (after six quarters). \\
Estimated impulse responses of output and tax receipts to news about a future tax cut for different values of $M$ and $N$ are displayed in Figure (ref). To make results quantitatively comparable across different specifications, responses are normalised such that the log change of government tax receipts is minus unity at its trough. Point estimates are shown together with 68% and 90% confidence intervals.
We start by investigating responses of output and government receipts for a shorter anticipation horizon. Setting $M=1$ and $N=4$ means that we use our prevalence measures to instrument tax news about policy changes that will have been implemented throughout the following four quarters. Results are displayed in Figures (ref)(a) and (ref)(b). Tax receipts react almost instantaneously, declining for several quarters and reaching the trough response one year after the arrival of the news. The associated response of output is insignificant and close to zero for about four quarters. After that, output increases significantly and steadily, peaking after 10 quarters. Comparing the timing of the decline in tax receipts and the increase in output indicates that the expansion in GDP is primarily driven by (post-)implementation effects. The results do not provide evidence for significant anticipation effects. Instead, the slightly delayed but large positive response in GDP is qualitatively and quantitatively comparable with findings in the literature on effects of aggregate tax changes RomerRomer10,mertens2012empirical, providing additional evidence for large output effects due to tax cuts.
As alluded to above, it could be that pre-implementation effects depend on the anticipation horizon and, thus, on the specification of the aggregation window in ((ref)). mertens2012empirical indeed find that the longer the anticipation horizon, the stronger the pre-implementation contraction in output. By setting $M=5$ and $N=8$ we use the tax prevalence measure to instrument news about tax changes in the more distant future, i.e. policy changes that will have been implemented throughout four consecutive quarters in one year from now.\footnote{This specification is not chosen randomly. mertens2012empirical find that the median anticipation horizon (the time frame from signing of the bill to implementation) of anticipated tax changes (i.e. tax changes that are not implemented in the quarter they are enacted) is six quarters.} Output and tax receipts responses for this specification are displayed in Figures (ref)(c) and (ref)(d).
The results are strikingly different compared to the findings discussed before. First, the contraction in tax receipts is delayed, with a trough response eight quarters after the arrival of the news (instead of four as in the previous experiment). The response of aggregate tax revenue is initially slightly positive (although barely significant) and only turns negative after a year. Overall, the response of tax receipts confirms that by setting $M=5$ and $N=8$ we indeed seem to successfully instrument policy changes in the more distant future, which means a longer anticipation horizon. Second, and most importantly, the delayed tax cut does not simply trigger a delayed output expansion with an otherwise similar pattern as observed in Figure (ref)(a). While GDP still does not react at impact, output starts to decline after two quarters for about a year and only starts to recover simultaneously with the implementation of the tax cut. We thus conclude that the initial contraction of output is due to anticipation (resp. pre-implementation) effects. Once the policy is implemented and taxes are cut, economic recovery follows a similar path as in the previous specification, but with a smaller peak expansion of output.
Although we use a conceptually fundamentally different approach, our results confirm the evidence about contractionary effects of anticipated tax cuts in mertens2012empirical. We investigate next whether we also find that the longer tax cuts are anticipated the more severe the decline in economic activity. To extend the anticipation horizon, we move the aggregation window (by increasing $M$ and $N$) and instrument tax changes that will have been implemented in an increasingly distant future. The results displayed in Figure (ref)(e) and (ref)(f) confirm indeed that: (i) the further in the future the tax cut is implemented the stronger the contractionary pre-implementation effects on output; and (ii) the smaller the post-implementation effect triggering economic expansion.\footnote{In Appendix (ref) we investigate to what extend the window size in the construction of $x_{1,t}$ (i.e. $N-M$) affects the results. We find that allowing for a larger window over which future tax changes are aggregated does not qualitatively alter the results. In particular we find that fixing $M=1$ and extending the time window over which we aggregate tax policy changes (i.e. increasing $N$) does not lead to results indicating any anticipation effects. It seems indeed that the stronger correlation of the tax prevalence measures with soon-to-be implemented policy changes dominates.}
We investigate next whether the findings above are sensitive to specific modelling choices. We present results of this robustness analysis for the arguably more interesting case of a longer anticipation horizon ($M=5$ and $N=8$) for which we find contractionary pre-implementation effects on output. For an easier comparison, output responses to news about a tax cut, as displayed in Figure (ref), always contain estimates (and confidence intervals) of the benchmark specification presented in Section (ref). As before, output responses are normalised, such that the trough response of tax receipts is minus one. Robustness checks for a shorter anticipation horizon ($M=1$ and $N=4$) can be found in Appendix (ref).\footnote{Results for a shorter anticipation horizon are qualitatively and quantitatively very similar to the findings discussed in Section (ref).}
It is plausible that political considerations drive the president's communication. One could imagine that Republican presidents would conduct different economic policies and that people are more likely to expect tax cuts from them than from Democratic presidents (or would deem initial announcements more credible). Similarly, upcoming presidential elections may affect our results if the president is campaigning for re-election with potentially less credible announcements. Moreover, presidents with low approval ratings may be inclined to propose policy reforms that increase popularity. And finally, the president's relation with Congress may influence the tone in his public communications, as well as the kind of policy reforms brought forward. Figure (ref)(a) shows output responses incorporating these different political controls and indicates that none of these aspects seem to matter.\footnote{To control for the president's popularity in public, we consider quarterly averages of Gallup's job approval ratings. The House and Senate concurrence is the percentage of members of congress who agree with the president's position on a roll call vote. Both series are obtained from UC Santa-Barbara's American Presidency Project \url{https://www.presidency.ucsb.edu/}. We add current observations and 12 lags of these additional controls to the benchmark set of regressors in the local projection regressions.}
As endogenous measure of future taxes we have used RR's narrative records which quantify, in monetary terms, all post-war legislated tax policy changes. RR's narrative series is an obvious choice due to the similarity of the underlying auxiliary data (presidential speeches and addresses). However, alternative measures of tax liability changes may be considered. A standard measure to proxy tax changes is cyclically adjusted revenue.\footnote{We use the revenue measure constructed by RR which is expressed in terms of change in revenues as a percent of GDP. Thus RR's narrative series as well as their measure of cyclically adjusted revenue show tax changes in the same unit and are hence comparable.} Figure (ref)(b) shows the response of output when using this endogenous series to construct the measure of future tax changes in ((ref)). Our previous finding of an initial pre-implementation decline in economic activity is robust to this alternative choice. Using cyclically adjusted revenue as endogenous variable results in an even more severe pre-implementation contraction. In addition, we observe a slightly smaller (and delayed) expansion in output after the tax cut is implemented.
The underlying identification assumption of the LP-IV approach is that our tax prevalence series are exogenous conditionally on the control variables. As motivated in (ref), we include two prevalence measures as contemporaneous controls; one series related to regulatory policies and one associated with polices aimed at improving long-run economic conditions. All other control variables in ${\bm x}_{2,t}$ are included as lagged values. This is akin to a recursive identification structure in a VAR where tax news is ordered after the prevalence measures associated with regulatory and long-term economic policies. Although we find that our tax prevalence measures cannot be predicted by economic conditions (see Section (ref)), it may still be the case that, for example, current recessionary shocks trigger an instantaneous discussion about tax cuts. In this case, the considered identification scheme would be flawed. To mitigate these concerns to some extent, we estimate the output response using the set of regressors of our benchmark specification but varying the variables which we include as contemporaneous regressors (thus essentially considering all possible ordering in the recursive identification scheme). The results, displayed in Figure (ref)(c), indicate that the shape of output responses is little affected by different orderings; at worst, our benchmark specification may somewhat underestimate expansionary effects due to implemented tax cuts.
Finally, Figure (ref)(d) shows the effect of using different model specifications. Including eight instead of 12 lags of all past regressors has little quantitative impact on the response of output. Excluding any trend leads - unsurprisingly - to a more persistent response.
Appendix (ref) gathers additional robustness checks. We investigate the role of our instruments by comparing the benchmark results presented in Figure (ref) with impulse responses based on the regression specified in equation (ref) estimated by OLS, using RR's exogenous tax changes to construct $x_{1,t}$. We find that anticipated future tax cuts do not lead to an initial contraction of economic activity when RR's narrative is used to construct future tax changes. There is no significant response of GDP prior to implementation of the tax change, suggesting that our tax prevalence measures contain important information regarding anticipation of future tax changes.
In this paper we analyse the public communications of U.S. presidents to identify information regarding planned tax reforms. Our semi-supervised topic modelling approach allows us to automatically determine what issues are mentioned in the texts, and in particular when the president discusses tax changes. Based on those results we create a measure of prevalence for the tax increase and tax decrease topics, which reflects the relative prominence across time of those policy issues on the president's agenda. We show that they are strong predictors for a variety of measures of tax changes, including those usually considered as unanticipated. This predictive power is not connected to other macroeconomic conditions, but is rather the effect of capturing the legislative process behind those changes.
To analyse the effects of tax news, we use our identified topics to instrument future changes in tax liabilities and estimate the reaction of output to news about a future tax cut. We find that output when tax cuts are implemented. However, a longer anticipation horizon leads to initial contractionary effects in output and also dampens (post-)implementation expansion.
There are several methodological implications of our findings. First, we show that presidential speeches contain signals about future tax changes, and as such might be crucial for solving the problem of fiscal foresight. Moreover, it is possible to meaningfully quantify those signals using an automated text-analytic approach. In particular, we propose a two-step estimation procedure for the LDA model, in which informative lexical priors are constructed to differentiate between similar topics. Our approach, while relatively simple, proves crucial for determining the direction of the discussed changes.
As an extension, our two-step approach could prove helpful in identifying news regarding different types of taxes (cf. mertens2013dynamic), however the identification of relevant lexicons would be more challenging than in this study. Our approach could in principle be used to identify news about government spending. Because text-analytic methods scale well, in both cases the analysis can be augmented by considering other channels of communication such as congressional records, news outlets or even Twitter.
From an econometric perspective an important extension would be to integrate estimation of the text model and the causal econometric model into a single procedure. While considering two separate steps has the advantage that one can build on established methods for both the text mining and the causal modeling, a single unified framework could potentially better exploit the causality that runs in both directions. By directly extracting only the exogenous components of the president's speeches, one could measure the (noisy) news content of speeches more accurately and thereby mimic economic agents' reception of potential news more closely. Such an approach, however, would require the development of new models and estimation methods combining text and economic data in a single approach. Given the current abundance of textual data sources, such an approach would be greatly beneficial to economic analysis, and therefore seems a promising future research agenda.