Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
24,405 characters · 6 sections · 0 citation commands
On the solution of the variational optimisation in the rational inattention framework
In two prominent papers (Sims 2003, 2006) Christopher A. Sims proposed to model decision under uncertainty as the optimal choice of the joint distribution of action $Y$ and external state $X$, under the constraint on the flow of information. It is assumed that the marginal distribution of $X$ is known, and the information flow is quantified as the mutual information of $X$ and $Y$, $I\left( X,Y\right) =H\left( X\right) +H\left( Y\right) -H\left( X,Y\right) =H\left( Y\right) -H\left( \left. Y\right \vert X\right) $, where for a random variable $W$ with distribution $p$, $H\left( W\right) \equiv -E\left[ \log _{2}p\right] $. This approach to optimisation under uncertainty belongs to a more general concept of rational inattention \ introduced by Sims, which within the last fifteen years has developed into a large literature, with applications to consumption, price and wage setting, and portfolio choice (Wiederholt, 2017).
Examples in Sims (2003, 2006) are maximisation of expected utility or minimisation of expected loss, with continuous distribution functions. The objective and the constraint are, therefore, definite integrals of unknown functions, and the optimisation problem is solved by finding an extremum of a functional. While in several follow-up applications the optimisation is carried out numerically, these two papers present analytical characterisation of the solution for several special cases. However, the analysis appears to have a fundamental flaw. Below, I outline the framework proposed by Sims and focus on two examples, a quadratic loss function (Sims 2003) and a two-period model of consumption and savings with logarithmic utility (Sims 2006).\footnote{ One of the working paper version of Sims (2006) is Sims (2005). The latter provides some details of analytical derivations of the results presented in the former.} The aim of my paper is twofold. First, it shows how the correct characterization of the solution can be obtained, using these two examples. Second, it demonstrates the restrictiveness of this framework, which suggests that it is unlikely to apply to a wider set of objective functions and distributions arising in economic models.
The rational inattention models are built on the assumption that an economic agent has a limited capacity for processing information when making a decision. An agent chooses an action taking into account an external state. The state cannot be perfectly observed, and both the action and the state are assumed to be random variables. The agent knows the distribution of the state which is fixed exogenously. The objective of the agent is to maximise some criterion function, $CF$, such as the expected utility or negative of the expected loss. Let $Y\in \mathcal{Y}$ be an action in the action space $ \mathcal{Y}$ and let $X\in \mathcal{X}$ be a state with distribution $ p\left( x\right) $ defined over space $\mathcal{X}$. Let $f\left( x,y\right) $ describe the joint distribution of $X$ and $Y$. The assumed limit on the agent's capacity to process information is modelled as the constraint on the mutual information between $X$ and $Y$. Thus, the agent solves
where $p\left( x\right) =\int \nolimits_{\mathcal{Y}}dy$ $f\left( x,y\right) $ and $g\left( y\right) =\int \nolimits_{\mathcal{X}}dx$ $f\left( x,y\right) $ are the marginal distribution. Sims (2003, 2006) suggested to use the joint distribution as the instrument of optimisation. Since $p\left( x\right) $ is fixed, this is equivalent to choosing the distribution of $Y$ conditional on $X$. When $X$ and $Y$ are continuous random variables, the agent's problem is
where $q\left( \left. y\right \vert x\right) =\frac{f\left( x,y\right) }{ p\left( x\right) }$ is the conditional distribution of action choice. This is a constrained optimisation problem of the calculus of variations (see, for example, Smirnov et al. 1933), since the unknown is a function, and the objective and the constraint are functionals. The problem in ((ref) ) is equivalent to the maximisation of a Lagrangean,
where $\widetilde{\lambda }\geq 0$ is the Lagrange multiplier, such that $ \widetilde{\lambda }>0$ when the constraint is binding (holds with equality) and $\widetilde{\lambda }=0$ otherwise. In addition, one needs to specify some boundary conditions for $q\left( \left. y\right \vert x\right) $. The natural boundary condition in this setting is the normalisation,
It is known from the calculus of variations that the necessary condition for an extremum of functional,
of function $y\left( x\right) $, with boundary condition $\left. y\left( x\right) \right \vert _{\left( x\right) \in \partial \left( \mathcal{X} \right) }=y_{0}\left( x\right) $, is given by $\delta \mathcal{F}=0$, leading to an Euler equation,
which, in general, can be rewritten as an ordinary differential equation of second order with respect to $x$. The general solution is a family of curves, and a particular solution is found from the boundary conditions. Similarly, the necessary condition $\delta \mathcal{F}=0$ for the extremum of functional
of function $z\left( x,y\right) $ of two variables, $x$ and $y$, with boundary condition $\left. z\left( x,y\right) \right \vert _{\left( x,y\right) \in \partial \left( \mathcal{X\times Y}\right) }=z_{0}\left( x,y\right) $, leads to the Euler equation given by
which, in general, is equivalent to a partial differential equation of second order. The general solution is a family of surfaces, and a particular solution is found from the boundary conditions. For a constrained optimisation the objective functional includes a term associated with the constraint with the Lagrange multiplier, and the corresponding first-order condition is known as the Euler-Lagrange equation.
When the objective function does not contain the derivatives of the unknown function, the necessary condition for the extremum, $\frac{\partial F}{ \partial y}=0$ for $y\left( x\right) $ in ((ref)), or $\frac{\partial F}{ \partial z}=0$ in ((ref)), is not a differential equation. The extremum in this case is described by $y=\varphi \left( x\right) $ (or, respectively, by $z=\varphi \left( x,y\right) $), and, in general, the solution does not exist, although the problem may have a solution in exceptional cases (Smirnov et al., 1933, p. 14). In other words, an extremum that satisfies the given boundary conditions may only exist for some exceptional boundary conditions.
One can see immediately that functional $\mathcal{L}$ in ((ref)) does not contain the derivatives of the unknown function. Therefore, the Euler-Lagrange equation for this optimisation problem is not a differential equation, and the solution does not, in general exist, -- in a sense that function $q\left( \left. x\right \vert y\right) =\varphi \left( x,y\right) $ that maximises $\mathcal{L}$ in ((ref)) may not satisfy condition ((ref)).
Suppose, however, that a solution exists for some exceptional case. Then it must satisfy the Euler-Lagrange equation, which for ((ref)) can be shown \footnote{ See Appendix for details.} to have the form
with boundary condition ((ref)), or, equivalently,
with boundary condition
where $\lambda \equiv \frac{\widetilde{\lambda }}{\ln 2}$, and natural logarithm is introduced for convenience in further derivations.
The potential solution is now analysed for two examples of $U\left( x,y\right) $ presented in Sims (2003, 2006).
Consider the problem of minimisation of the expected value of a linear-quadratic loss function\footnote{ This example can also be interpreted as maximisation of the expected value of a linear-quadratic utility in a two-period model of consumption and saving, allowing for negative consumption and wealth; see Sims (2005).},
This is a generalisation of the quadratic loss function ($\varphi =\theta =1$ , $b=c=0$) considered in Sims (2003), where it is stated that `when the $X$\ distribution is Gaussian, it is not too hard to show that the optimal form for $q$\ is also Gaussian, so that $Y$\ and $X$\ \ end up jointly normaly distributed' (p. 670). As I show below, Gaussian $q$ as a solution of ((ref)) given Gaussian $p$ only exists and satisfies the properties of a distribution function under certain restrictions on all but one of the loss function parameters.
Let $X\sim N\left( \mu _{x},\sigma _{x}^{2}\right) $. With $N\left( \mu _{\left. x\right \vert y},\sigma _{\left. x\right \vert y}^{2}\right) $ as a guess for $h\left( \left. x\right \vert y\right) $, we have
and, setting $I\left( X,Y\right) =\kappa $ gives
Next, using the properties of the conditional and marginal densities of the bivariate Gaussian distribution\footnote{ For the conditional distribution the mean and the variance are given by $\mu _{\left. x\right \vert y}=\mu _{x}+\rho \frac{\sigma _{x}}{\sigma _{y}} \left( y-\mu _{y}\right) $ and $\sigma _{\left. x\right \vert y}^{2}=\sigma _{x}^{2}\left( 1-\rho ^{2}\right) $.} we obtain from ((ref)) the expression for the Lagrange multiplier,
and the following set of relationships among the model parameters (see Appendix for details):
where $\mu _{y}$ is determined from
Equations ((ref)) and ((ref))-((ref)) effectively restrict three out of four parameters of the loss function, given $\kappa $ and $\left( \mu _{x},\sigma _{x}^{2}\right) $, for the optimisation problem to have conditional Gaussian distribution as a solution. Suppose, we fix $\theta $; this, along with ((ref)), determines $\varphi $ in ((ref)), and with $\mu _{y}$ calculated from ((ref)), determines $b$ and $c$ by ((ref)) and ((ref)). One can see that restrictions $\theta =\varphi =1$ and $b=c=0$ cannot hold simultaneously, and so the solution for $q$ in the case of quadratic loss function $-\left( Y-X\right) ^{2}$ analysed in Sims (2003) does not exist.
For $\mu _{x}=0$ and $\theta =1$ we have $\mu _{y}=\sqrt{\kappa \ln 2}$ and
In this case the optimal $q\left( \left. y\right \vert x\right) $ is Gaussian with
but this solution only exists for
This example is different in one important way which highlights how restrictive the variational approach is in the rational inattention framework. In the previous example the distributions of the state and action variables allow, in principle, for an unbounded support, and so a solution could be constructed for a suitable, albeit restricted, choice of the model parameters. When the nature of economic variables dictates the bounds on the support of the distribution (for example, non-negativity), the solution may not exist for any configuration of the remaining model parameters, -- the existence of bounds, in effect, poses additional restrictions that cannot be met simultaneously.
The following example of a two-period consumption-savings model with logarithmic utility was analysed in Sims (2006).\footnote{ In Sims (2005, 2006) $\beta =1$ and the notations correspond to $w=x$ and $ c=y$.} An individual with random endowment $X>0$ chooses how to allocate $X$ between consumption, $Y\leq X$, in the first period, and savings, $X-Y$, to be consumed in the second period. The objective is to maximise the expected utility function, $E\left[ U\left( X,Y\right) \right] $, where $U\left( X,Y\right) =\ln Y+\beta \ln \left( X-Y\right) $. The distribution of $X$ is given by $p\left( x\right) $, and the individual chooses $q\left( \left. y\right \vert x\right) $ under the constraint on the information flow.
A potential solution for $q\left( \left. y\right \vert x\right) $, if it exists, must be consistent with ((ref)):
This can be rewritten as
Because the support of the distribution is bounded, in order to satisfy ((ref)) it must be the case that
This can be verified directly:
Thus,
This formally resembles the expression obtained by Sims (2006)\footnote{ Sims (2005) derives the expression for the conditional density which contains a Lagrange multiplier on the marginal density constraint. This appears to be incorrect; see Appendix. Furthermore, the integrand in equation (7) in Sims (2006) (the same in Sims 2005) is not proportional to the density of $F\left( 2\alpha +2,2\alpha \right) $ distribution, -- contrary to what is stated in the paper (p. 161). For this to be the case the term in parentheses in the integrand should be $\left( v+\frac{\alpha }{ 1+\alpha }\right) $, rather than $\left( v+1\right) $.} with $\beta =1$. The conditional mean of $X$ exists for $\alpha >1$ (that is, for $\lambda <1$, so $\widetilde{\lambda }<\ln 2$) and is given by
Thus, $\frac{E\left[ \left. X\right \vert Y=y\right] }{y}=\left( 1+\beta \right) \frac{\alpha }{\alpha -1}>1+\beta $, whereas the certainty solution is $x/y=1+\beta $, -- consistent with the argument that the rational inattention solution is closer to the certainty solution, the lower is the shadow price of the information constraint.
As shown above, this solution for $h\left( \left. x\right \vert y\right) $ exists if $p\left( x\right) $ is a power law distribution ((ref)) with support $\mathcal{X}=\left[ x_{0},\infty \right) $ for some $x_{0}>0$. The normalisation condition,
determines the Lagrange multiplier, $\widetilde{\lambda }=\frac{\ln 2}{ \alpha }$, implicitly as a function of the model parameters:
However, it is impossible to construct the solution for $q\left( \left. y\right \vert x\right) $ that satisfies boundary condition ((ref)). Formally,
and
since the integrand is non-negative on $\left[ x_{0},x\right] $ and is strictly positive at least on some subinterval of $\left[ x_{0},x\right] $. However, ((ref)) implies $\frac{d}{dx}\int \limits_{\mathcal{Y}}dy$ $ q\left( \left. y\right \vert x\right) =0$. Therefore, the Euler-Lagrange equation in this example does not have a solution that would satisfy this condition.
The rational inattention framework has gained popularity as an alternative to the rational expectations approach to the decision under uncertainty. It is based on a plausible assumption that an economic agent has a limited amount of attention and allocates it optimally among available bits of information. However, formalisation of the solution as the optimal choice of conditional distribution of action given the exogenous distribution of the external state is not a well-posed problem, and the solution, in general, does not exist. This paper demonstrates that this approach may lead to a solution in one special case of the linear-quadratic objective with the Gaussian distribution of the state, under a specific choice of the model parameters that has no obvious interpretation. It is unlikely to be applicable to a wider set of problems that are of interest for economists. Other solution concepts used in the rational inattention literature can prove more fruitful in further developments.