EconBase
← Back to paper

Optimal Treatment Allocation under Constraints

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

144,001 characters · 15 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Treatment Allocation under Constraints

abstractIn optimal policy problems where treatment effects vary at the individual level, optimally allocating treatments to recipients is complex even when potential outcomes are known. We present an algorithm for multi-arm treatment allocation problems that is guaranteed to find the optimal allocation in strongly polynomial time, and which is able to handle arbitrary potential outcomes as well as constraints on treatment requirement and capacity. Further, starting from an arbitrary allocation, we show how to optimally re-allocate treatments in a Pareto-improving manner. To showcase our results, we use data from Danish nurse home visiting for infants. We estimate nurse specific treatment effects for children born 1959-1967 in Copenhagen, comparing nurses against each other. We exploit random assignment of newborn children to nurses within a district to obtain causal estimates of nurse-specific treatment effects using causal machine learning. Using these estimates, and treating the Danish nurse home visiting program as a case of an optimal treatment allocation problem (where a treatment is a nurse), we document room for significant productivity improvements by optimally re-allocating nurses to children. Our estimates suggest that optimal allocation of nurses to children could have improved average yearly earnings by USD 1,815 and length of education by around two months.

Introduction

Across a number of settings decision makers wish to make informed choices regarding provision of treatments to recipients in a way that maximizes some score.\footnote{Our use of treatment is broad and may include medical treatments, adverting, provision of services such as education, and more. Likewise, our use of recipient is broad and may include patients, customers, students, and more.} Examples include kim2011battle, chan2012optimizing, ozanne2014development, athey2017beyond, swaminathan2017off, fukuoka2018objectively, zhou2018evaluating, bastani2018interpreting, ban2019big, farias2019learning, bastani2020sequential, bertsimas2020predictive, bastani2020online, ranging from healthcare to digital advertising and public policies. A common theme across these settings is heterogeneous treatment effects,\footnote{For example: Given two patients, one might respond better to drug A while the other responds better to drug B.} in turn leading to benefits from “personalized treatment”.\footnote{This is also known as “personalized medicine” jain2002personalized, hamburg2010path, schork2015personalized, and can be seen as a subset of the broader setting of personalized treatment, by which is meant a broader definition of treatments than exclusively medical treatments.} However, to exploit this information requires insight into (estimates of) potential outcomes of recipients under different treatments and a policy mapping that, based on potential outcomes, is able to optimally allocate treatments to recipients, potentially while respecting constraints related to recipient treatment rights, treatment capacities, and equity concerns.

In this paper, we tackle the issue of optimally allocating treatments to recipients under constraints. Related to our work is the literature on policy learning dudik2011doubly, zhang2012estimating, zhao2012estimating, zhao2015doubly, swaminathan2015batch, zhou2017residual, kallus2018confounding, kitagawa2018should, athey2021policy, zhou2023offline. While most of the existing literature focuses on binary treatments, exceptions include swaminathan2015batch, zhou2017residual, kallus2018balances, zhou2023offline. What sets this study apart, however, is our focus on optimal treatment allocation, and we attempt to solve the computational challenges related to optimal policy design given potential outcomes that arise due to large numbers of treatments and recipients and various constraints related to treatment allocation. As such, our contribution is in developing a general method for finding policies that exactly maximize some objective function within a class of policies.\footnote{Or minimize some objective function, depending on whether a larger value signifies a better outcome.} Thus, our work is related to the task of finding the policy that exactly minimizes the approximate value function within the class of tree-based policies discussed in zhou2023offline, but where zhou2023offline use a general mixed integer program (MIP) which quickly becomes computationally infeasible, our approach, based on a graph representation of the problem, is more suitable for larger problems.\footnote{Recall we abstract from any complications related to forming an approximate value function using estimators, and thus our objective is exclusively related to finding the optimal policy given a value function. zhou2023offline consider the more general task of multi-action policy learning with observational data. Our contribution could help with certain aspect of this problem but is not an alternative to the full task.}

Our contribution is in proving that the general case of optimally allocating treatments to recipients taken potential outcomes as given is strongly polynomial, even when (1) treatment effects are allowed to vary arbitrarily between recipients, (2) there are arbitrarily many distinct types of treatments available, and (3) each treatment must respect capacity constraints. In an extension, we show that our results apply also when only Pareto-improvements are allowed, i.e., in settings where a re-allocation of treatments is not allowed to make any individual worse off.\footnote{Our approach lends itself well to extensions such as differential capacities by treatment and respecting constraints related to certain types of treatments not being eligible for some recipients.} For our proofs, we represent the optimal treatment allocation problem as a flow problem in an appropriately constructed network.\footnote{See the textbooks by ford1962flows, ahuja1993network, bang2008digraphs for introductions to flows in networks.} In addition to proving our complexity results, this representation also lends itself well to the construction of an algorithm for solving the problem.

To showcase our method, we consider nurse home visiting (NHV) in Denmark, where we consider each individual nurse a treatment and each child a recipient. NHV for infants started in 1937 in Denmark, and previous work has documented large, positive impacts of NHV and center care in Denmark wust2012early, hjort2017universal as well as in Norway buetikofer2019infant, Sweden bhalotra2017infant, and the US hoehn2021long. We combine a rich dataset consisting of nurse journals following infants during their first year of life with Danish administrative data, providing information on education and labor market outcomes.\footnote{In joint, concurrent work, baker2023universal, bjerregaard2023cohort have transcribed the contents of the nurse records using machine learning, based mostly on manual transcriptions from andersen2012weight, bjerregaard2014bmi.} The nurse journals originate from the Danish NHV program and cover all children born between 1959-1967 and living in Copenhagen during their first year of life. They contain information on birth characteristics (including birth weight, birth length, and preterm status), detailed data from visits taking place at child age two weeks and one, two, three, four, six, nine, and 12 months (including weight development, nutrition, and home and parent characteristics), and, crucially for this study, the name of the nurse performing the visits for each child. Linking the journals to Danish administrative data, we can follow the development of these children until around age 50-60 (depending on cohort and register), from which we obtain information on education and labor market outcomes.

The NHV setting lends itself well to our proposed method: First, using causal machine learning combined with detailed pre-treatment data for each child and her family, following children throughout their lives using Danish register data, and exploiting quasi-random allocation of nurses to children within nurse districts (and years), we can obtain individual-level treatment effects for all children of the Copenhagen NHV program, in turn allowing us to obtain estimates of potential outcomes under different treatments at the individual level. Second, the setting is complex in that there are large numbers of distinct treatments and recipients, each treatment is capacity-constrained (for example, allocating all children to the same nurse is not a valid solution), each recipient must be treated by exactly one treatment (exactly one nurse allocated each child), and we may want to impose Pareto-improvement constraints to make sure a re-allocation does not leave any child worse off.\footnote{Or even that some children must not be left worse off, e.g., those from families at the bottom of the socioeconomic distribution.}

We start by documenting that while the initial conditions of children (as measured by at-birth characteristics, such as birth weight, and parent characteristics) vary significantly between nurses, this is largely driven by differences between nurse districts as well as birth years. Once we remove these effects, differences in pre-treatment variables are much less pronounced, supporting our identification strategy relying on quasi-random allocation of nurses to children within district-by-year groups. However, clear differences in outcomes remain between children to whom different nurses were allocated. This suggests that nurses vary in terms of their treatment effects.\footnote{We also directly estimate differences in nurse treatment effects, results which are of independent interest. As we show later, the average (absolute) difference (beyond that which can occur due to noise) in length of education between groups of children for two separate nurses is around one month and in average income during ages 25-50 about 1.4%.}

We next turn to measuring child-specific treatment effects for each nurse. Here, we once again exploit the quasi-random allocation of nurses to children within district-by-year groups, and now estimate heterogeneous treatment effects using causal machine learning while exploiting our information on pre-treatment child and family characteristics.\footnote{Specifically, we use multi-arm causal forest wager2018estimation, athey2019generalized, nie2021quasi combined with child and family characteristics from the nurse records transcribed using machine learning, as well as information obtained from Danish registers by linking the transcribed nurse records to the registers.} This allows us to obtain estimates of potential outcomes under different treatments, and using these we are able to estimate counterfactual distributions of child outcomes under different treatment allocation policy rules. Using our method for finding optimal policy rules, we are able to estimate the maximum benefit from re-allocating nurses in a more efficient manner, as well as able to estimate the share of total theoretical benefits from re-allocation attained through other policy rules. Our estimates suggest that optimal allocation of nurses to children could have improved average yearly earnings by USD 1,815 (4%) and total length of education by around two months.

We contribute to two broad strands of literature. First, we contribute to the literature on optimal policy learning manski2004statistical, hirano2009asymptotics, dudik2011doubly, zhang2012estimating, zhao2012estimating, zhao2015doubly, hirano2020asymptotic, athey2021policy, zhou2023offline, specifically on the problem of finding optimal policies taken potential outcomes as given. Finding optimal policies is important in the sense of optimally allocating resources and thus achieving full efficiency, but the complex allocation mechanism to achieve this might not always be desirable; however, even in cases where simpler allocation mechanisms are desired, comparing such mechanisms with optimal rules has the benefit of allowing policy makers to correctly weigh the benefits of simplicity against its costs in terms of efficiency losses. In doing so, we also showcase how results from graph theory, and specifically network flows ford1956network, ford1962flows, murty1992network, ahuja1993network, dolan1993networks, are helpful in solving complex economic problems in settings where these can be represented by appropriately constructed graphs.

Second, we contribute to the literature on the impact of early-life policies, and in particular the role of treatment providers. The role of early-life circumstances on short- and long-run outcomes of individuals has long been an active area of research across disciplines, including economics Forsdahl1979, AlmondCurrieDuque2018. Within the realm of research investigating the enduring significance of early-life health policies, two primary currents emerge:\footnote{In addition to studies on the positive role of early-life investments, other studies have investigated the negative effects of early-life shocks, such as extreme weather events, and their interactions with early-life investments policies adhvaryu2018helping, duque2018early, garg2020temperature, aguilar2022nino.} The first strand consists of randomized trials that implement high-intensity targeted model programs such as the U.S. Nurse Family Partnership and the Perry Preschool Program, interventions which have underscored the substantial and lasting dividends reaped from precisely targeted investments in the well-being and development of socioeconomically disadvantaged children Oldsetal1986, Oldsetal1998, olds2019prenatal, belfield2006high, heckman2010rate. The second strand exploits naturally occurring variations in access to early-life health policies, and noteworthy examples include diverse interventions such as nutritional and income support programs hoynes2016long, barr2022investing, bailey2023safetynet, barr2023fighting, healthcare insurance and services wherry2018childhood, goodman2018public, miller2019long, noghanibehambari2022intergenerational, east2023multi, early educational initiatives rossin2020preschool, bailey2021prep, anders2023effect, as well as infant home visiting and center-based care hjort2017universal, bhalotra2017infant, buetikofer2019infant, hoehn2021long.\footnote{In concurrent work, baker2023universal combine some of the strengths of these two strands by evaluating a large-scale government trial of NHV in Copenhagen during the 1960s that quasi-randomly allocated some children to three vs. one year of NHV. This combines individual-level data on early-life circumstances with long-run administrative data on outcomes.} Where we depart from much of this existing literature is on our focus on the role of treatment providers, rather than the effect of a program. Understanding the role of providers is likely a key element in improving our understanding of optimal design and targeting of such programs, and we contribute with novel evidence on the magnitude of between-provider treatment effects and heterogeneity with respect to child characteristics. Key to facilitating our application is the development and use of machine learning (ML) for handwritten text recognition, which we use to transcribe the handwritten nurse records and link those with administrative data.\footnote{In concurrent work, bjerregaard2023cohort, baker2023universal transcribe handwritten nurse records from the 1960s Copenhagen NHV program and link those to Danish administrative data to construct a large cohort of children (92,902) to track from birth throughout their life. baker2023universal use this data to study the impact of extended NHV, exploiting a 1960s trial that quasi-randomly allocated children to three vs. one year of NHV. This cohort also forms the population of children we use for our application.} In doing so, we also contribute to the literature on layout detection clinchant2018tables, shen2021layoutparser, dahl2023bdad, dahl2023tableparser and scene, optical character, and handwritten text recognition goodfellow2013multi, bluche2014htr, lee2016recursive, bluche2017scan, bartz2021htr, geetha2021htr, kang2022htr, dahl2022dare, dahl2023hana.

The paper proceeds as follows: In Section (ref), we derive our allocation algorithm for solving the optimal allocation problem. Section (ref) then presents the institutional background of the 1960s Copenhagen infant NHV program and Section (ref) presents our data. Section (ref) describes our approach for estimation and Section (ref) presents our results. Section (ref) concludes.

Optimal Allocation

We define the optimal allocation problem as the problem of optimally allocating treatments to recipients (where a number of treatments are available), under constraints regarding recipients (receive exactly one treatment) and treatments (at most $m$ recipients can receive given treatment).\footnote{This encompasses the nurse allocation problem discussed in the paper, but it is not specific to it and may encompass a host of other such types of problems. It is straightforward to extend the problem to settings where capacities vary by treatments.} Formally, we define the problem as:

definition[Optimal allocation problem] Let $Y_i^j \in \mathbb{R}$ denote the outcome of recipient $i$ under treatment $j$ and let $D_i^j \in \{0, 1\}$ be one if recipient $i$ receives treatment $j$. The optimal allocation problem is then the problem of choosing $D_i^j$ such that the maximum average realized outcome is achieved, while respecting that each recipient must be allocated exactly one treatment and no treatment is allocated to more individuals than its capacity allows.

Noting that $D_i^j\in\{0, 1\}$ is $1$ if individual $i$ is under treatment $j$ and $0$ otherwise, we see that assignment of everyone to exactly one treatment ($\sum_j D_i^j = 1$) and respecting capacities ($\sum_i D_i^j \leq m$) leads to a binary program:\footnote{If capacities of different treatments vary, we may instead write $\sum_i D_i^j \leq m_j$ as the second set of conditions. Otherwise the problem remains unchanged.}

align*[align* omitted — 260 chars of source]

where $n$ is the number of recipients, $k$ the number of treatments, and $m$ the “capacity” of each treatment, i.e., the maximum number of recipients a given treatment may be allocated to. The solution to this binary program is exactly the optimal solution to the allocation problem of Definition (ref).

We claim there exists a strongly polynomial algorithm for solving the allocation problem (despite its potential $\mathcal{NP}$-completeness given its structure as a binary program). For our proof, we shall make use of flows in networks by appropriately constructing a network and representing the optimal allocation problem as an integer minimum cost feasible flow-problem in the constructed network.\footnote{See the textbooks by ford1962flows, ahuja1993network, bang2008digraphs for introductions to flows in networks.}

Let $V_1$ be the set of treatments and $V_2$ the set of recipients, and let $s$ and $t$ be special vertices. Define now $V := V_1 \cup V_2 \cup \{s, t\}$ as the set of vertices of a digraph and let $A := \{sv_1 \, : \, v_1 \in V_1\} \cup \{v_1v_2 \, : \, v_1 \in V_1, v_2 \in V_2\} \cup \{v_2t \, : \, v_2 \in V_2\}$ be the arcs of the digraph, i.e., $D = (V, A)$. Define now the network $\mathcal{N} := (V, A, l, u, b, c)$, where $b: V \rightarrow \mathbb{R}$ is a function on the vertices and $l: A \rightarrow \mathbb{R}$, $u: A \rightarrow \mathbb{R}$, and $c: A \rightarrow \mathbb{R}$ are functions on the arcs. By appropriately choosing $b$ and $l$, $u$, and $c$, it turns out that an integer minimum cost feasible flow in the network exactly corresponds to the optimal solution of the allocation problem. Formally, we define the network as:

definition[Network representation] Let $\mathcal{N} := (V, A, l, u, b, c)$ be defined as above and let, for all $v_1 \in V_1$ and $v_2 \in V_2$, $l_{v_1v_2} = 0$, $u_{v_1v_2} = 1$, and $c_{v_1v_2} = -Y_{v_2}^{v_1}$.\footnote{In some algorithmic settings non-negative costs are desired. This is easily guaranteed by adding $\max \{Y_i^j\}$ to all costs (leaving the problem unchanged). We refrain from this for notational simplicity.} Further, let, for all $v_1 \in V_1$, $l_{sv_1} = 0$, $u_{sv_1} = m$, and $c_{sv_1} = 0$ and let, for all $v_2 \in V_2$, $l_{v_2t} = 1$, $u_{v_2t} = 1$, and $c_{v_2t} = 0$. Finally, let $b(v) = 0$ for all $v \in V \setminus \{s, t\}$ and $b(s) = -b(t)$.\footnote{This, in turn, implies that $\sum_{v \in V}b(v) = 0$.} We shall call the network constructed this way the network representation of an instance of the optimal allocation problem of Definition (ref).

A flow in a network is a function on the arcs $x: A \rightarrow \mathbb{R}_+$, and we call it an integer flow if $x: A \rightarrow \mathbb{N}_0$. For a given flow, the balance vector of $x$ is

align*[align* omitted — 97 chars of source]

We say a flow $x$ is feasible if, for all $uw \in A$, $l_{uw} \leq x_{uw} \leq u_{uw}$ and, for all $v \in V$, $b(v) = b_x(v)$. Further, we define the cost of a flow as $\sum_{uw \in A}c_{uw}x_{uw}$. We are now in a position to prove the correspondence between the optimal allocation problem (Definition (ref)) and its network representation (Definition (ref)).

theoremAny integer feasible flow $x$ in $\mathcal{N}$ defined as in Definition (ref) corresponds to a valid solution to the treatment allocation problem defined as in Definition (ref). Further, the cost of any feasible flow $x$ is exactly the negative value of its corresponding value in the treatment allocation problem.
proofWe sketch the main points of the proof here. Appendix (ref) provides further details. Any integer feasible flow $x$ in $\mathcal{N}$ makes use of at most $m$ arcs from each $v_1 \in V_1$ (i.e., respects capacities) and exactly one arc from each $v_2 \in V_2$ (i.e., allocates exactly one treatment to each recipient). Let $D_{v_2}^{v_1} = x_{v_1v_2}$ for all $v_1 \in V_1$ and $v_2 \in V_2$. Then the set of $D_{v_2}^{v_1}$s satisfies the constraints of the binary program formulation of Definition (ref) and is thus a valid solution to the allocation problem, with cost $\sum_{uw \in A}c_{uw}x_{uw} = -\sum_{v_2 \in V_2}\sum_{v_1 \in V_1}Y_{v_2}^{v_1}D_{v_2}^{v_1}$.

It follows straightforwardly that the allocation problem has a solution if and only if an integer feasible flow exists in its network representation. Further, any integer minimum cost feasible flow in an allocation problem's corresponding network representation corresponds to the optimal value of the allocation problem.

corollaryThe allocation problem as defined in Definition (ref) has a solution if and only if $m|V_1| \geq |V_2|$, and any integer minimum cost feasible flow in its network representation (Definition (ref)) corresponds to an optimal solution of the associated allocation problem.
proofClearly, $\sum_j D_i^j = 1 \implies \sum_i\sum_j D_i^j = |V_2|$ and $\sum_iD_i^j \leq m \implies \sum_i\sum_j D_i^j \leq m|V_1|$. Hence, $m|V_1| \geq |V_2|$. Since, by Theorem (ref), the cost of any integer feasible flow in the network representation of the allocation problem is equal to the negative value of the optimal solution to the allocation problem, finding an integer minimum cost feasible flow in the allocation problem's network representation corresponds to finding an optimal solution to the allocation problem.

To prove that a strongly polynomial algorithm for solving the optimal allocation problem exists, it suffices to show that an integer minimum cost feasible flow in its corresponding network can be found in strongly polynomial time.\footnote{It also requires that the network representation of an allocation problem can be found in strongly polynomial time, but this may be easily verified.}

theoremLet $\mathcal{N} := (V, A, l, u, b, c)$ denote the network representation (Definition (ref)) of the allocation problem (Definition (ref)) and let $n_1$ denote the number of treatments and $n_2$ the number of recipients. Given the existence of an integer feasible flow in $\mathcal{N}$, an integer minimum cost feasible flow $x$ in $\mathcal{N}$ (and, in turn, an optimal solution to the allocation problem) can be found in time $\mathcal{O}\left[(n_1^3n_2^2 + n_1^2n_2^3) \log^2 (n_1 + n_2)\right]$.
proofWe sketch the main points of the proof here. Appendix (ref) provides further details. Let $n = |V|$ denote the number of vertices and $m = |A|$ the number of arcs in the underlying digraph $D$ of $\mathcal{N}$. Noting that capacities (and lower bounds and balance vectors) are integers, using the cancel-and-tighten algorithm of goldberg1989finding allows us to find a minimum cost feasible flow in $\mathcal{N}$ in time $\mathcal{O}(nm^2 (\log n)^2)$ (even when costs are arbitrary real-valued). Noting that the cancel-and-cut algorithm is a variant of the cycle cancelling algorithm klein1967primal, we may use a convenient integrality property of minimum cost flows, namely that given all integer lower bounds, capacities, and balance vectors, there exists an integer minimum cost flow: At each iteration, we augment our flow along the cycle $C$ by $\delta(C)$, by which we mean the minimum residual capacity of any arc on $C$ in $\mathcal{N}$, and thus our new flow after any iteration is $x^{t+1} = x^t \oplus \delta(C)$. Since the residual capacity on any arc is always an integer, $x^{t+1}$ is an integer flow so long as $x^t$ is, and thus by induction we have an integer minimum cost flow. Noting that $n = 2 + n_1 + n_2$ and $m = n_1 + n_2 + n_1n_2$, we have $\mathcal{O}(n m^2 (\log n)^2) = \mathcal{O}\left[(n_1^3n_2^2 + n_1^2n_2^3) \log^2 (n_1 + n_2)\right]$.

For our application of allocating nurses to infants, our ultimate goal is to provide guidance for a re-allocation of nurses in such a way as to maximize their “treatment effects”. In cases where nurses differentially affect children, i.e., there exists $i_1, i_2, j_1, j_2, i_1 \neq i_2, j_1 \neq j_2$ such that $Y_{i_2}^{j_1} - Y_{i_1}^{j_1} > Y_{i_1}^{j_2} - Y_{i_2}^{j_2}$, re-allocating nurse $j_1$ from child $i_2$ to $i_1$ (and nurse $j_2$ from child $i_1$ to $i_2$) will improve average outcomes.\footnote{However, there is no guarantee that neither child is made worse off, an issue we turn to in Section (ref).} While we do not observe all $Y_i^j$, we may use sample estimates of these, as may be obtained by estimating heterogeneous nurse treatment effects by child pre-treatment characteristics. Thus, our approach is designed to solve allocation problems once (estimates of) potential outcomes are obtained, and is then guaranteed to always be able to find the (non-parametric) optimal allocation of treatments to recipients in strongly polynomial time.

In a setting such as the one above, the number of treatments is required to grow as the number of recipients does, but this might not be the case if treatment is, e.g., a drug or an advertisement campaign. In such cases we obtain complexity $\mathcal{O}\left[n_2^3 (\log n_2)^2\right]$.

Pareto-Improvement Guarantees

While our approach taken thus far does not guarantee that no child is made worse off, it is easily extendable to cover scenarios where only Pareto improvements are allowed. Given some “initial solution” (allocation of treatments to recipients), we call this new problem the Pareto-guaranteed optimal allocation problem and formally define it as:

definition[Pareto-guaranteed optimal allocation problem] Let $Y_i^j \in \mathbb{R}$ denote the outcome of recipient $i$ under treatment $j$ and let $D_i^j \in \{0, 1\}$ be one if recipient $i$ receives treatment $j$. The Pareto-guaranteed optimal allocation problem is then the problem of choosing $D_i^j$ such that the maximum average realized outcome is achieved, while respecting that each recipient must be allocated exactly one treatment, that no treatment is allocated to more individuals than its capacity allows, and that no re-allocation leads to any recipient be made worse off compared to the baseline choices of $D_i^j$.

We shall make use of a slight modification to the network representation of Definition (ref) to prove our extension:

definition[Pareto-guaranteed network representation] Let $\mathcal{N} := (V, A, l, u, b, c)$ be defined as in Definition (ref). Remove arcs from $\mathcal{N}$ as follows: For each recipient $v_2^* \in V_2$, let $\bar{Y}_{v_2^*}^{v_1^*}$ denote the realized outcome under no re-allocation, i.e., nurse $v_1^*$ is allocated $v_2^*$ under no re-allocation. Now, for each $v_1 \in V_1$ and $v_2 \in V_2$, remove the arc $v_1v_2$ if $Y_{v_2}^{v_1} < \bar{Y}_{v_2^*}^{v_1^*}$ to construct the network $\mathcal{N}_P$.\footnote{Alternatively, one could set $u_{v_1v_2} = 0$ in these cases.} We shall call $\mathcal{N}_P$ constructed this way the Pareto-guaranteed network representation of an instance of the Pareto-guaranteed optimal allocation problem of Definition (ref).

With our modified network definition, we are in a position to prove that our results extend to settings where only Pareto improvements are allowed.

theoremLet $\mathcal{N_P}$ as defined in Definition (ref) be the Pareto-guaranteed network representation of the Pareto-guaranteed optimal allocation problem of Definition (ref), and let $n_1$ denote the number of treatments and $n_2$ the number of recipients. Given the existence of an integer feasible flow in $\mathcal{N_P}$, an integer minimum cost feasible flow $x$ in $\mathcal{N_P}$ is guaranteed to make no recipient worse off and to correspond to the (non-parametric) optimal allocation given the constraints. The solution can always be found in at most time $\mathcal{O}\left[(n_1^3n_2^2 + n_1^2n_2^3) \log^2 (n_1 + n_2)\right]$.
proofUnder the constraints of Definition (ref), no flow $x$ with value more than $0$ may pass through one of the arcs deleted when modifying $\mathcal{N}$ to $\mathcal{N}_P$. Thus, we may without loss of generality remove these arcs. The correspondence of an integer minimum cost feasible flow in $\mathcal{N}_P$ and the optimal solution to the Pareto-guaranteed optimal allocation problem then follows from Theorem (ref). Using Theorem (ref) (which puts no restriction on $\mathcal{N}$), we can always find an integer minimum cost feasible flow (if one exists) $x$ in $\mathcal{N}_P$ in at most time $\mathcal{O}\left[(n_1^3n_2^2 + n_1^2n_2^3) \log^2 (n_1 + n_2)\right]$.

Generally, our approach for solving the optimal allocation problem is readily generalizable to a broad range of alternative settings. In fact, any alternative formulation of the problem which can be represented by the type of network we introduce is immediately covered (with one such example being the Pareto-guaranteed optimal allocation problem). Other settings include the possibility of multiple (additive in effect) treatments available for some recipients (by appropriately changing capacities of the type $u_{v_2t}$) and treatments being differentially available to recipients (by appropriately removing arcs of the type $v_1v_2$).

Institutional Background of the 1960s Copenhagen Nurse Home Visiting Program

In this section, we give an overview of the 1960s Copenhagen NHV program that serves as the background of our empirical application of the results derived in Section (ref). The 1960s Copenhagen NHV program provides the scene for the study by baker2023universal on the role of extended NHV, and we refer the interested reader to that paper for additional details on the program, and in particular the Copenhagen trial on extended NHV.\footnote{The cohort profile we have created as part of our efforts of transcribing and linking the NHV records is described in more detail in bjerregaard2023cohort.}

Universal home visiting for families with infants in Denmark has a rich history, dating back to 1937 when the Danish National Board of Health (DNBH) launched a program to address high infant mortality rates. This initiative aimed to combat infant mortality rates of around six percent in the early 1930s. Utilizing staggered introductions across municipalities from 1937 to 1949, previous research has highlighted both short- and long-term health benefits of program participation wust2012early, hjort2017universal.

By the 1960s, with significant improvements in living conditions and a decline in infant mortality rates to around two percent, the DNBH revised the program to emphasize broader health monitoring and encourage relevant parental health investments dodelighed19311960. This shift in focus aligns with the evolving landscape of early childhood programs, such as the US Head Start program, which also underscored parental investments during toddler years BarrGibbs2017. With midwife-assisted home births being the norm and limited formal childcare options, interventions in the family home, especially during toddler years, gained prominence.

Amid these developments, the DNBH conducted experiments with extended home visiting, including the “Copenhagen trial” studied by baker2023universal, which encompassed children born between 1959 and 1967.\footnote{Exploiting quasi-random allocation to a three (vs. baseline one) year NHV program in Copenhagen in the 1960s, baker2023universal document positive long-run health (and to a lesser extend employment) effects, and notably significant heterogeneity with respect to characteristics such as birth weight (with low birth weight children much more positively affected than the average child).} In this trial, nurses offered additional follow-up during existing first-year visits to families residing in Copenhagen. While the “Copenhagen trial” aimed at evaluating the efficacy of a longer visiting schedule, it also meticulously tracked child development during the baseline period (that is, the first year of a child's life) for everyone living in Copenhagen and born between 1959-1967. Table (ref) shows an overview of topics covered at different visits during a child's first year visits.\footnote{The position as an infant health nurse required education beyond that of an ordinary nurse, within maternity care, specialized paediatric care, epidemic or tuberculosis care, or care for patients with mental illness. All infant health nurses had to pass a course at Aarhus University kuhn1939vejledning, dha1954vejledning, dha1961vejledning.} Topics include objective development measurements (e.g., child weight), self-reported parenting decisions (e.g., type of nutrition), and nurse-assessed family characteristics (e.g., socioeconomic status and mother well-being).\footnote{While more than the eight visits indicated in Table (ref) could take place (with an average of 13 visits per child taking place), structured information collection took place at these eight visits, with the nurse records containing pre-printed fields to be filled in at these specific visits.}

table[table omitted — 3,287 chars of source]

While the “Copenhagen trial” altered the landscape of child home visiting, it did so in a way unlikely to contaminate the study of differences in the impact of NHV by nurse, namely by quasi-random allocation to the prolonged program by day of the month of birth, assigning everyone born the first three days of any of the months of the nine year trial to the extended schedule. While the number of visits each child received was reduced as a consequence, in order to compensate nurses for the additional visits to the children enrolled into the extended schedule,\footnote{The average number of first year visits was reduced to around 13, compared to the pre-trial average of around 14 (stadsarkiv, various years).} this was done in a “homogeneous” way across children and should not interfere with our design.

Data

Our study combines two primary sources of data, those being handwritten nurse records from the 1960s and Danish administrative data, to allow us to combine detailed information on childhood development with long-run outcomes from administrative register data, in total following individuals from their birth to when they are around 50-60 years old (depending on their year of birth). Additionally, we use archive material detailing which nurse district each nurse of the Copenhagen infant nurse program worked in. Combining these sources, we are able to identify which nurse visited each child and follow that child throughout her first year of life and from her adulthood until current time.

Copenhagen Nurse Records

Nurse records with information on childhood development is available for all children born in Copenhagen between 1959-1967.\footnote{Years before or after are not available, and we speculate that these years were archived due to those being the years of the “Copenhagen trial” studied in baker2023universal.} Figure (ref) shows the scan of a nurse journal of a child, with parts blackened for confidentially reasons (in the source material we have available, these black patches are not present).

figure[figure omitted — 891 chars of source]

In joint, concurrent work, baker2023universal, bjerregaard2023cohort use ML to transcribe the contents of these records and link them to Danish administrative register data.\footnote{See dahl2023bdad for details on an unsupervised ML approach we use to identify types of journal pages and, as a consequence, identify which children were part of the treatment arm of the “Copenhagen trial”.} While the Danish unique personal identifier was introduced in 1968 -- i.e., after the birth of all the children of the records -- we nevertheless are able to obtain their personal identification number by using the name and date of birth of the child and her parents, linking them to the Danish Central Person Registry (CPR).

We transcribe the collection of nurse records by using timmsn, a Python library for image-to-text translation tsdj2023timmsn. The library is based on the PyTorch Image Models library by rw2019timm, an image recognition Python library. The full transcription code is available upon request and will be made available open-source at \url{https://github.com/TorbenSDJohansen/cihvr-transcription}. For the interested reader we refer to Appendix (ref) for additional details. We transcribe most of the contents of the nurse records with between 95%-99% accuracy, with slightly lower transcription accuracy of around 93% for nurse names (see Appendix Table (ref) for details).

\paragraph{Data Linkage and Coverage} We link the collection of nurse records to administrative data using the unique CPR number (similar to the US Social Security Number) of the children of the records. Figure (ref) shows the process from the raw, scanned nurse records to the link with outcomes from Danish registers. While the full collection of scanned nurse records consist of 95,323 documents, some of these we identify as duplicates and others as non-records (e.g., notice of movement), leading to a total of 92,902 records (of which 92,279 contain date of birth, which we need to obtain the children's CPR numbers). Some records we are unable to link to the Danish CPR, which may be due to poor scan quality or death prior to the establishment of the CPR in 1968, leaving us with 88,808 records we are able to link to the CPR. Note, however, that 808 of these children do not result in a match with our administrative data (which starts in 1977), primarily due to emigration or death, leaving us with a sample of 88,000 records (children) we can link to administrative outcome data. For the majority of our analyses, we shall make use of information on the nurse visiting the child and the district of the nurse; constraining the sample to only those records (children) with this information brings the sample size down to 75,318.\footnote{Here, we also restrain the sample to records on which the name of the nurse occur at least 100\xspace times across the entire collection of nurse records. This is done for two reasons: First, for confidentially reasons we are not allowed to report estimates of small group sizes, and once we combine nurse name information with year and district information, subgroups otherwise become too small. Second, nurse names that occur rarely are more likely to be artefacts from transcription, such as an otherwise valid name having been transcribed slightly incorrectly.}

figure[figure omitted — 3,151 chars of source]

\paragraph{Child-Nurse Matches} Our study relies on identifying the specific nurse responsible for visiting each child, and further on being able to identify which nurse district each nurse belonged to, in order for us to be able to compare children within the same nurse district with each other. To do this, we use our transcriptions of nurse name for each journal, and merge this with archive material on which district each nurse belonged to. We have been able to find information listing for each nurse the district they served in for the years 1957, 1963, 1965 (two separate accounts), and 1968, and it is this information we use to associate to each child the nurse district to which they belonged. This is a potential limitation, as nurses may change district over time, and while the data on nurse districts cover the entire time span we consider, we lack information for certain years (if, e.g., a certain nurse in 1964 served in a different district than in 1963 and 1965) and may have incomplete information on all nurses (if, e.g., a nurse was employed only in 1964, thus not showing up in any of the years we observe).

We match nurse records to nurse districts by identifying “most likely” matches between the nurse name transcriptions of the collection of nurse records with the archive material listing nurses and their districts. As the number of different nurses is limited (112\xspace), often encountered issues in linking based on names pertaining to non-unique names is limited. However, we note the following potential issues: First, while our transcription accuracy is high (see Appendix Table (ref)), it is not perfect, and this will result in a number of children with an incorrect nurse name transcription, leading to missed matches.\footnote{However, this is very unlikely to lead to a wrong match, as it would require a name to be transcribed incorrectly as another name. Further, we observe that incorrect transcriptions are primarily related to poor scan quality or source material degradation, both of which are unlikely to lead to any systematic bias. As such, we mainly view this weakness as one decreasing our precision, not as one introducing potential bias.} Second, nurses would not always write their name fully out. Specifically, we often observe that nurses would write their initial (of their first name) as well as their last name, rather than fully writing out their name. For this reason, we iteratively match nurses, looking first at exact matches where both first and last name of nurses are fully written, and then gradually relax this requirement over six steps.\footnote{In the second step, we -- for those we did not match in the previous step -- relax the requirements to allow use of initial in place of full first name. We then (3) match on first name initial and full last name, (4) match on full first name, (5) match on full last name, and (6) match on just first name initial. In cases where any of these steps leads to more than one match, we select the match closest in time, i.e., if a child is born in 1963 and we match the nurse on her record to our list of nurse districts for both 1963 and 1965, we select the match from the 1963 nurse district data.} Third, as we only observe certain “snapshots” with respect to the nurses' districts, we potentially assign a wrong district to some children, particularly for the years furthest away from when we have information on nurse districts. For this reason, we experiment with samples where we limit the maximum time span between a child's birth and one of our snapshots.

In total, we are able to match 43.1-98.4%\xspace of the children with a nurse district, depending on the strictness of our matching approach.\footnote{This refers to the share of nurse records with a non-rare nurse name transcription, of which there are 76,579 in total. This number is slightly higher than the final sample size shown in Figure (ref) due to missing district information for a few of the 76,579 records.} We focus on the larger sample that potentially includes lower quality matches, but also show robustness of our results to stricter versions of our matching approach. In total, 112\xspace different nurses appear across the collection of nurse records, and using our most lenient method of matching information on nurse district allows us to obtain this information for 111\xspace of the nurses.

Danish Administrative Data

We combine data from Danish registers to obtain information on long-run outcomes in the form of education and labor market outcomes. We obtain data from education and labor market registers for the years 1980-2018/2019, respectively. From education registers, we obtain information on the years of completed education of our focal individuals as well as their relatives, including information on highest completed educational level (e.g., mandatory, university, etc.). From labor market registers, we obtain information on employment and earnings of our focal individuals and their relatives. For each individual, we obtain the share of time in employment during each year for which we have data, as well as their earnings each year.\footnote{We inflation-adjust earnings to reflect 2015 values and winsorize one percent of each tail, the latter which we do separately for each age. } Using our employment and income data, we obtain average share of time in employment and income during ages 25-50 for each individual.\footnote{If information is missing for one or multiple ages of an individual during these 25 years, we take the average of the ages with non-missing employment/income information.}

Empirical Methods

Directly comparing outcomes of children for whom different nurses were allocated is likely to lead to false conclusions given significant differences in resources across different areas of Copenhagen (such as a family's available resources). To account for such differences, we make use of nurse districts and information on year of birth, exploiting that children born within a district-by-year group are likely to be “as good as randomly” assigned a nurse from that district.\footnote{In Section (ref), we empirically verify this by comparing pre-treatment characteristics of children allocated different nurses but born within the same district-by-year group.} Given that each nurse was supposed to be responsible for around 160 children, the birth of a new child in a district is likely to be assigned to the nurse from that district with the smallest current workload. Given that the Copenhagen nurse program followed all children during their first year of life, the allocation of a child to a nurse within a district is likely to be as good as random.

Cast in terms of the potential outcomes framework, we assume the potential outcomes of a child are independent of the nurse visiting the child, at least conditionally on the nurse district-by-year.\footnote{Conditioning on the district might be important if, e.g., the potential outcomes of a child changes if the child lived in another district where, e.g., the schools were better. Further, the parents of children in one district might vary systematically from the parents of children in another district, for instance reflecting socioeconomic differences.} Letting $Y_{id}^j$ denote the potential outcome of child $i$ in district-by-year $d$ under the “treatment” (visits) of nurse $j$, we assume:

align[align omitted — 66 chars of source]

where $D_i^j \in \{0, 1\}$ is one if child $i$ was visited by nurse $j$ and $G_{id}$ symbols the nurse district-by-year groups. The realized outcome of child $i$ is then $Y_{id} = \sum_j Y_{id}^j D_i^j$, and our objective is to draw inference on “treatment effects” of the type $\tau_{ij_1j_2} = Y_{id}^{j_1} - Y_{id}^{j_2}$, where now both $j_1$ and $j_2$, $j_1 \neq j_2$, represent nurses. This, then, represents the effect of assigning nurse $j_1$ rather than $j_2$ to child $i$, i.e., what is the change in the outcome of child $i$ by re-assigning nurses in such a way that child $i$ now receives visits form nurse $j_1$ rather than nurse $j_2$.

If we knew $Y_{id}^j$, we could directly calculate all $\tau_{idj_1j_2}$, but in the absence of this, where only the realized outcome is observed, we turn our attention first to the more coarse treatment effects of the type $\tau_{dj_1j_2} = \mathbb{E}\left[Y_d^{j_1}\right] - \mathbb{E}\left[Y_d^{j_2}\right]$, i.e., where we no longer subscript with $i$, instead averaging over children to obtain average treatment effects. Given (conditional) random allocation of nurses to children, this can be estimated by plugging in sample equivalents of $\mathbb{E}\left[Y_d^{j}\right]$, namely $\frac{1}{\sum_i D_i^j} \sum_i D_i^jY_{id}$.

Due to potential differences between district-by-year groups we consider the simplest case, focusing on some specific district-by-year group (and then leaving out subscripts $d$ for notational simplicity). Here, we may regress an outcome $Y_i$ of individual $i$ on dummy variables $D_i^j$, where $j$ enumerates the different nurses and where $D_i^j = 1$ if child $i$ was assigned nurse $j$:

align[align omitted — 49 chars of source]

where $\beta_j$ then denotes the expected outcome of children assigned nurse $j$; thus, our estimate of $\tau_{j_1j_2}$ is then:

align[align omitted — 174 chars of source]

The above approach allows us to compare the average outcomes of children by nurse within any specific district-by-year group. A further challenge arises, however, if we try to compare treatment effect estimates between district-by-year groups. The reason is as follows: Any treatment effect is a difference in average potential outcomes between two groups of children, and thus the size of any treatment effect depends not only on the visiting nurse, but also on the counterfactual alternative nurse you use as the base for your comparison. For that reason, large treatment effects might arise from either or both of (1) the nurse $j_1$ being in the right tail of the skill distribution and/or (2) the nurse $j_2$ being in the left tail of the skill distribution. For those reasons, directly comparing treatment effects between district-by-year groups is not possible

When we then turn to heterogeneity of nurse effects depending on child pre-treatment characteristics, we let $X_i \in \mathbb{R}^l$ denote the vector of characteristics for child $i$.\footnote{For example, $X_i$ might include low birth weight status or absence of father.} We once again consider the simplest case of one district-by-year group, and now include $X_i$ to account for nurse-by-child-characteristic differences:

align[align omitted — 61 chars of source]

where we may use causal ML to estimate heterogeneous treatment effects, such as generalized random forests wager2018estimation, athey2019generalized.\footnote{We use extensions by nie2021quasi to allow estimation in our setting of multi-arm treatments.} The heterogeneity we are able to capture this way reflects the cross-product of differences between nurses and children and, if present, allow us to consider welfare effects of nurse reallocation.

Our identification strategy relies on quasi-random allocation of children to nurses within nurse district-by-year groups. If this is not so, and nurses are selectively allocated to children as might be the case if children from families the least well off are more likely to be allocated a specific nurse, conditional independence between potential outcomes and treatment no longer holds. To mitigate such concerns, we perform several tests of differences in pre-treatment variables between the children of different nurses.

Results

Descriptive Statistics

We start by documenting the number of nurses and the number of children per year, district, and nurse. We can estimate the number of nurses in two ways, based on either statistics from archives or directly from the collection of nurse records we transcribe. From archive material, we know the number of nurses for the years from which we have data (1957, 1963, 1965-1969). This includes their names, which allow us to track them over time, allowing us to calculate the total number of unique nurses. From the collection of nurse records, we use our transcriptions of nurse names for each record to obtain statistics on the number of nurses for each year as well as the total number of unique nurses during the period 1959-1967. Due to imperfect transcriptions, however, we err on the side of caution and define a nurse only when the name of the nurse appears on a sufficient number nurse records, to avoid a small error in a transcription resulting in a name with one letter off now occurring as a unique nurse with just one record; we require 100\xspace occurrences of a name to include it.\footnote{Further, due to confidentiality reasons we are required to aggregate statistics up to include a certain minimum number of children.} Table (ref) shows the number of nurses for each year as well as the total number of nurses, as estimated from either archive material or the collection of nurse records. From the years for which we have data from both archive material and the nurse records, we generally see slightly fewer nurses from the archive material than from the nurse records (with 1965 being an exception). This is expected, as the first reports a snapshot and the second includes nurses present just for parts of the year.\footnote{For example, a snapshot of nurses for May may not include a nurse that stopped working in that year before May or one that started later than May.} In total, however, we identify more nurses from the archive material than from the nurse records; this is also not surprising as the archive material stretches over a longer period of time (1957-1969 vs. 1959-1967). In our main analyses, we continue with the sample of 112\xspace unique nurses, the name of which we know occurred on at least 100\xspace nurse records.\footnote{In our analyses exploiting information on district we continue with 111\xspace nurses, due to incomplete district information.}

table[table omitted — 1,332 chars of source]

Our identification strategy relies on the assumption of random allocation of children to nurses within each nurse district. From archive material, we know that the aim of the Copenhagen nurse visiting program was to have each nurse be responsible for around 160 children, and we would thus expect each nurse at any point in time to be responsible for somewhere around 160 children (stadsarkiv, various years). To assess the number of children by year, district, and nurse, Figure (ref) shows the number of children in our primary sample by year of birth (Panel (ref)), district (Panel (ref)), and nurse (Panel (ref)). Note that only nurse names occurring at least 100\xspace times are included.

figure[figure omitted — 969 chars of source]

As is evident from Figure (ref), the number of children by nurse varies substantially. While this at first appears problematic for our design, this is explained by differences in how many years the individual nurses were present: A nurse present in all of the years 1959-1967 will occur on more journals than one only present in 1959. To gauge the real number of children present at the same time per nurse, we therefore calculate the number of children by nurse-year. We do this by calculating the number of children born in each calendar year for each nurse, and only include a nurse-year if at least one child born in each month of the given year was assigned to the specific nurse.\footnote{This is done to eliminate the issue of a nurse only being present for part of a year.} We expect this to lead to a density with high mass centered at close to but below 160 (given that our sample includes close to but not all of the potential children of the nurse program and the program aimed at each nurse being allocated around 160 children). Indeed, Figure (ref), which shows the density of nurse-years as well as district-years, exhibits this pattern: Panel (ref) shows the density of the number of children by nurse-year, which exhibits a tight center of mass, with the average number of children by nurse-year being around 140. Turning to Panel (ref), which shows the density of the number of children by district-year, there is more dispersion, indicating that some districts were larger than others in terms of number of children in the district. Appendix Figure (ref) shows the number of children within each district-by-year combination, showcasing that districts that were larger at the start of the period (1959) generally continue to be so over all the available years (1959-1967). Further, as we would expect, the number of nurses is larger in those districts with more children.

figure[figure omitted — 745 chars of source]

Table (ref) shows means, standard deviations, and number of observations for pre-treatment and outcome variables for our primary sample, with the last column showing p-values from a Kruskal-Wallis H-test for independent samples kruskal1952use, comparing the population of nurses against each other with a null-hypothesis of no difference in the median outcome of children of different nurses.\footnote{Testing for differences in means (i.e., an ANOVA test) leads to the same picture, but we prefer the Kruskal-Wallis test due to its milder assumptions.} If children were randomly assigned nurses, we would expect the p-values of the pre-treatment variables (Panel A) to be large, while the p-values of the outcome variables (Panel B) could still be small if nurses differed with respect to their treatment effect. However, as is clear from the table there are statistically significant differences between nurses in all measures, with the largest p-value being for sex (0.069). This is not surprising, given that nurses served in different districts (and time periods) with different populations of children, due to, e.g., different levels of socioeconomic status of the parents of our focal individuals between nurse districts.

table[table omitted — 2,723 chars of source]

While Table (ref) document statistically significant differences between the groups of children of different nurses, it does little to report on the magnitudes of the differences. Table (ref) shows statistics for the variables of Table (ref), but where, for each variable, results are grouped by “nurse rank”:\footnote{This is done separately for each variable, meaning that the order of nurses vary to some degree between different rows of the table. However, the rank-order correlation between variables is high, and thus those nurses for whom the average outcome of one variable of their allocated children is high also tend to score highly on other variables.} Nurses for whom the average outcome of the given variable of the children they visit is below the first quantile forms the first group, the second group consists of those between the first and second quantile, the third group of those between the second and third quantile, and the final group of those above the third quantile. If nurses were randomly allocated across all districts, we would expect to see relatively little difference between the different groups for the variables in Panel A, and large differences in Panel B only if significant heterogeneity in nurse treatment effects is present.\footnote{Some variation is still expected due to finite sample sizes.} However, as is clear even for the pre-treatment variables, differences are also economically large between nurses, implying differences due to, e.g., variation in socioeconomic status between nurse districts (e.g., the difference between the lowest ranked quarter of nurses vs. the highest ranked quarter of nurses in mother years of education is over one year).

table[table omitted — 4,191 chars of source]

To get a more detailed look into the distribution of variables by nurse, Figure (ref) shows mean values (and associated 95% confidence intervals) of selects pre-treatment (i.e., determined before visit by nurse) variables by nurse, ranked such that nurses are sorted in ascending order. As expected following the results in Panel A of Table (ref), the figure shows significant differences in averages of pre-treatment variables by nurse, indicating that the population of children allocated different nurses vary substantially; as these variables are determined before the visits take place, the differences cannot be explained by differences in nurse treatment effects, but must instead occur due to non-random allocation.

figure[figure omitted — 1,039 chars of source]

Another way of assessing the degree of sorting that happens between nurses and children is to ask what the rank-order correlations between pre-treatment and outcome variables are. In a setting of no sorting, this would be (asymptotically) zero: While children that score “poorly” on pre-treatment characteristics (e.g., someone born low birth weight) are expected to score poorer on outcome variables (e.g., income), we would expect zero rank-order correlation between the average value of a pre-treatment and an outcome variable between nurses. We therefore calculate these averages for each nurse for select variables and then use a Spearman rank-order correlation (a non-parametric measure of monotonicity) to compare nurses fieller1957tests.\footnote{With 112 nurses, the asymptotic approximation of the p-value may yet be inaccurate, and the exact p-value is computationally impossible to derive. For these reasons, we use a permutation test that randomly draws 100,000 permutations, calculating the rank-order correlation coefficient of each and then compares it to the rank-order correlation coefficient of the non-permuted sample, letting the p-value be the share of permutations that results in a more extreme value of the rank-order correlation coefficient. } Appendix Figure (ref) shows the correlation coefficients and p-values for pairs of select pre-treatment and outcome variable means by nurses, verifying that substantial rank-order correlation is present, also when comparing pre-treatment and outcome variable pairs.

Taken at face value, the above results could be interpreted as detrimental to our identification strategy. However, the differences we observe between nurses turn out to be explained largely by differences between nurse districts (and to a smaller extent differences between cohorts). Figure (ref) shows the equivalent of Figure (ref) but now by nurse district rather than nurse. Notably, the variability of pre-treatment mean outcomes between districts is nearly as large as the variability of mean outcomes between nurses.\footnote{Note that the minimum and maximum mean outcomes of a district are bounded by the most extreme values of mean outcomes of any nurse.} The same pattern emerges when we consider the rank-order correlation in means of variables between districts (similarly to what we did above for nurses): Appendix Figure (ref) shows the correlation coefficients and p-values for pairs of select pre-treatment and outcome variable means by nurse districts, verifying that substantial rank-order correlation is present, also when comparing pre-treatment and outcome variable pairs.\footnote{With 16 nurse districts, the asymptotic approximation of the p-value will be inaccurate, and the exact p-value is computationally impossible to derive. For these reasons, we use a permutation test that randomly draws 100,000 permutations, calculating the rank-order correlation coefficient of each and then compares it to the rank-order correlation coefficient of the non-permuted sample, letting the p-value be the share of permutations that results in a more extreme value of the rank-order correlation coefficient.}

figure[figure omitted — 1,095 chars of source]

Differences in outcomes between districts persist throughout life: Figure (ref) shows income and share of time in employment for our focal individuals during ages 25-50 by nurse district. While there is some variability, those districts “doing well” in the early years of the focal individuals' lives tend to do so throughout their life-cycle.\footnote{We end at age 50 as that is about the highest age for which we observe the outcomes for all our focal individuals.}

figure[figure omitted — 916 chars of source]

We claim these differences between districts (and to a smaller extent cohorts) largely drive the apparent non-random allocation of children to nurses, and that properly handling this leads to close-to random allocation within district-by-year groups (see Appendix Figure (ref) for district-by-year group sample sizes). To show this, we once again turn to the rank-order correlation coefficient tests we performed above for nurses, but now do so within each district-by-year group of our sample. This leads to 137 (instead of 144, due to a few district-by-year groups containing exactly one nurse) separate district-by-year groups, and for each we obtain the p-value for the Spearman rank-order correlation coefficient test.\footnote{The number of nurses within each district-by-year group vary, and we thus apply a flexible inference approach that calculates the exact p-value whenever this is possible in fewer than 10,000 permutations and otherwise sample 10,000 permutations at random.} Figure (ref) shows the empirical cumulative density function (CDF) of the p-values obtained this way (colored lines). The black, dashed line indicates the asymptotic CDF given no rank-order correlation; given random allocation of children to nurses within these district-by-year groups, we expect the pairs with (at least) one pre-treatment variable to lie close to this line. Indeed, we also see this pattern, while also observing that the rank-order correlation between average income during ages 25-50 and years of education is statistically significant, which implies that while nurses did appear to have been allocated children at random, the children they visited systematically differ in how well they fare during their life.

figure[figure omitted — 1,330 chars of source]

To further investigate whether some selection between nurses could still take place within the district-by-year groups, we return to our earlier strategy to non-parametrically compare groups of children allocated different nurses by means of a Kruskal-Wallis H-test for independent samples kruskal1952use. Now, however, we do so separately for each district-by-year group and then plot the empirical CDF (as well as the theoretical CDF given no selection). Figure (ref) shows the results for select pre-treatment variables, with each colored line representing a pre-treatment variable and the black, dashed line indicating the theoretical CDF given no selection. Reassuringly, all empirical CDFs lie relatively close to the dashed line, indicating limited differences between groups of children allocated different nurses (within district-by-year groups); however, differences do not vanish entirely for all variables.

figure[figure omitted — 689 chars of source]

Estimation Results

Having established what appears to be close-to random allocation of children to nurses within nurse district-by-year groups, we next turn to estimating the “treatment effects” of nurses. Since children appear to have been more or less randomly allocated a nurse from the population of nurses in the district-by-year group they belong to, we can estimate the average potential outcome of children assigned a specific nurse by simply calculating the sample average for that nurse. First, however, we verify that such differences are indeed present.

Figure (ref) shows empirical CDFs for p-values from a Kruskal-Wallis H-test for independent samples between nurses in the same district-by-year group for select outcome variables (i.e., done similarly to Figure (ref) but now for outcomes). Panel (ref) shows results for childhood outcomes (obtained from the nurse records within the first year of a child's life) and Panel (ref) for adulthood outcomes (obtained from registers). The colored lines indicate outcomes and the dashed, black lines the theoretical CDFs in a setting of no differences between samples. As is evident from both panels, and particularly so for the childhood outcomes of Panel (ref), the empirical CDFs lie far above the black, dashed lines, indicating statistically significant differences between groups of children allocated different nurses, even when only comparing against nurses within the same district-by-year. Compared to the empirical CDFs for pre-treatment variables shown in Figure (ref), where the CDFs were close to the dashed line indicating no differences between groups, there now emerge differences far more statistically significant.

figure[figure omitted — 1,052 chars of source]

While the evidence above indicates strong, statistically significant differences between nurses in terms of their treatment effects, it does little to gauge the magnitudes of these differences. To illustrate the magnitudes of the differences, Figure (ref) shows box-plots of average outcomes by nurse-year for each nurse district. Evidently, there is significant heterogeneity both within and between districts in terms of averages of children allocated different nurses. Further, as supported by Appendix Figure (ref), districts that do well on one measure tends to also do well on others. However, while there are large differences between districts, differences within districts (i.e., between nurse-years within the same district) are in some cases larger still.

figure[figure omitted — 1,618 chars of source]

While Figure (ref) provides an overview of differences within and between nurse districts, it “aggregates” across cohorts, meaning that each box plot contains averages for different years, meaning that some of the variation comes from differences between cohorts rather than nurses. For this reason, Appendix Figures (ref) and (ref) show equivalents of Panels (ref) and (ref) of Figure (ref), respectively, but where each panel now refers to a different year of birth of the children. While the figures are more noisy, significant heterogeneity is still present both between and within nurse districts, documenting that the differences of Figure (ref) are not solely driven by cohort effects.

A potential concern regarding the differences between nurses documented above is that they may arise through noise: Since raw averages constitute the points of the box plots, the variance of these estimates are not accounted for. For this reason, we plot both the point estimate and its associated 95% confidence interval for each nurse, split by district and year, in Appendix Figure (ref) (for years of education) and Appendix Figure (ref) (for average earnings during ages 25-50). While some estimates are relatively imprecise (with wide confidence intervals), there is nonetheless clear differences between the outcomes of different nurses, even within the same district-by-year group, and these differences are economically significant (e.g., cases of more than one year of education on average between the children of two different nurses within the same district and year).

While the evidence above supports the hypothesis of differences between nurse skills leading to differences in outcomes between children allocated different nurses, its magnitude and interpretation is hampered by the complex setting of many nurses as well as district-by-year groups. We therefore turn to a simpler method of assessing what differences can be expected between two groups of otherwise similar children which happen to be allocated different nurses. Specifically, we calculate, for each district-by-year group, the average values of our outcomes by each nurse, and then take all distinct pairs of nurses and calculate the absolute value of the difference in the the average values of their children. The average of these absolute differences is then an estimate of the expected difference between two otherwise similar groups of children, but which happen to have been allocated different nurses.

Figure (ref) shows the empirical distributions of these absolute differences in the averages of children with different nurses but within the same district-by-year group, along with the average value of these absolute differences (black, solid line). As shown, these differences are economically large: For example, two otherwise similar groups of children that happened to be allocated different nurses differ, on average, by nearly half a year of education and around DKK 27,000 in annual earnings (note, however, that large parts of these differences arise due to considering absolute differences; nevertheless, as we show later, differences beyond those that might arise due to noise are present). Further, these differences are not driven by few, extreme differences, but rather by relatively large differences by a large share of nurse-pairs.

figure[figure omitted — 2,006 chars of source]

While the findings in Figure (ref) document large differences in outcomes of children allocated different nurses, the measure is by design bounded below at zero, and so the average differences will be larger than zero, even asymptotically and with no actual differences between nurses. Therefore, we want to ensure that these differences are not artefacts due to noise but represent real differences between nurses. However, the complex nature of the setup does not lend itself well to normal inference, and we therefore turn to a permutation-based strategy for inference.

Suppose nurses do not differentially affect children, but that potential outcomes differ by districts and cohorts. Then we could, separately for each district-by-year group, permute the children, leading us to -- in the counterfactual setting -- assign at random $k$ children within the same district-by-year group to some nurse with $k$ children in the real setting.\footnote{We need to keep the number of children for each nurse fixed to not bias our results to show as always significant: Since we measure un-weighted differences between pairs, removing noise in number of children by nurse would lead us to over-estimate the statistical significance of our findings, as this equalization in sample size would reduce statistical noise.} Doing so for each nurse within each district-by-year group, we obtain a permutation which, if nurses did not differentially affect children, would follow the same distribution as our actual sample. Repeating this would then lead to a counterfactual distribution of average absolute differences, and we can then get an estimate of the exact p-value by calculating the share of permutations that led to a larger average absolute difference than the one in the actual data.\footnote{Exact p-values are computationally infeasible, which is why we only obtain an estimate. The test is designed such that a more extreme value always corresponds to a greater value, which leads to the right-tail focus.} Specifically, we perform 1000 permutations and obtain for each a counterfactual estimate of the average absolute difference in the permuted sample.

Table (ref) shows the average absolute differences in outcomes between children of different nurses and, to take potential differences arising due to noise into account, the “excess” difference, i.e., the absolute average difference beyond those obtained on average in the permutation tests. Further, mean values of outcomes are reported (MDV), and the table also shows p-values obtained from the permutation tests, i.e., each p-value is the share of times a counterfactual estimate based on a permutation attained a more extreme value than the one from the non-permuted sample. As is evident, all the differences are highly statistically significant, as are they economically significant (e.g., differences of more than one month in length of education and income differences of around 1.4%). Appendix Figure (ref) shows, for each variable in Table (ref), histograms of the averages obtained from the permutation tests, clearly showcasing that the averages obtained in the actual sample are larger than those which could appear due to random noise.

table[table omitted — 1,846 chars of source]

Taken together, our results showcase large, economically as well as statistically significant, differences in how well children fare as a consequence of which nurse they got allocated for their visits during their first year of life. Further, recall that these estimates are not estimates of the impact of the nurse visiting program; rather, they are specifically the differences within the program as a consequence of the visiting nurse. As such, our findings suggest that an integral part of understanding the impact of such policies lies in understanding the role of the treatment provider, in addition to the program itself.

When comparing nurses and nurse “treatment” effects in this setting it is important, however, to apply some caution: Any treatment effect is a difference in averages between two nurses within the same nurse district, and so a large treatment effects may be any one or both of (1) a skilled nurse and (2) a poor comparison nurse. Further, it is only possible to obtain these estimates within district-by-year group without imposing stronger assumptions on the underlying potential outcomes, and thus comparing differences across district-by-year groups must be done with caution.

Child-Specific Treatment Effects & Allocation Mechanisms

While our results documenting important differences in how well children allocated different nurses fare are interesting by themselves, and may have policy implications in terms of added focus on improving the skill of the worst nurses, some differences are likely impossible to erase, due either to skills attained only through years of experience or some inherent differences between nurses which are non-trivial to mitigate, another avenue for potential gains is in optimally “matching” children and nurses. Our hypothesis posits that certain children are particularly likely to benefit from a highly skilled nurse compared to others; for example, perhaps children with poor initial health or born to parents in the lower end of the socioeconomic distribution are more likely to benefit from a high rather than low skill nurse compared to a child born in perfect health to parents from the higher end of the socioeconomic distribution. If this is the case, there is potentially room for improved impacts of the nurse visiting program by exploiting such information to allocate children to nurses in a better way.

For it to be possible to identify potential gains from improving the allocation mechanism of the nurse visiting program, we require estimates of treatment effects that take into account such heterogeneity. Further, for any such allocation mechanism to be feasible, it needs to rely on readily available information of the child and/or her family.

In terms of relevant information readily available for potential (re)allocation, the nurse records provide information available immediately from birth in terms of birth weight and length, sex, and number of weeks born prior to due date (if relevant), and from administrative data we obtain other information visible directly at birth in the form of parent characteristics such as age and education.\footnote{While our use of register data here obviously was not available at the time of the sample we study, we only use information from the registers that would have been readily available from the parents and which, today, is immediately available. Nonetheless, we choose to drop parents' years of education, as we observe this only later in the registers and parents may potentially have obtained additional education between the birth of their child and the point at which we observe their education.} Specifically, we use birth weight, birth length, born prior to due date (and number of weeks when applicable), an indicator for born one of the first three days of a month,\footnote{This is potentially an important factor to consider given the “Copenhagen trial”, which selected children born during the first three days of a month to an extended, three year NHV program baker2023universal.}, parity, an indicator for sex, mother and father age at birth, and an indicator for nurse assessment of socioeconomic status of the family (low, average, high).\footnote{Nurse assessment of the socioeconomic status of the home is not strictly speaking observed before the first visit, as the nurse records it during her visit at child age one month. However, it is unlikely to be affected by the nurse during such a short time span. Our results are robust to excluding this variable from the list of variables we use to estimate heterogeneous treatment effects.}

With this information, we estimate individual-level treatment effects by means of a series of multi-arm causal forest wager2018estimation, athey2019generalized, nie2021quasi:\footnote{These effects are “individual” in the sense of being specific with respect to the at-birth information we have available on the child and her family. Two children which are identical with respect to all our measures will not differ in terms of their estimated treatment effects.} For each district-by-year group, we treat each nurse as a separate treatment arm and estimate heterogeneous treatment effects (for each pair of arms, i.e., comparing pairs of nurses) with respect to the information on the child and her family available to us at her birth.\footnote{Here, we add a new sample restriction in the form of requiring at least 100 children per nurse (within our district-by-year groups) for inclusion, as we otherwise risk imprecise estimates related to very low estimated propensity scores for some nurses.}

Having estimated a multi-arm causal forest for each district-by-year group, we can use the out-of-bag predictions of the forests to estimate counterfactual outcome distributions under different treatment allocation mechanisms. Limiting the counterfactual treatment allocation rules to those feasible within the nurse program, we can then estimate potential gains in the nurse visiting program from alternative allocation mechanisms. Further, we can estimate the total room for gains from an alternative allocation rule by estimating the gains from the optimal allocation mechanism, something which we proved, in Section (ref), is always possible in strongly polynomial time.

Table (ref) shows estimated average values of our outcome variables under the current as well as under two alternative allocation mechanisms: One uses a greedy heuristic to identify gains from re-allocation (but which is not guaranteed to achieve the optimal allocation) and one finds the optimal allocation. The results indicate economically significant room for improvements: Our estimates for average earnings during ages 25-50 suggest that improved allocation of nurses to children could have resulted in additional yearly income of around (2023) USD 1,815.\footnote{We arrive at this number by adjusting for Danish inflation between 2015 and 2023 (17%) and using an exchange rate of 0.14 DKK/USD.} Turning to education, we estimate the total gains in terms of length of education to be around two months on average.

table[table omitted — 1,622 chars of source]

How much of the gains from improved allocation is attainable through a simpler re-allocation algorithm? To answer this question, Table (ref) reports results from a re-allocation mechanism that attempts to find re-allocations using a greedy heuristic (the column “Greedy”). Specifically, it keeps track of the number of children allocated each nurse and then goes through each child in a district-by-year group and allocates the nurse to that child that (1) is not yet at capacity and (2) results in the highest potential outcome for that child (thus obtaining an $\mathcal{O}(n_2)$ re-allocation mechanism). However, as is clear from our estimates, using this heuristic algorithm leaves significant room for improvements. This highlights the importance of an exact algorithm for complex tasks like optimal treatment allocation, emphasizing the value of deriving a strongly polynomial algorithm for this task.

While Table (ref) documents significant room for average improvements through an optimized nurse allocation mechanism, it potentially hides important ways in which these average improvements are obtained: If those children currently the worst off are negatively impacted (e.g., through complementarity of initial conditions and nurse skill leading to matching of the best nurses to the best off children), welfare may not be improved under the re-allocation (in terms of a social planner with preferences for equality). Conversely, if those children the worst off are (particularly) positively impacted (e.g., through substitutability of initial conditions and nurse skill leading to matching of the best nurses to the worst off children), welfare may be more positively impacted under re-allocations compared to what the average improvement suggests (in terms of a social planner with preferences for equality).

To study the impact of nurse re-allocation on the distribution (rather than just its expectation) of outcomes, Figure (ref) shows the distribution of our non-binary outcomes under the current as well as under the optimal allocation mechanisms.\footnote{Binary outcomes are left out as Table (ref) (i.e., their expectations) fully captures the effects on these.} Across the outcomes, we observe a shift to the right across the entire distribution when moving from the actual to the optimal allocation, suggesting neither strong complementarity nor substitutability.

figure[figure omitted — 1,362 chars of source]

Robustness Checks

The biggest concern to our empirical design is selection in the allocation mechanism: If children with relatively good potential outcomes are selectively allocated to one nurse, and children with relatively poor potential outcomes to another, then our estimates of the differences in treatment effects between nurses will be biased.

To mitigate such potential concerns, Table (ref) shows average differences in children pre-treatment variables by nurse (with associated p-values from permutation tests), i.e., similarly to Table (ref) but now for pre-treatment rather than outcome variables. A necessary condition for our assumption of no selective allocation within district-by-year groups is that differences in pre-treatment variables should not differ between groups of children allocated different nurses. Indeed, with the exception of “born prior to due date”, the estimated differences are at most marginally significant.\footnote{The “excess” difference in the share of children born prior to due date between nurses is around one percent. However, the number of weeks born prior to due date does not systematically differ between children. We interpret this as differences between nurses in terms of when they would note a child as being born prior to due date, with some nurses being stricter than others.} While these results do not rule out potential selection, they are nevertheless reassuring for our design. Further, Appendix Figure (ref) shows the full distribution of counterfactual differences from our permutation tests (similarly to Appendix Figure (ref), but for pre-treatment rather than outcome variables).

table[table omitted — 1,619 chars of source]

Could our relatively lenient strategy for matching nurse records and nurse district information introduce some potential bias? To answer this question, we re-estimate the average differences in outcomes of children between nurses, similarly to Table (ref), but now for a subsample where records are only included if we can match it to nurse district information using nurse first name initial and full last name. Table (ref) shows our findings, indicating that our results are robust to this additional sample restriction. Magnitudes remain relatively stable and our estimated differences are still highly statistically significant.

table[table omitted — 1,993 chars of source]

Another potential concern is misclassification of children to districts as a result of nurses changing districts over time and us only observing nurse district for a subset of our sample years (1963 and 1965). To mitigate such concerns, we re-estimate the average differences in outcomes of children between nurses, but now for a subsample where records are only included if the child is born in one of the years for which we have data on nurse district. Table (ref) shows our findings, indicating that our results are robust to this additional sample restriction. The magnitudes of our estimates for this restricted sample are close to the results of Table (ref), though less precisely estimated.

table[table omitted — 2,020 chars of source]

What role might spillover effects between siblings introduce, and could nurses be selectively allocated children from families of which they already were assigned previous children? Table (ref) shows our estimates of average differences in outcomes of children between nurses, but now for a sample of only firstborn children. Here, some of our results are slightly attenuated (when compared to Table (ref)), but the same overarching pattern remains, again statistically significantly estimated.

table[table omitted — 1,932 chars of source]

Our above placebo tests and results for alternative samples, combined with our earlier tests for rank-order correlation between pre-treatment and outcome variables by nurse within district-by-year groups (see Figure (ref)) and our non-parametric tests for differences in pre-treatment variables by nurse within district-by-year groups (see Figure (ref)), support our identification strategy. As far as pre-treatment variables are concerned, there does not appear to have been significant selection, as long as we only compare nurses within the same district-by-year group.

Discussion

How do our estimated treatment effects and effects of re-allocation compare to other early-life interventions? In particular, how much can be gained by optimizing an existing program -- as is the case of the nurse re-allocation -- vs. adding or changing a program with associated increases in required resources?

The most similar early-life interventions to our setting are the introduction of NHV programs and center care in Scandinavia wust2012early, hjort2017universal, bhalotra2017infant, buetikofer2019infant, as well as the study by baker2023universal which examines the impacts of extending the Danish NHV program from one to three years. What sets our setting apart from the studies above is that we measure the impact of an, in principal, “free” intervention, which instead of adding resources considers the impact of improved allocation of existing resources.

Compared to the study on the introduction of NHV in Denmark (1937) by hjort2017universal, which finds health effects but none for education and labor market outcomes, we document positive impacts on education and labor market outcomes. In Sweden, bhalotra2017infant find that the introduction of a health intervention (pioneered in 1931-1933) that provided information to mothers and infant monitoring through home visits and clinics led to substantial health benefits and increased likelihood for secondary schooling enrolment for females by around 3-4%-points (15%).\footnote{They find no effects on likelihood of secondary schooling enrolment for males. They interpret this as being due to relatively fixed supply of schooling during that time.} Our estimates suggest that optimal re-allocation could have improved the probability of completing at least secondary education by around 3%-points. In Norway, buetikofer2019infant study increased access to mother and child health care centers during a child's first year of life (introduced during the 1930s), finding that access to these centers improved length of schooling by around 1.8 months and earnings by around 2%, as well as improved health. Our estimates suggest that optimal re-allocation could have improved the length of education by around two months and income by around 4%.

Compared to the above studies on the introduction of early-life policies in Scandinavia, our results show that, in terms of education and labor market outcomes, optimally allocating providers (nurses) to recipients (children) may have effects of similar magnitude. In the US, hoehn2021long studies the impact of county-level health departments (instituted from 1908-1933) on long-run outcomes, showing that affected men's later-life earnings were improved by 2-5%, again similar in magnitude to our results. However, our study is set at a later point in time and focuses not on the introduction but rather the provision of early-life investments. The closest study to ours in terms of timing and target group is the study by baker2023universal on the impact of extended NHV (three vs. one year) for the same group of children comprising our study. Their findings suggest positive and persistent health effects of enrolment into the extended NHV program, with some effects on labor market participation of women (around 1.4%-points for share of time in employment during ages 30-50) but none for education or income. In comparison, our results suggest that optimal re-allocation could have improved share of time in employment during ages 25-50 by around 2%-points.

How come our estimates are of similar magnitudes to impacts of introduction of similar services in Scandinavia and the US? A likely answer lies in findings across these studies of significant effect heterogeneity, with disadvantaged children disproportionately positively impacted. Unlike the studies above, which focus on universal programs, our counterfactual policy focuses precisely on those children most likely to be positively impacted. We view this as yet another motivation for studying the role played by treatment providers when evaluating and designing policies.

Conclusions

We tackle the issue of solving optimal treatment allocation problems of the type where each recipient may be differentially affected by each treatment, and where there are constraints on treatment capacities. We prove that the problem is solvable in strongly polynomial time and present an algorithm based on flows in networks for problems of this type. We also extend our algorithm to cases where only Pareto-improving re-allocations of treatments to recipients are allowed by modifying the underlying network to ensure that only re-allocations that make no individual worse off are selected.

To showcase our method, we study NHV in the 1960s Copenhagen. Earlier work in Denmark, as well as other countries, has documented positive impacts of the introduction of NHV and center care wust2012early, hjort2017universal, bhalotra2017infant, buetikofer2019infant, hoehn2021long and extended NHV baker2023universal, but we go beyond these studies by estimating differences in treatment effects between individual nurses. We do so by transcribing and linking data from historical nurse records to Danish administrative data, identifying the nurse allocated each child and following the child throughout her life using Danish register data.\footnote{This transcription work has been done concurrently and in collaboration with the work of bjerregaard2023cohort, baker2023universal.}

We show that outcomes of children vary significantly by the visiting nurse, even when comparing children in the same nurse district and born the same year. We show that these differences are not artefacts due to selection (once operating within district-by-year groups) by comparing pre-treatment variables between children allocated different nurses. Further, using causal machine learning, we show that children are heterogeneously affected by nurses. This, in turn, allows us to obtain estimates of potential outcomes under different allocation mechanisms, and we document that a re-allocation of the nurses in the Copenhagen NHV program could have resulted in significant efficiency improvements. Our estimates suggest that optimal allocation of nurses to children could have improved average yearly earnings by USD 1,815 (4%) and length of education by around two months.

We contribute with new knowledge within optimal policy and the literature on the role of early-life conditions and investments for long-run outcomes. Within optimal policy, we introduce a strongly polynomial algorithm for optimal treatment allocation problems under constraints. Within the literature on early-life investments, we add novel evidence on the role of treatment providers for the effects of such policies. Further, we show that optimal allocation of such investments to recipients may play a crucial role in fully exploiting potential benefits of early-life investments.

While our approach for solving optimal treatment allocation problems is applicable to settings with various constraints, we assume that potential outcomes can be sufficiently precisely estimated. Optimal policy learning is a rapidly evolving field with a number of challenges related to identification. Directly incorporating policy learning with our approach for solving allocation problems could prove a fruitful avenue for new research.

For our empirical application, two challenges remain for proper integration of our optimal allocation design. First, our results for long-run outcomes are not observable before a significant time-gap, and thus these results are not directly applicable in practice for estimating potential outcomes. However, given the strong rank-order correlation between short- and long-run outcomes, using estimates for short-run outcomes are likely sufficient to obtain re-allocation rules to improve also long-run outcomes. Second, if all nurses are re-allocated based on our algorithm, it would no longer be possible to obtain causal estimates of potential outcomes if the underlying distribution changes (e.g., as a result of the population of children changing or the population of nurses changing) since the algorithm by design exploits selection. Implementing an approach in practice would thus require only partial selection, leaving some children to be randomly allocated to update estimates of potential outcomes as its underlying distribution changes. This leaves a complex “exploration vs. exploiting” problem sutton2018reinforcement, where the optimal share of children to randomly allocate to nurses (in order to update estimates of potential outcomes, i.e., explore) must be decided in a way to optimize average outcomes over time (i.e., exploit our knowledge to improve allocation and, in turn, outcomes).

\singlespacing \printbibliography

\onehalfspacing \FloatBarrier