EconBase
← Back to paper

Inferring hidden potentials in analytical regions: uncovering crime suspect communities in Medellín

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,037 characters · 12 sections · 35 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inferring hidden potentials in analytical regions: uncovering crime suspect communities in Medell\'in

\pubMonth\pubYear\pubVolume\pubIssue \Keywords{Bayesian econometrics, crime, hidden populations, neighbourhood, one-sided errors, spatial effects.}

abstractThis paper proposes a Bayesian approach to perform inference regarding the size of hidden populations at analytical region using reported statistics. To do so, we propose a specification taking into account one-sided error components and spatial effects within a panel data structure. Our simulation exercises suggest good finite sample performance. We analyze rates of crime suspects living per neighborhood in Medellín (Colombia) associated with four crime activities. Our proposal seems to identify hot spots or “crime communities”, potential neighborhoods where under-reporting is more severe, and also drivers of crime schools. Statistical evidence suggests a high level of interaction between homicides and drug dealing in one hand, and motorcycle and car thefts on the other hand.

Introduction

We introduce in this paper an inferential framework within a Bayesian paradigm to incorporate that reported rates of individuals with unknown features in analytical regions might be lower bounds. For instance, positive rate tests of individuals having particular diseases, drug consumers, cheating behavior or police report rates of criminal activity in analytical regions are under reports of the total population size. All these are examples of what we called “hidden populations”. Our proposal controls for spatial effects and unobserved heterogeneity on panel data settings.

We depart from the stochastic frontier analysis (composed error models) Aigner1977,Meeusen1977, and consider inefficiency, which is a one-sided error term from a statistical point of view, as the conditional percentage of the total population actually reported per analytical region. Accordingly, we incorporate structural and transitory lower bounds as one-sided error terms. In particular, we extend tsionas2014firm proposal by including spatial effects; this component induces heteroscedasticity, and as a consequence, omitting one-sided error terms potentially causes biased estimators Wang2002.

Crime and arrest rates at a regional level based on reported statistics are lower bounds Kirk2006. Consistently, we apply our proposal to model the rate of crime suspects living per neighborhood in Medellín (Colombia). In particular, we include spatial effects as our cross-section units are neighborhoods, and permanent and transient one-sided errors to handle time-varying lower bound issues. In this way, we can perform inference about potential under reports of crime suspects living in a region; this would be valuable for policymakers. Particularly, it would suggest regional hot spots to implement structural interventions and efficient police operations to reduce criminality. This is especially relevant in a city like Medellín, which is a natural experiment field due to its history of violence associated with drug trafficking and urban wars. Nevertheless, it is worth mentioning that Medellín is also an exceptional example of social transformation based on public and private interventions.

It seems that mainstream crime literature neither has taken into account potential bias due to omitting lower bound issues nor has implemented strategies to infer the population of criminals. A remarkable exception for the latter is Heijden2014, who based their strategy on count models, but requires data at the individual level, whereas ours requires data at an aggregated regional level. andresen2006spatial,kakamu2008spatial,kikuchi2010neighborhood,arnio2012demography,He2015 take into account spatial effects modeling crime data. However, these authors do not take the lower bound issue into account, which may have consequences on the statistical properties of estimators. Observe that unlike most of the crime literature, which focuses on crime rates, we model place of residence of crime suspects, that is, we consider that neighborhoods have socioeconomic conditions that might promote “schools of criminality”. So, following the spirit of stochastic frontier analysis, we can consider these conditions as production factors of potential criminals.

On the other hand, druska2003generalized,schmidt2009spatial,mastromarco2016modelling,tsionas2016spatial,glass2016spatial,gude2018heterogeneous propose stochastic frontier models including spatial effects, but focusing on “standard production functions” such as GDP at state level. Additionally, the tendency is to incorporate both effects in the same error: spatial dependence and inefficiency, whereas we separate these components. Also, our proposal differs from previous stochastic frontier proposals with spatial effects due to including a random conditional autoregressive spatial (CAR) effect. schmidt2009spatial includes CAR effects in the inefficiency component, whereas most of the stochastic frontier literature includes spatial autoregressive (SAR) components. The former proposal does not induce explicitly heteroscedasticity when inefficiency is omitted due to being present precisely in this omitted part. On the other hand, the SAR process is not Markovian, so it generates global spatial patterns, and it seems that criminal activity in Medellín is controlled by gangs with influence in specific areas Collazos2020. Therefore, we use the CAR specification, which is a Markovian process in space ramirez2017welfare, to control for local spatial effects.

We aim to contribute to crime and stochastic frontier literature, but mainly, considering one minus the (exponential) one-sided errors as the percentage of covered (uncaptured) criminals helps to build a link between these two well-developed areas of knowledge.

It seems from our simulation exercises that the sampling performance of our proposal is sound, allowing us to capture both the one-sided error terms and the spatial dependency. Additionally, it allows us to obtain good estimates for the location parameters in the presence of a five--way error component model. We also find that predictive inference regarding hidden populations has good predictive interval coverage.

Our empirical analysis suggests that homicide and drug dealing are strongly linked as both seem to generate local urban displacement, share some determinants and same hot spots (crime communities), which are mainly located in the most west analytical region in Medell\'in. On the other hand, there are common links between car and motorcycle thefts. For instance, both crime activities are associated with high local unemployment rates, and there is a crime community in the central-west part of the city. Although, there are some isolated crime communities specialized in each crime activity with their own determinants.

The remainder of this paper proceeds as follows: Section (ref) outlines our econometric model, the conditional posterior distributions, and the results of simulation exercises. Section (ref) presents our application. In particular, construction of the analytical regions, unconditional spatial analysis based on hypothesis tests using standardized rates, and posterior inferential results. Section (ref) concludes.

Econometric approach

The model

The point of departure is the observed under-reported ratio of the number of target individuals per inhabitants (${Y}_{it}$) at analytical region $i=1,2,\dots,N$ and time period $t=1,2\dots,T$, where target individuals belong to the “hidden population” ($P_{it}$), which is also standardized,

align[align omitted — 71 chars of source]

where $\text{R}_{it}$ is the report rate, that is, the percentage of individuals belonging to the target population that have been observed.

We can think about $P_{it}$ as depending on environmental variables that promote or discourage the number of individuals belonging to the target population as well as spatial effects reflecting spatial clusters, and also unobserved regional heterogeneity and idiosyncratic stochastic errors. So, we propose $P_{it}=f(\bm X_{it},\bm \beta)\times \exp\{\alpha_i+v_i+\epsilon_{it}\}=\prod_{k=1}^K X_{kit}^{\beta_k}\times \exp\{\alpha_i+v_i+\epsilon_{it}\}$ where $X_{kit}$ are $k$ potential drivers which may include spatial lags, that is, given a set of controls ($\bm{z}_{it}$), their spatial lags are $\sum_{j=1}^N w_{ij}\bm{z}_{ijt}$, where $w_{ij}$ is the $ij$-th element of the contiguity matrix $\mathbf{W}_N$, $\bm{x}_{it}=\left[\bm{z}_{it}' \: \left(\sum_{j=1}^N w_{ij}\bm{z}_{ijt}\right)' \right]'$. The location parameters are given by $\beta_k$. In addition, $\alpha_i$ is the unobserved stochastic heterogeneity, $v_{i}$ is the spatial random effect, and $\epsilon_{it}$ is the idiosyncratic stochastic error.

On the other hand, we specify the report rate as $\text{R}_{it}=\exp\{-\eta_{i}^{+}-u_{it}^{+}\}$ where $\eta_i^+$ and $u_{it}^+$ are one-sided positive stochastic errors to account for unobserved persistent and transient lower bound issues, that is, if $\eta_{i}^{+}=u_{it}^{+}=0$, then $R_{it}=1$, and $Y_{it}=P_{it}$, otherwise we observe just a lower bound of $P_{it}$.

Observe that this setting follows the statistical framework of stochastic frontier analysis with permanent and transient one-sided components, unobserved heterogeneity and spatial effects. Therefore, we extend tsionas2014firm proposal including spatial effects,

align[align omitted — 114 chars of source]

such that the reduce form equation (ref) is in $\log$-$\log$ form.

Following tsionas2014firm we assume that the i.i.d random components have the following distributions:

align[align omitted — 234 chars of source]

We assume that each $v_{i}$ has an improper (intrinsic) conditionally autoregressive structure besag1991:

align[align omitted — 168 chars of source]

where $\bm{v}_{i\sim j}$ is a vector of stochastic spatial errors for the neighbors $j$ of $i$ ($i\sim j$).

The joint distribution of the improper CAR is $\bm{v}\sim \mathcal{N}(\bar{\bm{v}},\sigma^2_v(\bm{D}_w-\bm{W}_N)^{-1})$, where $\bm{D}_w=\text{diag}(\sum_{i\sim j}w_{ij})$ banerjee2014hierarchical. The $ij$-th element of $\bm{W}_N$ is equal 1 if region $i$ and $j$ are neighbors, and 0 otherwise. By definition the elements of the main diagonal are set equal to zero.

Likelihood and priors

Set $\tau_{it}=\alpha_i+\epsilon_{it}$ such that taken assumptions in ((ref)) into account, $\bm{\tau}_i\sim \mathcal{N}(\bm{0},\bm{\Sigma})$, $\bm{\Sigma}=\sigma_{\epsilon}^2\bm{I}_T+\sigma_{\alpha}^2\bm{i}_T\bm{i}_T'$, where $\bm{i}_T$ is a $T$-dimensional vector of 1's and $\bm{I}_T$ is a $T$-dimensional identity matrix. Given our model specification (equations ((ref)), ((ref)) and ((ref))), the joint conditional distribution function, given spatial random effects, is the product over individuals of a $T$-variate closed skew normal distributions Dominguez2003. Working directly with this distribution is demanding given that it is not readily available in closed form. So, we follow Sanchez2020 who in similar settings use data augmenting Tanner1987. In particular, set $\bm{\theta}=(\bm{\bm{\beta}}',\sigma^2_\alpha,\sigma^2_\epsilon,\sigma^2_v,\sigma^2_u,\sigma^2_\eta)'$, and the augmented vector $\bm{\Theta}=(\bm{\theta}',u_{it}^+,\eta_{i}^+,v_i)$, then taking into account equations ((ref)), ((ref)) and ((ref)), the “augmented” likelihood is

align*[align* omitted — 1,004 chars of source]

where we stack information by individual such that $\bm{X}_i$ is a $T\times dim\left\{\bm\beta\right\}$ dimensional matrix with information of individual $i$.

We follow standard practice in Bayesian econometrics with conditional independent priors such that $\bm{\bm{\beta}}\sim \mathcal{N}(\bm{\beta}_0,\bm{B}_0)$, where $\bm{\beta}_0=\bm{0}$ and $\bm{B}_0=1000\bm{I}$ which implies vague prior information. For the scale parameters, $\frac{\bar{Q}_k}{\sigma^2_k}\sim \chi^2(\bar{N}_k)$, $k=\epsilon,\alpha,v$, where $\bar{N}_k=1$ and $\bar{Q}_k=10^{-4}$ tsionas2014firm. We follow makiela2017bayesian for the priors of $\sigma_u^2$ and $\sigma_\eta^2$, that is, we use $\sigma_u^{2}\sim \mathcal{I}\mathcal{G}(v_{0u}/2,2v_{0u}\log^2(r_u^*)/2)$ and $\sigma_{\eta}^{2}\sim \mathcal{I}\mathcal{G}(v_{0\eta}/2,2v_{0\eta}\log^2(r_{\eta}^*)/2)$ where the prior medians of the transient and persistent one-sided errors are equal to $r^*_u=0.85$ and $r^*_{\eta}=0.70$, and $v_{0u}=v_{0\eta}=10$.\footnote{We perform robustness analysis regarding these hyperparameters. Available upon request.} Even though $\bm{v}$ has an improper distribution, Theorem 2 in sun1999posterior guarantees that a proper posterior distribution exists if $\bm{D}_w-\bm{W}_N$ is nonnegative definite, the precision parameters have gamma prior distributions, and the intercepts have diffuse prior distributions (we fulfill all these requirements).

Conditional posterior distributions

The conditional posterior distribution for the location parameters is

align*[align* omitted — 101 chars of source]

where $\bar{\bm B}=\bigg(\sum_i \bm{X}_i'\bm{\Sigma}^{-1}\bm{X}_i+\bm{B}_0^{-1}\bigg)^{-1}$, $\bar{\bm{\beta}}=\bar{\bm{B}}\bigg(\sum_i \bm{X}_i'\bm{\Sigma}^{-1}\tilde{\bm y}_i+\bm{B}_0^{-1}\bm{\beta}_0\bigg)$, and $\tilde{\bm y}_i=\bm{y}_i-\bm{u}_i^+ -v_i\bm{i}_T-\eta_i^+\bm{i}_T$. The notation $\bm\Theta_{-\psi}$ indicates all elements in $\bm\Theta$ except $\psi$. The conditional posterior distribution for $\bm{u}_i^+$ is

align*[align* omitted — 199 chars of source]

where $\bm{\Omega}=\left(\bm{\Sigma}^{-1}+\bm{I}_T\frac{1}{\sigma_{u}^{2}}\right)^{-1}$ and $\bm{\mu}=\bm{\Omega}\bigg(\bm{\Sigma}^{-1}(\bm{y}_i-\bm{X}_i\bm{\beta}-v_i\bm{i}_T-\eta_i^+\bm{i}_T)\bigg)$. Given the multivariate condition $I(\bm{u}_i^+>0)$, which can be difficult to meet in high dimensional settings, we sample from $\pi(u_{it}^+|\bm{\Theta}_{-\bm{u}_{it}^+},\bm{y},\bm{X})$ using the result from a conditional multivariate normal distribution eaton1983multivariate. Let $\bm{u_i}^+=(u^+_{i1},\hdots,u^+_{iT})'$,

align*[align* omitted — 309 chars of source]

then, the conditional distribution of $u^+_{1t}$ given $\bm{u}^+_{2t}$ is

align*[align* omitted — 140 chars of source]

where $\bar{\mu}=\mu_{1}+\bm{\Omega}_{12}\bm{\Omega}_{22}^{-1}(\bm{u}_{2t}^+-\bm\mu_{2})$ and $\bar{\omega}=\omega_{11}-\bm{\Omega}_{12}\bm{\Omega}_{22}^{-1}\bm{\Omega}_{21}$.

The conditional posterior distribution for $\eta_i^+$ is:

align*[align* omitted — 54 chars of source]

where $\psi^2=\sigma_\eta^2 \bigg(1+\sigma_\eta^2\bm{i}_T'\Sigma^{-1}\bm{i}_T\bigg)^{-1}$ and $m_i=\psi^2\bm{i}_T'\Sigma^{-1}(\bm{y}_i-\bm{X}_i\bm{\bm{\beta}}-\bm{u}_{i}^+-v_i\bm{i}_T)$. The conditional posterior distribution for $v_i$ is

align*[align* omitted — 92 chars of source]

where $\bar{\sigma}^2_{vi}=\bigg(\bm{i}_{T}'\bm\Sigma^{-1}\bm{i}_{T}+\frac{\sum_{i\sim j}w_{ij}}{\sigma^2_v}\bigg)^{-1}$ and $\bar{v}_i=\bar{\sigma}^2_{vi}\bigg(\bm{i}_{T}'\bm\Sigma^{-1}(\bm y_{i}-\bm{X}_{i}\bm{\beta}-\bm{u}_{i}^+-\eta_i^+\bm{i}_T)+\sum_{i\sim j}\frac{w_{ij}v_j}{\sigma^2_v}\bigg)$. The conditional posterior distribution for $\sigma^2_v$ is

align*[align* omitted — 173 chars of source]

The conditional posterior distribution for $\sigma_{u}^2$ is

align*[align* omitted — 184 chars of source]

The conditional posterior distribution for $\sigma_{\eta}^2$ is

align*[align* omitted — 194 chars of source]

Up to this point we have standard conditional posterior distributions, so we can use the Gibbs sampling algorithm. However, the conditional posterior distributions for $\sigma_{\alpha}^2$ and $\sigma_{\epsilon}^2$ do not have standard form. We use the Metropolis-Hastings algorithm for these two parameters.

In particular, the conditional posterior distribution for $\sigma_{\alpha}^2$ is

align*[align* omitted — 580 chars of source]

The conditional posterior distribution for $\sigma_{\epsilon}^2$ is

align*[align* omitted — 593 chars of source]

We use as proposal distributions scaled Chi-squared distributions with one degree of freedom.

Simulations

We consider the following data generating process:

align[align omitted — 130 chars of source]

where the contiguity criterion is queen, and $z_{it}$ is drawn from a standard normal distribution. Following tsionas2014firm, $\sigma_\alpha=0.1,\sigma_\eta=0.5,\sigma_u=0.2,$ and $\sigma_\epsilon=0.1$, and we set $\sigma_v=0.4$. We perform 20,000 iterations, a burn-in equal to 10,000, and a thinning parameter equal 5.

Population parameters: point estimate results

Table (ref) displays sampling properties regarding point estimates of our Bayesian proposal using different combinations of $N$ an $T$, where one of them closely matches our application. We can see that our Bayesian proposal has good sampling properties as the highest density intervals (HDI) are relatively narrow and contain the population scale and location parameters. Comparing the first scenario ($N=49$, $T=5$) with the last scenario ($N=196$, $T=10$), we observe that the HDIs get narrower as the sample size increases. This is particularly relevant for the scale parameters. In general, the expected value of the correlation of the posterior draws of one-sided errors and spatial effects, and the Population values ($\hat{\mathbf{E}}(\rho)$) is higher than 0.5 except in one case. Also, the descriptive statistics (mean and median) of these unobserved stochastic components are similar.

We produce another two sets of simulation results (see Appendix (ref), Tables (ref) and (ref)). The first shows the consequences of varying $\lambda=\frac{\sigma_{\eta}+\sigma_u}{\sigma_{\epsilon}}$. This parameter has attracted a lot of attention in the stochastic frontier community. olson1980monte identifies two main issues when $\lambda \to \infty$: two-step estimators have very unstable empirical moments, and negative bias (constant term and $\hat{\sigma}^2_u$). On the other hand, the wrong skew problem, $\lambda \to 0$, implies opposite direction bias. waldman1982stationary,horrace2019stationary prove the existence of a stationary point in this case, then the probability of the wrong skew problem converges to zero when the sample size converges to infinity. However, simar2009inferences show that finite sample problems remain. Robustness checks for $\lambda$ are reported in Table (ref) in the Appendix. The sampling properties of our proposal seems to follow previous studies, that is, good performance as far as $\lambda$ is higher than one but bounded. Posterior estimates in our application suggests $\lambda$ is between 2 and 9, which seems a safe ground.

We also perform robustness checks regarding the presence of heavy tails in the stochastic error ($\epsilon_{it}$). In particular, we assume normality to perform statistical inference, but simulating equation (ref) using a Student's t-distribution with four degrees of freedom. It seems from results in Table (ref) in the Appendix that our inferential procedure is robust to heavy tails presence.

Hidden population: coverage results

One of the main purposes of our proposal is being able to make inference about “hidden populations”. Given $Y_{it}=P_{it}\exp\left\{-(\eta_i^+ + u_{it}^+)\right\}$, then $\left\{-(\eta_i^+ + u_{it}^+)\right\}$ is the percentage of reported cases of our target population, and as a consequence, 1-$\exp(-(\eta_i^+ + u_{it}^+))$ represents the percentage that is still covered. Therefore, we would expect that a sensible $1-\alpha$ predictive interval for the target population ($P_{it}$) is $\mathcal{P}_{\alpha}=\left\{Y_{it}\times\exp(\eta_i^+ + u_{it}^+):\pi(Y_{it}\times\exp(\eta_i^+ + u_{it}^+)|\bm\Theta,\bm Y, \bm X)\geq k(\alpha)\right\}$, $k(\alpha)$ is the largest constant such that $P(\mathcal{P}_{\alpha})\geq 1-\alpha$, that is, the $(1-\alpha)$ highest density (predictive) interval (HDI).

landscape\begin{table}[p] \caption{Sampling properties of Bayes estimators} \resizebox{1.48\textheight}{!}{ \begin{tabular}{cccccccccccccccccccccccccccccc} \hline & \multicolumn{3}{c}{$\eta^+$} & \multicolumn{3}{c}{$\sigma_\eta$}& \multicolumn{3}{c}{$u^+$}& \multicolumn{3}{c}{$\sigma_u$} & \multicolumn{2}{c}{$v$}& \multicolumn{3}{c}{$\sigma_v$} & \multicolumn{3}{c}{$\sigma_\alpha$} & \multicolumn{3}{c}{$\sigma_\epsilon$} & \multicolumn{3}{c}{$\beta_1$} & \multicolumn{3}{c}{$\beta_2$} \\ \cmidrule(r){2-4} \cmidrule(r){5-7} \cmidrule(r){8-10} \cmidrule(r){11-13} \cmidrule(r){14-15} \cmidrule(r){16-18} \cmidrule(r){19-21} \cmidrule(r){22-24} \cmidrule(r){25-27} \cmidrule(r){28-30} & Mean & Median & $\hat{\operatorname{\mathbb{E}}}(\rho)$ & Mean & \multicolumn{2}{c}{HDI} & Mean & Median & $\hat{\operatorname{\mathbb{E}}}(\rho)$ & Mean & \multicolumn{2}{c}{HDI} & Median & $\hat{\operatorname{\mathbb{E}}}(\rho)$ & Mean & \multicolumn{2}{c}{HDI} & Mean & \multicolumn{2}{c}{HDI} & Mean & \multicolumn{2}{c}{HDI} & Mean & \multicolumn{2}{c}{HDI}& Mean & \multicolumn{2}{c}{HDI} \\ \hline Population & -0.367 & -0.314 & & 0.500 & & & -0.152 & -0.132 & & 0.200 & & & 0.005 & & 0.400 & & & 0.100 & & & 0.100 & & & 0.500 & & & -0.500 & & \\ Estimate & -0.362 & -0.262 & 0.744 & 0.478 & 0.387 & 0.576 & -0.164 & -0.135 & 0.527 & 0.209 & 0.156 & 0.249 & -0.028 & 0.558 & 0.364 & 0.185 & 0.513 & 0.057 & 0.003 & 0.131 & 0.092 & 0.059 & 0.131 & 0.493 & 0.471 & 0.517 & -0.496 & -0.505 & -0.487 \\ N=49, T=5 & & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\ Population & -0.338 & -0.266 & & 0.500 & & & -0.172 & -0.145 & & 0.200 & & & -0.031 & & 0.400 & & & 0.100 & & & 0.100 & & & 0.500 & & & -0.500 & & \\ estimate & -0.355 & -0.316 & 0.564 & 0.455 & 0.373 & 0.556 & -0.158 & -0.133 & 0.590 & 0.199 & 0.164 & 0.232 & 0.008 & 0.726 & 0.566 & 0.366 & 0.781 & 0.065 & 0.003 & 0.156 & 0.096 & 0.073 & 0.120 & 0.509 & 0.494 & 0.524 & -0.498 & -0.503 & -0.492 \\ N=49, T=10 & & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\ Population & -0.391 & -0.328 & & 0.500 & & & -0.154 & -0.120 & & 0.200 & & & 0.005 & & 0.400 & & & 0.100 & & & 0.100 & & & 0.500 & & & -0.500 & & \\ estimate & -0.374 & -0.289 & 0.707 & 0.487 & 0.407 & 0.569 & -0.183 & -0.153 & 0.588 & 0.231 & 0.172 & 0.275 & -0.005 & 0.573 & 0.447 & 0.280 & 0.608 & 0.073 & 0.003 & 0.150 & 0.086 & 0.045 & 0.134 & 0.493 & 0.477 & 0.507 & -0.497 & -0.503 & -0.490 \\ N=100, T=5 & & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\ Population & -0.396 & -0.291 & & 0.500 & & & -0.163 & -0.137 & & 0.200 & & & 0.005 & & 0.400 & & & 0.100 & & & 0.100 & & & 0.500 & & & -0.500 & & \\ estimate & -0.388 & -0.305 & 0.687 & 0.490 & 0.414 & 0.564 & -0.173 & -0.148 & 0.611 & 0.218 & 0.184 & 0.250 & 0.000 & 0.352 & 0.354 & 0.175 & 0.519 & 0.147 & 0.058 & 0.214 & 0.092 & 0.070 & 0.120 & 0.493 & 0.482 & 0.503 & -0.501 & -0.504 & -0.497 \\ N=100, T=10 & & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\ Population & -0.399 & -0.312 & & 0.500 & & & -0.162 & -0.137 & & 0.200 & & & 0.008 & & 0.400 & & & 0.100 & & & 0.100 & & & 0.500 & & & -0.500 & & \\ estimate & -0.416 & -0.358 & 0.735 & 0.521 & 0.461 & 0.581 & -0.154 & -0.134 & 0.542 & 0.194 & 0.153 & 0.224 & -0.012 & 0.565 & 0.474 & 0.350 & 0.600 & 0.081 & 0.006 & 0.144 & 0.099 & 0.076 & 0.125 & 0.504 & 0.498 & 0.511 & -0.500 & -0.503 & -0.498 \\ N=196, T=10 & & & & & & & & & & & & & & & & & & & & & & & & & & & & & \\ \hline \end{tabular}} \begin{tablenotes} • \tiny Note: By Population value of $\eta^+$, $u^+$ and $v$ we mean the average and median values as they were generated from the Monte Carlo experiment. We report posterior means, medians and highest density intervals from posterior draws using 20,000 iterations, 10,000 burn-in and thinning equal to 5. $\hat{\mathbf{E}}(\rho)$ is the expected value of the correlation of the posterior draws of one-sided errors and spatial effects, and the Population one-sided errors and spatial effects coming from the Monte Carlo experiment, respectively. \end{tablenotes} \end{table}

We check the performance of this predictive interval calculating its coverage ($cov$). In particular, we propose a Beta-Binomial model for this coverage using as prior a non-informative Beta distribution, that is, $\pi(cov)\sim B(1,1)$, therefore its posterior distribution is $cov|\bm\Theta,\bm Y,\bm X\sim B(a,b)$ where $a=1+\sum_{i=1}^N\sum_{t=1}^T I_{it}$, $b=1+N\times T-\sum_{i=1}^N\sum_{t=1}^T I_{it}$, $I_{it}=\mathbbm{1}\left[P_{it}\in \mathcal{P}_{\alpha}\right]$. Observe that $\mathbb{E}(cov|\bm\Theta,\bm Y,\bm X)\approx \frac{\sum_{i=1}^N\sum_{t=1}^T I_{it}}{N\times T}$ due to using a non-informative prior. Algorithm (ref) shows details.

It seems from Table (ref) that our proposal has a good coverage as all means are close to $(1-\alpha)$ credibility levels with narrow 95% HDIs, where most of them embrace the credibility levels. To have an idea of how close is our estimate of $P_{it}$ to the real value, we calculate the mean absolute percentage error (MAPE) for each observation as

align*[align* omitted — 131 chars of source]

Table (ref) reports summary statistics corresponding the MAPE's for each of the sample sizes of Table (ref). Results suggest that our point estimate for $P_{it}$ tends to be very close to the real values. On average, our estimate diverges from $P_{it}$ in 23 percentage points. Additionally, half of the estimates diverge in no more than 9 percentage points from the real values.

algorithm[algorithm omitted — 1,071 chars of source]
table[table omitted — 1,469 chars of source]
table[table omitted — 1,141 chars of source]

Crime suspect communities: the Medellín case

Medellín is a natural experimental field to analyze crime. In particular, this city had a slightly increasing homicide rate from the mid-1960s to 1980, then there is a remarkable increase between 1981 to 1991 passing from 52 to 388 homicides per one hundred thousand inhabitants Garcia2012; this is mainly explained by the drug trafficking war of the local cartel against its competitors and the state. Then, there is a significant decrease of the homicide rate until 2008, reaching 46, that may be explained by the peace agreements with the guerrilla groups M-19 and EPL in 1990 and 1992, respectively, the dismantling of the local cartel, and death of its main figure in 1993, the intervention of the police and army (Orion operation) in Comuna 13, a neighborhood characterized by no state law enforcement, and the demobilization of the paramilitary group Cacique Nutibara in 2003, and other groups until 2006 Giraldo2015. However, there is a sudden upturn in 2009, the homicide rate was 94, this may be explained by disputes to get control of the drug trafficking business after the extradition of former leaders. Since then, Medellín has experienced a steady decrease in the homicide rate achieving 20 in 2015, which is a historical low record for the city in the last 50 years. We got confidential information from Colombian police reports about residential address at capture moment of suspects associated with four criminal activities: homicide, drug dealing, motorcycle and car thefts (see Appendix (ref), Table (ref)). We do not have the socioeconomic characteristics of these individuals. So, we calculate crime suspects rates per one hundred thousand inhabitants at the analytical region level in Medell\'in between 2011-2014.\footnote{However, we use rates per one million inhabitants for regressors in our models for coefficient scale purposes.} Then, as we consider environmental factors at a neighborhood level as potential drivers of “crime communities”, we calculated for these analytical regions averages of socioeconomic characteristics that have been considered previously in crime literature Nunez2003,andresen2006spatial,hipp2007block,kakamu2008spatial,kikuchi2010neighborhood,arnio2012demography,He2015 using information from annual Living Standards Surveys (See Appendix (ref), Table (ref)). However, these surveys are not representative at a neighborhood level (325 units) in Medellín. Therefore, we use the max-p-region algorithm duque2012max obtaining 175 analytical regions which are more representative. This algorithm merges adjacent neighborhoods to create new analytical regions such that the algorithm minimizes within attribute heterogeneity, but maximizes this heterogeneity between the new analytical regions (see duque2012max for details).

Table (ref) in Appendix (ref) shows descriptive statistics. The main take away from this table is the high level of heterogeneity regarding control variables between the analytical regions; this would suggest that Medellín is characterized by a very high level of social inequality. We also observe in this table a mean homicide rate equal to 49 per one hundred thousand inhabitants; there are many analytical regions without any homicide, whereas the central business district analytical region has a remarkable 2,062 rate, explained by relatively no many people living in this area. Observe that this is a common flaw when modeling crime rates. On the other hand, we model crime suspects residence per region where we also found that there are many analytical regions without any, but others with very high figures. Homicide suspects rate is the highest on average, followed by drug dealing, motorcycle thefts, and car thefts, respectively.

Unconditional analysis

We perform unconditional analysis to identify potential “crime communities” using standardized incidence ratios banerjee04, $SIR_{it}=\frac{S_{it}}{E_{it}}$, where $S_{it}$ is number of crime suspects living in analytical region $i$ at time $t$ based on capture police reports, and $E_{it}=n_{it}\frac{\sum_{i=1}^N S_{it}}{\sum_{i=1}^N n_{it}}$ is the expected number of crime suspects, $n_{it}$ is the number of inhabitants in analytical region $i$ at time $t$.

It is possible to find a $SIR$'s estimator by maximum likelihood; assuming $S_{it}|\eta_{it}$ distributes poisson, that is, $S_{it}|\eta_{it}\sim P(E_{it}{\eta}_{it})$. Then, the maximum likelihood (ML) estimator is $\hat{\eta}_{it} = SIR_{it}$. Nonetheless, assuming equi-dispersion may be non-realistic clayton87. To get a more flexible model we assume that $S_{it}|\eta_{it}\sim P(E_{it}\eta_{it})$ such that $\eta_{it}$ distributes gamma, $\eta_{it} \sim G(\nu, \alpha)$; this implies that $\eta_{it}| S_{it}=s_{it}\sim G(s_{it} + \nu, E_{it} + \alpha)$. So, at the end, we get smoother ratios, through a prior distribution on $\eta_i$, overcoming equi-dispersion.

We are interested on the probability of a $SIR$ being higher than an observed value, that is, $H_0:S_{it}=E_{it}$ versus $H_1:S_{it}>E_{it}$. Then, if $S_{it}\sim P(E_{it}\eta_{it})$ under the null hypothesis $\eta_{it}=1$, $P(\eta_{it}> 1|s_{it},E_{it})$ is the posterior probability against the null hypothesis of evenly distributed rates across space, $P(\eta_{it}> 1|s_{it},E_{it})=1-P(\eta_{it}<1|s_{it},E_{it})=1-\int_0^{s_{it}}\frac{\eta_{it}^{s_{it}+\nu-1}\exp^{-\eta_i(E_{it} + \alpha)}(E_{it} + \alpha)^{s_{it} + \nu}}{\Gamma(s_{it} + \nu)}d\eta_{it}$ where $\Gamma(.)$ is the gamma function, and $\nu=\alpha=0.01$ to have non informative priors.

Figure (ref) shows probability maps, $P(\eta_{it}> 1|s_{it},E_{it})$ choynowski59. This help to easily identify “crime communities”. In particular, we observe that the central business district (map center) is a potential hot spot for homicide, drug dealing and car theft suspects. On the other hand, it seems that there are some other specialized “crime communities”. Homicide suspects are located in the western (see Figure (ref)), drug dealing suspects at central-eastern (see Figure (ref)), car thefts at north-western (see Figure (ref)), and motorcycle thefts at north-eastern (see Figure (ref)).

figure[figure omitted — 1,077 chars of source]

Conditional posterior results

Table (ref) shows posterior estimates of our econometric proposal for four criminal activities: homicides, drug dealing, motorcycle and car thefts. Our dependent variable is $\log(1+Y_{it})\approx Y_{it}$ where $Y_{it}=\frac{S_{it}}{(n_{it}/100,000)}$. There are some interesting results regarding home location of crime suspects. It seems that homicide and drug dealing suspect communities are positively associated with high proportions of young males, and low population densities, but surrounded by neighborhoods with a high population density. This would suggest local urban displacement associated with these crime communities, although, this displacement seems not to be explicitly forced. Observe that these characteristics do not play any statistical significant role in motorcycle and car thefts suspect communities, which on the other hand, are positively associated with young unemployment rates. This suggests that local focused employment policies may reduce these criminal activities. Additionally, it seems that there is a kind of optimal location regarding these communities as they are located near middle income neighborhoods.

There are other specific statistical significant variables to each crime suspect community. For instance, homicide communities are positively associated with less immigrants, low male education and household expenditures, high household sizes and neighborhoods with a lower safety perception. Drug dealers communities are associated with less proportion of Caucasians, but higher levels of safety perception. The latter is also positively associated with motorcycle suspect communities, which in turn, is also positively associated with neighbors with high forced displacement and low safety perception. It seems that motorcycle thieves travel to close neighborhoods to commit their crimes (average travel time is 10 minutes from crime location to home location). Finally, car thieves communities is positively associated with a higher proportion of divorced males, more immigrants, less people per household locally and in surrounding neighborhoods.

Table (ref) reports posterior mean estimates of error components (one-sided and two-sided) associated with homicide, drug dealing, motorcycle and car thefts. We notice that $\lambda$ is approximately between 2 and 8, which implies a safe ground for inference in one-sided error models as shown from simulation exercises in Table (ref) and previous studies olson1980monte,simar2009inferences. In addition, the percentage of total variability due to spatial effects is between 20% (homicides) and 13% (car thefts), where $\frac{\sigma_v}{0.7\left(\sum_{i\sim j}w_{ij}\right)^{Ave}}$ is the marginal standard deviation due to spatial effects ramirez2017welfare, $\left(\sum_{i\sim j}w_{ij}\right)^{Ave}$ is the average number of neighbors. This highlights the relevance of this effect. Finally, the posterior mean estimate of permanent percentage of potential covered (uncaptured) crime suspects ($\hat{\mathbb{E}}(1-\exp(\eta_i^+))$) fluctuates between 18% (car thefts) and 26% (drug dealing), and the total (permanent plus transient, $\hat{\mathbb{E}}(1-\exp(\eta_i^+ + u_{it}^+))$) is between 32% (car thefts) and 57% (homicides and drug dealing). This means that on average the highest transient effect were associated with homicides (33%).

However, the former figures have a lot of heterogeneity through time and space. Figure (ref) shows analytical regions specific total percentages (permanent and transient) of potentially still covered suspects by crime and time. Regarding homicide (top-left panel), it seems that between 2011 and 2013 there was a hot spot of potential covered crime communities in the most western area (analytical region 121) with percentages over 90%. However, this situation drastically changed in 2014 for this area, it seems that these hot spots moved a little bit to east (analytical regions 121 and 177). In addition, there is a cluster of homicide crime communities from the central east (analytical region 80) to the north limit of the central business center (analytical region 163) for this last year.

Drug dealing have similar pattern to homicides regarding the hot spot in the most western area between 2011 and 2013 (top-right panel). However, analytical region 121 still seems to be a drug dealers community in 2014. Observe that similar patters regarding statistically relevant variables was also found in estimation results, and coincides with the violent history of Medell\'in due to the drug trafficking war of gangs for business control. Both crime activities (homicide and drug dealing) have uncaptured percentage rates as high as 90%.

Car thefts communities are located in the central-north area near the west riverside of the Medell\'in river (bottom-left panel). This river is a geographical barrier between the west and the east of the city, and plays an important role regarding crime communities. It seems that there is a hot spot composed by analytical regions 44 to 47. Another hot spot is composed by analytical regions 95, 71 and 73 located on the central-west. It seems that this is also a community of motorcycle thieves. Observe that the percentage of uncaptured car thieves is as high as 50%, whereas this figure is as high as 70% for motorcycle thieves.

table[table omitted — 2,720 chars of source]
table[table omitted — 1,099 chars of source]
figure[figure omitted — 1,237 chars of source]

Concluding remarks

We propose a Bayesian approach to perform inference regarding “hidden populations” at analytical region level such as criminal activity. We extend a generalized random effects model including spatial effects where “hidden populations” are taken into account using one-sided errors. Simulation exercises suggest that our proposal has good sampling properties regarding point estimates and “hidden population” predictions.

Our application based on home place of crime suspects suggest that there is association between homicide and drug dealing which has caused potential urban displacement to neighbourhoods near these crime communities. This is also supported by historical facts and higher levels of uncaptured crime suspects, which are as high as 90% in the hot spots of both activities. On the other hand, motorcycle and car thefts have lower uncaptured rates (70% and 50%, respectively), and both activities are associated with high local unemployment rates, which would suggest that focused employment policies would mitigate these activities.

Due to the elapsed time between crime moment and reported time using CCTV monitoring, we suggest that a potentially good strategy to capture crime suspects is to lock down the potential destination neighborhood of criminals. Our modelling strategy would help to predict this potential destination neighborhoods as we identified potential “crime communities”. On the other hand, focused policy interventions targeting these communities with specific education and employment objectives would help to structurally reduce crime activity.

Future research should take into account sensitivity of our proposal to spatial contiguity criteria, where contiguity matrices can be selected based on Bayes factors or performing Bayesian model average to take into account this uncertainty source.