Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,202 characters · 6 sections · 15 citation commands
Combinatorial Models of Cross-Country Dual Meets: What is a Big Victory?
\graphicspath{{CCfinal/}} Cross-country racing is a varsity sport in many high schools with more than 15 thousand schools offering cross-country NFHS. It is the fourth most popular sport for boys and fifth most for girls, as defined by the number of school teams. Almost half a million high school students compete on cross-country teams every year, and hundreds of thousands of middle school students also compete. Scoring is an application of nonparametric statistics: all that matters is the rank of the runner as he/she crosses the finish line. The team score is the sum of the ranks of the runners whose score counts. This scoring is termed “Rank Sum Scoring.” In this article, we calculate the distribution of scores under a variety of assumptions on the relative speed of the runners. First, we address the distribution of scores when all runners have the same distribution of running times.
To the best of our knowledge, the distribution of outcomes has never been analyzed from probabilistic point of view. There is much excellent research on cross-country scoring from a game theoretical perspective Huntington, Hammond, Sze,MixonKing, BESW2014, BS, BERS. The basic discrepancy/inequity is that in a multi-team meet, Team A can beat Team B under dual-meet scoring and lose to it under the multi-team scoring rule. These papers analyze the multi-team scoring and propose alternative scoring rules that eliminate these paradoxes. In Hammond and BESW2014, distributions of scores are calculated for meets involving three or more teams and the probability of a “social choice principle violation” is calculated. In BERS, actual multi-team empirical data is analyzed to see how often these scoring paradoxes occur in practice. They find that combinatorial analyses are helpful in predicting the probability of a paradox. Their results lend credence to our combinatorial analysis.
The standard scoring method is well established and unlikely to change. dual-meet scoring is based purely upon the order that the runners finish the race. Our study examines the standard rank-based scoring in a cross-country dual-meet. We derive the actual scoring distributions for the standard scoring rules for dual-meets using combinatorics and rank statistics. Indeed, we believe that the nonparametric (rank-based) nature of cross-country scoring makes it an ideal example for teaching stochastic and combinatorial analysis.
Definition: An $(M,N)$ dual meet consists of a team of $M$ runners. Only the $N$ fastest runners on each team are counted. The meet is scored as follows: the $k$th runner to finish receives a a score of $k$. The placement of runners from $N+1$ to $M$ are not included in the score but do increase the score of the other team. The $N$ lowest scores are added up for each team. The lowest score wins. If the runners $N+1$ to $M$ do not raise the score of the other team, we refer to the event as a cross-country meet with {\it no displacement}.
Cross-country meets occur with and without displacement scoring, but displacement scoring is definitely more common. Thus we make displacement scoring the default and explicitly mention {\it no displacement} in the cases where we study this alternative. For an $(M,N)$ event, the sum of the two scores is between $N*(2N+1)$ and $N*(M+N+1)$. The upper bound score, $N*(M+N+1)$, occurs when all $M$ runners of one team are faster than any of the other team. For the no displacement case, the sum of the two scores is $N*(2N+1)$.
The standard American dual meet is a $(7,5)$ with seven runners and lowest five scores are counted. An international dual meet has six runners and four are counted. We will consider two teams, Team A and Team B. In certain cases, we will assume that Team A has the fastest runner or even the fastest two runners. We will denote the score of Team A by $s_A$ and the score of Team B by $s_B$. The margin of victory for Team A will be denoted by $m_{A,B}= s_B -s_A$.
The following results are well known to the cross-country community.
1) The best possible score is $s_A= N*(N+1)/2$. The worst possible score is $s_A= MN + N*(N+1)/2$
2) The biggest possible victory margin is $MN$. For $M=N=4$, this is 16, and for $M=N=5$, the maximum victory margin is $25$. For $(M,N)=(7,5)$, the maximum victory margin is $35$. For $(M,N)=(6,4)$, the maximum victory margin is $24$.
3) If $N=5$ and $M \le 7$ and Team A has the fastest three runners, Team A has won. Proof: $(1+ 2+ 3 +11+12)< (4+5+6+7+8)$.
For no displacement scoring, we have the following results
A) The sum of the two teams' scores satisfies $s_A +s_B= N *(2N+1)$. For $N=4$ and $N=5$, $s_A +s_B= 36$ and $55$ respectively.
B) Ties can occur only if $N$ is even. In particular, ties can occur for $(6,4)$, but not for $(7,5)$.
C) The score difference is always odd for $(7,5)$. The score difference for $(6,4)$ is always even.
D) If $M=6$, $N=4$ and Team A has the fastest two runners, Team A cannot lose. Proof: $1+2+7+8=3+4+5+6$
As pointed out by the anonymous referee, the no displacement case is a reparameterization of Wilcoxon rank-sum distribution (Mann-Whitney distribution) Wilcoxon, MannWhitney, Lehman, Conover. We calculate the distribution of $s_B-s_A = 2 s_B - N *(2N+1)$. The case with displacement seems to be a related but new distribution. In the next sections, we calculate the distribution of victory margins, both one-sided and two sided. We define a large victory to be a score difference (positive or negative) in the .9 quantile range or greater. This means ten percent of the dual-meets will result in a “large” victory by our definition.
\graphicspath{{./CCfinal}}
We begin by considering the case where all runners are equally likely to come in $k$th place in the competition for $k$ in $1, \ldots ,2M$. In other words, the runners' times are independently, identically distributed (IID). We label the places from $1$ to $2M$. We then draw $M$ of these without replacement to determine the the placement of the runners on Team A. There are ${2M \choose M}$ possible outcomes.
To illustrate the distribution and scoring, we consider the three runner case with two scores counting i.e. $(3,2)$. The list of possible outcomes for Team A is $[(1, 2, 3), (1, 2, 4), (1, 2, 5)$, $(1, 2, 6), (1, 3, 4)$, $(1, 3, 5), (1, 3, 6), (1, 4, 5)$, $(1, 4, 6)$, $(1, 5, 6)$, $(2, 3, 4)$, $(2, 3, 5)$, $(2, 3, 6), (2, 4, 5)$, $(2, 4, 6), (2, 5, 6)$, $(3, 4, 5)$, $(3, 4, 6)$, $(3, 5, 6), (4, 5, 6)]$. By assumption, all cases are equally likely. Since only the top two runners count, the corresponding list of scores is distribution of scores for top 2 of each team. When there is no displacement, the distribution of scores is $\{(1, 2):4, (1, 3):3, (1, 4):3, (2, 3):3, (2, 4):3, (3, 4):4 \}$. Thus probability of a tie is $0.3$, the probability of a two point victory is $0.3$ and the probability of a four point victory is $0.4$. The team with the fastest runner cannot lose and ends up winning $70$ percent of the time. Now consider the case with displacement. We present the scores as a tuple, $(s_A,s_B)$ for each scenario, $[(3,9),(3, 8), (3,7), (3,7), (4,7), (4,6), (4,6), (5,5), (5,5)$, $(6,5),(5,6),(5,5),(5,5)$, $(6,4),(6,4),(7,4),(7,3),(7,3),(8,3),(9,3)]$. The distribution of score differences, $s_B-s_A$ is symmetric and satisfies $\{0:4,1:1,2:2,3:1, 4:2, 5:1,6:1\}$. Not surprisingly, displacement increases the diversity of outcomes and can have larger victories.
Let us consider the alternative case of two runners on each team and both runners count, $(2,2)$. Now, all six cases are equally likely: $[(1, 2), (1, 3), (1, 4), (2, 3), (2, 4), (3, 4)]$. Here the probability of a tie is $1/3$, the probability of winning by two points is $1/3$ and the probability of a four point victory is $1/3$. Finally, consider the case of $(M,2)$ for large $M$. As $M$ increases, the distribution of scores converges to the distribution corresponding to drawing with replacement: the probability that one team has the two fastest runners is $0.5$. If there is no displacent, the probability of a tie is $0.25$ and the probability of a two point margin is $0.25$. With displacement, let Team A have the fastest runner. Team A has a score distribution:$\{ 3:.5,4:.25,5:.125,6:.0625\ldots \}$ for a mean score of 4. Team B has a score distribution:$\{5:.25,6:.125,7:.1875 \ldots \}$ for an expected score of $8$. (The first Team B runner has an expected score of three while the second Team B runner has an expected score of five.) This illustrates a common theme: the larger $M$ for a given $N$, the larger the tail event scores are.
These three cases illustrate how the distribution of scores vary as the assumption of number of runners varies. We now focus on the two real-world situations: international dual-meets, $(6,4)$ and American dual-meets, $(7,5)$. Rather than trying to derive the score distribution analytically, we write a very short program which evaluates the score for each of the ${2M \choose M}$ possible outcomes. In Python, the outcomes are given by $map(tuple,itertools.combinations(range(2*M), M))$,
We tabulate the score difference, $s_B-s_A$, conditional on Team A having the fastest runner. Remember the team with the lowest score wins, so $s_B-s_A$ is the victory margin for Team A. A negative score means that Team A lost. We begin by computing the distribution with no score displacement in Table (ref). There are only 13 distinct values of score difference conditioned on Team A having the fast runner and the maximum victory margin being 16. Given the fastest runner, Team A wins with probability $0.6948$ and ties with probability $0.0974$. Conditional on Team A winning and having the fastest runner, the mean victory margin is $8.11215$. Team A loses with probability $0.20779$ and the average loss is $4.45833$.
To get the unconditional distribution, $|s_A-s_B|$ (either Team A or Team B may have the fastest runner), we just symmetrize Table (ref). The symmetrized distribution is a reparameterization of the Wilcoxon rank-sum distribution. Unconditionally, the mean score difference is $6.56277 \pm 4.52355$ and median victory is $6$. The $75$th percentile victory is $10$. A reasonable definition of a large victory is the $90\%$ quantile of the score difference. We see that this quantile value is $14$.
Table (ref) has the more realistic case of displacement. There are now 39 possible values of the score difference. The probability of an odd score difference is nonzero but clearly smaller than the probability of an even score difference.
Table (ref) has four rows of score differences to accomodate all the possible outcomes. In Table (ref), we assume that Team A has the fastest runner. Given the fastest runner, Team A wins with probability $0.6991$ and ties with probability $0.0584$. Conditional on Team A winning and having the fastest runner, the mean victory margin is $9.08$. Team A loses with probability $0.30$ and the average loss is $4.973$. Thus, displacement makes the importance of the fastest runner in determining which team wins significantly less, because the slower runners have more influence.
To get the unconditional distribution (either Team A or Team B may have the fastest runner) for $(6,4)$ with displacement, we just symmetrize Table (ref) to get the distribution of $|s_A-s_B|$. Unconditionally, the mean score difference is $7.554 \pm 5.334$ and median score difference is $6.5$. The $75$th percentile of the score difference is $11$. A reasonable definition of a large victory is the $90\%$ quantile of the absolute score difference. We see that this quantile value is $15$.
As a variant of the basic distribution, let us consider the scoring distribution of $(6,4)$ conditional on Team A having the fastest two runners. This result is given in Table (ref) for the case of no displacement. In this case, the probability of a tie is $0.0714$. The mean score differential is $9.162\pm 4.711$ and the median difference is $10$.
Since this is an American journal, we now evaluate the distribution victory margins under the IID outcome assumption. The result is in Table (ref) for the case of no displacement. Conditional on Team A having the fastest runner, the probability of victory is $72.14\%$ with an average victory of $10.37$. The probability of a loss is $27.86\%$ with a mean loss of $5.837$. There are no ties.
For the symmetric distribution of $|s_A-sB|$, (not conditioning on Team A having the fastest runner), the mean score difference is $9.108 \pm 6.283$ with a median value of $9$ and $.9$ quantile of $19$. As previously mentioned, the no displacement case is equivalent to the Wilcoxon rank-sum distribution. The distributions conditional on Team A having the fastest runner appear not to be in the literature.
The distribution for $(7,5)$ with displacement ranges between $-23$ and $35$ conditional on Team $A$ having the fastest runner. Table (ref) displays this broad distribution. Odd score differences, $s_B-s_A$, are more likely than even differences. Conditional on Team A having the fastest runner, the probability of victory is $70.22\%$ with an average victory of $11.66$. The probability of a loss is $28.15\%$ with a mean loss of $6.882$.
Overall (not conditional on Team A having the fastest runner), the mean score difference is $10.12 \pm 7.100$ with a median value of $9$. The $75$th percentile of the score difference is $15$ and the $90$th percentile victory is $21$ for $(7,5)$ .
Similarly, we evaluate the distribution of scoring differences from the IID outcome model when Team A has the fastest two runners. Table (ref) presents these results for $(7,5)$ with displacement. Team A wins with a $90.4\%$ probability with an average victory margin of $14.06$ Team A loses with a $8.71\%$ probability with average loss of $3.90$. Ties occur with probability $0.9\%$.
If one team has the fastest runner and the other team has the second fastest runner, the distribution of scores for $(M,N)$ then coresponds to a shifted version of $(M,N)$. In particular for $(7,5)$, let Team A be the team with the third fastest runner. The distribution of victory margins is given by Table (ref) shifted by one. If Team A has the first and third fastest runners, add one to the score differences in Table (ref). If Team A has the second and third fastest runners, subtract one to the score differences in Table (ref).
Our first model postulates that all runners are equally likely to win or come in last. This models the situation where there is a large amount of variation in each runner's time relative to the skill difference between runners. If, on the other hand, the standing of the runners did not vary and the results were predetermined before the start of the race, what would matter is the size of the population from which the runners are chosen.
Assume that School A has a population of $M_A$ potential runners and School B has a population of $M_B$ potential runners. The speeds of the potential runners are randomly distributed over the population of $M_A + M_B$ potential runners. We assume that each coach can perfectly identify the best $M$ runners.
In this random population model, the probability that Team A has the fastest runner (or any other runner) is $M_A/(M_A +M_B)$. One proxy for $M_A$ is the number of students at a school times some constant that corresponds to the participation rate. The participation rate is larger for smaller schools, but we can assume that good coaches always recruit the fastest runners and the recruiting population is proportional to school size. Each coach then chooses the best $M$ runners to compete.
This random population model is parameterized by the tuple $(M_A,M_B,M,N)$. The random outcome model is a special case with $M_A = M_B = M$. Of course, in the random population model, the rankings are predetermined before the race starts by the abilities of each runner. For each value of $(M_A,M_B,M,N)$, we could calculate the distribution of results. Instead we consider the limit of $M_A >> M$ and $M_B >>M$ such that the ratio $M_A/(M_A +M_B)$ tends to a constant fraction, $r$. In this limit, the score distribution converges to the distribution which corresponds to drawing runners with replacement.
One very convenient property of cross-country scoring is that the final score depends only on the results of the runners who come in in the first $2N -1$ places. Thus we need to compute probabilities of the $2^{2N-1}$ racing scenarios. As $M_A$ and $M_B$ increase, the probability of any of the scenarios converges to $r^k (1-r)^{2N-k-1}$ where $k$ is the number of runners from Team A.
In the large population limit of the random population model with no displacement, the results depend only on $N$ and $r$. We tabulate the score difference distribution as a function of $r =M_A/(M_A +M_B)$ for $N=4$ and $N=5$. Once again, we consider both the case of displacement and the simpler case of no displacement. Tables (ref)-(ref) present the summary statistics for the large population limit. In Tables (ref)- (ref), {\it Prob of Win} is the probability that Team A wins, {\it Mn score Diff} is the mean score difference i.e. the expectation of $s_A -s_B$, {\it Std Dev Diff} is the standard deviation of $s_A -s_B$, {\it Mn Win Diff} is the expectation of $s_A -s_B$ conditional on Team A winning and {\it Mn Loss Diff} is the expectation of $s_A -s_B$ conditional on Team A losing. The final row is our large victory margin, the $quantile(.9)$. For this, we are computing the quantile for the absolute victory margin. Ten percent of the time the there will be a victory margin of this size, but it may be a loss for Team A.
Both the mean victory margin and the standard deviation are larger for the case of displacement as seen by comparing Tables (ref) and (ref).
For completeness, we give the score difference distribution as a function of $r$ in Tables (ref) and (ref). Note that $s_A -s_B=0$ is the probability of a tie in Table (ref).
For international dual-meets, Table (ref) shows that more often than not the victory margin is divisible by four. Specifically, for $(6,4)$, the probability that the victory margin is divisible by four is .5625. For the random population problem with ratio $r=.5$, the probability that the victory margin is divisble by four is $.59375$.
Since a picture is worth a thousand words, Figure (ref)-(ref) plot the distributions of Table (ref)-(ref). For $N=4$, $r=.55$, $m_{A,B}=8$ is the most probable value. As $r$ increases, the primary maximum shifts to the largest possible value, $16$, with a secondary maximum at 8.
Figure (ref) appears smooth while Figure (ref) appears highly oscillatory. This is an illusion as the odd numbered values are identically zero for the no displacement case. These same values are small but nonzero for the case with displacement. Our input into the plotting package for Figure (ref) does not give the zero values at the odd integers.
Figure (ref) appears smooth because the even numbered values are identically zero for the no displacement case with $M=5$. For $M=5$, the secondary relative maximum of the density occurs at $m_{A,B}=15$.
We now discuss possible generalizations and alternative models. We can model the effect of an injury during the race such as a pulled hamstring by calculating the score distribution of $(M,M-1,N)$ such as $(7,6,5)$.
The models of Sections II and III are simple to explain and can be computed easily numerically. To get models that can be easily evaluated analytically, we can make more restrictive assumptions on the distribution of speeds. A very simple assumption is that the top runners on both teams are of equal ability and much faster than all other runners. We make this assumption iteratively on the remaining runners on the team. In this scenario, the distribution of scores corresponds to $N$ independent races. The score differences are then a shifted scaling binomial distribution. (The shifting and scaling occur because the score differences are $\{-1, 1\}$, not $\{0,1\}$.) Another tractable model is to assume that the top two runners of each team are much faster than all other runners and next two runners of each team are much faster than the rest of the runners. This essentially decomposes the scores for a $(M,N)$ race into two independent $(2,2)$ races plus the residual scoring from the remainder of the team. These last two examples illustrate the ease with which one can create new scenarios for cross-country models
Given actual race time data for each runner, we can create probabilistic models for distribution of race times for each runner. We continue to assume that the runners' times are independent. Of course, this ignores the competitive nature of racing: If another runner is running fast, you are likely to run faster even at the risk of burning yourself out and possibly finishing much worse. We could fit each runner's race times to a density, $p_i(t; c_i)$ where the free parameters $c_i$ are fit to the data. The probability that runner $i$ is faster than runner $j$ is given by $Prob(T_i < T_j) = \int P_i(t; c_i) p_j(t; c_j) \mathrm{d}t $ where $P_i$ is the integral of $p_i$: $P_i(t; c_i) \equiv \int_0^t p_i(s; c_i) ds$. The simplest model is that the $i$th runner's time is uniformly distributed on the interval, $[b_i,B_i]$. Given the values of $b_i$ and $B_i$ for each runner, one can compute the probability of each possible ordering either evaluating the integrals explicity (I assume using a symbolic manipulation program) or using Monte Carlo. I believe that the running times are probably more peaked about the mode with a long tail corresponding to “off” days. Thus a shifted stretched beta distribution with $a \approx 1.5$ and $b \approx 3$ may be a more realistic model. The shifted stretched beta distribution has four free parameters and this is often too many.
Not all races or race courses are equally easy. Hilly race courses and unpleasant race conditions will slow down all runners. Any attempt to fit empirical data will likely benefit from standardizing the times by race or at least by race course. I have personal experience with a runner who only wins on hilly race courses. Of course, modelling this would again add an extra free parameter to each density, $p_i(t; c_i)$, and is likely not worthwhile.
The two models presented above, the IID Outcome model and the Large Population model are baseline models models from which we get interesting and very testable score distributions. The population model gives a probability of victory when two schools of different sizes compete. Given a database of dual-meet scores and school sizes, it would be of interest to see how well the predictions in Tables 5 - 8 are supported by the data. Our tabulated distributions show that as $M$ increases at fixed $N$, the probability of a large victory increases. Figure (ref) shows the distributions for $N=5$ as M increases from $M=5$ to $M=7$ to $M=\infty$ for the case of no no displacement..
We defined a big victory to be a victory in the 90{\it th} percentile of victories. For the IID Outcome model, a big victory is $|m_{A,B}|=15$ with displacement and $|m_{A,B}|=14$ without for $(6,4)$ (International dual-meets). For American dual-meets, $(7,5)$, the $.9$ quantile is $|m_{A,B}|=21$ with displacement and $|m_{A,B}|=19$ without.
We have analyzed the statistical properties of rank sum scoring with and without scoring. The true argument in favor of scoring with displacement is to motivate the slower runners so that they can favorably impact the success of the team.
We thank the anonymmous referee for pointing out the Wilcoxon rank-sum statistic as well as many other references.