Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,143 characters · 30 sections · 8 citation commands
Partial Identification from LLM Prompts
\baselineskip=1.15\baselineskip
Large language models (LLMs) are now routinely used as binary classifiers. They label text as toxic or non-toxic, factual or non-factual, policy-violating or safe, relevant or irrelevant, and so on. In many applications the target label is not observed in the main sample. The econometrician observes only LLM reports and wants to learn the latent prevalence \[ \theta=\mathbb{P}(X^*=1), \] where $X^*\in\{0,1\}$ is the true label.
The central problem is not only that LLMs make mistakes. It is that there is more than one way to replicate an LLM measurement, and different replications produce different observable objects. This paper distinguishes three designs given in Table (ref) below.
If the same LLM is asked repeated exchangeable versions of the same binary question, a count is natural. By contrast, if GPT-4, GPT-4 Turbo, Claude, and open-weight models are each asked once, replacing the named vector by a vote count throws away information: the response patterns $(1,1,0,0)$ and $(0,0,1,1)$ have the same count but different implications if the first two models are more reliable. If each named model is asked repeated prompt variants, neither the simple count nor the single-prompt named vector is adequate; the analyst observes a two-way measurement panel and must decide whether model identity, prompt identity, or both should be preserved.
\paragraph{Paper Contributions} The baseline nonidentification result and the use of sensitivity or specificity-style restrictions to bound a prevalence are applications of partial identification techniques (see manski1999identification, tamer2010partial). The paper's contribution is to adapt, organize, and extend that literature to handle LLM measurement panels, where errors are plausibly dependent across models and prompts, and to add results that are, to our knowledge, new in this setting:
An augmented empirical illustration quantifies the practical payoff: in a toxicity-labeling application with three named LLMs, reporter-specific calibration on the named vector roughly halves the width of the identified set relative to the count coarsening used in earlier drafts.
We connect our results to important literatures.
Latent-class and multiple-rater models. Estimating rater error rates without a gold standard goes back at least to DawidSkene1979. That tradition typically obtains point identification through conditional independence of raters given the truth (or low-order dependence corrections). Our setting deliberately drops conditional independence: LLMs share training corpora, benchmarks, synthetic data, distillation pipelines, and alignment procedures, so their errors can be arbitrarily dependent given $X^*$. Without independence, the Dawid--Skene identification route is unavailable, and the model becomes a two-component mixture with unrestricted components---hence partial identification.
Partial Identification with Misclassification, corrupted data and Mixtures. Bounds on parameters under misclassified or contaminated outcomes are classical HorowitzManski1995,Molinari2008,Hu2008. Our calibrated-score bounds are recognizably of this family: bounds on sensitivity and specificity translate linearly into bounds on prevalence. What we add is (i) the score/event organization that nests reporter-specific accuracy, thresholds, unanimity, weighted ensembles, and matrix events as one assumption type rather than many; and (ii) the sharpness analysis of Section (ref), which distinguishes what is sharp given the score mean from what is sharp given the score law. Finally HenryKitamuraSalanie2014 study partial identification of finite mixtures using observable variation in mixture weights. Our degeneracy result (Proposition (ref)) is the mixture-identification observation specialized to LLM panels; its value is the discipline it imposes on applied work, not mathematical novelty. Our group-level extension (Appendix (ref)) connects directly to the exclusion-restriction logic of that literature.
Classifier calibration and LLM-based labeling. A growing literature calibrates classifier and LLM confidence and uses calibrated predictions for prevalence estimation SilvaFilho2023,Hovsepian2024,Multical2026. We treat external calibration as the source of identification: validated lower confidence bounds on score sensitivity/specificity or event predictive values are exactly the inputs our bounds require. The division of labor is deliberate: that literature supplies calibrated constants; this paper says what those constants identify under arbitrary dependence.
For each item, let $X^*\in\{0,1\}$ be the latent truth. Let $j=1,\dots,J$ index named LLMs and $m=1,\dots,M$ index repeated questions, prompt variants, stochastic completions, or elicitation templates. The binary response from model $j$ under prompt $m$ is $R_{jm}\in\{0,1\}$, and the full response matrix is \[ R=(R_{jm})_{j\le J,\,m\le M}\in\{0,1\}^{J\times M}. \] The parameter of interest is $\theta=\mathbb{P}(X^*=1)$. In the spirit of the partial identification literature, we develop bounds on $\theta$ using minimal plausible assumptions.
\ \ \
Write $\pi_R(r)=\mathbb{P}(R=r)$ and, conditional on the latent state, $f_z(r)=\mathbb{P}(R=r\mid X^*=z)$ for $z\in\{0,1\}$. Then
No factorization of $f_z$ is assumed: entries of $R$ may be arbitrarily dependent within each latent state.
The analyst may work with a finite summary $U=g(R)\in\mathcal{U}$: a count, a named response vector, a model-count vector, a prompt-count vector, a column-pattern histogram, or the matrix itself. Let $p_U(u)=\mathbb{P}(U=u)$ and $q_{z,U}(u)=\mathbb{P}(U=u\mid X^*=z)$. Every summary satisfies
For restrictions $\mathcal{A}_U$ on $(q_{0,U},q_{1,U})$, define the identified set
\ \
Weak shape restrictions do not alter this conclusion. For count summaries, first-order stochastic dominance (FOSD), monotone likelihood ratio (MLR) ordering, and weak mean ordering formalize the idea that positive items should produce more positive reports; for vector or matrix summaries, coordinatewise stochastic orders play the same role. These restrictions are often plausible, but they permit $q_0=q_1$, so the degenerate decomposition remains feasible.
For clarity, in a count experiment $S\in\{0,\ldots,M\}$, FOSD means \[ \sum_{t=s}^M q_1(t)\ge \sum_{t=s}^M q_0(t),\qquad s=1,\ldots,M. \] MLR means that $q_1/q_0$ is increasing in the usual cross-product sense: for $s>s'$, $q_1(s)q_0(s')\ge q_1(s')q_0(s)$, with the standard conventions at zeros. Weak mean ordering means $\mathbb{E}[S\mid X^*=1]\ge \mathbb{E}[S\mid X^*=0]$. For vector or matrix summaries, coordinatewise FOSD means $\mathbb{E}[\varphi(U)\mid X^*=1]\ge \mathbb{E}[\varphi(U)\mid X^*=0]$ for every bounded coordinatewise increasing function $\varphi$; equivalently, the inequality holds for all increasing upper sets. All of these are weak orders. Hence $q_0=q_1=p_U$ satisfies them with equality.
The nonidentification result has an inferential counterpart. Let $\mathcal{M}_U$ be a class of joint laws for $(X^*,U)$. Suppose that for every observable law $p\in\Delta(\mathcal{U})$ and every $t$ in a set $\Theta_0\subseteq[0,1]$, the uninformative experiment $U\sim p$, $X^*\sim\mathrm{Bernoulli}(t)$, $X^*\perp U$ belongs to $\mathcal{M}_U$.
\ \
Section (ref) sharpens this in the matrix design: Proposition (ref) characterizes exactly when a coarsening is lossless for the truth, and Theorem (ref) gives a design symmetry under which a specific coarsening is lossless for the identified set.
This section contains the paper's maintained identifying content. Everything else---reporter-specific accuracy, threshold classifiers, relaxed support, unanimity, weighted votes, matrix-agreement rules---is a special case obtained by choosing the summary $U$, the score $w$, or the events $A,B$.
Let $w:\mathcal{U}\to[0,1]$ be a pre-specified score and write \[ \bar{w}=\mathbb{E}[w(U)]=\sum_{u\in\mathcal{U}}w(u)\,p_U(u). \] The score may be a count fraction, a threshold rule, a named-model report, a weighted ensemble, or a matrix-agreement statistic.
The interval (ref) uses only the mean $\bar{w}$ of the score. When the analyst observes the full law of $w(U)$, the sharp identified set can be strictly smaller, and it admits an exact characterization through a rearrangement (concentration) function.
For $t\in[0,1]$ define
$W^+(t)$ is the largest possible contribution to $\mathbb{E}[w(U)]$ from a sub-population of mass $t$: it is computed by a greedy fill, allocating mass to the values of $u$ with the largest $w(u)$ first (a Hardy--Littlewood rearrangement bound). $W^+$ is concave and nondecreasing, with $W^+(0)=0$ and $W^+(1)=\bar{w}$, and satisfies $W^+(t)\le\min\{t,\bar{w}\}$.
Let $A\subseteq\mathcal{U}$ be a high-confidence positive event and $B\subseteq\mathcal{U}$ a high-confidence negative event.
Rich summaries admit many candidate high-agreement events. Rather than choosing one arbitrarily, the analyst can pre-specify a finite class and calibrate the whole class on a validation sample, treating event selection explicitly as a multiple-testing problem.
Let $\mathcal{A}^+$ be a finite class of positive events $A\subseteq\mathcal{U}$ (row-threshold events, column-pattern events, trusted-model events, weighted-score threshold events). Suppose validation data deliver simultaneous lower confidence bounds $\widehat\rho_A$ such that, with probability at least $1-\alpha$,
Then with the same probability all lower bounds hold simultaneously, so
and symmetrically $\theta\le\min_{B\in\mathcal{A}^-}\{1-\widehat\lambda_B\mathbb{P}(U\in B)\}$ for a simultaneously calibrated negative class $\mathcal{A}^-$.
The paper's baseline identification comes from calibrated scores and calibrated events. Two additional restrictions are useful as sensitivity analyses, but they should not be presented as maintained assumptions unless independently justified. For a count or vote score $C\in\{0,\ldots,J\}$, directional asymmetry with parameter $\gamma>0$ imposes \[ \mathbb{E}[C\mid X^*=0]\le \gamma\,\mathbb{E}[J-C\mid X^*=1], \] so false-positive votes are bounded relative to false-negative votes. For a one-LLM count $S\in\{0,\ldots,M\}$ with observed mean $\bar s=\mathbb{E}[S]$, anchored separation with tolerance $\varepsilon>0$ imposes \[ \mathbb{E}[S\mid X^*=1]\ge \bar s+\varepsilon, \qquad \mathbb{E}[S\mid X^*=0]\le \bar s-\varepsilon, \] which implies $\theta\ge \varepsilon/(\varepsilon+M-\bar s)$ and $\theta\le \bar s/(\bar s+\varepsilon)$. These restrictions are informative because they impose cross-state separation, not because they use replication by itself.
A single LLM is used repeatedly to measure the same latent binary truth. For one item, write the repeated reports as $R_1,\dots,R_M\in\{0,1\}$. If the repeated prompts are exchangeable probes of the same truth, their labels carry no structural content and the natural observable is the count \[ S=\sum_{m=1}^M R_m\in\{0,\dots,M\},\qquad p(s)=(1-\theta)q_0(s)+\theta q_1(s). \] No conditional independence is imposed across prompts: repeated prompts from one model may share the same systematic errors.
The count design has two natural calibrated scores: $w_1(S)=S/M$ and $w_k(S)=\mathbf{1}\{S\ge k\}$.
\paragraph{Average repeated-prompt accuracy.} With $w=S/M$ in Proposition (ref) and $\bar r=\mathbb{E}[S/M]$,
Because $S/M$ is a graded score, Theorem (ref) applies with content: the sharp set computed from $W^+$ can be strictly inside (ref), and is obtained by sorting the support of $S$.
\paragraph{Threshold accuracy.} For a threshold $k$, let $D_k=\mathbf{1}\{S\ge k\}$ and $p_k^+=\mathbb{P}(S\ge k)$. If validation data support $\mathbb{P}(D_k=1\mid X^*=1)\ge a_k$ and $\mathbb{P}(D_k=0\mid X^*=0)\ge b_k$, then \[ \max\Bigl\{0,\frac{p_k^++b_k-1}{b_k}\Bigr\}\le\theta\le\min\Bigl\{1,\frac{p_k^+}{a_k}\Bigr\}, \] and by Corollary (ref)(ii) these bounds are exactly sharp.
\paragraph{High- and low-count events.} With $A=\{S\ge k\}$ and $B=\{S\le\ell\}$, Proposition (ref) gives $\theta\ge\rho_k\mathbb{P}(S\ge k)$ and $\theta\le 1-\lambda_\ell\mathbb{P}(S\le\ell)$; Proposition (ref) gives the tail-error bounds \[ \theta\ \ge\ \max\Bigl\{0,\frac{\mathbb{P}(S\ge k)-\alpha_{0k}}{1-\alpha_{0k}}\Bigr\},\qquad \theta\ \le\ \min\Bigl\{1,\frac{1-\mathbb{P}(S\le\ell)}{1-\alpha_{1\ell}}\Bigr\}. \] Relaxed support is the special case $k=M$, $\ell=0$: $q_0(M)\le\alpha_0$, $q_1(0)\le\alpha_1$. Exact support ($\alpha_0=\alpha_1=0$) gives $p(M)\le\theta\le 1-p(0)$ but is rarely credible for LLMs; relaxed, validation-calibrated tail bounds are usually preferable.
\paragraph{Implication.} Design I is useful only to the extent that the count distribution can be calibrated. Report a small number of validation-calibrated score or event bounds---ideally the sharp set of Theorem (ref) for the graded score---rather than many sensitivity restrictions.
$J$ named LLMs each answer the same binary question once. The response vector is $Y=(Y_1,\dots,Y_J)\in\{0,1\}^J$ with law $\pi(y)$ and class-conditionals $f_z(y)$ satisfying $\pi(y)=(1-\theta)f_0(y)+\theta f_1(y)$. No conditional independence is assumed across named LLMs. The vote count $\sum_j Y_j$ is a coarsening, not the primitive object.
The named-vector design matters because it permits calibrated scores and events that use model identity: the single-reporter score $w_j(Y)=Y_j$, the weighted score $w(Y)=\sum_j c_jY_j$ with $c_j\ge0$, $\sum_j c_j=1$, and events $A\subseteq\{0,1\}^J$ encoding agreement by a trusted subset rather than simple majority.
\paragraph{Reporter-specific accuracy.} Suppose validation data provide
and let $m_j=\mathbb{P}(Y_j=1)$. Applying Proposition (ref) to each $w_j(Y)=Y_j$ and intersecting,
Each one-reporter bound is sharp (binary score); the intersection is valid and is the sharp set based on the marginals alone. Using the joint law $\pi$ with the restrictions (ref) simultaneously can tighten further; this is the LP of Section (ref).
\paragraph{Named events.} If $A\subseteq\{0,1\}^J$ satisfies $\mathbb{P}(X^*=1\mid Y\in A)\ge\rho_A$ then $\theta\ge\rho_A\pi(A)$; if $B$ satisfies $\mathbb{P}(X^*=0\mid Y\in B)\ge\lambda_B$ then $\theta\le 1-\lambda_B\pi(B)$. Relaxed unanimity is only the special case $A=\{\mathbf 1_J\}$, $B=\{\mathbf 0_J\}$ and should not be the default event unless validation evidence supports it.
\paragraph{The exact cost of counting votes.} If only the vote count $S=\sum_jY_j$ is stored, the reporter-specific marginals and named events are not recoverable; the count identifies only $\bar m=\frac1J\mathbb{E}[S]=\frac1J\sum_j m_j$. With averaged calibration constants $\beta_1=\frac1J\sum_j a_j$, $\beta_0=\frac1J\sum_j b_j$, the count-only analogue is
which is generally strictly weaker than (ref); Section (ref) quantifies the gap in the toxicity application. The same loss applies to weighted scores and trusted-subset events.
\paragraph{Implication.} Analyze Design II with the named vector. Count-based analysis is a robustness check or a fallback when only counts were stored. The identifying content should come from reporter-specific or named-event calibration, not from generic independence assumptions.
The third design is a two-way measurement panel: for one item, observe $R=(R_{jm})\in\{0,1\}^{J\times M}$ with mixture law (ref). No conditional independence is assumed across rows, columns, or cells. Design III is valuable not because it makes independence credible, but because it records where agreement occurs.
Useful lower-dimensional summaries: the model-count vector $T=(T_1,\dots,T_J)$, $T_j=\sum_m R_{jm}$; the prompt-count vector $S=(S_1,\dots,S_M)$, $S_m=\sum_j R_{jm}$; the total count $C=\sum_{j,m}R_{jm}$; and, when prompt labels are exchangeable but model labels are not, the column-pattern histogram
which counts the prompts on which the named response pattern equals $y$ (with $J=3$, $N_{110}$ counts prompts where models 1 and 2 say one and model 3 says zero). The histogram preserves cross-model agreement within prompts while discarding prompt labels, and refines the model-count vector via $T_j=\sum_y y_jN_y$. The coarsening hierarchy is \[ R\ \longrightarrow\ N\ \longrightarrow\ T\ \longrightarrow\ C,\qquad R\ \longrightarrow\ S\ \longrightarrow\ C, \] with support sizes $2^{JM}$, $\binom{M+2^J-1}{2^J-1}$, $(M+1)^J$, $(J+1)^M$, and $JM+1$ respectively. The full matrix distinguishes patterns with the same total count: one weak model saying yes on every prompt is not equivalent to several trusted models each saying yes repeatedly.
Which coarsenings lose information about $X^*$? The answer is a likelihood-ratio invariance condition.
Let $\Omega=\{0,1\}^{J\times M}$, let $U=h(R)$ be any finite summary with fibers $F_u=\{r:h(r)=u\}$, and let $q_z(u)=\sum_{r\in F_u}f_z(r)$.
Some reductions are not losses: they remove labels that have no design meaning. In Design III, prompt labels may be exchangeable even when model labels are not.
Let $S_M$ be the group of permutations of prompt labels; for $\sigma\in S_M$, $(\sigma r)_{jm}=r_{j,\sigma(m)}$. The histogram $N(R)$ indexes the orbits of this action: two matrices have the same $N$ iff one is obtained from the other by permuting prompt columns. The action extends to distributions by $(\sigma q)(r)=q(\sigma^{-1}r)$.
The reduction from $R$ to $N$ is lossless only when prompt exchangeability is part of the maintained design. To avoid ambiguity, let $\Theta_R^{\mathrm{ex}}(\pi_R;\mathcal{A}_R)$ denote the full-matrix identified set when admissible component distributions $q_0,q_1$ are required to be prompt-exchangeable in addition to satisfying the matrix-level restrictions $\mathcal{A}_R$. The required correspondence between matrix-level and histogram-level restrictions is explicit below.
Design III should not be analyzed as a large vote count unless both dimensions are intentionally treated as exchangeable. The natural calibrated objects are matrix scores and matrix events.
A weighted matrix score has the form \[ w(R)=\frac{\sum_{j}\sum_{m}c_{jm}R_{jm}}{\sum_{j}\sum_{m}c_{jm}},\qquad c_{jm}\ge0,\ \ \sum_{j,m}c_{jm}>0, \] with special cases $w=T_j/M$, $w=S_m/J$, and $w=C/(JM)$; if (ref) holds for $w$, Proposition (ref) and Theorem (ref) apply. Under prompt exchangeability a score may be written as a function of $N$, e.g.\ placing weight on column patterns where trusted models agree.
Matrix events encode repeated agreement across both dimensions, e.g. \[ A^+(K,L)=\Bigl\{R:\ \sum_{j=1}^J\mathbf{1}\{T_j\ge L\}\ \ge K\Bigr\},\qquad A^-(K,L)=\Bigl\{R:\ \sum_{j=1}^J\mathbf{1}\{T_j\le L\}\ \ge K\Bigr\}, \] the events that at least $K$ named models each produce at least (at most) $L$ positive prompt responses. The histogram also suggests events invisible from $T$: with a trusted subset $J_0\subseteq\{1,\dots,J\}$, \[ \sum_{y:\,y_j=1\ \forall j\in J_0} N_y\ \ge\ L \] requires the trusted models to agree positively on at least $L$ prompts, regardless of the other models, while remaining invariant to prompt relabeling. Calibrated predictive values or wrong-state errors for any of these events feed Propositions (ref)--(ref); searches over event classes use Section (ref).
The model-count vector is a tractable compromise when $R$ or $N$ is too large. If for each named model $j$ \[ \mathbb{E}[T_j/M\mid X^*=1]\ge a_j,\qquad \mathbb{E}[(M-T_j)/M\mid X^*=0]\ge b_j, \] then with $r_j=\mathbb{E}[T_j/M]$, \[ \max_{1\le j\le J}\max\Bigl\{0,\frac{r_j+b_j-1}{b_j}\Bigr\}\ \le\ \theta\ \le\ \min_{1\le j\le J}\min\Bigl\{1,\frac{r_j}{a_j}\Bigr\}, \] and symmetrically for calibrated prompt-count scores $S_m/J$. Model-specific bounds suit stable, calibratable model error profiles; prompt-specific bounds suit stable prompt families. If prompt labels are exchangeable but within-prompt agreement matters, $N$ is preferable to $T$ because it retains more agreement geometry.
Full cell-level calibration of every pair $(j,m)$ may be too parameter-rich. A practical compromise groups models and prompts: with model groups $g(j)$ and prompt families $h(m)$, calibrate $\mathbb{P}(R_{jm}=1\mid X^*=1)\ge a_{g(j),h(m)}$ and $\mathbb{P}(R_{jm}=0\mid X^*=0)\ge b_{g(j),h(m)}$, with the grouping fixed before analysis or justified by validation data.
Design III allows diagnostics that add no identifying assumptions.
\paragraph{Coarsening loss.} For $U\in\{R,N,T,S,C\}$ and compatible restrictions, report widths of identified sets and their differences, e.g. \[ \mathrm{Loss}(N\to T)=\mathrm{wid}(\Theta_T)-\mathrm{wid}(\Theta_N),\qquad \mathrm{Loss}(T\to C)=\mathrm{wid}(\Theta_C)-\mathrm{wid}(\Theta_T). \] A large loss from $N$ to $T$ means within-prompt cross-model agreement matters; from $T$ to $C$, that model identity matters. If prompt exchangeability is credible and restrictions are invariant and matched, Theorem (ref) predicts no loss from $R$ to $N$ at the population level---a checkable implication of the design symmetry.
\paragraph{Row and column influence.} Let $\Theta^{(-j)}$ and $\Theta^{(-m)}$ be the identified sets after dropping model $j$ or prompt $m$. Large movements in either bound show that conclusions hinge on a particular model or prompt; this is often more informative than another structural assumption.
\paragraph{Dependence and redundancy.} On a validation sample, estimate within-model dependence $\mathrm{Corr}(R_{jm},R_{jm'}\mid X^*=z)$ and across-model dependence $\mathrm{Corr}(R_{jm},R_{j'm}\mid X^*=z)$. High correlations indicate redundant measurements; low correlations indicate nonredundant variation. These are diagnostics, not independence assumptions.
\paragraph{Linear programming.} For any finite summary $U$, fixed-$\theta$ feasibility is a linear program. With joint masses $h_z(u)=\mathbb{P}(X^*=z,U=u)$: \[ h_0(u)+h_1(u)=p_U(u),\qquad \sum_u h_1(u)=\theta,\qquad h_z\ge0, \] and conditional restrictions become linear after multiplying through, e.g.\ $\mathbb{E}[w(U)\mid X^*=1]\ge a$ becomes $\sum_u w(u)h_1(u)\ge a\sum_u h_1(u)$. Sharp bounds under the baseline linear score and event restrictions are two LPs: minimize and maximize $\sum_u h_1(u)$. For fixed $\theta$, FOSD restrictions are linear in the conditional component probabilities and can be included in fixed-$\theta$ feasibility checks; exact MLR restrictions are bilinear in the components and should be handled only through explicit relaxations or nonlinear optimization. Prompt exchangeability is implemented either by using $N$ directly or by orbit-equality constraints on matrix probabilities.
\paragraph{The sharp score set by sorting.} $W^+(t)$ in (ref) is computed greedily: order support points by descending $w(u)$ and fill mass to total $t$. $W^+$ is piecewise linear and concave, so $\Theta_w$ in (ref) is found by intersecting a concave piecewise-linear function with two affine functions---no LP solver needed.
\paragraph{Sampling uncertainty.} When $p_U$ is estimated from $N$ items, replace mixture equalities by bands $|(1-\theta)q_0(u)+\theta q_1(u)-\widehat p_U(u)|\le\Delta_u$. A simple finite-sample choice is the Hoeffding-union bound
with $\Delta_u=\varepsilon_N(\alpha)$ for every $u$, yielding a conservative outer confidence set; less conservative alternatives use the multinomial bootstrap, empirical Bernstein bands, or moment-inequality methods. For optimized event bounds, sampling uncertainty enters twice---through $\widehat p_U$ in the main sample and through validation estimates of predictive values---and validity requires the simultaneous calibration (ref).
\paragraph{Specification testing by minimum tolerance.} Let $\mathcal{Q}$ be the set of distributions over $\mathcal{U}$ implied by the maintained restrictions for some $\theta$, and define \[ \Delta^*=\min_{q\in\mathcal{Q}}\|\widehat p_U-q\|_\infty, \] the minimum $\ell_\infty$ tolerance restoring feasibility. If the restrictions are correct at the population law $p_U^0$, then $\Delta^*\le\|\widehat p_U-p_U^0\|_\infty$, so a finite-sample test rejects at level $\alpha$ if $\Delta^*>\varepsilon_N(\alpha)$, and \[ \Delta_0\in\bigl[\max\{0,\Delta^*-\varepsilon_N(\alpha)\},\ \Delta^*+\varepsilon_N(\alpha)\bigr] \] is a confidence interval for the population misspecification distance. The test detects out-of-distribution collapse: if an LLM panel becomes uninformative, restrictions calibrated in distribution may become infeasible. Emptiness of the sharp score set $\Theta_w$ is the same idea specialized to a single calibrated score.
Set $M=5$ and $\theta_0=0.65$, with class-conditional counts $S\mid X^*=1\sim\mathrm{BetaBin}(5,4,1.5)$ and $S\mid X^*=0\sim\mathrm{BetaBin}(5,1.5,4)$; the observed histogram is $p=(1-\theta_0)q_0+\theta_0q_1$. The design has $\mathbb{E}[S\mid X^*=0]=1.36$, $\mathbb{E}[S\mid X^*=1]=3.64$, average repeated-prompt accuracy $0.727$, and observed mean $\mathbb{E}[S]=2.84$, so $\bar{w}=\mathbb{E}[S/M]=0.568$ for the graded score $w=S/M$. Without restrictions, or under FOSD/MLR/weak mean ordering alone, the identified set is $[0,1]$ because the degenerate decomposition $q_0=q_1=p$ remains feasible (Proposition (ref)); Table (ref) records this once as a benchmark and then isolates the paper's main computational comparison: the mean-only linear bounds (ref) versus the sharp rearrangement set of Theorem (ref), computed from $W^+$ by sorting the six support points of $S$.
The table mirrors the paper's identification message in miniature. Weak ordering restrictions leave $[0,1]$ untouched, so they are recorded in a single benchmark row. Calibration is what moves the set, and the shape of the score distribution carries identifying content beyond its mean, exactly as Theorem (ref) predicts: at $a=b=0.60$ the rearrangement constraint $W^+(\theta)\ge \bar{w}-(1-b)(1-\theta)$ binds before the linear lower bound, raising $\theta_L$ from $0.280$ to $0.317$; at $a=b=0.70$ the sharp set $[0.479,0.784]$ removes $29\%$ of the linear interval's width ($0.429$ to $0.305$), tightening both endpoints. Because $w(S)$ is a graded score on six support points, this gain is obtained by a sort and costs nothing computationally. Exact support tightens only the weakly calibrated case ($a=b=0.60$, where it caps $\theta_U$ at $0.882$) and is redundant once calibration is strong; this ordering---calibration first, support as a benchmark---is the recommended reporting style of Section (ref).
This section revisits the toxicity-response classification task studied by ChengEtAl2024SoftLabel. The task is to bound the prevalence of toxic responses when the analyst observes binary labels from LLM annotators and, for a validation sample, expert-adjudicated truth (which gives us a benchmark to allow us to validate our bounds).
In particular, the data involve $J=3$ named LLM annotators---GPT-4, GPT-4 Turbo, and Claude-2---and three human annotators. This is a Design II setting. Following the taxonomy of this paper, we now report the named-vector analysis in three increasingly disciplined steps---plug-in marginal bounds, the full named-vector LP on the joint law of $Y=(Y_1,Y_2,Y_3)$, and an honest version that replaces plug-in calibration constants by one-sided lower confidence bounds---and we tag every bound with the response object that must be stored to compute it.
Table (ref) shows why named-vector analysis matters: GPT-4 is much more accurate than Claude-2, and all three LLMs have higher specificity than sensitivity. A count-only analysis collapses this heterogeneity into an exchangeable average. The validation class-conditionals also show that exact support is empirically violated with only three LLMs: among truly toxic items, 14.2% receive zero positive LLM votes; among truly non-toxic items, 4.3% receive unanimous positive votes. Exact support can be reported as a benchmark, but relaxed support or validation-calibrated score bounds are more credible.
Table (ref) records the observable objects that the bounds below are built from. The top panel gives the count histograms---$\widehat p(s)$ with the validation class-conditionals $\widehat q_z(s)$ for $N=1{,}000$, and the count-only training histogram for $N=28{,}194$---while the lower panel reports the full named-vector law $\widehat\pi(y)$ over $\{0,1\}^3$. The two are not interchangeable: the count is the coarsening $s=\sum_j y_j$ of $\widehat\pi$, so patterns that disagree on which model flagged the item are merged. Here $(1,1,0)$, where GPT-4 and GPT-4 Turbo flag and Claude-2 abstains, has mass $0.250$, whereas the equal-count pattern $(0,1,1)$ has mass only $0.010$; counting pools them into $\widehat p(2)=0.295$ and discards exactly the reporter identity that the named-vector bounds will exploit.
Table (ref) reports reporter-specific bounds (ref) in two versions. Panel A is the plug-in benchmark: calibration constants $(a_j,b_j)$ are the validation sensitivities and specificities of Table (ref), estimated on the same sample as the observed rates $m_j$, so the panel illustrates the identification logic but is not honest inference. Panel B is the honest calibration version recommended in Section (ref): $(a_j,b_j)$ are replaced by one-sided Clopper--Pearson lower confidence bounds $(a_j^L,b_j^L)$ at Bonferroni level $\alpha/6$ with $\alpha=0.05$, so all six constraints hold simultaneously with probability at least $0.95$ and the resulting bounds are valid by Proposition (ref).
Table (ref) is the section's central exhibit. It reports four identified sets for the same population, ordered by the richness of the stored response object, including the full named-vector LP of Section (ref), which imposes the mixture equation on the joint law $\widehat\pi(y)$ over $\{0,1\}^3$ together with all reporter-specific calibration restrictions simultaneously.
Three facts stand out. First, the cost of counting: moving from the named vector to the count widens the set from $[0.497,0.707]$ to $[0.348,0.731]$---the width nearly doubles, from $0.209$ to $0.383$---purely because counting discards reporter identity, quantifying Theorem (ref). Second, the full LP on the joint law coincides with the marginal intersection to four decimals. This is informative rather than disappointing: it certifies that, given reporter-specific calibration alone, the marginal intersection is already sharp at this observed law, so no further information can be extracted from these restrictions; the joint law is nevertheless the object to store, because only it makes the sharpness check possible and because pattern-level calibrated events (trusted-subset agreement, Section (ref)) require it whenever validation evidence supports them. Third, honest calibration costs about $0.07$ of width relative to the plug-in benchmark but preserves the storage ranking: the honest named-vector set (width $0.277$) remains far narrower than the honest count set (width $0.480$).
The binding reporter on both sides is GPT-4, the most accurate model; the influence diagnostic of Section (ref) applies verbatim: dropping GPT-4 widens the plug-in intersection to the Turbo bounds $[0.391,0.730]$. For the training set ($N=28{,}194$), only the count histogram was stored, so every named-vector row of Table (ref) is unavailable by construction: with $\bar m=1.847/3=0.616$ and the exchangeable-average constants, (ref) gives $[0.554,\,1.000]$. The storage choice has real consequences: on the training set the count bound is $[0.554,1.000][0.554, 1.000] [0.554,1.000]$, whose upper endpoint is the trivial value 1, so counting leaves the prevalence bounded only from below. Had the named vector been stored, both endpoints would be informative.
So far the target has been the scalar prevalence $\theta=\mathbb{P}(X^*=1)$. In many applications $X^*$ is not the final object but a latent regressor: the analyst wants the partial association between an observed outcome and the latent label, holding covariates fixed. This section shows that the calibration bounds feed directly into such a regression, with prevalence recovered as the special case of a constant outcome.
For each item let $(V,Z,U)$ be observed, where $V\in\mathbb{R}$ is a scalar outcome, $Z\in\mathbb{R}^{d}$ a vector of observed covariates, and $U=g(R)$ the LLM summary of Sections (ref)--(ref); $X^*\in\{0,1\}$ remains latent. (We write the outcome as $V$ to avoid collision with the named-LLM vector $Y$ of Section (ref).) Stack the regressors as $W=(1,Z',X^*)'\in\mathbb{R}^{d+2}$ and consider the best linear predictor (BLP) coefficient \[ \beta \;=\; \big(\mathbb{E}[WW']\big)^{-1}\mathbb{E}[WV], \] assuming $\mathbb{E}[WW']$ is nonsingular. The coefficient on $X^*$, denoted $\gamma$, is typically the parameter of interest; if $\mathbb{E}[V\mid Z,X^*]$ is linear, $\beta$ is the conditional-expectation coefficient.
\paragraph{Reduction to latent cross-moments.} Because $(V,Z)$ are observed and $X^*$ is binary (so $(X^*)^2=X^*$), every entry of $\mathbb{E}[WW']$ and $\mathbb{E}[WV]$ is identified from the data except the blocks that pair $X^*$ with an observed variable: \[ \theta=\mathbb{E}[X^*],\qquad \mathbb{E}[X^*Z]\in\mathbb{R}^{d},\qquad \mathbb{E}[X^*V]\in\mathbb{R}. \] Collect these in $\mu=(\theta,\mathbb{E}[X^*Z]',\mathbb{E}[X^*V])'$. Then $\beta=\Phi(\mu;m_{\mathrm{obs}})$ for a known map $\Phi$ that is rational in $\mu$---a matrix inverse times a vector, both affine in $\mu$---where $m_{\mathrm{obs}}$ collects the observed moments. Identifying $\beta$ thus reduces to identifying $\mu$.
\paragraph{Bounding the latent cross-moments.} Since $(U,Z,V)$ are jointly observed, write the unknown attachment of the latent label to the data as \[ \eta(u,z,v)=\mathbb{P}(X^*=1\mid U=u,Z=z,V=v)\in[0,1]. \] Every latent cross-moment is a linear functional of $\eta$: for any observed $h$, \[ \mathbb{E}[X^*\,h(U,Z,V)]=\mathbb{E}[h(U,Z,V)\,\eta(U,Z,V)]. \]
The calibration inequalities enter exactly as in Section (ref): a score restriction $\mathbb{E}[w(U)\mid X^*=1]\ge a$ becomes $\sum_u w(u)\,\mathbb{E}[\eta\mid U=u]\,p_U(u)\ge a\,\mathbb{E}[\eta]$ after writing $\mathbb{E}[X^*w(U)]=\mathbb{E}[w(U)\eta]$, and an event restriction becomes the corresponding linear inequality; both are linear in $\eta$. The lemma is therefore the paper's fixed-$\theta$ LP with the objective $\mathbb{E}[\eta]$ replaced by $\mathbb{E}[h\,\eta]$.
Replacing $\mathcal M$ by the box of component-wise intervals from Lemma (ref) gives valid but generally conservative outer bounds; the joint program is sharp. This is the same distinction drawn for the named-vector LP versus the marginal intersection in Section (ref).
\paragraph{A transparent special case.} With no covariates and target the mean contrast $\gamma=\mathbb{E}[V\mid X^*=1]-\mathbb{E}[V\mid X^*=0]$, \[ \gamma=\frac{\mathbb{E}[VX^*]-\theta\,\mathbb{E}[V]}{\theta(1-\theta)}, \] so $\gamma$ is pinned down once $\theta$ and $\mathbb{E}[VX^*]$ are, each bounded by Lemma (ref); the identified set is obtained by ranging the numerator and $\theta$ jointly over their calibrated region. If, in addition, $V$ is itself a calibrated high-confidence label, the contrast inherits the sharpness of Section (ref).
\paragraph{Inference.} Sampling uncertainty enters through the observed law of $(U,Z,V)$ and, for honest calibration, through the validation estimates of the calibration constants. Propagating the bands (ref) of Section (ref) (or a multiplier bootstrap) through the LPs of Lemma (ref) yields a valid outer confidence set for each coordinate of $\beta$; moment-inequality methods apply verbatim because all restrictions are linear in $\eta$.
\paragraph{Use calibration as the main identifying content.} The maintained assumptions should be calibrated score and event restrictions. Report the source of calibration, preferably lower confidence bounds from validation data on a split disjoint from the main sample.
\paragraph{Treat weak ordering as regularity, not identification.} FOSD, MLR, and weak mean ordering are plausible but do not rule out $q_0=q_1$ and should not be presented as identifying assumptions.
\paragraph{Report sharp sets where they are cheap.} For binary scores and events, the linear bounds are already sharp. For graded scores, report the sharp set of Theorem (ref); it costs a sort, and the gap from the linear bounds measures the information in the score's shape.
\paragraph{Exploit design symmetries, not independence assumptions.} If prompt labels are exchangeable by design, reduce the matrix to the column-pattern histogram $N$ under matched, invariant restrictions (Assumption (ref)) rather than imposing conditional independence.
\paragraph{Store the richest object possible.} For Design I, store all repeated responses even if the count is analyzed. For Design II, store the named vector. For Design III, store the full matrix whenever feasible; under prompt exchangeability, store or construct $N$; at minimum store the model-count vector. The training-set panel of Section (ref) shows what is lost otherwise.
\paragraph{Report coarsening and influence diagnostics.} Compare bounds under richer and coarser summaries; report whether results depend on particular models or prompts; quantify losses along $R\to N\to T\to C$ where relevant.
\paragraph{Keep sensitivity assumptions separate.} Directional asymmetry, anchored separation, and independence-style restrictions can be useful, but label them as sensitivity analysis unless independently justified.
The identifying content of LLM measurement panels depends on the source of replication. Asking one LLM repeated exchangeable questions creates a count experiment; asking several named LLMs once creates a named-reporter experiment; asking several named LLMs repeated questions creates a two-way response matrix. These designs should not be collapsed into a single vote-count model at the outset.
The general theory is simple, and we present it as an adaptation of known mixture and misclassification logic to a setting where dependence among reporters is the rule. Every finite summary yields a two-component mixture for the latent truth; without restrictions, prevalence is completely unidentified, and weak ordering restrictions do not help because they permit equality of the latent components. Useful identification comes from calibrated scores and calibrated events, of which reporter-specific accuracy, threshold rules, relaxed support, unanimity, weighted ensembles, and matrix-agreement events are special cases. The sharpness analysis clarifies exactly what each calibrated object delivers: linear bounds that are sharp for binary scores and for any analysis recording only the score mean, and a rearrangement characterization of the strictly sharper set available from the full score law.
The matrix design adds one further lesson. Even when entries of the matrix are arbitrarily correlated, design symmetries justify lossless reductions: under prompt exchangeability and matched invariant restrictions, the column-pattern histogram preserves all identifying information in the full matrix. Calibration identifies; storage and symmetry determine what can be calibrated without loss.