EconBase
← Back to paper

Partial Identification from LLM Prompts

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

76,143 characters · 30 sections · 8 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Identification from LLM Prompts

abstractLarge language models are increasingly used as binary classifiers when the true label is latent. We study partial identification of the prevalence $(\theta=P(X^*=1))$ from panels of LLM reports whose errors may be arbitrarily dependent given the truth. The design of replication determines the observable, and hence the identifying content: repeated prompts to one model yield a count, several named models a response vector, and both a response matrix. Cast as a two-component finite mixture, the problem makes the identification failure transparent—absent restrictions that separate the latent components, the prevalence $\theta$ is completely unidentified, and weak stochastic-ordering restrictions (first-order dominance, monotone likelihood ratio, mean ordering) leave the identified set at $[0,1]$. Identifying power comes instead from externally calibrated scores and events, which discipline the mixture in the spirit of the misclassification and corrupted-data literature. We characterize the resulting bounds, establishing validity and sharpness, and give an exact account of the identifying information in the full score distribution beyond its mean. When named models are asked repeated versions of the same question, what identifies $\theta$ is not the number of positive answers but which models agree across prompts—a feature a vote count discards. An extension derives implied bounds on regression coefficients when $X^*$ is a regressor of interest that is not directly observed.

\baselineskip=1.15\baselineskip

Introduction

Large language models (LLMs) are now routinely used as binary classifiers. They label text as toxic or non-toxic, factual or non-factual, policy-violating or safe, relevant or irrelevant, and so on. In many applications the target label is not observed in the main sample. The econometrician observes only LLM reports and wants to learn the latent prevalence \[ \theta=\mathbb{P}(X^*=1), \] where $X^*\in\{0,1\}$ is the true label.

The central problem is not only that LLMs make mistakes. It is that there is more than one way to replicate an LLM measurement, and different replications produce different observable objects. This paper distinguishes three designs given in Table (ref) below.

table[table omitted — 1,301 chars of source]

If the same LLM is asked repeated exchangeable versions of the same binary question, a count is natural. By contrast, if GPT-4, GPT-4 Turbo, Claude, and open-weight models are each asked once, replacing the named vector by a vote count throws away information: the response patterns $(1,1,0,0)$ and $(0,0,1,1)$ have the same count but different implications if the first two models are more reliable. If each named model is asked repeated prompt variants, neither the simple count nor the single-prompt named vector is adequate; the analyst observes a two-way measurement panel and must decide whether model identity, prompt identity, or both should be preserved.

\paragraph{Paper Contributions} The baseline nonidentification result and the use of sensitivity or specificity-style restrictions to bound a prevalence are applications of partial identification techniques (see manski1999identification, tamer2010partial). The paper's contribution is to adapt, organize, and extend that literature to handle LLM measurement panels, where errors are plausibly dependent across models and prompts, and to add results that are, to our knowledge, new in this setting:

enumerate[leftmargin=1.6em,itemsep=1pt] • A design taxonomy (Table (ref)) that maps the source of replication---repeated prompts, named models, or both---into the correct observable object, with an exact characterization of when a coarsening of the response matrix is truth-sufficient (Proposition (ref)). • A unified calibration theory: all of our identifying restrictions are calibrated scores or calibrated events. We prove validity of the resulting bounds, give an exact sharpness characterization of the score bounds via a rearrangement function of the observed score distribution (Theorem (ref)), show the simple linear bounds are sharp when only the score mean is recorded and exactly sharp for binary scores (Corollary (ref)), and prove sharpness of the event bounds. • A symmetry result for the two-way design: under prompt exchangeability and an explicit restriction-compatibility condition (Assumption (ref)), the full matrix has a lossless reduction to the column-pattern histogram $N$ (Theorem (ref)). No independence is assumed anywhere. • Diagnostics---coarsening loss, row/column influence, dependence audits, and a minimum-tolerance specification test---that add transparency without adding identifying assumptions, together with a multiple-testing-aware treatment of optimized calibrated events.

An augmented empirical illustration quantifies the practical payoff: in a toxicity-labeling application with three named LLMs, reporter-specific calibration on the named vector roughly halves the width of the identified set relative to the count coarsening used in earlier drafts.

Relation to the literature

We connect our results to important literatures.

Latent-class and multiple-rater models. Estimating rater error rates without a gold standard goes back at least to DawidSkene1979. That tradition typically obtains point identification through conditional independence of raters given the truth (or low-order dependence corrections). Our setting deliberately drops conditional independence: LLMs share training corpora, benchmarks, synthetic data, distillation pipelines, and alignment procedures, so their errors can be arbitrarily dependent given $X^*$. Without independence, the Dawid--Skene identification route is unavailable, and the model becomes a two-component mixture with unrestricted components---hence partial identification.

Partial Identification with Misclassification, corrupted data and Mixtures. Bounds on parameters under misclassified or contaminated outcomes are classical HorowitzManski1995,Molinari2008,Hu2008. Our calibrated-score bounds are recognizably of this family: bounds on sensitivity and specificity translate linearly into bounds on prevalence. What we add is (i) the score/event organization that nests reporter-specific accuracy, thresholds, unanimity, weighted ensembles, and matrix events as one assumption type rather than many; and (ii) the sharpness analysis of Section (ref), which distinguishes what is sharp given the score mean from what is sharp given the score law. Finally HenryKitamuraSalanie2014 study partial identification of finite mixtures using observable variation in mixture weights. Our degeneracy result (Proposition (ref)) is the mixture-identification observation specialized to LLM panels; its value is the discipline it imposes on applied work, not mathematical novelty. Our group-level extension (Appendix (ref)) connects directly to the exclusion-restriction logic of that literature.

Classifier calibration and LLM-based labeling. A growing literature calibrates classifier and LLM confidence and uses calibrated predictions for prevalence estimation SilvaFilho2023,Hovsepian2024,Multical2026. We treat external calibration as the source of identification: validated lower confidence bounds on score sensitivity/specificity or event predictive values are exactly the inputs our bounds require. The division of labor is deliberate: that literature supplies calibrated constants; this paper says what those constants identify under arbitrary dependence.

Framework and nonidentification

Latent truth, response matrices, and summaries

For each item, let $X^*\in\{0,1\}$ be the latent truth. Let $j=1,\dots,J$ index named LLMs and $m=1,\dots,M$ index repeated questions, prompt variants, stochastic completions, or elicitation templates. The binary response from model $j$ under prompt $m$ is $R_{jm}\in\{0,1\}$, and the full response matrix is \[ R=(R_{jm})_{j\le J,\,m\le M}\in\{0,1\}^{J\times M}. \] The parameter of interest is $\theta=\mathbb{P}(X^*=1)$. In the spirit of the partial identification literature, we develop bounds on $\theta$ using minimal plausible assumptions.

\ \ \

Write $\pi_R(r)=\mathbb{P}(R=r)$ and, conditional on the latent state, $f_z(r)=\mathbb{P}(R=r\mid X^*=z)$ for $z\in\{0,1\}$. Then

equation[equation omitted — 112 chars of source]

No factorization of $f_z$ is assumed: entries of $R$ may be arbitrarily dependent within each latent state.

The analyst may work with a finite summary $U=g(R)\in\mathcal{U}$: a count, a named response vector, a model-count vector, a prompt-count vector, a column-pattern histogram, or the matrix itself. Let $p_U(u)=\mathbb{P}(U=u)$ and $q_{z,U}(u)=\mathbb{P}(U=u\mid X^*=z)$. Every summary satisfies

equation[equation omitted — 104 chars of source]

For restrictions $\mathcal{A}_U$ on $(q_{0,U},q_{1,U})$, define the identified set

equation[equation omitted — 172 chars of source]

\ \

Degeneracy and weak ordering

proposition[Nonidentification of $\theta$] Fix any finite summary $U=g(R)$. Without restrictions that rule out equality of the latent component distributions, $\Theta_U(p_U)=[0,1]$: for every $\theta\in[0,1]$, the choice $q_{0,U}=q_{1,U}=p_U$ satisfies (ref) and the data contain no information about $\theta.$
proofFor any $u$, $(1-\theta)p_U(u)+\theta p_U(u)=p_U(u)$, for every $\theta\in[0,1]$.

Weak shape restrictions do not alter this conclusion. For count summaries, first-order stochastic dominance (FOSD), monotone likelihood ratio (MLR) ordering, and weak mean ordering formalize the idea that positive items should produce more positive reports; for vector or matrix summaries, coordinatewise stochastic orders play the same role. These restrictions are often plausible, but they permit $q_0=q_1$, so the degenerate decomposition remains feasible.

For clarity, in a count experiment $S\in\{0,\ldots,M\}$, FOSD means \[ \sum_{t=s}^M q_1(t)\ge \sum_{t=s}^M q_0(t),\qquad s=1,\ldots,M. \] MLR means that $q_1/q_0$ is increasing in the usual cross-product sense: for $s>s'$, $q_1(s)q_0(s')\ge q_1(s')q_0(s)$, with the standard conventions at zeros. Weak mean ordering means $\mathbb{E}[S\mid X^*=1]\ge \mathbb{E}[S\mid X^*=0]$. For vector or matrix summaries, coordinatewise FOSD means $\mathbb{E}[\varphi(U)\mid X^*=1]\ge \mathbb{E}[\varphi(U)\mid X^*=0]$ for every bounded coordinatewise increasing function $\varphi$; equivalently, the inequality holds for all increasing upper sets. All of these are weak orders. Hence $q_0=q_1=p_U$ satisfies them with equality.

A confidence-set implication

The nonidentification result has an inferential counterpart. Let $\mathcal{M}_U$ be a class of joint laws for $(X^*,U)$. Suppose that for every observable law $p\in\Delta(\mathcal{U})$ and every $t$ in a set $\Theta_0\subseteq[0,1]$, the uninformative experiment $U\sim p$, $X^*\sim\mathrm{Bernoulli}(t)$, $X^*\perp U$ belongs to $\mathcal{M}_U$.

theorem[Distribution-free impossibility] Let $\widehat\Theta_{n,\alpha}\subseteq[0,1]$ be a possibly randomized confidence set for $\theta$, constructed from an i.i.d.\ sample $U_1,\dots,U_n$. If $\sup_{P\in\mathcal{M}_U}\mathbb{P}_P\{\theta(P)\notin\widehat\Theta_{n,\alpha}\}\le\alpha$, then for every observable law $p$ and every $t\in\Theta_0$, $\mathbb{P}_p\{t\notin\widehat\Theta_{n,\alpha}\}\le\alpha$. Consequently \[ \mathbb{E}_p\bigl[\lambda(\widehat\Theta_{n,\alpha}\cap\Theta_0)\bigr]\ge(1-\alpha)\lambda(\Theta_0), \] where $\lambda$ is Lebesgue measure. If $\Theta_0=[0,1]$ then $\mathbb{E}_p[\lambda(\widehat\Theta_{n,\alpha})]\ge 1-\alpha$, and if the confidence set is always an interval contained in $[0,1]$ then $\mathbb{P}_p\{\widehat\Theta_{n,\alpha}=[0,1]\}\ge 1-2\alpha$.
proofFix $p$ and $t\in\Theta_0$. The law with $U\sim p$, $X^*\sim\mathrm{Bernoulli}(t)$, $X^*\perp U$ belongs to $\mathcal{M}_U$; under it the sample has distribution $p^n$ and the prevalence is $t$, so coverage implies $\mathbb{P}_p\{t\notin\widehat\Theta_{n,\alpha}\}\le\alpha$. Integrating over $t\in\Theta_0$ (Tonelli) gives the expected-length bound. If $\Theta_0=[0,1]$ and the set is always an interval, coverage of both endpoints implies the interval is $[0,1]$; the final claim follows by the union bound.

\ \

Coarsening and information loss

theorem[Coarsening weakens identification under compatible restrictions] Let $U_2=h(U_1)$. Suppose every feasible $(\theta,q_{0,U_1},q_{1,U_1})$ under restrictions $\mathcal{A}_{U_1}$ projects to a feasible $(\theta,q_{0,U_2},q_{1,U_2})$ under restrictions $\mathcal{A}_{U_2}$. Then $\Theta_{U_1}(p_{U_1};\mathcal{A}_{U_1})\subseteq\Theta_{U_2}(p_{U_2};\mathcal{A}_{U_2})$.
proofGiven a feasible point under $U_1$, define $q_{z,U_2}(u_2)=\sum_{u_1:h(u_1)=u_2}q_{z,U_1}(u_1)$. Aggregating (ref) over fibers of $h$ gives the mixture equation for $U_2$; compatibility gives the projected restrictions.

Section (ref) sharpens this in the matrix design: Proposition (ref) characterizes exactly when a coarsening is lossless for the truth, and Theorem (ref) gives a design symmetry under which a specific coarsening is lossless for the identified set.

Identification by calibration

This section contains the paper's maintained identifying content. Everything else---reporter-specific accuracy, threshold classifiers, relaxed support, unanimity, weighted votes, matrix-agreement rules---is a special case obtained by choosing the summary $U$, the score $w$, or the events $A,B$.

Calibrated score bounds: validity

Let $w:\mathcal{U}\to[0,1]$ be a pre-specified score and write \[ \bar{w}=\mathbb{E}[w(U)]=\sum_{u\in\mathcal{U}}w(u)\,p_U(u). \] The score may be a count fraction, a threshold rule, a named-model report, a weighted ensemble, or a matrix-agreement statistic.

proposition[Calibrated score bounds] Suppose \begin{equation} \mathbb{E}[w(U)\mid X^*=1]\ge a,\qquad \mathbb{E}[1-w(U)\mid X^*=0]\ge b,\qquad a,b\in(0,1]. \end{equation} Then every admissible prevalence satisfies \begin{equation} \max\Bigl\{0,\ \frac{\bar{w}+b-1}{b}\Bigr\}\ \le\ \theta\ \le\ \min\Bigl\{1,\ \frac{\bar{w}}{a}\Bigr\}. \end{equation}
proofLet $\mu_z=\mathbb{E}[w(U)\mid X^*=z]$, so $\bar{w}=(1-\theta)\mu_0+\theta\mu_1$. Since $\mu_1\ge a$ and $\mu_0\ge 0$, $\bar{w}\ge\theta a$, giving the upper bound. Since $\mu_0\le 1-b$ and $\mu_1\le 1$, $\bar{w}\le(1-\theta)(1-b)+\theta=1-b+b\theta$, giving the lower bound. Intersect with $[0,1]$.

Sharpness: score mean vs score law

The interval (ref) uses only the mean $\bar{w}$ of the score. When the analyst observes the full law of $w(U)$, the sharp identified set can be strictly smaller, and it admits an exact characterization through a rearrangement (concentration) function.

For $t\in[0,1]$ define

equation[equation omitted — 152 chars of source]

$W^+(t)$ is the largest possible contribution to $\mathbb{E}[w(U)]$ from a sub-population of mass $t$: it is computed by a greedy fill, allocating mass to the values of $u$ with the largest $w(u)$ first (a Hardy--Littlewood rearrangement bound). $W^+$ is concave and nondecreasing, with $W^+(0)=0$ and $W^+(1)=\bar{w}$, and satisfies $W^+(t)\le\min\{t,\bar{w}\}$.

theorem[Sharp identified set under score calibration] Maintain (ref) and suppose the analyst observes $p_U$ (hence the law of $w(U)$) but imposes no other restriction. The identified set for $\theta$ is \begin{equation} \Theta_w\ =\ \Bigl\{\theta\in[0,1]\ :\ W^+(\theta)\ \ge\ a\theta\ \ and\ \ W^+(\theta)\ \ge\ \bar{w}-(1-b)(1-\theta)\Bigr\}, \end{equation} and $\Theta_w$ is a (possibly empty) closed interval.
proofWork with the joint mass $h_1(u)=\mathbb{P}(X^*=1,U=u)$ and $h_0=p_U-h_1$. Feasibility of a prevalence $\theta$ is equivalent to the existence of $h_1$ with \[ 0\le h_1\le p_U,\qquad \sum_u h_1(u)=\theta,\qquad \sum_u w(u)h_1(u)\ge a\theta,\qquad \sum_u w(u)h_0(u)\le(1-b)(1-\theta), \] where the third inequality is $\mathbb{E}[w\,\mathbf{1}\{X^*=1\}]\ge a\,\mathbb{P}(X^*=1)$, i.e.\ (ref) for $\mu_1$, and the fourth is (ref) for $\mu_0$. The fourth inequality rewrites as $\sum_u w(u)h_1(u)\ge\bar{w}-(1-b)(1-\theta)$. The feasible set for $h_1$ given the first two constraints is a nonempty polytope (for $\theta\in[0,1]$), and the achievable values of the linear functional $s=\sum_u w(u)h_1(u)$ over that polytope form a closed interval $[W^-(\theta),W^+(\theta)]$. Hence a feasible $h_1$ exists if and only if $W^+(\theta)\ge\max\{a\theta,\ \bar{w}-(1-b)(1-\theta)\}$. Both $\theta\mapsto a\theta$ and $\theta\mapsto\bar{w}-(1-b)(1-\theta)$ are affine and $W^+$ is concave, so each constraint defines an interval and $\Theta_w$ is their intersection.
corollary[Relation to the linear bounds; exact sharpness for binary scores] (i) $\Theta_w$ is contained in the interval (ref). (ii) If $w(U)\in\{0,1\}$ almost surely, then $\Theta_w$ equals (ref): the linear bounds are exactly sharp for binary scores (in particular, for all event indicators and threshold classifiers). (iii) The interval (ref) is sharp in the class of observed laws with score mean $\bar{w}$: for every $\theta$ in (ref) there exists a law of $w(U)$ with mean $\bar{w}$ and a feasible decomposition supporting $\theta$. Hence (ref) cannot be improved using $\bar{w}$ alone.
proof(i) From $W^+(\theta)\le\bar{w}$ and $W^+(\theta)\ge a\theta$ we get $\theta\le\bar{w}/a$; from $W^+(\theta)\le\theta$ and $W^+(\theta)\ge\bar{w}-(1-b)(1-\theta)$ we get $b\theta\ge\bar{w}+b-1$. (ii) For binary $w$, $W^+(\theta)=\min\{\theta,\bar{w}\}$. If $\theta\le\bar{w}$ the first constraint in (ref) holds since $a\le1$, and the second reduces to $\theta\ge(\bar{w}+b-1)/b$; if $\theta>\bar{w}$ the second holds automatically and the first reduces to $\theta\le\bar{w}/a$. Since $(\bar{w}+b-1)/b\le\bar{w}\le\bar{w}/a$, the union of the two regimes is exactly (ref). (iii) Given $\bar{w}$, take $w(U)$ binary with $\mathbb{P}(w(U)=1)=\bar{w}$ and apply (ii).
remark[Interpretation] The gap between $\Theta_w$ and (ref) is an exact measure of the information in the shape of the score distribution beyond its mean. For binary scores the shape carries nothing extra; for graded scores (e.g.\ $w=S/M$ with $M$ large) the rearrangement constraint $W^+(\theta)\ge a\theta$ can bind strictly earlier than $\theta=\bar{w}/a$, tightening the upper bound. $W^+$ is computed by sorting, so the sharp set costs no more than the linear bounds in practice (Section (ref)). Emptiness of $\Theta_w$ is a specification rejection of the calibration constants $(a,b)$ at the observed law.

Calibrated events: validity and sharpness

Let $A\subseteq\mathcal{U}$ be a high-confidence positive event and $B\subseteq\mathcal{U}$ a high-confidence negative event.

proposition[Posterior event calibration; sharp] If $\mathbb{P}(X^*=1\mid U\in A)\ge\rho$ then $\theta\ge\rho\,\mathbb{P}(U\in A)$. If $\mathbb{P}(X^*=0\mid U\in B)\ge\lambda$ then $\theta\le 1-\lambda\,\mathbb{P}(U\in B)$. Each one-sided bound is sharp under the corresponding single event-calibration restriction.
proofThe lower bound follows from \[ \theta\ge\mathbb{P}(X^*=1,U\in A)=\mathbb{P}(X^*=1\mid U\in A)\mathbb{P}(U\in A)\ge\rho\mathbb{P}(U\in A), \] and the upper bound is analogous. For sharpness of the lower bound, under the single restriction involving $A$, set $h_1(u)=\rho p_U(u)$ for $u\in A$ and $h_1(u)=0$ otherwise, with $h_0=p_U-h_1$. Then $0\le h_1\le p_U$, the calibration constraint holds with equality whenever $\mathbb{P}(U\in A)>0$, and $\theta=\sum_u h_1(u)=\rho\mathbb{P}(U\in A)$. The upper bound is symmetric, assigning mass $h_0(u)=\lambda p_U(u)$ on $B$ and $h_0(u)=0$ off $B$.
proposition[Wrong-state event errors; sharp] If $\mathbb{P}(U\in A\mid X^*=0)\le\alpha_A$ with $\alpha_A\in[0,1)$, then \[ \theta\ \ge\ \max\Bigl\{0,\ \frac{\mathbb{P}(U\in A)-\alpha_A}{1-\alpha_A}\Bigr\}. \] If $\mathbb{P}(U\in B\mid X^*=1)\le\alpha_B$ with $\alpha_B\in[0,1)$, then \[ \theta\ \le\ \min\Bigl\{1,\ \frac{1-\mathbb{P}(U\in B)}{1-\alpha_B}\Bigr\}. \] Each one-sided bound is sharp under the corresponding single wrong-state error restriction.
proofLet $p_A=\mathbb{P}(U\in A)$. Since $p_A\le(1-\theta)\alpha_A+\theta$, rearrangement gives the lower bound. To see sharpness, if $p_A\le\alpha_A$, the bound is zero and is attained by $\theta=0$ with $q_0=p_U$. If $p_A>\alpha_A$, set \[ \theta_A=\frac{p_A-\alpha_A}{1-\alpha_A}. \] Allocate positive-state joint mass only inside $A$, proportionally to $p_U$ on $A$, with total mass $\theta_A$, and set $h_0=p_U-h_1$. Then $h_0(A)=p_A-\theta_A=\alpha_A(1-\theta_A)$, so $\mathbb{P}(U\in A\mid X^*=0)=\alpha_A$ and the lower bound is attained. The upper bound is symmetric: if $\mathbb{P}(U\in B)\le\alpha_B$, take $\theta=1$; otherwise set $\theta_B=(1-\mathbb{P}(U\in B))/(1-\alpha_B)$, allocate all mass outside $B$ to the positive state and allocate additional positive-state mass inside $B$ so that $\mathbb{P}(U\in B\mid X^*=1)=\alpha_B$.
remark[One source of identification] Propositions (ref)--(ref) are the maintained identifying content of the paper. Reporter-specific accuracy, threshold classifiers, relaxed support, unanimity, weighted votes, and matrix-agreement rules are not separate assumptions; they are obtained by choosing different $U$, $w$, $A$, and $B$. Proposition (ref) is the one-sided binary-score analogue of Proposition (ref): the lower bound uses $w=\mathbf{1}\{U\in A\}$ and the specificity-type constraint $\mathbb{E}[1-w(U)\mid X^*=0]\ge1-\alpha_A$, while the upper bound uses $w=\mathbf{1}\{U\in B\}$ and the false-negative constraint $\mathbb{E}[w(U)\mid X^*=1]\le\alpha_B$. The direct construction above gives sharpness.

Optimized calibrated events and multiplicity

Rich summaries admit many candidate high-agreement events. Rather than choosing one arbitrarily, the analyst can pre-specify a finite class and calibrate the whole class on a validation sample, treating event selection explicitly as a multiple-testing problem.

Let $\mathcal{A}^+$ be a finite class of positive events $A\subseteq\mathcal{U}$ (row-threshold events, column-pattern events, trusted-model events, weighted-score threshold events). Suppose validation data deliver simultaneous lower confidence bounds $\widehat\rho_A$ such that, with probability at least $1-\alpha$,

equation[equation omitted — 125 chars of source]

Then with the same probability all lower bounds hold simultaneously, so

equation[equation omitted — 110 chars of source]

and symmetrically $\theta\le\min_{B\in\mathcal{A}^-}\{1-\widehat\lambda_B\mathbb{P}(U\in B)\}$ for a simultaneously calibrated negative class $\mathcal{A}^-$.

remark[Pre-specification and honest calibration] The event class must be fixed before examining the main unlabeled sample, or the calibration step must account for selection. Simultaneity in (ref) can be obtained by Bonferroni corrections, split-sample validation, or conformal-style calibration. The same discipline applies to the calibration constants $(a_j,b_j)$ used anywhere in the paper: for honest inference they should be lower confidence bounds estimated on a sample (or sample split) disjoint from the one used to compute observed rates. This adds no structural assumption; it only ensures the validity statements survive the search over events.

Optional sensitivity restrictions

The paper's baseline identification comes from calibrated scores and calibrated events. Two additional restrictions are useful as sensitivity analyses, but they should not be presented as maintained assumptions unless independently justified. For a count or vote score $C\in\{0,\ldots,J\}$, directional asymmetry with parameter $\gamma>0$ imposes \[ \mathbb{E}[C\mid X^*=0]\le \gamma\,\mathbb{E}[J-C\mid X^*=1], \] so false-positive votes are bounded relative to false-negative votes. For a one-LLM count $S\in\{0,\ldots,M\}$ with observed mean $\bar s=\mathbb{E}[S]$, anchored separation with tolerance $\varepsilon>0$ imposes \[ \mathbb{E}[S\mid X^*=1]\ge \bar s+\varepsilon, \qquad \mathbb{E}[S\mid X^*=0]\le \bar s-\varepsilon, \] which implies $\theta\ge \varepsilon/(\varepsilon+M-\bar s)$ and $\theta\le \bar s/(\bar s+\varepsilon)$. These restrictions are informative because they impose cross-state separation, not because they use replication by itself.

table[table omitted — 1,670 chars of source]

Design I: one LLM, repeated binary questions

A single LLM is used repeatedly to measure the same latent binary truth. For one item, write the repeated reports as $R_1,\dots,R_M\in\{0,1\}$. If the repeated prompts are exchangeable probes of the same truth, their labels carry no structural content and the natural observable is the count \[ S=\sum_{m=1}^M R_m\in\{0,\dots,M\},\qquad p(s)=(1-\theta)q_0(s)+\theta q_1(s). \] No conditional independence is imposed across prompts: repeated prompts from one model may share the same systematic errors.

remark[When the count is appropriate] The count is appropriate only if the repetitions are exchangeable measurements of the same latent truth. If prompts ask substantively different questions, there is no single scalar $X^*$ behind all responses. If prompt identities have known reliability differences, keep the prompt-response vector rather than counting.

The count design has two natural calibrated scores: $w_1(S)=S/M$ and $w_k(S)=\mathbf{1}\{S\ge k\}$.

\paragraph{Average repeated-prompt accuracy.} With $w=S/M$ in Proposition (ref) and $\bar r=\mathbb{E}[S/M]$,

equation[equation omitted — 128 chars of source]

Because $S/M$ is a graded score, Theorem (ref) applies with content: the sharp set computed from $W^+$ can be strictly inside (ref), and is obtained by sorting the support of $S$.

\paragraph{Threshold accuracy.} For a threshold $k$, let $D_k=\mathbf{1}\{S\ge k\}$ and $p_k^+=\mathbb{P}(S\ge k)$. If validation data support $\mathbb{P}(D_k=1\mid X^*=1)\ge a_k$ and $\mathbb{P}(D_k=0\mid X^*=0)\ge b_k$, then \[ \max\Bigl\{0,\frac{p_k^++b_k-1}{b_k}\Bigr\}\le\theta\le\min\Bigl\{1,\frac{p_k^+}{a_k}\Bigr\}, \] and by Corollary (ref)(ii) these bounds are exactly sharp.

\paragraph{High- and low-count events.} With $A=\{S\ge k\}$ and $B=\{S\le\ell\}$, Proposition (ref) gives $\theta\ge\rho_k\mathbb{P}(S\ge k)$ and $\theta\le 1-\lambda_\ell\mathbb{P}(S\le\ell)$; Proposition (ref) gives the tail-error bounds \[ \theta\ \ge\ \max\Bigl\{0,\frac{\mathbb{P}(S\ge k)-\alpha_{0k}}{1-\alpha_{0k}}\Bigr\},\qquad \theta\ \le\ \min\Bigl\{1,\frac{1-\mathbb{P}(S\le\ell)}{1-\alpha_{1\ell}}\Bigr\}. \] Relaxed support is the special case $k=M$, $\ell=0$: $q_0(M)\le\alpha_0$, $q_1(0)\le\alpha_1$. Exact support ($\alpha_0=\alpha_1=0$) gives $p(M)\le\theta\le 1-p(0)$ but is rarely credible for LLMs; relaxed, validation-calibrated tail bounds are usually preferable.

\paragraph{Implication.} Design I is useful only to the extent that the count distribution can be calibrated. Report a small number of validation-calibrated score or event bounds---ideally the sharp set of Theorem (ref) for the graded score---rather than many sensitivity restrictions.

Design II: many named LLMs, one binary question each

$J$ named LLMs each answer the same binary question once. The response vector is $Y=(Y_1,\dots,Y_J)\in\{0,1\}^J$ with law $\pi(y)$ and class-conditionals $f_z(y)$ satisfying $\pi(y)=(1-\theta)f_0(y)+\theta f_1(y)$. No conditional independence is assumed across named LLMs. The vote count $\sum_j Y_j$ is a coarsening, not the primitive object.

The named-vector design matters because it permits calibrated scores and events that use model identity: the single-reporter score $w_j(Y)=Y_j$, the weighted score $w(Y)=\sum_j c_jY_j$ with $c_j\ge0$, $\sum_j c_j=1$, and events $A\subseteq\{0,1\}^J$ encoding agreement by a trusted subset rather than simple majority.

\paragraph{Reporter-specific accuracy.} Suppose validation data provide

equation[equation omitted — 135 chars of source]

and let $m_j=\mathbb{P}(Y_j=1)$. Applying Proposition (ref) to each $w_j(Y)=Y_j$ and intersecting,

equation[equation omitted — 175 chars of source]

Each one-reporter bound is sharp (binary score); the intersection is valid and is the sharp set based on the marginals alone. Using the joint law $\pi$ with the restrictions (ref) simultaneously can tighten further; this is the LP of Section (ref).

\paragraph{Named events.} If $A\subseteq\{0,1\}^J$ satisfies $\mathbb{P}(X^*=1\mid Y\in A)\ge\rho_A$ then $\theta\ge\rho_A\pi(A)$; if $B$ satisfies $\mathbb{P}(X^*=0\mid Y\in B)\ge\lambda_B$ then $\theta\le 1-\lambda_B\pi(B)$. Relaxed unanimity is only the special case $A=\{\mathbf 1_J\}$, $B=\{\mathbf 0_J\}$ and should not be the default event unless validation evidence supports it.

\paragraph{The exact cost of counting votes.} If only the vote count $S=\sum_jY_j$ is stored, the reporter-specific marginals and named events are not recoverable; the count identifies only $\bar m=\frac1J\mathbb{E}[S]=\frac1J\sum_j m_j$. With averaged calibration constants $\beta_1=\frac1J\sum_j a_j$, $\beta_0=\frac1J\sum_j b_j$, the count-only analogue is

equation[equation omitted — 154 chars of source]

which is generally strictly weaker than (ref); Section (ref) quantifies the gap in the toxicity application. The same loss applies to weighted scores and trusted-subset events.

\paragraph{Implication.} Analyze Design II with the named vector. Count-based analysis is a robustness check or a fallback when only counts were stored. The identifying content should come from reporter-specific or named-event calibration, not from generic independence assumptions.

Design III: many named LLMs, repeated questions each

The third design is a two-way measurement panel: for one item, observe $R=(R_{jm})\in\{0,1\}^{J\times M}$ with mixture law (ref). No conditional independence is assumed across rows, columns, or cells. Design III is valuable not because it makes independence credible, but because it records where agreement occurs.

Summaries and coarsenings

Useful lower-dimensional summaries: the model-count vector $T=(T_1,\dots,T_J)$, $T_j=\sum_m R_{jm}$; the prompt-count vector $S=(S_1,\dots,S_M)$, $S_m=\sum_j R_{jm}$; the total count $C=\sum_{j,m}R_{jm}$; and, when prompt labels are exchangeable but model labels are not, the column-pattern histogram

equation[equation omitted — 140 chars of source]

which counts the prompts on which the named response pattern equals $y$ (with $J=3$, $N_{110}$ counts prompts where models 1 and 2 say one and model 3 says zero). The histogram preserves cross-model agreement within prompts while discarding prompt labels, and refines the model-count vector via $T_j=\sum_y y_jN_y$. The coarsening hierarchy is \[ R\ \longrightarrow\ N\ \longrightarrow\ T\ \longrightarrow\ C,\qquad R\ \longrightarrow\ S\ \longrightarrow\ C, \] with support sizes $2^{JM}$, $\binom{M+2^J-1}{2^J-1}$, $(M+1)^J$, $(J+1)^M$, and $JM+1$ respectively. The full matrix distinguishes patterns with the same total count: one weak model saying yes on every prompt is not equivalent to several trusted models each saying yes repeatedly.

Truth-sufficient reductions of the response matrix

Which coarsenings lose information about $X^*$? The answer is a likelihood-ratio invariance condition.

Let $\Omega=\{0,1\}^{J\times M}$, let $U=h(R)$ be any finite summary with fibers $F_u=\{r:h(r)=u\}$, and let $q_z(u)=\sum_{r\in F_u}f_z(r)$.

proposition[Truth-sufficient coarsenings] Assume $0<\mathbb{P}(X^*=1)<1$. The following are equivalent, up to null sets. \begin{enumerate}[leftmargin=1.6em,itemsep=1pt] • $U$ is sufficient for the latent truth: $\mathbb{P}(X^*=1\mid R=r)=\mathbb{P}(X^*=1\mid U=h(r))$ for all $r$ with positive probability. • The residual matrix information given the summary is independent of the truth: $\mathbb{P}(R=r\mid U=u,X^*=1)=\mathbb{P}(R=r\mid U=u,X^*=0)$ for $r\in F_u$. • The likelihood ratio $L(r)=f_1(r)/f_0(r)$ is constant on each fiber $F_u$ (usual conventions for zeros). • There is a Markov kernel $K(r\mid u)$, independent of $z$, with $f_z(r)=q_z(h(r))\,K(r\mid h(r))$ for $z=0,1$. \end{enumerate} When these conditions hold, $R$ and $U$ contain the same information about $X^*$; when they fail, the coarsening discards information about the latent truth.
proof(2)$\Leftrightarrow$(4): take $K(r\mid u)=\mathbb{P}(R=r\mid U=u,X^*=z)$, which is independent of $z$ exactly when (2) holds. (2) is equivalent to $f_1(r)/q_1(u)=f_0(r)/q_0(u)$ on $F_u$, i.e.\ to $f_1(r)/f_0(r)=q_1(u)/q_0(u)$ constant on $F_u$, which is (3). By Bayes' rule, $\mathbb{P}(X^*=1\mid R=r)=\theta f_1(r)/\{(1-\theta)f_0(r)+\theta f_1(r)\}$ depends on $r$ only through $h(r)$ iff $f_1/f_0$ does, giving (1)$\Leftrightarrow$(3).
remark[Common matrix summaries] $C$ is truth-sufficient only if $f_1/f_0$ is constant across matrices with the same total count; $T$ only if constant across matrices with the same row-count vector; $N$ only if constant across matrices with the same pattern histogram. Prompt exchangeability of both $f_0$ and $f_1$ is sufficient for the last property but not necessary: likelihood-ratio invariance within prompt-permutation orbits is enough.

Prompt exchangeability and a lossless matrix reduction

Some reductions are not losses: they remove labels that have no design meaning. In Design III, prompt labels may be exchangeable even when model labels are not.

Let $S_M$ be the group of permutations of prompt labels; for $\sigma\in S_M$, $(\sigma r)_{jm}=r_{j,\sigma(m)}$. The histogram $N(R)$ indexes the orbits of this action: two matrices have the same $N$ iff one is obtained from the other by permuting prompt columns. The action extends to distributions by $(\sigma q)(r)=q(\sigma^{-1}r)$.

The reduction from $R$ to $N$ is lossless only when prompt exchangeability is part of the maintained design. To avoid ambiguity, let $\Theta_R^{\mathrm{ex}}(\pi_R;\mathcal{A}_R)$ denote the full-matrix identified set when admissible component distributions $q_0,q_1$ are required to be prompt-exchangeable in addition to satisfying the matrix-level restrictions $\mathcal{A}_R$. The required correspondence between matrix-level and histogram-level restrictions is explicit below.

assumption[Restriction compatibility] (i) The matrix-level restrictions $\mathcal{A}_R$ are invariant to prompt relabeling: if $(q_0,q_1)\in\mathcal{A}_R$ and the $q_z$ are prompt-exchangeable, then relabeling prompts does not change the truth of the restrictions. (ii) The histogram-level restriction set is exactly the projection of the exchangeable part of $\mathcal{A}_R$: \[ \mathcal{A}_N=\Bigl\{(q_0^N,q_1^N):\ (q_0,q_1)\in\mathcal{A}_R,\ q_z(\sigma r)=q_z(r)\ \forall\sigma\in S_M,\ q_z^N\ \text{is the law of }N\text{ under }q_z\Bigr\}. \]
theorem[Lossless reduction under prompt exchangeability] Suppose prompt exchangeability is a maintained design restriction, the observed population law $\pi_R$ is prompt-exchangeable, and Assumption (ref) holds. Then the full matrix and the column-pattern histogram give the same identified set: \[ \Theta_R^{\mathrm{ex}}(\pi_R;\mathcal{A}_R)\ =\ \Theta_N(p_N;\mathcal{A}_N). \]
proof($\subseteq$) Let $(\theta,q_0,q_1)$ be feasible for $\Theta_R^{\mathrm{ex}}(\pi_R;\mathcal{A}_R)$. Projecting each $q_z$ along $N$ gives $(q_0^N,q_1^N)$ satisfying the histogram mixture equation because projection commutes with mixing. Assumption (ref)(ii) gives $(q_0^N,q_1^N)\in\mathcal{A}_N$. ($\supseteq$) Let $(\theta,q_0^N,q_1^N)$ be feasible for $\Theta_N(p_N;\mathcal{A}_N)$. For each histogram value $n$, lift $q_z^N(n)$ uniformly over the finite orbit $\{r:N(r)=n\}$. The lifted $q_z$ are prompt-exchangeable and project back to $q_z^N$. Because $\pi_R$ is prompt-exchangeable, it is itself the uniform orbit lift of $p_N$, so the lifted distributions satisfy the matrix mixture equation for $\pi_R$. Assumption (ref)(ii) ensures that the lifted pair satisfies $\mathcal{A}_R$. Thus the same $\theta$ is feasible for $\Theta_R^{\mathrm{ex}}(\pi_R;\mathcal{A}_R)$.
remark[No independence is used] The theorem uses only a design symmetry and matched restrictions. It does not assume prompts are independent conditional on $X^*$, nor that cells are weakly correlated; entries of $R$ may be arbitrarily dependent within each latent state.
remark[Why compatibility matters] Both directions of Assumption (ref) have bite. If the analyst imposes additional restrictions after reducing to $N$ (so $\mathcal{A}_N$ is strictly smaller than the projection), then only $\Theta_N\subseteq\Theta_R^{\mathrm{ex}}$ is guaranteed. Conversely, a matrix-level restriction that is not prompt-invariant---for example, a calibrated accuracy bound for prompt $m=1$ only---cannot be expressed through $N$ at all; dropping it in the reduction gives only $\Theta_R^{\mathrm{ex}}\subseteq\Theta_N$. Equality is a statement about matched restriction classes, not about the statistic alone.
remark[General orbit reduction] The argument applies to any finite group of design symmetries: if $G$ acts on the matrix space, admissible $q_z$ are $G$-invariant, and the restrictions are $G$-invariant and matched as in Assumption (ref), the orbit statistic gives the same identified set as the full matrix. Prompt exchangeability gives $N$; independent prompt relabeling within each model gives $T$; full exchangeability of all cells gives $C$, but that last symmetry is rarely credible for named LLMs.

Matrix scores and matrix events

Design III should not be analyzed as a large vote count unless both dimensions are intentionally treated as exchangeable. The natural calibrated objects are matrix scores and matrix events.

A weighted matrix score has the form \[ w(R)=\frac{\sum_{j}\sum_{m}c_{jm}R_{jm}}{\sum_{j}\sum_{m}c_{jm}},\qquad c_{jm}\ge0,\ \ \sum_{j,m}c_{jm}>0, \] with special cases $w=T_j/M$, $w=S_m/J$, and $w=C/(JM)$; if (ref) holds for $w$, Proposition (ref) and Theorem (ref) apply. Under prompt exchangeability a score may be written as a function of $N$, e.g.\ placing weight on column patterns where trusted models agree.

Matrix events encode repeated agreement across both dimensions, e.g. \[ A^+(K,L)=\Bigl\{R:\ \sum_{j=1}^J\mathbf{1}\{T_j\ge L\}\ \ge K\Bigr\},\qquad A^-(K,L)=\Bigl\{R:\ \sum_{j=1}^J\mathbf{1}\{T_j\le L\}\ \ge K\Bigr\}, \] the events that at least $K$ named models each produce at least (at most) $L$ positive prompt responses. The histogram also suggests events invisible from $T$: with a trusted subset $J_0\subseteq\{1,\dots,J\}$, \[ \sum_{y:\,y_j=1\ \forall j\in J_0} N_y\ \ge\ L \] requires the trusted models to agree positively on at least $L$ prompts, regardless of the other models, while remaining invariant to prompt relabeling. Calibrated predictive values or wrong-state errors for any of these events feed Propositions (ref)--(ref); searches over event classes use Section (ref).

Model-specific and prompt-specific bounds; structured calibration

The model-count vector is a tractable compromise when $R$ or $N$ is too large. If for each named model $j$ \[ \mathbb{E}[T_j/M\mid X^*=1]\ge a_j,\qquad \mathbb{E}[(M-T_j)/M\mid X^*=0]\ge b_j, \] then with $r_j=\mathbb{E}[T_j/M]$, \[ \max_{1\le j\le J}\max\Bigl\{0,\frac{r_j+b_j-1}{b_j}\Bigr\}\ \le\ \theta\ \le\ \min_{1\le j\le J}\min\Bigl\{1,\frac{r_j}{a_j}\Bigr\}, \] and symmetrically for calibrated prompt-count scores $S_m/J$. Model-specific bounds suit stable, calibratable model error profiles; prompt-specific bounds suit stable prompt families. If prompt labels are exchangeable but within-prompt agreement matters, $N$ is preferable to $T$ because it retains more agreement geometry.

Full cell-level calibration of every pair $(j,m)$ may be too parameter-rich. A practical compromise groups models and prompts: with model groups $g(j)$ and prompt families $h(m)$, calibrate $\mathbb{P}(R_{jm}=1\mid X^*=1)\ge a_{g(j),h(m)}$ and $\mathbb{P}(R_{jm}=0\mid X^*=0)\ge b_{g(j),h(m)}$, with the grouping fixed before analysis or justified by validation data.

Diagnostics unique to the matrix design

Design III allows diagnostics that add no identifying assumptions.

\paragraph{Coarsening loss.} For $U\in\{R,N,T,S,C\}$ and compatible restrictions, report widths of identified sets and their differences, e.g. \[ \mathrm{Loss}(N\to T)=\mathrm{wid}(\Theta_T)-\mathrm{wid}(\Theta_N),\qquad \mathrm{Loss}(T\to C)=\mathrm{wid}(\Theta_C)-\mathrm{wid}(\Theta_T). \] A large loss from $N$ to $T$ means within-prompt cross-model agreement matters; from $T$ to $C$, that model identity matters. If prompt exchangeability is credible and restrictions are invariant and matched, Theorem (ref) predicts no loss from $R$ to $N$ at the population level---a checkable implication of the design symmetry.

\paragraph{Row and column influence.} Let $\Theta^{(-j)}$ and $\Theta^{(-m)}$ be the identified sets after dropping model $j$ or prompt $m$. Large movements in either bound show that conclusions hinge on a particular model or prompt; this is often more informative than another structural assumption.

\paragraph{Dependence and redundancy.} On a validation sample, estimate within-model dependence $\mathrm{Corr}(R_{jm},R_{jm'}\mid X^*=z)$ and across-model dependence $\mathrm{Corr}(R_{jm},R_{j'm}\mid X^*=z)$. High correlations indicate redundant measurements; low correlations indicate nonredundant variation. These are diagnostics, not independence assumptions.

Computation and inference

\paragraph{Linear programming.} For any finite summary $U$, fixed-$\theta$ feasibility is a linear program. With joint masses $h_z(u)=\mathbb{P}(X^*=z,U=u)$: \[ h_0(u)+h_1(u)=p_U(u),\qquad \sum_u h_1(u)=\theta,\qquad h_z\ge0, \] and conditional restrictions become linear after multiplying through, e.g.\ $\mathbb{E}[w(U)\mid X^*=1]\ge a$ becomes $\sum_u w(u)h_1(u)\ge a\sum_u h_1(u)$. Sharp bounds under the baseline linear score and event restrictions are two LPs: minimize and maximize $\sum_u h_1(u)$. For fixed $\theta$, FOSD restrictions are linear in the conditional component probabilities and can be included in fixed-$\theta$ feasibility checks; exact MLR restrictions are bilinear in the components and should be handled only through explicit relaxations or nonlinear optimization. Prompt exchangeability is implemented either by using $N$ directly or by orbit-equality constraints on matrix probabilities.

\paragraph{The sharp score set by sorting.} $W^+(t)$ in (ref) is computed greedily: order support points by descending $w(u)$ and fill mass to total $t$. $W^+$ is piecewise linear and concave, so $\Theta_w$ in (ref) is found by intersecting a concave piecewise-linear function with two affine functions---no LP solver needed.

\paragraph{Sampling uncertainty.} When $p_U$ is estimated from $N$ items, replace mixture equalities by bands $|(1-\theta)q_0(u)+\theta q_1(u)-\widehat p_U(u)|\le\Delta_u$. A simple finite-sample choice is the Hoeffding-union bound

equation[equation omitted — 105 chars of source]

with $\Delta_u=\varepsilon_N(\alpha)$ for every $u$, yielding a conservative outer confidence set; less conservative alternatives use the multinomial bootstrap, empirical Bernstein bands, or moment-inequality methods. For optimized event bounds, sampling uncertainty enters twice---through $\widehat p_U$ in the main sample and through validation estimates of predictive values---and validity requires the simultaneous calibration (ref).

\paragraph{Specification testing by minimum tolerance.} Let $\mathcal{Q}$ be the set of distributions over $\mathcal{U}$ implied by the maintained restrictions for some $\theta$, and define \[ \Delta^*=\min_{q\in\mathcal{Q}}\|\widehat p_U-q\|_\infty, \] the minimum $\ell_\infty$ tolerance restoring feasibility. If the restrictions are correct at the population law $p_U^0$, then $\Delta^*\le\|\widehat p_U-p_U^0\|_\infty$, so a finite-sample test rejects at level $\alpha$ if $\Delta^*>\varepsilon_N(\alpha)$, and \[ \Delta_0\in\bigl[\max\{0,\Delta^*-\varepsilon_N(\alpha)\},\ \Delta^*+\varepsilon_N(\alpha)\bigr] \] is a confidence interval for the population misspecification distance. The test detects out-of-distribution collapse: if an LLM panel becomes uninformative, restrictions calibrated in distribution may become infeasible. Emptiness of the sharp score set $\Theta_w$ is the same idea specialized to a single calibrated score.

Simulation: count-based Beta-Binomial experiment

Set $M=5$ and $\theta_0=0.65$, with class-conditional counts $S\mid X^*=1\sim\mathrm{BetaBin}(5,4,1.5)$ and $S\mid X^*=0\sim\mathrm{BetaBin}(5,1.5,4)$; the observed histogram is $p=(1-\theta_0)q_0+\theta_0q_1$. The design has $\mathbb{E}[S\mid X^*=0]=1.36$, $\mathbb{E}[S\mid X^*=1]=3.64$, average repeated-prompt accuracy $0.727$, and observed mean $\mathbb{E}[S]=2.84$, so $\bar{w}=\mathbb{E}[S/M]=0.568$ for the graded score $w=S/M$. Without restrictions, or under FOSD/MLR/weak mean ordering alone, the identified set is $[0,1]$ because the degenerate decomposition $q_0=q_1=p$ remains feasible (Proposition (ref)); Table (ref) records this once as a benchmark and then isolates the paper's main computational comparison: the mean-only linear bounds (ref) versus the sharp rearrangement set of Theorem (ref), computed from $W^+$ by sorting the six support points of $S$.

table[table omitted — 391 chars of source]
table[table omitted — 1,416 chars of source]

The table mirrors the paper's identification message in miniature. Weak ordering restrictions leave $[0,1]$ untouched, so they are recorded in a single benchmark row. Calibration is what moves the set, and the shape of the score distribution carries identifying content beyond its mean, exactly as Theorem (ref) predicts: at $a=b=0.60$ the rearrangement constraint $W^+(\theta)\ge \bar{w}-(1-b)(1-\theta)$ binds before the linear lower bound, raising $\theta_L$ from $0.280$ to $0.317$; at $a=b=0.70$ the sharp set $[0.479,0.784]$ removes $29\%$ of the linear interval's width ($0.429$ to $0.305$), tightening both endpoints. Because $w(S)$ is a graded score on six support points, this gain is obtained by a sort and costs nothing computationally. Exact support tightens only the weakly calibrated case ($a=b=0.60$, where it caps $\theta_U$ at $0.882$) and is redundant once calibration is strong; this ordering---calibration first, support as a benchmark---is the recommended reporting style of Section (ref).

Empirical illustration: toxicity classification

This section revisits the toxicity-response classification task studied by ChengEtAl2024SoftLabel. The task is to bound the prevalence of toxic responses when the analyst observes binary labels from LLM annotators and, for a validation sample, expert-adjudicated truth (which gives us a benchmark to allow us to validate our bounds).

In particular, the data involve $J=3$ named LLM annotators---GPT-4, GPT-4 Turbo, and Claude-2---and three human annotators. This is a Design II setting. Following the taxonomy of this paper, we now report the named-vector analysis in three increasingly disciplined steps---plug-in marginal bounds, the full named-vector LP on the joint law of $Y=(Y_1,Y_2,Y_3)$, and an honest version that replaces plug-in calibration constants by one-sided lower confidence bounds---and we tag every bound with the response object that must be stored to compute it.

table[table omitted — 591 chars of source]

Table (ref) shows why named-vector analysis matters: GPT-4 is much more accurate than Claude-2, and all three LLMs have higher specificity than sensitivity. A count-only analysis collapses this heterogeneity into an exchangeable average. The validation class-conditionals also show that exact support is empirically violated with only three LLMs: among truly toxic items, 14.2% receive zero positive LLM votes; among truly non-toxic items, 4.3% receive unanimous positive votes. Exact support can be reported as a benchmark, but relaxed support or validation-calibrated score bounds are more credible.

Table (ref) records the observable objects that the bounds below are built from. The top panel gives the count histograms---$\widehat p(s)$ with the validation class-conditionals $\widehat q_z(s)$ for $N=1{,}000$, and the count-only training histogram for $N=28{,}194$---while the lower panel reports the full named-vector law $\widehat\pi(y)$ over $\{0,1\}^3$. The two are not interchangeable: the count is the coarsening $s=\sum_j y_j$ of $\widehat\pi$, so patterns that disagree on which model flagged the item are merged. Here $(1,1,0)$, where GPT-4 and GPT-4 Turbo flag and Claude-2 abstains, has mass $0.250$, whereas the equal-count pattern $(0,1,1)$ has mass only $0.010$; counting pools them into $\widehat p(2)=0.295$ and discards exactly the reporter identity that the named-vector bounds will exploit.

table[table omitted — 1,424 chars of source]

Named-vector bounds: plug-in benchmark and honest calibration

Table (ref) reports reporter-specific bounds (ref) in two versions. Panel A is the plug-in benchmark: calibration constants $(a_j,b_j)$ are the validation sensitivities and specificities of Table (ref), estimated on the same sample as the observed rates $m_j$, so the panel illustrates the identification logic but is not honest inference. Panel B is the honest calibration version recommended in Section (ref): $(a_j,b_j)$ are replaced by one-sided Clopper--Pearson lower confidence bounds $(a_j^L,b_j^L)$ at Bonferroni level $\alpha/6$ with $\alpha=0.05$, so all six constraints hold simultaneously with probability at least $0.95$ and the resulting bounds are valid by Proposition (ref).

table[table omitted — 1,591 chars of source]

The key comparison: what storing richer objects buys

Table (ref) is the section's central exhibit. It reports four identified sets for the same population, ordered by the richness of the stored response object, including the full named-vector LP of Section (ref), which imposes the mixture equation on the joint law $\widehat\pi(y)$ over $\{0,1\}^3$ together with all reporter-specific calibration restrictions simultaneously.

table[table omitted — 1,101 chars of source]

Three facts stand out. First, the cost of counting: moving from the named vector to the count widens the set from $[0.497,0.707]$ to $[0.348,0.731]$---the width nearly doubles, from $0.209$ to $0.383$---purely because counting discards reporter identity, quantifying Theorem (ref). Second, the full LP on the joint law coincides with the marginal intersection to four decimals. This is informative rather than disappointing: it certifies that, given reporter-specific calibration alone, the marginal intersection is already sharp at this observed law, so no further information can be extracted from these restrictions; the joint law is nevertheless the object to store, because only it makes the sharpness check possible and because pattern-level calibrated events (trusted-subset agreement, Section (ref)) require it whenever validation evidence supports them. Third, honest calibration costs about $0.07$ of width relative to the plug-in benchmark but preserves the storage ranking: the honest named-vector set (width $0.277$) remains far narrower than the honest count set (width $0.480$).

The binding reporter on both sides is GPT-4, the most accurate model; the influence diagnostic of Section (ref) applies verbatim: dropping GPT-4 widens the plug-in intersection to the Turbo bounds $[0.391,0.730]$. For the training set ($N=28{,}194$), only the count histogram was stored, so every named-vector row of Table (ref) is unavailable by construction: with $\bar m=1.847/3=0.616$ and the exchangeable-average constants, (ref) gives $[0.554,\,1.000]$. The storage choice has real consequences: on the training set the count bound is $[0.554,1.000][0.554, 1.000] [0.554,1.000]$, whose upper endpoint is the trivial value 1, so counting leaves the prevalence bounded only from below. Had the named vector been stored, both endpoints would be informative.

Extension: regression on a latent label

So far the target has been the scalar prevalence $\theta=\mathbb{P}(X^*=1)$. In many applications $X^*$ is not the final object but a latent regressor: the analyst wants the partial association between an observed outcome and the latent label, holding covariates fixed. This section shows that the calibration bounds feed directly into such a regression, with prevalence recovered as the special case of a constant outcome.

For each item let $(V,Z,U)$ be observed, where $V\in\mathbb{R}$ is a scalar outcome, $Z\in\mathbb{R}^{d}$ a vector of observed covariates, and $U=g(R)$ the LLM summary of Sections (ref)--(ref); $X^*\in\{0,1\}$ remains latent. (We write the outcome as $V$ to avoid collision with the named-LLM vector $Y$ of Section (ref).) Stack the regressors as $W=(1,Z',X^*)'\in\mathbb{R}^{d+2}$ and consider the best linear predictor (BLP) coefficient \[ \beta \;=\; \big(\mathbb{E}[WW']\big)^{-1}\mathbb{E}[WV], \] assuming $\mathbb{E}[WW']$ is nonsingular. The coefficient on $X^*$, denoted $\gamma$, is typically the parameter of interest; if $\mathbb{E}[V\mid Z,X^*]$ is linear, $\beta$ is the conditional-expectation coefficient.

\paragraph{Reduction to latent cross-moments.} Because $(V,Z)$ are observed and $X^*$ is binary (so $(X^*)^2=X^*$), every entry of $\mathbb{E}[WW']$ and $\mathbb{E}[WV]$ is identified from the data except the blocks that pair $X^*$ with an observed variable: \[ \theta=\mathbb{E}[X^*],\qquad \mathbb{E}[X^*Z]\in\mathbb{R}^{d},\qquad \mathbb{E}[X^*V]\in\mathbb{R}. \] Collect these in $\mu=(\theta,\mathbb{E}[X^*Z]',\mathbb{E}[X^*V])'$. Then $\beta=\Phi(\mu;m_{\mathrm{obs}})$ for a known map $\Phi$ that is rational in $\mu$---a matrix inverse times a vector, both affine in $\mu$---where $m_{\mathrm{obs}}$ collects the observed moments. Identifying $\beta$ thus reduces to identifying $\mu$.

\paragraph{Bounding the latent cross-moments.} Since $(U,Z,V)$ are jointly observed, write the unknown attachment of the latent label to the data as \[ \eta(u,z,v)=\mathbb{P}(X^*=1\mid U=u,Z=z,V=v)\in[0,1]. \] Every latent cross-moment is a linear functional of $\eta$: for any observed $h$, \[ \mathbb{E}[X^*\,h(U,Z,V)]=\mathbb{E}[h(U,Z,V)\,\eta(U,Z,V)]. \]

lemma[Calibrated bounds on latent moments] Let $h$ be any observed, bounded function. Over all $\eta$ consistent with the observed law of $(U,Z,V)$ and with the maintained calibration restrictions of Section (ref), the cross-moment $\mathbb{E}[X^*h]$ ranges over a closed interval whose endpoints solve the linear programs \[ \min_{\eta}\ /\ \max_{\eta}\ \mathbb{E}[h(U,Z,V)\,\eta(U,Z,V)] \quad\text{s.t.}\quad 0\le\eta\le1,\ \text{marginal consistency},\ \text{calibration}. \] Taking $h\equiv1$ returns the prevalence bounds of Section (ref) exactly; taking $h\in\{Z_1,\dots,Z_d,V\}$ bounds the remaining components of $\mu$.

The calibration inequalities enter exactly as in Section (ref): a score restriction $\mathbb{E}[w(U)\mid X^*=1]\ge a$ becomes $\sum_u w(u)\,\mathbb{E}[\eta\mid U=u]\,p_U(u)\ge a\,\mathbb{E}[\eta]$ after writing $\mathbb{E}[X^*w(U)]=\mathbb{E}[w(U)\eta]$, and an event restriction becomes the corresponding linear inequality; both are linear in $\eta$. The lemma is therefore the paper's fixed-$\theta$ LP with the objective $\mathbb{E}[\eta]$ replaced by $\mathbb{E}[h\,\eta]$.

proposition[Identified set for the regression coefficient] Let $\mathcal M\subseteq\mathbb{R}^{d+2}$ be the set of $\mu$ attainable by some feasible $\eta$; $\mathcal M$ is convex and compact, and is a polytope under score/event (linear) calibration. The sharp identified set for the BLP coefficient is the image \[ B=\{\Phi(\mu;m_{\mathrm{obs}}):\mu\in\mathcal M\}. \] Each coordinate of $B$ is computed by the fixed-value method of Section (ref). By Frisch--Waugh--Lovell, $\gamma=\mathbb{E}[\tilde V\,\eta]\,/\,D(\theta,\mathbb{E}[X^*Z])$, where $\tilde V$ is the residual of $V$ on $(1,Z)$ (observed) and $D=\mathbb{E}[\tilde X^{*2}]\ge0$ depends only on the first-stage moments $(\theta,\mathbb{E}[X^*Z])$. Fixing those first-stage moments at any value in their calibrated region makes $\gamma=c$ linear in $\eta$, so $B$ is traced by a parametric family of LPs indexed by that low-dimensional region.

Replacing $\mathcal M$ by the box of component-wise intervals from Lemma (ref) gives valid but generally conservative outer bounds; the joint program is sharp. This is the same distinction drawn for the named-vector LP versus the marginal intersection in Section (ref).

remark[Conditional versus marginal calibration] The bounds are valid under the calibration of Section (ref), which constrains $\eta$ only through its $U$-marginal and so permits $\eta$ to vary freely with $(Z,V)$ within report cells. If instead calibration is validated within covariate strata---$\mathbb{E}[w(U)\mid X^*=1,Z=z]\ge a(z)$, the natural design when validation data carry $Z$---then within-stratum prevalences and outcome cross-moments are bounded separately and the coefficient bounds tighten accordingly. Conditional calibration is to this section what reporter-specific calibration was to Design II: it is where the covariates earn their keep.
remark[Relation to misclassified-regressor bounds] With $U$ a single noisy report, this is the partial-identification problem for a regression with a misclassified binary regressor Bollinger1996,Mahajan2006,Hu2008,Molinari2008. Our contribution is not a new bound for that problem but the observation that the same externally calibrated scores and events that bound prevalence also bound the regression coefficient, through the single channel of the latent cross-moments $\mu$; no independence or instrument is invoked.

\paragraph{A transparent special case.} With no covariates and target the mean contrast $\gamma=\mathbb{E}[V\mid X^*=1]-\mathbb{E}[V\mid X^*=0]$, \[ \gamma=\frac{\mathbb{E}[VX^*]-\theta\,\mathbb{E}[V]}{\theta(1-\theta)}, \] so $\gamma$ is pinned down once $\theta$ and $\mathbb{E}[VX^*]$ are, each bounded by Lemma (ref); the identified set is obtained by ranging the numerator and $\theta$ jointly over their calibrated region. If, in addition, $V$ is itself a calibrated high-confidence label, the contrast inherits the sharpness of Section (ref).

\paragraph{Inference.} Sampling uncertainty enters through the observed law of $(U,Z,V)$ and, for honest calibration, through the validation estimates of the calibration constants. Propagating the bands (ref) of Section (ref) (or a multiplier bootstrap) through the LPs of Lemma (ref) yields a valid outer confidence set for each coordinate of $\beta$; moment-inequality methods apply verbatim because all restrictions are linear in $\eta$.

Practical recommendations

\paragraph{Use calibration as the main identifying content.} The maintained assumptions should be calibrated score and event restrictions. Report the source of calibration, preferably lower confidence bounds from validation data on a split disjoint from the main sample.

\paragraph{Treat weak ordering as regularity, not identification.} FOSD, MLR, and weak mean ordering are plausible but do not rule out $q_0=q_1$ and should not be presented as identifying assumptions.

\paragraph{Report sharp sets where they are cheap.} For binary scores and events, the linear bounds are already sharp. For graded scores, report the sharp set of Theorem (ref); it costs a sort, and the gap from the linear bounds measures the information in the score's shape.

\paragraph{Exploit design symmetries, not independence assumptions.} If prompt labels are exchangeable by design, reduce the matrix to the column-pattern histogram $N$ under matched, invariant restrictions (Assumption (ref)) rather than imposing conditional independence.

\paragraph{Store the richest object possible.} For Design I, store all repeated responses even if the count is analyzed. For Design II, store the named vector. For Design III, store the full matrix whenever feasible; under prompt exchangeability, store or construct $N$; at minimum store the model-count vector. The training-set panel of Section (ref) shows what is lost otherwise.

\paragraph{Report coarsening and influence diagnostics.} Compare bounds under richer and coarser summaries; report whether results depend on particular models or prompts; quantify losses along $R\to N\to T\to C$ where relevant.

\paragraph{Keep sensitivity assumptions separate.} Directional asymmetry, anchored separation, and independence-style restrictions can be useful, but label them as sensitivity analysis unless independently justified.

Conclusion

The identifying content of LLM measurement panels depends on the source of replication. Asking one LLM repeated exchangeable questions creates a count experiment; asking several named LLMs once creates a named-reporter experiment; asking several named LLMs repeated questions creates a two-way response matrix. These designs should not be collapsed into a single vote-count model at the outset.

The general theory is simple, and we present it as an adaptation of known mixture and misclassification logic to a setting where dependence among reporters is the rule. Every finite summary yields a two-component mixture for the latent truth; without restrictions, prevalence is completely unidentified, and weak ordering restrictions do not help because they permit equality of the latent components. Useful identification comes from calibrated scores and calibrated events, of which reporter-specific accuracy, threshold rules, relaxed support, unanimity, weighted ensembles, and matrix-agreement events are special cases. The sharpness analysis clarifies exactly what each calibrated object delivers: linear bounds that are sharp for binary scores and for any analysis recording only the score mean, and a rearrangement characterization of the strictly sharper set available from the full score law.

The matrix design adds one further lesson. Even when entries of the matrix are arbitrarily correlated, design symmetries justify lossless reductions: under prompt exchangeability and matched invariant restrictions, the column-pattern histogram preserves all identifying information in the full matrix. Calibration identifies; storage and symmetry determine what can be calibrated without loss.