EconBase
← Back to paper

An Adversarial Approach to Identification, Computation, and Inference in Models with a Linear-in-Measures Representation

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

102,749 characters

An Adversarial Approach to Identification, Computation, and Inference in Models with a Linear-in-Measures Representation


\title{
  An Adversarial Approach to Identification, Computation, and Inference in Models with a Linear-in-Measures Representation
  \thanks{Botosaru: Department of Economics, McMaster University, \texttt{[email removed]}.
  Loh: Department of Economics, University of North Carolina Wilmington, \texttt{[email removed]}.
  Muris: Department of Economics, McMaster University, \texttt{[email removed]}.
  We thank Andr\'es Aradillas-L\'opez, Tim Christensen, Inga Deimen, Jiaying Gu, Bo Honor\'e, Hide Ichimura, Hiro Kasahara, Vadim Marmer, Francesca Molinari, Adam Rosen, Rami Tabri, Alex Torgovitsky, and Victoria Zinde-Walsh for discussions and suggestions.
  We also thank audiences at various seminars and conferences for questions and comments.
  Botosaru gratefully acknowledges financial support from the Canada Research Chairs Program.
  Muris gratefully acknowledges financial support from the Social Sciences and Humanities Research Council of Canada (Insight Grant 435-2025-1380).
  }
}
\author{Irene Botosaru \and Isaac Loh \and Chris Muris}
\date{September 25, 2026}
\maketitle

\begin{abstract}
We develop a framework for identification, computation, and inference in econometric models with a \emph{linear-in-measures representation}.
These models express maintained restrictions as moment conditions linear in the joint probability measure of observed and latent inputs, and map that measure linearly to the distribution of outputs, even with nonlinear outcome equations.
We construct an \emph{adversarial discrepancy function} whose zeros characterize the identified set for structural and counterfactual parameters.
With finite output support, finite linear programs compute the discrepancy function or provide certified bounds even when latent inputs have infinite support, and a penalized bootstrap yields confidence sets with uniform per-point coverage.
We apply the framework to two open cases in binary choice panels with fixed effects and discrete covariates: sequential exogeneity with unspecified conditional marginal error distributions, and known conditional marginal error distributions with unrestricted serial dependence.
In an entry game with multiple equilibria, the framework recovers the known sharp identification region.
\end{abstract}

\smallskip
\noindent\textbf{Keywords:} identification; latent-variable models; nonlinear panel models; sequential exogeneity; linear programming; uniform inference.\par
\noindent\textbf{JEL:} C12, C14, C23, C61.

\section{Introduction}
\label{sec:intro}

\subsection{Motivation and approach}

We develop a new approach to identification, computation, and inference for a broad class of econometric models.
A model in this class specifies two objects: (i) moment conditions that restrict the joint probability measure of its \emph{inputs}, the set of observed and latent variables from which it generates its \emph{outputs};
\footnote{``Input'' and ``output'' refer to roles in the representation.
Observed inputs may be covariates, instruments, initial conditions, or the endpoints of an interval-censored covariate.
These variables are also outputs, which the model returns unchanged.
Outcomes are outputs generated from the observed and latent inputs.}
and (ii) the conditional distribution of the outputs given the inputs.
Averaging this conditional distribution over the \emph{input measure} yields the output distribution.
Thus, at each parameter value, the model has a \emph{linear-in-measures representation} (LIMR) because both the moment restrictions and the mapping to the output distribution are linear in the input measure.
A parameter value belongs to the identified set if some input measure satisfies the moment restrictions at that value and reproduces the observed output distribution.

Using the LIMR and a separation argument in the space of distributions, we construct the \emph{adversarial discrepancy function} (ADF), a criterion function whose zeros characterize the identified set.
This places the adversarial approach in the criterion function tradition of \textcite{manskiInferenceRegressionsInterval2002} and \textcite{Chernozhukovetal2007}.
Here the ADF has a zero-sum-game form: a test function seeks to separate the observed output distribution from those generated by the admissible input measures, while the model seeks an input measure that minimizes the resulting discrepancy.
Identification requires neither an explicit characterization of the set of model-implied output distributions nor a separate characterization of its observable implications.

When the outputs take finitely many values, even when the latent inputs have infinite support, finite linear programs (LPs) either compute the ADF exactly or yield certified bounds.
The bounds are valid at every certified iterate, so no \textit{a priori} convergence guarantee is required.
The test statistic is the sample analog of the ADF, computed by the same LPs.
Inference is by test inversion, with critical values from a penalized bootstrap, and the resulting confidence sets have uniform per-point coverage.

Within the LIMR class, specifications differ in the choice of inputs, the restrictions on their joint probability measure, and the conditional distribution of the outputs given the inputs.
The ADF, and hence the LPs and the bootstrap procedure, is constructed from these objects.
Thus, changes in the outcome equation or maintained restrictions alter the ingredients of the same construction rather than requiring a new identification, computation, or inference method.

We demonstrate the scope of the LIMR in short nonlinear panel models with fixed effects.
We address open questions in binary choice panel models under sequential exogeneity without parametric restrictions on the error distributions, and under parametrically specified marginal error distributions with unrestricted serial dependence.
We analyze additional specifications that allow unrestricted dependence between the fixed effect and the error terms.
We also find that the framework recovers the point identification benchmark for the panel logit model with fixed effects.
The LIMR class also includes models with interval-censored covariates and dynamic discrete choice, and extends beyond panels to entry games with multiple equilibria and randomized trials with imperfect compliance.
Section~\ref{sec:examples} constructs a LIMR for each model mentioned above.

We call the map from input measures to output distributions the \emph{output operator}.
Linearity in the input measure restricts neither the outcome equation nor the moment functions, which may be nonlinear or nonseparable.
The moment conditions may be unconditional or conditional on any subset of the inputs and may form a continuum, as when a restriction holds at every value of a continuously distributed fixed effect.
They include zero-mean and median restrictions, parametric distributional restrictions, stationarity restrictions, and definitions of counterfactual parameters such as the average structural function (ASF).
Adding an equilibrium selector accommodates incomplete entry games, while a change of inputs yields a LIMR under sequential exogeneity in a nonlinear panel model with fixed effects.

In nonlinear panel models with fixed effects and few time periods, repeated observations do not in general reveal the individual fixed effect.
Hence, its unrestricted distribution given the covariates is an infinite-dimensional latent object.
The fixed effect also enters counterfactual parameters that average over its distribution, such as the ASF.
Under special parametric structure, the fixed effect can be eliminated through a sufficient statistic or a functional differencing transformation for the identification of some structural coefficients.
More generally, restrictions on the errors are often stated conditional on the fixed effect and must therefore hold at every value of its unknown support.
Structural and counterfactual parameters may consequently be only partially identified.

A substantial literature has developed powerful identification methods for important specifications of these models, both for point identification \parencite[e.g.,][]{cham1980,Manski1987,honorePanelDataDiscrete2000,Bonhomme2012} and for partial identification of structural coefficients and counterfactual parameters \parencite[e.g.,][]{honoreBoundsParametersPanel2006,ChernValHahnNewey2013,daveziesIdentificationEstimationAverage2022,botosaruMuris2025}.
These results exploit various restrictions on the errors and use constructions tailored to the outcome equation and the parameter of interest.
For example, stationarity restrictions on error distributions yield model-specific inequalities that depend only on observables \parencite[e.g.,][]{khanIdentificationDynamicBinary2023,gaoIdentificationNonlinearDynamic2024,pakesPorterMomentInequalities,mbakopIdentificationSomeDiscrete2023}, while with parametric errors functional differencing yields moment equalities free of the fixed effect \parencite{Bonhomme2012,honoreWeidnerMomentConditions2025}, which need not exhaust the observable implications \parencite{dobronyiGuKim2021}.

Our two binary choice applications depart from familiar benchmark restrictions in different ways.
Fully parametric specifications obtain a likelihood by specifying the joint error distribution \parencite{cham1980,Chamberlain2010}, whereas semiparametric approaches exploit restrictions such as equality of error distributions across periods \parencite{Manski1987,ChernValHahnNewey2013}.
Under sequential exogeneity, feedback from past outcomes to future covariates makes the conditioning information in the restrictions on the errors evolve over time.
This has long been a difficult case in nonlinear panel models \parencite{arellanoPanelDataModels2001,chamberlainFeedbackPanelData2022,bonhommeDanoGraham2025}.
With finitely supported covariates, we leave the conditional marginal error distributions unspecified and characterize the identified set of the coefficient and the ASF.
\footnote{Existing identification results for the coefficient rely on a parametric error distribution \parencite{arellanoBinaryChoicePanel2003,piginiConditionalInferenceBinary2022,bonhommeDanoGraham2023} or a special regressor \parencite{honoreSemiparametricBinaryChoice2002}, while \textcite{ChernValHahnNewey2013} derive bounds for the ASF.}
Our second specification fixes the conditional marginal distribution of each error term given the fixed effect and covariates but leaves serial dependence unrestricted.
It therefore lies between the two benchmarks: the marginal distributions are specified, but the joint error distribution, and hence the likelihood, is not.
\footnote{Specified serial dependence can instead be incorporated into a likelihood \parencite{heckmanStatisticalModelsDiscrete1981,Hyslop1999}.}
In both specifications, the LIMR retains the fixed effect among the inputs, so restrictions stated conditional on it and counterfactual parameters defined through its distribution can be imposed directly.

Random set, optimal transport, and entropic latent-variable methods also provide general identification procedures for partially identified models.
Section~\ref{sec:related_literature} compares those approaches with the adversarial approach.

\subsection{Contributions}
\label{sec:contributions}

Our first contribution is an adversarial characterization of the identified set.
Fix $\theta$ and let $\Gamma_\theta$ denote the set of admissible input measures.
Let $\Lgen$ denote the output operator, so that $\Lgen\gamma$ is the output measure, the output distribution generated by $\gamma\in\Gamma_\theta$.
Let $\mu^*$ denote the true output measure.
Under the conditions of Theorem~\ref{thm:main_sharpness}, we construct the ADF below
\begin{equation}
T(\theta)
\equiv
\sup_{\substack{\phi\ \mathrm{measurable}\\ 0\leq\phi\leq 1}} \;
\inf_{\gamma\in\Gamma_\theta} \;
\Big(
\EE{\mu^*}{\phi}
-
\EE{\Lgen\gamma}{\phi}
\Big),
\label{eq:discrepancy_function}
\end{equation}
whose zeros characterize the identified set.
The outer problem searches over test functions $\phi$ and the inner problem chooses the admissible input measure that minimizes the resulting discrepancy.
\footnote{\textcite{Kaji2023} use a generator--discriminator minimax construction for simulation-based estimation of parametric structural models.
The adversary here instead ranges over test functions that separate $\mu^*$ from the set of output measures generated by the model.}
Thus $T(\theta)>0$ whenever some $\phi$ separates $\mu^*$ from every output measure at $\theta$.
By the separation argument in Section~\ref{sec:identification}, $T(\theta)$ equals the total variation (TV) distance from $\mu^*$ to the TV closure of the set of output measures.\footnote{Hence $T(\theta)=0$ means that the model can approximate $\mu^*$ arbitrarily closely in TV. When the output measure set is TV closed, exact and approximate compatibility coincide. The set of output measures is convex by the LIMR. See Section~\ref{sec:discussion_discrepancy} for discussion.}
The max--min order in \eqref{eq:discrepancy_function} is intentional: by the LIMR, the inner problem is linear in $\gamma$, while the outer problem is linear in $\phi$.
This linear structure yields the LP formulations used for computation and inference.

Our second contribution is computation: when the output space is finite, an auxiliary finite linear program computes $T(\theta)$ exactly under two checkable conditions and, under one of them together with a row bound, encloses it.
The auxiliary program keeps finitely many input values as rows and finitely many moment restrictions as columns.
Omitting input values weakly raises the auxiliary value, whereas omitting moment restrictions weakly lowers it, so the auxiliary value need not bound $T(\theta)$ in either direction.
\emph{Column certification} verifies that omitted moment restrictions leave the auxiliary value unchanged, making it an upper bound on $T(\theta)$.
\emph{Row certification} verifies that no omitted input value lowers the auxiliary value.
When both hold, the auxiliary value equals $T(\theta)$ (Theorem~\ref{thm:exact_computation}).
Column certification requires no optimization in any of our examples, whereas row certification requires a global optimization over the input space.
Deciding whether $\theta$ belongs to the identified set needs less than equality.
A \emph{row bound}, a certified bound on how much omitted input values can lower the auxiliary value, gives a lower bound, and with column certification it encloses $T(\theta)$ (Theorem~\ref{thm:enclosure}).
An upper bound of zero settles inclusion, and a positive lower bound settles exclusion.
If neither bound decides, column-and-row generation adds as a new row an input value that lowers the auxiliary value, constructs a column-certified set of moment restrictions for the enlarged support, and solves the program again.
We make no general convergence claim, and a decision does not need one: every column-certified iterate carries a valid enclosure.


Our third contribution is inference.
Replacing $\mu^*$ by the empirical output measure gives $T_n(\theta)$, the sample analog of \eqref{eq:discrepancy_function}.
With finite output support, the pointwise limiting distribution of $\sqrt{n}(T_n(\theta)-T(\theta))$ depends on the \emph{contact set}, the set of test functions attaining the population supremum.
We use a penalized bootstrap whose diverging penalty localizes the bootstrap supremum to the contact set.
Uniform validity does not require uniform estimation of the contact set.
For $\theta$ in the identified set, a finite-sample inequality reduces uniform size control to a uniform empirical-process approximation on the finite output space.
Test inversion then gives confidence sets that cover each point of the identified set with asymptotic probability at least $1-\alpha$, uniformly over $\mu^*$ and over $\theta$ in its identified set, including output measures with zero-probability cells.
\footnote{The tests add a fixed $\varepsilon>0$ to the bootstrap critical value on the $\sqrt{n}T_n(\theta)$ scale, as in Corollary~\ref{C:inf_finite}.}
For alternatives satisfying $T(\theta)\geq\Delta/\sqrt n$, the worst-case asymptotic acceptance probability converges to zero uniformly over the sampling distribution and the parameter value as $\Delta\to\infty$.
The population ADF, its sample analog, and the penalized bootstrap statistic are computed with the same LP structure.


Our fourth contribution concerns binary choice panels with fixed effects.
Under the baseline specification, the error terms have a known distribution, are independent across periods, and are independent of the fixed effect and covariates.
In two-period designs based on six error distributions, logit reproduces the classical point identification result \parencite{cham1980,Chamberlain2010}, whereas the other five distributions produce interval-identified coefficients.
If serial dependence is unrestricted while each error term retains its known conditional distribution given the fixed effect and covariates, the coefficient set widens but its sign remains identified in all six designs.
If instead the fixed effect and the errors may be arbitrarily dependent while the errors remain serially independent conditional on the covariates, the identified sets contain zero in these designs.
For the probit design, we also compute joint identified sets for the coefficient and the ASF.
Additional periods contract these sets, but more slowly under either relaxation than under the baseline.
When the fixed effect and errors may be arbitrarily dependent, five periods are required in this design to identify the sign of the coefficient.
With two periods and interval-censored covariates, the coefficient sign remains identified even when each covariate is observed only through three equal-width bins; see Additional Appendix~\ref{app:interval_results}.
In the probit simulations, rejection rates at evaluated coefficients in the identified set do not exceed the nominal level, while rejection rates outside the set increase with sample size.

We also characterize the identified set under sequential exogeneity.
For binary choice panels with finite covariate support, this gives, to our knowledge, the first identified-set characterization for the coefficient without a parametric error distribution or a special regressor.
In our baseline numerical design, adding a third period produces an upper bound on the coefficient, although its sign remains unidentified with either two or three periods.
Feedback from past outcomes to future covariates may violate conditional stationarity \parencite{Manski1987}, which conditions on the complete covariate history, while leaving sequential exogeneity correctly specified.

Finally, the framework applies beyond panel models.
We study the two-player complete-information entry game of \textcite{Tamer2003} under the specification of \textcite{beresteanuSharpIdentificationRegions2011} (BMM).
Without covariates and with bivariate standard normal error terms, augmenting the inputs with an unrestricted equilibrium selector yields a LIMR whose ADF has BMM's identified set as its zero set.
Thus an incomplete model can be handled by completing it through the input measure rather than first deriving a random set characterization.
The two criteria agree on membership, and finite linear programs approximate the adversarial criterion to arbitrary accuracy; see Supplemental Appendix~\ref{app:entry_game}.

\subsection{Related literature}
\label{sec:related_literature}

\paragraph{Identification.}
A useful distinction among identification methods for latent-variable models is whether they eliminate the latent variables before characterizing compatibility or retain a latent object in the characterization.

Random set methods characterize compatibility in observable space \parencite{BeresteanuMolinari2008,beresteanuSharpIdentificationRegions2011}.
BMM represent the model-implied moment set as an Aumann expectation and characterize membership by support functions.
For their entry-game specification, which fixes the distribution of the error terms and leaves equilibrium selection unrestricted, the completed LIMR of Example~\ref{ex:entry_game_main} generates the same set of outcome probability vectors as the Aumann expectation of their equilibrium-outcome random set; see Supplemental Appendix~\ref{app:entry_game}.
This equivalence relies on unrestricted equilibrium selection.
Restrictions on equilibrium selection change the set of admissible selections, and BMM note that the resulting moment set need not remain convex \parencite[p.~1788]{beresteanuSharpIdentificationRegions2011}.
The input measure fixes the marginal distribution of the error terms.
A restriction that is linear in the conditional distribution of the equilibrium selector given the error terms is therefore linear in the input measure, and the LIMR can impose it.
The same applies to restrictions on errors conditional on other latent variables whenever those restrictions are linear in the input measure.
Additionally, the approach here does not require an explicit characterization of the set of model-implied output distributions or its observable implications.

\textcite{chesherRosenZhang2026} project out the fixed effect and then apply random set methods to characterize compatibility in the resulting incomplete model.
Projection permits unrestricted dependence between the fixed effect and the error terms.
The LIMR can allow unrestricted dependence while retaining the fixed effect among the inputs, as in Example~\ref{ex:binary_parametric}.
Restrictions or counterfactual parameters involving the fixed effect or its distribution can then be imposed directly.
After projection, they must instead admit an equivalent representation in terms of the projected model.
For discrete-outcome specifications, the projected characterization can require containment inequalities indexed by a core-determining collection of sets.
With finite observable support, the LIMR evaluates compatibility through the linear programming procedure of Section~\ref{sec:computation} without constructing such a collection.

A second class of methods retains a latent object in the compatibility problem.
Optimal transport methods use couplings of observed and latent variables, with the latent marginal specified or restricted through finitely many moments \parencite{ekelandOptimalTransportationFalsifiability2010,galichonSetIdentificationModels2011}.
\textcite{schennachEntropicLatentVariable2014} profiles out the latent distribution by entropic tilting when the model is represented by finitely many moment restrictions, and treats countably many restrictions through an increasing sequence of finite systems.
\textcite{li2026} uses a support-function criterion to characterize the moment closure of the identified set and treats structural and counterfactual parameters jointly.
The LIMR also accommodates models in which the distribution of the fixed effect is unrestricted and an uncountable family of restrictions is imposed conditional on it.
Those restrictions enter directly through the joint input measure, while the closure relevant for our identification result is taken in the space of output probability measures rather than in moment space.

Other methods retain type probabilities, cell probabilities, a subdistribution, or a latent function \parencite{balkePearlBoundsTreatmentEffects1997,laffersIdentificationModelsDiscrete2019,MogstadSantosTorgovitsky2018,torgovitskyPartialIdentificationExtending2019,Tebaldi2023,guCounterfactualIdentificationLatent2025}.
In nonlinear panels, related methods optimize over latent heterogeneity or exploit model-specific reductions to characterize coefficients or average effects \parencite{honoreBoundsParametersPanel2006,ChernValHahnNewey2013,daveziesIdentificationEstimationAverage2022,bonhommeDanoGraham2023}.
\textcite{guCounterfactualIdentificationLatent2025} instead construct an exact finite representation for discrete-outcome models with indices linear in latent variables and characterize additional restrictions that can be imposed in that representation.
These reductions depend on the outcome equation, restrictions on the latent variables, and the parameter of interest, and need not preserve arbitrary restrictions indexed by continuously distributed latent heterogeneity.

\textcite{christensenCounterfactualSensitivityRobustness2023} study sensitivity of counterfactuals to a parametric specification of the distribution of the unobservables, optimizing over a $\varphi$-divergence neighborhood subject to finitely many moment restrictions.
Their nonparametric case also yields a membership criterion whose dual contains an optimization over latent values; finite-radius neighborhoods replace this optimization by a convex expectation (their Section~2.5 and Remark~2.7).
With finite observable support, we instead bound the analogous optimization over input values and obtain certified lower and upper bounds on $T(\theta)$ (Section~\ref{sec:computation}).

\paragraph{Computation.}
Finite computation in latent-variable models typically follows either from an exact finite representation \parencite{balkePearlBoundsTreatmentEffects1997,honoreBoundsParametersPanel2006,KitamuraStoye2018,laffersIdentificationModelsDiscrete2019,Tebaldi2023,guCounterfactualIdentificationLatent2025} or from a finite approximation to an infinite-dimensional latent object \parencite{ChernValHahnNewey2013,MogstadSantosTorgovitsky2018,bonhommeDanoGraham2023}.
For their finite linear programming formulation, \textcite{honoreBoundsParametersPanel2006} impose finite support on the fixed effect, while \textcite{ChernValHahnNewey2013}, \textcite{bonhommeDanoGraham2023}, and \textcite{botosaru2024} implement continuous-support problems on finite grids without formal bounds on the discretization error.
\textcite{pakelBoundsAverageEffects2026} instead obtain tractable outer bounds.
Our finite input support and retained moment restrictions define an auxiliary problem.
The target remains the unrestricted $T(\theta)$, and the certificates in Section~\ref{sec:computation} translate the finite problem into valid bounds on that target.
The resulting procedure is related to cutting-plane, row-exchange, and column-generation methods for semi-infinite and large-scale linear programs \parencite{Kelley1960,HettichKortanek1993,Luebbecke2011,Muter2013}.
In econometrics, \textcite{SmeuldersCramaSpieksma2021} generate rational types for the finite program of \textcite{KitamuraStoye2018}.

\paragraph{Inference.}
We use the per-point coverage criterion of \textcite{im2004} and \textcite{stoya2009}, uniformly over the sampling distribution and the parameter value under test, rather than simultaneous coverage of the identified set \parencite{Chernozhukovetal2007,romano2010inference}.
The nonregularity is analogous to that in moment inequality models, where the binding restrictions can change along sequences of data-generating processes and motivate moment selection \parencite{RomanoShaikh2008,andrews2010}.
There, moment selection operates on observable moment inequalities.
Here the corresponding object is the contact set of test functions generated by the adversarial criterion, and uniform validity does not require uniform estimation of that set.
\textcite{GalichonHenry2009} and \textcite{loh2024} use related minimax statistics, while the pointwise analysis uses results for directionally differentiable functionals \parencite{fang2018,HongLi2018}.

\subsection{Notation}

We use $\one\{\cdot\}$ for the indicator function and $\overset{d}{=}$ for equality in distribution.
For each measurable space $\mathcal S$, let $\mathcal B(\mathcal S)$ denote its $\sigma$-algebra.
Product spaces carry the product $\sigma$-algebra.
We write $\Pa(\mathcal S)$ for the set of probability measures on $\mathcal B(\mathcal S)$ and $\delta_s$ for the Dirac measure at $s\in\mathcal S$, defined by $\delta_s(B)=\one\{s\in B\}$.
A probability kernel from $\mathcal S$ to $\mathcal S'$ is a map $K(\cdot\mid\cdot)$ such that $K(\cdot\mid s)$ is a probability measure on $\mathcal S'$ for each $s\in\mathcal S$ and $K(B\mid\cdot)$ is measurable for each $B\in\mathcal B(\mathcal S')$.
For a measurable map $f\colon \mathcal S\to\mathcal S'$ and $\mu\in\Pa(\mathcal S)$, the pushforward measure $f_*\mu\in\Pa(\mathcal S')$ is defined by $(f_*\mu)(B)=\mu(f^{-1}(B))$ for $B\in\mathcal B(\mathcal S')$.
For an index set $\mathcal I$, $\R^{(\mathcal I)}$ denotes the set of vectors in $\R^{\mathcal I}$ with finite support, and $\R^{(\mathcal I)}_+$ its nonnegative cone.
For a finite signed measure $\nu$, write $\inner{\phi,\nu}\equiv\int\phi\,\d\nu$ and, when $\nu=\mu$ is a probability measure, $\EE{\mu}{\phi}\equiv\inner{\phi,\mu}$.
For probability measures, we use the convention
$\norm{\mu-\mu'}_{\mathrm{TV}}\equiv\sup_{B\in\mathcal B(\mathcal S)}|\mu(B)-\mu'(B)|$
for total variation distance.
For a vector $X$, $X'$ denotes its transpose.




\section{Model}
\label{sec:model}

This section defines what it means for an econometric model to have a \emph{linear-in-measures representation} (LIMR).
Section~\ref{sec:examples} gives six examples of models with a LIMR.

Let $X \in \mathcal{X}$ and $U \in \mathcal{U}$ denote observable and unobservable inputs, respectively, and collect them into the input $W = (X, U) \in \mathcal{W} = \mathcal{X} \times \mathcal{U}$.
Let $Y \in \mathcal{Y}$ denote the outcome, and define the output as $Z = (Y, X) \in \mathcal{Z} = \mathcal{Y} \times \mathcal{X}$, so that the observable $X$ serves as both an input and an output.

In our leading binary choice panel models, $X = (X_1,\dots,X_T)$ collects time-varying covariates, $Y = (Y_1,\dots,Y_T)$ collects binary outcomes, and $U$ contains a time-invariant fixed effect and, in some specifications, idiosyncratic error terms.
In other applications, $U$ may contain random coefficients, true values of partially observed covariates, equilibrium selectors, potential outcomes, or other structural primitives.

\begin{asm}
\label{asm:measurable}
The supports $\mathcal Y$, $\mathcal X$, and $\mathcal U$ are measurable spaces.
\end{asm}
A probability measure $\gamma\in\Pa(\mathcal W)$ on the inputs is called an \emph{input measure}.
A parameter $\theta\in\Theta$ collects structural parameters, such as regression coefficients, and counterfactual parameters defined by restrictions linear in $\gamma$, such as ASFs.
\begin{asm}
\label{asm:moment_conditions}
For each $\theta\in\Theta$, the set of admissible input measures $\Gamma_\theta \subseteq \Pa(\mathcal W)$ is characterized by a system of moment restrictions:\footnote{These restrictions implicitly require integrability of the moment functions with respect to $\gamma$.}
$$
\Gamma_\theta = \left\{ \gamma \in \Pa(\mathcal W) : \begin{aligned}
\EE{\gamma}{g_{1,j}(W;\theta)} &= 0,   &&\quad\text{for all } j \in \mathcal J \\
\EE{\gamma}{g_{2,k}(W;\theta)} &\le 0, &&\quad\text{for all } k \in \mathcal K
\end{aligned} \right\},
$$
where $\mathcal J$ and $\mathcal K$ are arbitrary index sets, and $\{g_{1,j}(\cdot;\theta)\}_{j \in \mathcal J}$ and $\{g_{2,k}(\cdot;\theta)\}_{k \in \mathcal K}$ are families of real-valued measurable functions on $\mathcal W$.
\end{asm}
In the examples, the equalities represent sequential exogeneity, parametric restrictions, and random assignment.
The inequalities may represent shape restrictions.
The equalities can also define the counterfactual parameters, as in Example~\ref{ex:binary_parametric}.
The index sets $\mathcal J$ and $\mathcal K$ may be infinite, which matters in nonlinear panel models because restrictions conditional on fixed effects typically generate a continuum of moment equalities.

Not every model assumption is a linear restriction in $\gamma$ in its natural parameterization.
For example, independence with unrestricted marginals imposes a nonlinear factorization of the joint measure.
Such restrictions may admit an equivalent linear representation after reparameterization or augmentation, as in Example~\ref{ex:binary_sequential}.

\begin{asm}
\label{asm:linear_operator}
For each $\theta\in\Theta$, the model specifies an \emph{output kernel}, a probability kernel $K_\theta(\cdot\mid\cdot)$ from $\mathcal W$ to $\mathcal Z$ that does not depend on $\gamma$ and that preserves the observable input:
\begin{equation}
\label{eq:probability_kernel}
(\pi_{\mathcal X})_*K_\theta(\cdot\mid x,u)=\delta_x
\quad\text{for all } (x,u)\in\mathcal W,
\end{equation}
where $\pi_{\mathcal X}\colon\mathcal Z\to\mathcal X$ denotes the coordinate projection.
\end{asm}
Integrating the output kernel with respect to $\gamma$ defines the \emph{output measure} $\mu_{\theta,\gamma} \equiv \Lgen\gamma$ by
\begin{equation}
\label{eq:kernel_operator}
(\Lgen\gamma)(B) \equiv \int_{\mathcal W} K_\theta(B\mid w)\,\d\gamma(w)
\quad\text{for all } B\in\mathcal B(\mathcal Z).
\end{equation}
We call $\Lgen$ the \emph{output operator}.
Because $K_\theta$ is fixed as $\gamma$ varies, the output operator is linear in the input measure.
\footnote{Formally, because $\Pa(\mathcal W)$ is not a vector space, the map $\Lgen\colon\Pa(\mathcal W)\to\Pa(\mathcal Z)$ is affine. Throughout, linearity refers to its extension to finite signed measures.}

Assumption~\ref{asm:linear_operator} does not require the outcome equation to be linear, separable, or parametric.
For example, consider a model with outcome equation $Y=h_\theta(W)$, where $h_\theta\colon\mathcal W\to\mathcal Y$ is measurable for each $\theta$, and define $\psi_\theta\colon\mathcal W\to\mathcal Z$ by $\psi_\theta(w)\equiv\bigl(h_\theta(w),x\bigr)$ for $w=(x,u)$.
The output kernel is
$K_\theta(B\mid w)=\one\{\psi_\theta(w)\in B\}$,
and the output measure $\Lgen\gamma=(\psi_\theta)_*\gamma$ is the pushforward of $\gamma$ through the map $w\mapsto\psi_\theta(w)$.

The representation is flexible in how the model assumptions are allocated among $\mathcal W$, $\Gamma_\theta$, and $K_\theta$.
Section~\ref{sec:examples} illustrates this flexibility: the parametric panel examples integrate out the error terms through $K_\theta$; the semiparametric panel example retains them in $W$ and restricts their distribution through $\Gamma_\theta$; the interval-censoring example builds the censoring restriction directly into $\mathcal W$; and the entry game augments $W$ with an equilibrium selector.

We call the family $\{(\Gamma_\theta,\Lgen):\theta\in\Theta\}$ a LIMR if, for each $\theta$, the set $\Gamma_\theta\subseteq\Pa(\mathcal W)$ is characterized by Assumption~\ref{asm:moment_conditions}, and the output operator $\Lgen\colon\Pa(\mathcal W)\to\Pa(\mathcal Z)$ is characterized by Assumption~\ref{asm:linear_operator}.
The name reflects that the restrictions characterizing $\Gamma_\theta$ and the output operator $\Lgen$ are linear in the input measure.
We say that an econometric model has a LIMR if there exists such a family for which, at every $\theta\in\Theta$, $\{\Lgen\gamma:\gamma\in\Gamma_\theta\}$ coincides with the set of probability measures of $Z$ that the model can generate at $\theta$.

\begin{asm}
\label{asm:domination}
For each $\theta\in\Theta$, there exists a $\sigma$-finite measure $\lambda_\theta$ on $\mathcal Z$ such that, for all $\gamma \in \Gamma_\theta$, $\Lgen\gamma$ is absolutely continuous with respect to $\lambda_\theta$.
\end{asm}

Assumption~\ref{asm:domination} ensures that every output measure has a Radon--Nikodym density in the common space $L^1(\lambda_\theta)$.
It supplies the dominating measure used in the separation argument of Section~\ref{sec:identification}.
This assumption is satisfied when $\mathcal Z$ is finite by taking $\lambda_\theta$ to be counting measure.
With continuous $X$ or $Y$, Assumption~\ref{asm:domination} holds, for example, when $\Gamma_\theta$ fixes the marginal of $X$ at some probability measure and, for every $(x,u)\in\mathcal W$, $K_\theta(\cdot\times\mathcal X\mid x,u)$ is dominated by a probability kernel from $\mathcal X$ to $\mathcal Y$ that does not depend on $u$.


\section{Examples}
\label{sec:examples}

This section illustrates the scope of the LIMR by constructing one for each of six econometric models.
Each example specifies the input $W$, the admissible set $\Gamma_\theta$, and the output kernel $K_\theta$.
Throughout, the outcomes and the observable input take finitely many values, so Assumption~\ref{asm:domination} holds with $\lambda_\theta$ equal to counting measure on $\mathcal Z$, and we can write the output kernel as a probability mass function.
The panel examples feature linear indices, additive fixed effects, and binary outcomes, although the framework does not require them.

\begin{example}[Parametric binary choice with fixed effects]
\label{ex:binary_parametric}
For $t=1,\ldots,T$, let
\begin{equation}
\label{eq:panel_binary_outcome}
Y_t=\one\{X_t'\beta+A-V_t\ge0\},
\end{equation}
where $A$ is a fixed effect whose distribution given the covariate path $X=(X_1,\ldots,X_T)$ is unrestricted, and $V=(V_1,\ldots,V_T)$ collects the error terms.
Let $H$ be a known continuous CDF.
We consider a baseline in which the errors are i.i.d.\ given $(A,X)$, and two relaxations:
\begin{equation}
\label{eq:parametric_three_specifications}
\begin{aligned}
V\mid(A,X)&\sim H^{\otimes T} && \text{(baseline)},\\
V_t\mid(A,X)&\sim H,\quad t=1,\ldots,T && \text{(serial dependence)},\\
V\mid X&\sim H^{\otimes T} && \text{(fixed effect--error dependence)}.
\end{aligned}
\end{equation}
The first relaxation allows serial dependence while retaining independence of each error from $(A,X)$.
The second allows dependence between the fixed effect and the errors while retaining serial independence given $X$.
Section~\ref{subsec:numerical_example1} compares their identifying information.
For the baseline specification, take $W=(X,A)$, $Z=(Y,X)$, and $\Gamma_\beta=\Pa(\mathcal X\times\R)$, leaving the input measure unrestricted.
Integrating out $V$ gives the output kernel
$$
K_\beta(\{(y,x)\}\mid x,a)=\prod_{t=1}^T H(x_t'\beta+a)^{y_t}\bigl(1-H(x_t'\beta+a)\bigr)^{1-y_t}.
$$
We incorporate the ASF $\tau_{\mathrm{ASF}}\equiv\Pr(\bar x'\beta+A-V_1\ge0)$ at a counterfactual covariate value $\bar x$ by setting $\theta=(\beta,\tau_{\mathrm{ASF}})$ and imposing the linear restriction $\EE{\gamma}{H(\bar x'\beta+A)-\tau_{\mathrm{ASF}}}=0$, which refines $\Gamma_\beta$ to $\Gamma_\theta$.
Similarly, an average treatment effect (ATE) between two counterfactual covariate values $\bar x^0$ and $\bar x^1$ can be incorporated with $\theta=(\beta,\tau_{\mathrm{ATE}})$ and $\EE{\gamma}{H\bigl((\bar x^1)'\beta+A\bigr)-H\bigl((\bar x^0)'\beta+A\bigr)-\tau_{\mathrm{ATE}}}=0$.
For the two relaxations, take $W=(X,A,V)$, use the kernel induced by~\eqref{eq:panel_binary_outcome}, and impose the corresponding distributional restriction through the admissible set.
Supplemental Appendix~\ref{app:computation_example1} develops the LIMR for each specification.
\end{example}

\begin{example}[Interval-censored covariates]
\label{ex:binary_interval}
For $t=1,\ldots,T$, let
$$
Y_t=\one\{(X_t^\star)'\beta+A-V_t\ge0\},
\qquad
X_{L,t}\le X_t^\star\le X_{U,t}.
$$
The covariate path $X^\star$ and the fixed effect $A$ are latent, and $X^\star$ is observed only through the endpoints $X=(X_L,X_U)$, the observable input, so the output is $Z=(Y,X_L,X_U)$.
Impose $V\mid(X_L,X_U,X^\star,A)\sim H^{\otimes T}$ for known $H$.
With $\theta=\beta$, take $W=(X_L,X_U,X^\star,A)$ on the input space $\mathcal W\equiv\{(x_L,x_U,x^\star,a):x_L\le x^\star\le x_U\}$ and set $\Gamma_\theta=\Pa(\mathcal W)$.
The support restriction carries the interval censoring, and the input measure is otherwise unrestricted.
Integrating out $V$ gives the output kernel: $Y$ has the probability mass function of Example~\ref{ex:binary_parametric} with $x^\star$ in place of $x$, and the endpoints in the output equal those in the input.
Additional Appendix~\ref{app:interval_results} reports identified sets for $\beta$ in a numerical example.
\end{example}

\begin{example}[Panel binary choice with sequential exogeneity]
\label{ex:binary_sequential}
Consider outcome equation~\eqref{eq:panel_binary_outcome}.
Denote by $X^t\equiv(X_1,\ldots,X_t)$ the covariate history through period $t$ and impose sequential exogeneity:
\begin{equation}
\label{eq:seq_exog}
V_t\mid (A,X^t)\overset{d}{=}V_1\mid (A,X_1),
\qquad
t=2,\ldots,T.
\end{equation}
This is the predetermined version of time homogeneity in \textcite[Assumption~3]{ChernValHahnNewey2013}: given the fixed effect, the error has the same distribution in every period, and that distribution may depend on the covariate history only through $X_1$.
Future covariates do not enter the conditioning set, so covariates may respond to past outcomes \parencite{ChernValHahnNewey2013,chamberlainFeedbackPanelData2022,bonhommeDanoGraham2023,Bonhomme2025BackToFeedback}.
The restriction is nonlinear in the distribution of $(X,V,A)$, but becomes linear once we replace $A$ by $\nu\equiv(\Pr(X=x\mid A))_{x\in\mathcal X}$, the distribution of the covariate path given the fixed effect, and $V_t$ by the composite error $\tilde V_t\equiv A-V_t$.
Write $\nu_t(x^t)\equiv\sum_{\tilde x\in\mathcal X:\,\tilde x^t=x^t}\nu(\tilde x)$ for the probability of the history $x^t$ under $\nu$.
Lemma~\ref{lem:nu_lifting} in Supplemental Appendix~\ref{app:semiparametric_binary_choice} shows that, for finite $\mathcal X$, the model has a LIMR with input $W=(X,\tilde V,\nu)$, output $Z=(Y,X)$, output operator the pushforward of $\gamma$ through the map $w\mapsto\bigl((\one\{x_t'\beta+\tilde v_t\ge0\})_{t=1}^T,x\bigr)$, and admissible set $\Gamma_\beta$ defined by
\begin{align}
&\EE{\gamma}{r(\nu)\bigl(\one\{X=x\}-\nu(x)\bigr)}=0,
\quad\text{for all }x,\ \text{and all bounded measurable }r,
\label{eq:nu_consistency}\\
&\EE{\gamma}{r(\nu)\bigl[f(\tilde V_t)\one\{X^t=x^t\}\nu_1(x_1)-f(\tilde V_1)\one\{X_1=x_1\}\nu_t(x^t)\bigr]}=0,
\label{eq:seq_lifted}\\
&\quad\text{for all }t=2,\ldots,T,\ \text{all }x^t,\ \text{and all bounded measurable }r,f.
\notag
\end{align}
Restriction~\eqref{eq:seq_lifted} is sequential exogeneity with $\nu$ in place of $A$, multiplied through by the history probabilities so that it is linear in $\gamma$.
We incorporate the ASF at a counterfactual covariate value $\bar x$ by setting $\theta=(\beta,\tau_{\mathrm{ASF}})$ and imposing $\EE{\gamma}{\one\{\bar x'\beta+\tilde V_1\ge0\}-\tau_{\mathrm{ASF}}}=0$ in addition to the restrictions defining $\Gamma_\beta$.
Section~\ref{subsec:numerical_semiparametric} compares the identified sets with those under conditional stationarity.
\end{example}

\begin{example}[Dynamic discrete choice and dynamic panels]
\label{ex:dynamic_state_dependence}
\textcite{honoreBoundsParametersPanel2006} study the dynamic panel model with outcome equation
$$
Y_t=\one\{X_t'\beta+\rho Y_{t-1}+A+V_t\ge0\}, \qquad t=1,\ldots,T,
$$
where the errors $V_t$ are i.i.d.\ standard normal variables independent of $(X,A,Y_0)$, and the binary initial condition $Y_0$ is latent.
With $\theta=(\beta,\rho)$, take $W=(X,A,Y_0)$, $Z=((Y_1,\ldots,Y_T),X)$, and $\Gamma_\theta=\Pa(\mathcal X\times\R\times\{0,1\})$, leaving the distribution of $(A,Y_0)$ given $X$ unrestricted.
Integrating out $V$ gives the output kernel
$$
K_\theta(\{(y,x)\}\mid x,a,y_0)=\prod_{t=1}^T \Phi(x_t'\beta+\rho y_{t-1}+a)^{y_t}\bigl(1-\Phi(x_t'\beta+\rho y_{t-1}+a)\bigr)^{1-y_t},
$$
where $\Phi$ denotes the standard normal CDF.

The same construction extends immediately to the habit persistence specification of \textcite{heckmanStatisticalModelsDiscrete1981}, in which the lagged latent index replaces $Y_{t-1}$ in the outcome equation: $Y_t=\one\{Y_t^*\ge0\}$ with $Y_t^*=X_t'\beta+\rho Y_{t-1}^*+A+V_t$.
The latent input component becomes $(A,Y_0^*)$, and the output kernel is given by multivariate normal orthant probabilities.
\end{example}

\begin{example}[Simultaneous entry game with multiple equilibria]
\label{ex:entry_game_main}
Firms $j=1,2$ simultaneously choose entry actions $y_j\in\{0,1\}$.
Firm $j$ receives payoff $y_j(\delta_jy_{-j}+V_j)$, where $\theta=(\delta_1,\delta_2)\in(-\infty,0)^2$ collects the spillover parameters.
Both firms observe the error terms $V=(V_1,V_2)$ before play, and $V$ follows a known joint CDF $H$ with continuous marginals.
The game has three Nash equilibria when $0<V_j<-\delta_j$ for both firms, so the model is incomplete \parencite{Tamer2003}.
We complete the model with an equilibrium selector $S\in\{1,2,3\}$ that indexes those equilibria.
Take $W=(V,S)$ and $Z=Y=(Y_1,Y_2)$, the realized action profile, since $\mathcal X$ is a singleton.
The output kernel $K_\theta(\cdot\mid v,s)$ is the outcome distribution of the equilibrium that $s$ selects at $v$.
The admissible set $\Gamma_\theta$ imposes the moment equalities $\EE{\gamma}{\one\{V\le c\}-H(c)}=0$ for all $c\in\R^2$, ensuring that $V$ has CDF $H$.
It leaves the conditional distribution of $S$ given $V$ unrestricted, thereby spanning every equilibrium selection mechanism \parencite{BerryTamer2007}.
Proposition~\ref{prop:entry_equivalence} in Supplemental Appendix~\ref{app:entry_game} shows, for bivariate standard normal errors, that the identified set coincides with the sharp identification region of \textcite{beresteanuSharpIdentificationRegions2011}.
\end{example}

\begin{example}[Randomized trial with imperfect compliance]
\label{ex:imperfect_compliance}
Let $R\in\{0,1\}$ be a randomized assignment, $D\in\{0,1\}$ the treatment received, and $Y\in\{0,1\}$ the outcome, allowing $D\ne R$ \parencite{balkePearlBoundsTreatmentEffects1997}.
Write $D(r)$ for treatment under assignment $r$ and $Y(d)$ for the outcome under treatment $d$, and collect the latent response type as $U=(Y(0),Y(1),D(0),D(1))\in\{0,1\}^4$.
Take the observable input to be $X=R$ and the model outcome to be $(Y,D)$, so $W=(R,U)$ and $Z=(Y,D,R)$.
The output kernel maps $(r,u)$, with $u=(y(0),y(1),d(0),d(1))$, to $(y(d(r)),d(r),r)$.
This mapping incorporates the exclusion restriction.
Let $\theta=\tau$ denote the ATE.
Random assignment and the target enter $\Gamma_\theta$ through the linear equalities
$$
\begin{aligned}
\EE{\gamma}{\bigl(\one\{R=1\}-p\bigr)h(U)}&=0
&&\text{for all bounded measurable }h,\\
\EE{\gamma}{Y(1)-Y(0)-\tau}&=0,
\end{aligned}
$$
where $p\equiv\Pr(R=1)$ is known by design.
Because $W$ takes finitely many values, the identified set for $\tau$ is the interval between the sharp bounds of \textcite{balkePearlBoundsTreatmentEffects1997}, which solve the same linear program over the distribution of $U$.
Restrictions beyond those of \textcite{balkePearlBoundsTreatmentEffects1997} enter directly: monotonicity $D(1)\ge D(0)$ is the moment equality $\EE{\gamma}{\one\{D(0)=1,\ D(1)=0\}}=0$, which rules out defiers.
\end{example}


\section{Identification}
\label{sec:identification}

We first define the identified set, then construct the ADF $T(\theta)$ via a separation argument.
Our main result shows that $\theta$ belongs to the identified set if and only if $T(\theta)=0$.

The econometrician observes the true probability measure $\mu^*$ of the output $Z$.
For each $\theta\in\Theta$, let $\mathcal M_\theta\subseteq\Pa(\mathcal Z)$ denote the set of probability measures of $Z$ that the model can generate at $\theta$.
If the model has a LIMR, this set can be written as
\begin{equation}
    \label{eq:model_probabilities}
    \mathcal M_\theta = \Lgen\Gamma_\theta = \{\mu_{\theta,\gamma}:\gamma\in\Gamma_\theta\}.
\end{equation}
Denote the TV closure of $\mathcal M_\theta$ by
\begin{equation}
    \label{eq:closure_definition}
    \Mtheta
    \equiv
    \Bigl\{
     m \in \Pa(\mathcal{Z}) :
     \forall \epsilon > 0,\ \exists\, \gamma \in \Gamma_\theta,
     \ \norm{\mu_{\theta,\gamma} - m}_{\mathrm{TV}} < \epsilon \Bigr\}.
\end{equation}
We define the identified set as
\begin{equation}
    \label{eq:identified_set}
    \Thetaw \equiv
    \{\theta \in \Theta :
    \mu^*\in\Mtheta\}.
\end{equation}

The closure adds probability measures that can be approximated arbitrarily closely in TV by measures in $\mathcal M_\theta$.
\footnote{Identification using closures also arises in \textcite{schennachEntropicLatentVariable2014} and \textcite{li2026}.
The former takes the closure of attainable expected moment values, while the latter characterizes the moment closure of the identified set.
Both closures are taken in finite-dimensional moment space rather than in the space of output measures used here.}
When $\mathcal M_\theta$ is already TV closed, $\mathcal M_\theta=\Mtheta$ and the closure is redundant.
This is the case in Example~\ref{ex:binary_sequential}, as shown in Proposition~\ref{prop:se_finite_support} in Supplemental Appendix~\ref{app:semiparametric_binary_choice}.
By contrast, $\mathcal M_\theta\subsetneq\Mtheta$ in the logit specification of Example~\ref{ex:binary_parametric}.
For a fixed covariate path $x$, the point mass at $(Y,X)=((1,\ldots,1),x)$ is the TV limit of output measures as the fixed effect diverges to $+\infty$, but no admissible input measure generates it (Remark~\ref{rem:closure_role}).

Under i.i.d.\ sampling, the boundary probability measures added by the TV closure $\Mtheta$ cannot be statistically distinguished from the measures in $\mathcal M_\theta$: any test whose size is controlled uniformly over $\mathcal M_\theta$ has power no greater than size against a measure in $\Mtheta\setminus\mathcal M_\theta$; see Remark~\ref{rem:tv_indistinguishability}.

We construct the ADF from the LIMR and Assumptions~\ref{asm:measurable} and~\ref{asm:domination}.
The LIMR makes $\mathcal M_\theta$ convex, so its TV closure $\Mtheta$ is closed and convex, while Assumption~\ref{asm:domination} represents every element of $\Mtheta$ by a density in $L^1(\lambda_\theta)$.
The associated separating hyperplane argument yields $\mu^*\notin\Mtheta$ if and only if there is a bounded measurable test function $\phi$ such that $\EE{\mu^*}{\phi}-\sup_{\mu\in\Mtheta}\EE{\mu}{\phi}>0$.
By $L^1$--$L^\infty$ duality, every continuous linear functional on $L^1(\lambda_\theta)$ is integration against a bounded measurable function, so no larger class of test functions is needed.
Because positive rescaling and translation preserve the sign of the separating gap, we normalize this full class as
\begin{equation}
    \Phi(\mathcal{Z})
    \equiv
    \{\phi\colon\mathcal{Z}\to[0,1]\text{ measurable}\}.
    \label{eq:phi_class}
\end{equation}

The ADF is then the value of this separation problem.
The outer supremum maximizes the separating gap over $\Phi(\mathcal Z)$, while the inner infimum searches over the admissible input measures $\gamma\in\Gamma_\theta$:
\footnote{We use the convention $\inf\varnothing=+\infty$, so parameter values $\theta$ with $\Gamma_\theta=\varnothing$ are automatically excluded.}
\begin{align}
    T(\theta)
    &\equiv
    \sup_{\phi \in \Phi(\mathcal{Z})} \inf_{\gamma \in \Gamma_\theta}
        \Bigl( \EE{\mu^*}{\phi} - \EE{\Lgen\gamma}{\phi} \Bigr)
    \label{eq:discrepancy_function_input}
    \\
    &=
    \sup_{\phi \in \Phi(\mathcal{Z})} \inf_{\mu \in \Mtheta}
        \Bigl( \EE{\mu^*}{\phi} - \EE{\mu}{\phi} \Bigr).
    \label{eq:discrepancy_function_model}
\end{align}
For each $\phi\in\Phi(\mathcal Z)$, the gap $\mu\mapsto\EE{\mu^*}{\phi}-\EE{\mu}{\phi}$ is continuous in TV.
Its infimum over $\mathcal M_\theta=\Lgen\Gamma_\theta$ therefore equals its infimum over the closure $\Mtheta$, which gives the second equality.

Let $\Thetam$ denote the zero set of $T(\theta)$:
$$\Thetam \equiv \{\theta \in \Theta : T(\theta)=0\}.$$
\begin{thm}
    \label{thm:main_sharpness}
    Under Assumptions~\ref{asm:measurable}--\ref{asm:domination}, the identified set is
    \begin{equation}
        \Thetaw = \Thetam.
    \end{equation}
\end{thm}
\begin{proof}
See Appendix~\ref{app:proofs}.
\end{proof}

Identification can be viewed as a separation problem in the space of probability measures on $\mathcal Z$.
By \eqref{eq:discrepancy_function_model}, $T(\theta)=0$ exactly when no bounded test function separates $\mu^*$ from $\Mtheta$; Theorem~\ref{thm:main_sharpness} shows that this happens exactly when $\theta\in\Thetaw$.\footnote{The value $T(\theta)$ is nonnegative because $\phi\equiv0$ is feasible.}
The separation argument uses only two properties of the set of output measures: convexity and common domination.
The LIMR supplies the first and Assumption~\ref{asm:domination} the second, and the LIMR also makes the model's response to each test function a linear problem in the input measure.
Evaluating $T(\theta)$ therefore does not require characterizing $\Mtheta$.
For computation and inference, we instead use \eqref{eq:discrepancy_function_input}, which works with the input measure $\gamma$.

\subsection{Discussion}
\label{sec:discussion_discrepancy}

\begin{remark}[Role of the closure]
\label{rem:closure_role}
Consider the baseline specification of
Example~\ref{ex:binary_parametric}, where $\theta=\beta$, $\Gamma_\beta=\Pa(\mathcal X\times\R)$, and $H$ has full support on $\R$, as in the logit case.
Fix a covariate path $x$ and let $\mu^*$ assign probability one to $(Y,X)=((1,\ldots,1),x)$.
For the admissible input measures $\gamma_a\equiv\delta_{(x,a)}$, the output measures satisfy $\mu_{\beta,\gamma_a}(\{((1,\ldots,1),x)\})=\prod_{t=1}^T H(x_t'\beta+a)\to1$ as $a\to+\infty$, so $\mu_{\beta,\gamma_a}\to\mu^*$ in TV and $\mu^*\in\overline{\mathcal M}_\beta$ for every $\beta$.
No admissible input measure generates $\mu^*$, because $\EE{\gamma}{\prod_{t=1}^T H(X_t'\beta+A)}<1$ for every $\gamma\in\Gamma_\beta$.
The limiting output measure here is the one that would be generated if $A$ were allowed to take the boundary value $+\infty$.
More generally, allowing the boundary values $\pm\infty$ would let $\Pr(Y=(1,\ldots,1)\mid X=x)$ range over $[0,1]$ rather than $(0,1)$.
This matters when $\mu^*$ is itself such a boundary point, as in this example: the closure yields $\Thetaw=\Theta$ rather than the empty set.
In this example, that inclusiveness is desirable, because a population with no outcome variation cannot distinguish values of $\beta$.

For this baseline specification, the closed set of output measures can be obtained by adjoining $\pm\infty$ to the support of $A$ and extending the output kernel continuously to those values.
We instead take the closure in output space, which avoids imposing model-specific conditions guaranteeing a closed set of output measures.
\end{remark}

\begin{remark}[Statistical indistinguishability of TV-closure points]
\label{rem:tv_indistinguishability}
Fix $\theta\in\Theta$ and $n\in\mathbb N$, and suppose that
$Z_1,\ldots,Z_n$ are i.i.d.
Let $\psi_n\colon\mathcal Z^n\to[0,1]$ be any possibly randomized test satisfying
\begin{equation}
    \sup_{\nu\in\mathcal M_\theta}
    \EE{\nu^{\otimes n}}{\psi_n}
    \leq \alpha .
    \label{eq:uniform_size_exact_model}
\end{equation}
Then, for every $\mu\in\Mtheta$,
\begin{equation}
    \EE{\mu^{\otimes n}}{\psi_n}
    \leq \alpha .
    \label{eq:closure_indistinguishability}
\end{equation}
Hence, if $\mu\in\Mtheta\setminus\mathcal M_\theta$, no test whose size is at most $\alpha$ uniformly over $\mathcal M_\theta$ rejects $\mu$ with probability above $\alpha$.\footnote{To see~\eqref{eq:closure_indistinguishability}, fix $\mu\in\Mtheta$ and choose
$\nu_j\in\mathcal M_\theta$ with
$\norm{\nu_j-\mu}_{\mathrm{TV}}\to0$. Then, for fixed $n$,
$$
    \norm{\nu_j^{\otimes n}-\mu^{\otimes n}}_{\mathrm{TV}}
    \leq
    n\norm{\nu_j-\mu}_{\mathrm{TV}}
    \longrightarrow0.
$$
Since $0\leq\psi_n\leq1$,
$\EE{\nu_j^{\otimes n}}{\psi_n}
\to\EE{\mu^{\otimes n}}{\psi_n}$, and
\eqref{eq:closure_indistinguishability} follows from
\eqref{eq:uniform_size_exact_model}.
The result applies only to $\mu\in\Mtheta$ and gives no impossibility
statement for $\mu\notin\Mtheta$.}
This TV-closure indistinguishability argument is standard in the literature on impossible inference; see, e.g., \textcite{BERTANHA2020247}.
Relatedly, \textcite{BaiPonomarevSantosShaikhTabordMeehanTorgovitsky2026} characterize the TV closure of the null hypothesis in their analysis of partially identified linear systems.
When $\mathcal Z$ is finite, $\Mtheta$ is simply the Euclidean closure of the feasible set of cell-probability vectors.
\end{remark}

\begin{remark}[Point-to-set integral probability metric]
\label{rem:ipm}
The ADF can be interpreted as an extension of an integral probability metric (IPM) from two fixed probability measures to a point-to-set discrepancy.
More generally, for any class $\mathcal F$ for which the expectations are well defined, define the corresponding point-to-set discrepancy:
$$
T_{\mathcal F}(\theta)
\equiv
\sup_{\phi\in\mathcal F}\inf_{\gamma\in\Gamma_\theta}
\Bigl(
\EE{\mu^*}{\phi}-\EE{\Lgen\gamma}{\phi}
\Bigr).
$$
For two fixed probability measures, different choices of $\mathcal F$ give familiar discrepancies, e.g., the unit ball of a reproducing kernel Hilbert space gives maximum mean discrepancy \parencite{Gretton2012}, while the $1$-Lipschitz functions give the $1$-Wasserstein distance.
The corresponding max--min criterion equals the infimum of these pairwise discrepancies over the model set when the relevant minimax conditions hold.
We establish this equality for TV under our maintained assumptions.

The separation argument above, which uses Assumptions~\ref{asm:measurable}--\ref{asm:domination}, \emph{selects} $\mathcal F = \Phi(\mathcal Z)$.
More regular classes require additional topological, metric, or kernel structure on $\mathcal Z$ that the construction of $T(\theta)$ does not use.
If $\mathcal F\subseteq\Phi(\mathcal Z)$ and $0\in\mathcal F$, then
$0\leq T_{\mathcal F}(\theta)\leq T(\theta)$, so restricting the class of test functions can only enlarge the zero set.

For $\Phi(\mathcal Z)$ in \eqref{eq:phi_class}, convexity of $\Mtheta$ and the minimax argument in the proof of Theorem~\ref{thm:main_sharpness} give:
$$
T(\theta)
=
\inf_{\mu\in\Mtheta}
\sup_{\phi\in\Phi(\mathcal Z)}
\Bigl(
\EE{\mu^*}{\phi}-\EE{\mu}{\phi}
\Bigr)
=
\inf_{\mu\in\Mtheta}
\norm{\mu^*-\mu}_{\mathrm{TV}}.
$$
Thus, $T(\theta)$ is the TV distance from $\mu^*$ to $\Mtheta$.
We do not use this representation for computation or inference.

The IPM representation shows that $T(\theta)$ is a supremum of linear expectation differences over the class of test functions.
Combined with the LIMR, this makes the objective bilinear in $(\phi, \gamma)$, leading to the finite LPs of Section~\ref{sec:computation}.
Since TV is an $f$-divergence, $T(\theta)$ is the distance from $\mu^*$ to $\Mtheta$ in an $f$-divergence.
Inference and model specification using $f$-divergences within the point-to-set geometry appear in, e.g., \textcite{KitamuraStutzer1997,kaidoMolinari2025}.
While $T(\theta)$ retains the connection to the divergence-based literature, it exploits the linear structure of both the IPM representation and the LIMR.
\end{remark}

\section{Computation}
\label{sec:computation}

Section~\ref{sec:identification} characterizes $\Thetaw$ through the zeros of $T(\theta)$, but evaluating $T(\theta)$ requires optimization over all test functions and all admissible input measures.
We construct an auxiliary finite linear program by restricting the input measure to finite support and retaining finitely many moment restrictions.
Because the program omits input values and moment restrictions, its value need not equal $T(\theta)$.
The program is auxiliary because $T(\theta)$ remains the computational target.

Section~\ref{subsec:computation_main} gives two conditions under which the auxiliary program is exact.
Column certification requires an optimal input measure of the program to satisfy the full family of moment restrictions, and row certification requires an optimal solution to satisfy its constraint at every input value, not only at the finitely many imposed.
Together they give $T_{\mathrm{LP}}(\theta)=T(\theta)$ (Theorem~\ref{thm:exact_computation}).
Section~\ref{subsec:computation_exchange} shows that a verdict on $\theta$ needs less: column certification alone bounds $T(\theta)$ from above, a bound on the row residual bounds it from below (Theorem~\ref{thm:enclosure}), and a column-and-row generation algorithm iteratively enlarges the finite sets, tightening the two bounds.

We maintain Assumptions~\ref{asm:measurable}--\ref{asm:linear_operator} and replace Assumption~\ref{asm:domination} by:
{\begin{asm}
\label{asm:finite_Z}
The observable space $\mathcal Z$ is finite and every subset of $\mathcal Z$ is measurable.
\end{asm}}\setcounter{asm}{4}
The assumption restricts only the observable space and imposes no restriction or topology on $\mathcal W$.
It makes $\Phi(\mathcal Z)=[0,1]^{\mathcal Z}$ a convex polytope, so the outer supremum in $T(\theta)$ is finite dimensional.
The remaining sources of infinite dimensionality are the input measure $\gamma$ and the possibly infinite family of moment restrictions defining $\Gamma_\theta$.
When both the outcomes and the covariates are discrete, point identification typically fails in nonlinear panels with fixed effects \parencite{Chamberlain2010}, so this is the setting in which a characterization of the identified set matters most.

\subsection{Main result}
\label{subsec:computation_main}

Fix $\theta\in\Theta$.
For a nonempty finite set $\mathcal W'\subseteq\mathcal W$ and finite sets
$\mathcal J'\subseteq\mathcal J$ and $\mathcal K'\subseteq\mathcal K$, define the \emph{restricted discrepancy}
\begin{align}
\TLD{\theta}{\mathcal W'}{\mathcal J',\mathcal K'}
&\equiv
\sup_{\phi\in[0,1]^{\mathcal Z}}
\inf_{\gamma_p\in\Gamma_\theta(\mathcal W',\mathcal J',\mathcal K')}
\Bigl(
\EE{\mu^*}{\phi}-\EE{\Lgen\gamma_p}{\phi}
\Bigr),
\label{eq:T_restricted}
\end{align}
where $\Gamma_\theta(\mathcal W',\mathcal J',\mathcal K')$ consists of finite-support input measures
$\gamma_p\equiv\sum_{w\in\mathcal W'}p_w\delta_w$, with probability weights
$p=(p_w)_{w\in\mathcal W'}$, satisfying the retained moment restrictions:
\begin{equation}
\label{eq:restricted_gamma}
\Gamma_\theta(\mathcal W',\mathcal J',\mathcal K')
\equiv
\left\{
\gamma_p:
\begin{aligned}
p&\in\R^{(\mathcal W')}_+, &&\textstyle\sum_{w\in\mathcal W'}p_w=1,\\
\EE{\gamma_p}{g_{1,j}(W;\theta)}&=0, &&\quad\text{for all } j\in\mathcal J',\\
\EE{\gamma_p}{g_{2,k}(W;\theta)}&\le0, &&\quad\text{for all } k\in\mathcal K'
\end{aligned}
\right\}.
\end{equation}
The definitions in~\eqref{eq:T_restricted} and~\eqref{eq:restricted_gamma} also apply to the full index sets $\mathcal J$ and $\mathcal K$.
Because $\mathcal W'$, $\mathcal J'$, and $\mathcal K'$ are finite, the inner infimum in~\eqref{eq:T_restricted} is a finite-dimensional linear program.
Dualizing it gives the auxiliary finite linear program
\begin{align}
\TLD{\theta}{\mathcal W'}{\mathcal J',\mathcal K'}
&=
\sup_{\substack{\phi\in[0,1]^{\mathcal Z},\ \zeta\in\R\\
u\in\R^{(\mathcal J')},\ v\in\R^{(\mathcal K')}_{+}}}
\left\{\EE{\mu^*}{\phi}-\zeta\right\}
\label{eq:TLD_restricted}
\\
\quad\text{subject to}\quad
q_{\phi,\theta}(w)
&\le
\zeta
+
\sum_{j\in\mathcal J'}u_j g_{1,j}(w;\theta) +
\sum_{k\in\mathcal K'}v_k g_{2,k}(w;\theta)
\quad\text{for all }w\in\mathcal W',
\nonumber
\end{align}
where the \emph{input-space payoff}
$
q_{\phi,\theta}(w)
\equiv
\EE{\Lgen\delta_w}{\phi}
=
\sum_{z\in\mathcal Z}\phi(z)\,K_\theta(\{z\}\mid w)
$
is the mean of $\phi$ implied by the model at input $w$.
For fixed $\mathcal W'$, $\mathcal J'$, and $\mathcal K'$, write
$$
T_{\mathrm{LP}}(\theta)\equiv\TLD{\theta}{\mathcal W'}{\mathcal J',\mathcal K'}.
$$

Restricting the input measure to $\mathcal W'$ weakly raises $T_{\mathrm{LP}}(\theta)$, whereas omitting moment restrictions weakly lowers it, so $T_{\mathrm{LP}}(\theta)$ need not bound $T(\theta)$ in either direction.
In~\eqref{eq:TLD_restricted}, each input value $w\in\mathcal W'$ contributes a constraint, or row, and each retained moment restriction contributes a multiplier, or column.

Column certification controls the error from omitting moment restrictions.
\begin{defi}[Column certification]
\label{def:column_certified}
The pair $(\mathcal J',\mathcal K')$ is \emph{column-certified for $\mathcal W'$} if $\Gamma_\theta(\mathcal W',\mathcal J,\mathcal K)$ is nonempty and
\begin{equation}
\label{eq:column_certificate}
\TLD{\theta}{\mathcal W'}{\mathcal J',\mathcal K'}
=
\TLD{\theta}{\mathcal W'}{\mathcal J,\mathcal K}.
\end{equation}
\end{defi}
Under column certification, imposing the omitted moment restrictions in $\mathcal J\setminus\mathcal J'$ and $\mathcal K\setminus\mathcal K'$ leaves the value at $\mathcal W'$ unchanged.
Equivalently, $(\mathcal J',\mathcal K')$ is column-certified for $\mathcal W'$ when some optimal $p^*$ of the inner program in~\eqref{eq:T_restricted} gives an input measure $\gamma_{p^*}$ that satisfies the full family of moment restrictions.
Because the admissible input measures supported on $\mathcal W'$ are a subset of $\Gamma_\theta$, column certification implies $T(\theta)\le T_{\mathrm{LP}}(\theta)$, leaving only the error from the omitted input values.
In Examples~\ref{ex:binary_parametric} and~\ref{ex:binary_sequential}, column certification can be established directly and requires no additional optimization (Section~\ref{subsec:computation_exchange}).

Write the slack of~\eqref{eq:TLD_restricted} at $w\in\mathcal W$ as
\begin{equation}
\label{eq:slack}
s(w;\phi,\zeta,u,v)
\equiv
\zeta
+
\sum_{j\in\mathcal J'}u_j g_{1,j}(w;\theta)
+
\sum_{k\in\mathcal K'}v_k g_{2,k}(w;\theta)
-
q_{\phi,\theta}(w),
\end{equation}
so that the constraint in~\eqref{eq:TLD_restricted} is
$s(w;\phi,\zeta,u,v)\ge0$ for all $w\in\mathcal W'$.
For an optimizer $(\phi^*,\zeta^*,u^*,v^*)$, define the \emph{row residual}
\begin{equation}
\label{eq:row_oracle_residual}
r^*
\equiv
\inf_{w\in\mathcal W}
s(w;\phi^*,\zeta^*,u^*,v^*) \leq 0.
\end{equation}
If $r^*=0$, the constraint holds at every input value in $\mathcal W$, not only on $\mathcal W'$.

\begin{defi}[Row certification]
\label{def:row_certified}
The support $\mathcal W'$ is \emph{row-certified for $(\mathcal J',\mathcal K')$} if the row residual~\eqref{eq:row_oracle_residual} equals zero at some optimizer of~\eqref{eq:TLD_restricted}.
\end{defi}

\begin{thm}[Exact computation]
\label{thm:exact_computation}
Under Assumptions~\ref{asm:measurable}--\ref{asm:linear_operator} and~\ref{asm:finite_Z}, if $(\mathcal J',\mathcal K')$ is column-certified for $\mathcal W'$ and $\mathcal W'$ is row-certified for $(\mathcal J',\mathcal K')$, then
\begin{equation}
\label{eq:exact_computation}
T_{\mathrm{LP}}(\theta)=T(\theta).
\end{equation}
\end{thm}
\begin{proof}
See Appendix~\ref{app:proofs}, where Theorem~\ref{thm:exact_computation} is obtained from Theorem~\ref{thm:enclosure} below by setting the row bound to zero.
\end{proof}

\subsection{Column-and-row generation}
\label{subsec:computation_exchange}

Under the conditions of Theorem~\ref{thm:exact_computation}, the auxiliary finite linear program gives a verdict on $\theta$: $\theta\in\Thetaw$ if $T_{\mathrm{LP}}(\theta)=0$, and $\theta\notin\Thetaw$ if $T_{\mathrm{LP}}(\theta)>0$.
Of the two certifications the theorem requires, row certification is the demanding one, and it is more than a verdict needs.
Inclusion needs only column certification, and for exclusion row certification can be replaced by a bound on the row residual.

The column side is straightforward: a finite, column-certified $(\mathcal J',\mathcal K')$ can be constructed directly, with no optimization, in all of our examples for any finite $\mathcal W'$ that supports an admissible input measure.
Supplemental Appendices~\ref{app:computation_example1} and~\ref{app:se_exact_pricing} do so for Examples~\ref{ex:binary_parametric} and~\ref{ex:binary_sequential}, and Lemma~\ref{lem:finite_column_reduction} in Supplemental Appendix~\ref{app:algorithm} gives a general construction whenever the moment inequalities are finite in number, allowing arbitrary moment equalities.
Column certification implies $T(\theta)\le T_{\mathrm{LP}}(\theta)$, so $T_{\mathrm{LP}}(\theta)=0$ settles inclusion.
A positive value does not settle exclusion, because the bound $0\le T(\theta)\le T_{\mathrm{LP}}(\theta)$ allows both zero and positive values.
Without a bound on the error from omitted input values, a column-certified discretization gives only this upper bound.

The row side is harder, because computing $r^*$ requires minimizing the slack over all of $\mathcal W$.
A verdict does not need $r^*$ itself, only a number $r_{\mathrm{low}}$ that is guaranteed to lie below it, which we call a \emph{row bound}.
In Example~\ref{ex:binary_parametric}, a row bound requires, for each covariate path, an upper bound on the model-implied mean of $\phi^*$ over the scalar fixed effect, which we obtain by bounding this mean on intervals of the real line and splitting the intervals until the bound is tight enough.
Under fixed effect--error dependence, no such calculation is needed (Additional Appendix~\ref{app:parametric_implementation}).
\footnote{For Example~\ref{ex:binary_sequential}, Supplemental Appendix~\ref{app:se_exact_pricing} obtains the row bound by a bilinear optimization over admissible conditional distributions.}
A row bound gives the lower bound $T_{\mathrm{LP}}(\theta)+r_{\mathrm{low}}$ on $T(\theta)$, so a positive value settles exclusion.

Column certification and a row bound together enclose $T(\theta)$:
\begin{thm}[Enclosure]
\label{thm:enclosure}
Under Assumptions~\ref{asm:measurable}--\ref{asm:linear_operator} and~\ref{asm:finite_Z}, let $(\mathcal J',\mathcal K')$ be column-certified for $\mathcal W'$.
Then~\eqref{eq:TLD_restricted} has an optimizer, and for the row residual $r^*$ of any optimizer, every $r_{\mathrm{low}}\le r^*$ satisfies
\begin{equation}
\label{eq:certified_enclosure}
\max\{0,T_{\mathrm{LP}}(\theta)+r_{\mathrm{low}}\}
\;\le\;
T(\theta)
\;\le\;
T_{\mathrm{LP}}(\theta).
\end{equation}
\end{thm}
\begin{proof}
See Appendix~\ref{app:proofs}.
\end{proof}

\begin{remark}[Row residual over input measures]
\label{rem:row_measures}
The proof of Theorem~\ref{thm:enclosure} uses the row inequality only through its integral against measures in $\Gamma_\theta$.
The theorem therefore holds with the row residual~\eqref{eq:row_oracle_residual} replaced by $\inf_{\gamma\in\Gamma_\theta}\EE{\gamma}{s(W;\phi^*,\zeta^*,u^*,v^*)}$, which is weakly larger.
This refinement is used under sequential exogeneity, in Example~\ref{ex:binary_sequential} and Supplemental Appendix~\ref{app:se_exact_pricing}.
\end{remark}

The enclosure at a given $(\mathcal W',\mathcal J',\mathcal K')$ may be too wide to decide, and we then enlarge the sets by column-and-row generation, which combines two standard devices in linear programming \parencite{Luebbecke2011,Muter2013}.
An input value with negative slack becomes a new row, and we then construct a column-certified $(\mathcal J',\mathcal K')$ for the enlarged support.
We re-solve the auxiliary program and repeat, as stated precisely in Algorithm~\ref{alg:master} of Supplemental Appendix~\ref{app:algorithm}.
A fixed discretization can serve as the initial support, which the algorithm then certifies and, if needed, refines.
Because the bounds of Theorem~\ref{thm:enclosure} are valid at every step, the algorithm may stop as soon as they determine whether $\theta\in\Thetaw$, or once they are narrow enough.

We impose no topology on $\mathcal W$, and hence no compactness or continuity, so the usual arguments for the convergence of a refined grid do not apply, and we make no general convergence claim.
Because the bounds are valid at every step, a verdict does not require convergence.
When the algorithm stops before the bounds decide, the parameter value is reported as undecided.
The implementation treats a value as zero only to linear programming accuracy, with no separate membership threshold (Additional Appendix~\ref{app:parametric_implementation}).
Model-specific compactness and continuity, or a finite reduction, may yield a general result through standard cutting-plane and semi-infinite programming results \parencite{Kelley1960,HettichKortanek1993}.

Figure~\ref{fig:enclosure_iterations} illustrates the algorithm in the baseline design of Section~\ref{subsec:numerical_example1}: the fixed effect probit model of Example~\ref{ex:binary_parametric} with $T=2$, a binary covariate correlated with the fixed effect, error terms independent of both, and $\beta_0=1$.
For two coefficient values only $1.6\times10^{-4}$ apart, on opposite sides of the boundary of the identified set, the algorithm reaches a verdict in eight iterations.
The figure shows the two enclosures, starting from a single fixed effect value at each covariate path, $\mathcal W'=\{(x,0):x\in\mathcal X\}$, with $\mathcal J'=\mathcal J=\mathcal K'=\mathcal K=\varnothing$.
Each iteration requires one solution of the auxiliary program and four one-dimensional interval branch and bound searches, one per covariate path, to bound the row residual (Additional Appendix~\ref{app:parametric_implementation}), and at the deciding iteration the program has $22$ rows, spanning nine distinct fixed effect values.
Across one hundred equally spaced coefficient values on $[0.01,2]$, the median number of iterations to a verdict is three, against eight for the two values above.
Deciding one value takes about ten milliseconds on a single core of an Intel Core i7-11370H.

\begin{figure}
    \centering
    \includegraphics[width=0.6\linewidth]{figures/s5-enclosure-trajectory}
    \caption{Bounds on $T(\beta)$ by iteration, at $\beta=1.050598$ (inside the identified set) and $\beta=1.050759$ (outside).
    Solid: $T_{\mathrm{LP}}(\beta)$. Dashed: $\max\{0,T_{\mathrm{LP}}(\beta)+r_{\mathrm{low}}\}$. Dotted: first decision.
    Logarithmic scale, with zero on the floor of each panel.}
    \label{fig:enclosure_iterations}
\end{figure}

\section{Inference}
\label{sec:inference}

We develop inference procedures for the finite output setting of Assumption~\ref{asm:finite_Z}.
We construct tests of candidate parameter values and, by test inversion, confidence sets that cover each point of the identified set $\Thetaw$ with asymptotic probability at least $1-\alpha$, uniformly over the data-generating process and the point under test.
The procedure uses the sample analog $T_n(\theta)$ of the ADF and a bootstrap approximation to the distribution of a penalized statistic that dominates $\sqrt nT_n(\theta)$ under the null.
Both may be obtained with the machinery of Section~\ref{sec:computation}: once the retained problem is column-certified and a certified lower bound on $T_n(\theta)$ has been computed, each bootstrap replicate requires only one finite linear program to obtain a conservative upper bound on the bootstrap statistic (Remark~\ref{rem:crit_val_LP}).
Supplemental Appendix~\ref{app:inference_computation} provides implementation details.

Under Assumption~\ref{asm:finite_Z}, the space of test functions $\Phi(\mathcal{Z}) = [0,1]^{\mathcal Z}$ is finite-dimensional, as in moment inequality problems, but our approach differs from standard procedures in two ways.
First, the same model-specific row bound calculation used for population identification evaluates the empirical distribution of output variables against the full set of input-space moment restrictions. This is possible because the constraint functions are always fully known to the researcher.
Second, our bootstrap critical values use a geometric contact-set penalty that focuses power on the separating hyperplanes most informative about whether $\theta$ belongs to the identified set.

\subsection{Setup and notation}

We denote the sample expectation by $\EE{n}{\cdot} \equiv n^{-1}\sumn$.
We observe $n$ i.i.d.\ realizations $\{Z_i\}_{i=1}^n$ from $\mu^*$.
Define the sample analog of the ADF as
\begin{align}
	\label{E:Tn_definition}
	T_n(\theta) \equiv \sup_{\phi \in \Phi(\mathcal{Z})}  \inf_{\mu \in \Mtheta} \left(\EE{n}{\phi} - \EE{\mu}{\phi}\right),
\end{align}
which replaces the population expectation in the ADF by its empirical counterpart.
Throughout this section, we subscript the population ADF of Section~\ref{sec:identification} by $\mu^*$ to make its dependence explicit:
\begin{align*}
	T_{\mu^*}(\theta) \equiv \sup_{\phi \in \Phi(\mathcal{Z})} \inf_{\mu \in \Mtheta} (\EE{\mu^*}{\phi} - \EE{\mu}{\phi}).
\end{align*}

We define the penalty function
\begin{align}
	\label{E:penalty}
	\eta_{\theta, \mu^*}(\phi) &\equiv \inf_{\mu \in \Mtheta} \left(\EE{\mu^*}{\phi} - \EE{\mu}{\phi}\right),
\end{align}
with empirical analog $\eta_{\theta, n}(\phi) \equiv \inf_{\mu \in \Mtheta} (\EE{n}{\phi} - \EE{\mu}{\phi})$.

This penalty function measures how well a test function $\phi$ discriminates between $\mu^*$ and the model set $\Mtheta$.
We likewise write $\Thetaw(\mu^*)$ for the identified set when the data are generated from $\mu^*$.
Note that $\eta_{\theta, \mu^*}(\phi) \le 0$ for all $\phi$ when $\theta \in \Thetaw(\mu^*)$.
By definition, $T_{\mu^*}(\theta) = \sup_{\phi \in \Phi(\mathcal{Z})} \eta_{\theta, \mu^*}(\phi)$ and $T_{n}(\theta) = \sup_{\phi \in \Phi(\mathcal{Z})} \eta_{\theta, n}(\phi)$.

Let $\mathbf{P} \subset \Pa(\mathcal{Z})$ denote a class of distributions over which we seek uniform results.
For $\mu^* \in \mathbf{P}$ and $\theta \in \Theta$, define the contact set
\begin{align}
\label{E:contact_set}
K_{\mu^*}(\theta) \equiv \left\{\phi \in \Phi(\mathcal{Z}): \eta_{\theta, \mu^*}(\phi) = T_{\mu^*}(\theta)\right\}.
\end{align}
The contact set $K_{\mu^*}(\theta)$ consists of the test functions attaining the population adversarial supremum.
When $\mu^*\in\Mtheta$, its elements define supporting directions at $\mu^*$; if $\mu^*$ lies in the relative interior of a set $\Mtheta$ that is full-dimensional relative to the probability simplex, these reduce to the constant functions.
When $\mu^*\notin\Mtheta$, the contact set consists of test functions attaining the maximal separating gap.

We say that a set of random variables $\{A_{n,\mu^*}: \mu^* \in \mathbf{P}\}$ is $o_p(1)$ uniformly in $\mu^*$ if, for all $c > 0$,
$
	\limsup_{n \to \infty} \sup_{\mu^* \in \mathbf{P}} \mu^*(\{|A_{n,\mu^*}| > c\}) = 0.
$
If instead
$
	\lim_{C \to \infty} \limsup_{n \to \infty} \sup_{\mu^* \in \mathbf{P}} \mu^*(\{|A_{n,\mu^*}| > C\}) = 0,
$
then the set is $O_p(1)$ uniformly in $\mu^*$.
Incorporating an additional supremum over $\theta$ extends these definitions to uniformity over $\theta$.

\subsection{Asymptotic distribution}

\begin{asm}[Random sampling]
	\label{A:finite_inference}
	The observations $\{Z_i\}_{i=1}^n$ are i.i.d.\ from $\mu^* \in \mathbf{P}$, where $\mathbf{P}$ is a collection of probability measures on the finite space $\mathcal{Z}$.
\end{asm}

In particular, $\mathbf{P}$ may be taken as all of $\Pa(\mathcal{Z})$: because $\mathcal{Z}$ is finite and test functions in $\Phi(\mathcal{Z})$ are uniformly bounded, no further regularity on $\mathbf{P}$ is required for the uniform results that follow.
\footnote{The argument does not divide by cell probabilities or require them to be bounded away from zero. Cells with zero probability produce degenerate coordinates of the empirical process, which is allowed.}
Define the centered empirical process $\mathbb{G}_{n,\mu^*}(\phi) = \sqrt{n}(\EE{n}{\phi} - \EE{\mu^*}{\phi})$.
Because $\mathcal{Z}$ is finite, the multivariate central limit theorem implies that $\mathbb{G}_{n,\mu^*}$ converges in distribution to a Gaussian process $\mathbb{G}_{\mu^*}$ indexed by $\phi \in \Phi(\mathcal{Z})$.

\begin{remark}[Finite-dimensional stochastic representation]\label{rem:remark_finitedimrep}
If $\mathcal Z=\{z_1,\ldots,z_{|\mathcal Z|}\}$, write $p$ and $\widehat p_n$ for the population and empirical probability vectors and set
$$
V_n\equiv\sqrt n(\widehat p_n-p)\in\mathbb R^{|\mathcal Z|}.
$$
Identifying $\phi$ with $(\phi(z_1),\ldots,\phi(z_{|\mathcal Z|}))'$, we have $\mathbb G_{n,\mu^*}(\phi)=\phi'V_n$.
Thus all stochastic variation in the empirical process is $|\mathcal Z|$-dimensional.
The representation also makes explicit that the parameter space for the observable distribution is the entire probability simplex.
We retain the empirical-process formulation below because the finite output space gives the required distribution-free entropy bound uniformly over this simplex.
\end{remark}

Let $\{Z_i^*\}_{i=1}^n$ denote a bootstrap sample drawn with replacement from the observed data, and define the bootstrap empirical process $\mathbb{G}_{n,\mu^*}^*(\phi) = \sqrt{n}(\mathbb{E}_n^*[\phi] - \EE{n}{\phi})$, where $\mathbb{E}_n^*[\cdot] = n^{-1}\sumn (\cdot)(Z_i^*)$ denotes the empirical bootstrap average.
We reserve $\mathbb E^*$ and $\mathbb P^*$ for conditional bootstrap expectation and probability:
$$
\mathbb P^*(A)\equiv
\mathbb P(A\mid Z_1,\ldots,Z_n),
\qquad
\mathbb E^*[X]\equiv
\mathbb E[X\mid Z_1,\ldots,Z_n].
$$

Let $d_{\mathrm{BL}}$ denote the bounded Lipschitz metric between probability distributions \parencite[\S 1.12]{VW1996}.
We write $\mathbb{G}_n \Rightarrow \mathbb{G}$ for random variables if $d_{\mathrm{BL}}(L_n,L)$ converges to $0$, where $L_n$ is the distribution of $\mathbb{G}_n$ and $L$ is the distribution of $\mathbb{G}$.
For a bootstrap quantity $\mathbb{G}_n^*$ whose distribution $L_n^*$ is itself random for each $n$, we write $\mathbb{G}_n^*  \overset{\mu^*}{\Rightarrow} \mathbb{G}$ if $d_{\mathrm{BL}}(L_n^*,L)$ converges in probability to $0$.

\begin{prop}
	\label{P:inf_finite}
	Let Assumptions~\ref{asm:measurable}, \ref{asm:moment_conditions}, \ref{asm:linear_operator}, \ref{asm:finite_Z}, and~\ref{A:finite_inference} hold with $\Gamma_\theta$ nonempty.
	For any sequence of positive penalty parameters $(\lambda_n)$,
	\begin{align}
		\sqrt{n} T_n(\theta) &\le \sup_{\phi \in \Phi(\mathcal{Z})} \left(\mathbb{G}_{n,\mu^*}(\phi) + \lambda_n \eta_{\theta,n}(\phi)\right) - \lambda_n T_n(\theta)\label{E:inf_bound}
	\end{align}
	for all $n$, $\mu^* \in \mathbf{P}$, and $\theta \in \Thetaw(\mu^*)$.
	If, in addition, $\lambda_n \to \infty$ and
	$\lambda_n=o(\sqrt n)$, then, for each fixed
	$\mu^* \in \mathbf{P}$ and $\theta \in \Theta$ such that $\Gamma_\theta\neq\varnothing$, the contact set $K_{\mu^*}(\theta)$ is nonempty and, as $n\to\infty$,
	\begin{align}
		\sqrt{n}(T_n(\theta) - T_{\mu^*}(\theta)) &\Rightarrow \sup_{\phi \in K_{\mu^*}(\theta)} \mathbb{G}_{\mu^*}(\phi), \label{E:inf_convergence}
	\end{align}
	and the bootstrap analog satisfies
	\begin{align}
		\sup_{\phi \in \Phi(\mathcal{Z})} \left(\mathbb{G}_{n,\mu^*}^*(\phi) + \lambda_n \eta_{\theta,n}(\phi)\right) - \lambda_n T_n(\theta) \overset{\mu^*}{\Rightarrow} \sup_{\phi \in K_{\mu^*}(\theta)} \mathbb{G}_{\mu^*}(\phi). \label{E:inf_bootstrap}
	\end{align}
\end{prop}

\begin{proof}
    See Appendix~\ref{app:proofs}.
\end{proof}

Proposition~\ref{P:inf_finite} bounds $\sqrt{n} T_n(\theta)$ by a penalized supremum that is consistently approximated by the bootstrap.
The empirical penalty term $\lambda_n\eta_{\theta,n}(\phi)$ is the sample counterpart of the population contact-set penalty and down-weights test functions $\phi$ outside $K_{\mu^*}(\theta)$, which do not enter the limit in~\eqref{E:inf_convergence}.
The $\sqrt n$ population drift localizes the sample criterion to the contact set in~\eqref{E:inf_convergence}.
Divergence of $\lambda_n$ produces the corresponding localization of the bootstrap criterion in~\eqref{E:inf_bootstrap}, while $\lambda_n=o(\sqrt n)$ makes estimation error in the empirical penalty asymptotically negligible.
The domination inequality~\eqref{E:inf_bound} is finite-sample and does not require either $\lambda_n\to\infty$ or $\lambda_n=o(\sqrt n)$.
These rate conditions are used only for the pointwise and bootstrap approximations in~\eqref{E:inf_convergence} and \eqref{E:inf_bootstrap}.

The requirement that $\Gamma_\theta$ is nonempty ensures that the penalty function $\eta_{\theta, n}$ and $T_n$ are well defined.
Nonemptiness of $\Gamma_\theta$ is easy to verify directly in many of our examples, and the feasibility check FEAS in Supplemental Appendix~\ref{app:algorithm} certifies it.
Parameter values for which FEAS certifies $\Gamma_\theta=\varnothing$ are not in the identified set and can be excluded before inference.

For every fixed null pair $(\mu^*,\theta)$ with $\theta\in\Thetaw(\mu^*)$, \eqref{E:inf_bound} holds for every $n$ and every positive $\lambda_n$.
Under $\lambda_n\to\infty$ and $\lambda_n=o(\sqrt n)$, the bootstrap statistic in \eqref{E:inf_bootstrap} consistently estimates the pointwise limiting distribution of this upper-bounding statistic.
This combination of an exact finite-sample null domination and a pointwise-consistent bootstrap approximation parallels the approach in \textcite{HongLi2018} \parencite[see also][]{fang2018}.

\begin{remark}[Bootstrap computation]
\label{rem:crit_val_LP}
Supplemental Appendix~\ref{app:inference_computation} shows that the bootstrap value $\sup_{\phi \in \Phi(\mathcal{Z})} \left(\mathbb{G}_{n,\mu^*}^*(\phi) + \lambda_n \eta_{\theta,n}(\phi)\right)$ used for estimating the asymptotic distribution of $T_n$ has an LP representation which exactly mirrors that of $T_n$, with a signed measure replacing the empirical measure.
This fact makes valid, if conservative, inference especially computationally tractable.
For example, under column certification of $(\mathcal{J}', \mathcal{K}')$ for $\mathcal{W}'$ obtained through equality of the feasible sets, as in Lemma~\ref{lem:finite_column_reduction}, Supplemental Appendix~\ref{app:inference_computation} bounds the left-hand side of~\eqref{E:inf_bootstrap} above by the direct bootstrap analog of $T_{\mathrm{LP}}(\theta)$ minus $\lambda_n$ times a certified lower bound on $T_n(\theta)$.
This upper bound is explicitly the value of a finite linear program and requires no row or column generation.
\end{remark}



\subsection{Confidence sets}

The corollary below gives tests whose inversion yields confidence sets with uniform asymptotic coverage of each point in the identified set.
It also considers power over the alternatives $\Theta_n^\Delta(\mu^*) \equiv \{\theta \in \Theta: T_{\mu^*}(\theta) \ge \Delta / \sqrt{n}\}$ for $\Delta > 0$.
\begin{cor}
	\label{C:inf_finite}
	Let the assumptions of Proposition~\ref{P:inf_finite} hold.
	Let $\varepsilon>0$ and let $\hat{c}_{1-\alpha}(\theta)$ denote the $1-\alpha$ quantile of the bootstrap distribution of $\sup_{\phi \in \Phi(\mathcal{Z})} (\mathbb{G}_{n,\mu^*}^*(\phi) + \lambda_n \eta_{\theta,n}(\phi)) - \lambda_n T_n(\theta)$.
	Then
	\begin{align}
		\liminf_{n \to \infty} \inf_{\substack{\mu^* \in \mathbf{P} \\ \theta \in \Thetaw(\mu^*)}} \PP{\mu^*}{\sqrt{n} T_n(\theta) \le \hat{c}_{1-\alpha}(\theta) + \varepsilon} \ge 1 - \alpha. \label{E:coverage}
	\end{align}
	On the other hand,
	\begin{align}
		\limsup_{\Delta \to \infty} \limsup_{n \to \infty} \sup_{\substack{\mu^* \in \mathbf{P} \\ \theta \in \Theta_{n}^\Delta(\mu^*)}} \PP{\mu^*}{\sqrt{n} T_n(\theta) \le \hat{c}_{1-\alpha}(\theta) + \varepsilon} = 0. \label{E:power}
	\end{align}
\end{cor}
\begin{proof}
See Supplemental Appendix~\ref{sec:app-additional-proofs}.
\end{proof}

The first result of Corollary~\ref{C:inf_finite} is that the critical values $\hat{c}_{1-\alpha}(\theta)+\varepsilon$ control asymptotic size uniformly over points in the identified set.
Define the confidence set
\begin{align}
    \mathrm{CS}_{1-\alpha} \equiv \left\{\theta \in \Theta: \sqrt{n} T_n(\theta) \le \hat{c}_{1-\alpha}(\theta) + \varepsilon\right\}.
\end{align}
Corollary~\ref{C:inf_finite} implies that $\mathrm{CS}_{1-\alpha}$ follows the paradigm of \textcite{im2004}: for every $\theta$ in the identified set, $\mathrm{CS}_{1-\alpha}$ covers $\theta$ with asymptotic probability at least $1-\alpha$ uniformly over both the parameter $\theta \in \Thetaw(\mu^*)$ and the underlying distribution $\mu^* \in \mathbf{P}$.
This is per-point rather than simultaneous coverage: the result does not require the confidence set to contain the entire identified set with probability at least $1-\alpha$.
Uniform coverage does not require uniform estimation of $K_{\mu^*}(\theta)$.
It follows from the exact null domination in \eqref{E:inf_bound}, the uniform empirical-process approximation implied by the distribution-free entropy bound on the finite output space, and the fact that the penalty-estimation error does not depend on $\theta$.

For alternatives with $T_{\mu^*}(\theta)\geq \Delta/\sqrt n$, the second result states that the worst-case asymptotic acceptance probability converges to zero as $\Delta\to\infty$.
The fixed positive tolerance $\varepsilon$ accommodates possible atoms in the limiting distribution $\sup_{\phi \in K_{\mu^*}(\theta)} \mathbb{G}_{\mu^*}(\phi)$.
For example, when $K_{\mu^*}(\theta)$ degenerates to the constant functions, the limit collapses to a point mass at zero.
Because uniform approximation of distributions does not in general imply uniform approximation of their quantiles at discontinuity points, the positive slack allows the bootstrap approximation to yield uniform coverage without imposing an anti-concentration condition.
Thus, for any fixed $\varepsilon>0$, Corollary~\ref{C:inf_finite} gives asymptotic coverage of at least $1-\alpha$.
Related positive-slack adjustments to bootstrap quantiles are used by~\textcite{andrews2013} and~\textcite{marcoux2024}.\footnote{The tolerance is fixed on the $\sqrt n T_n(\theta)$ scale and is therefore part of the rejection rule, rather than a numerical approximation tolerance.}

\begin{remark}[Specification test]
\label{rem:specification_test}
Corollary~\ref{C:inf_finite} yields a specification test as a
by-product.
If the model is correctly specified, then $\Thetaw(\mu^{*})$ is nonempty, and any fixed $\theta^{*}\in\Thetaw(\mu^{*})$ is covered with asymptotic probability at least $1-\alpha$.
Because $\{\mathrm{CS}_{1-\alpha}=\varnothing\}\subseteq\{\theta^{*}\notin\mathrm{CS}_{1-\alpha}\}$, it follows that $\limsup_{n\to\infty}\PP{\mu^*}{\mathrm{CS}_{1-\alpha}=\varnothing}\le\alpha$.
Rejecting correct specification when the confidence set is empty is therefore an asymptotically level-$\alpha$ test.
Tests of this form are studied for moment inequality models by
\textcite{bugni2015}, who show that such by-product tests are valid but generally conservative relative to dedicated specification tests.
\end{remark}
\section{Binary choice models with fixed effects}
\label{sec:numerical}

We report identified sets and inference in the models of Examples~\ref{ex:binary_parametric} and~\ref{ex:binary_sequential}.
Results for Example~\ref{ex:binary_interval} are in Additional Appendix~\ref{app:interval_results}.

\subsection{Parametric binary choice}
\label{subsec:numerical_example1}

We begin from the design in Figure~2 of \textcite{ChernValHahnNewey2013}:
\begin{equation}
\label{eq:cfhn_dgp}
Y_{t} = \one\{\beta_0\, X_{t} + A - V_{t} \ge 0\},
\qquad
V \mid (A,X) \sim H^{\otimes T},
\qquad
X_{t} = \one\{A - \eta_{t} \ge 0\},
\end{equation}
with $\eta_{t}$ and $A$ standard normal, the $\eta_t$ independent across $t$ and independent of $A$, so the binary covariate is correlated with the fixed effect.
We set $\beta_0=1$ and consider six known links $H$, each standardized to have mean zero and variance one: the normal (probit) link, and the logistic (logit), Laplace, Gumbel, uniform, and asymmetric truncated normal links.
The restriction $V\mid(A,X)\sim H^{\otimes T}$ makes the error terms independent across periods and independent of $(A,X)$.
The serial dependence specification of Example~\ref{ex:binary_parametric} keeps each error term's distribution given $(A,X)$ but leaves the dependence across periods unrestricted.
The fixed effect--error dependence specification instead keeps $V\mid X\sim H^{\otimes T}$ but lets the error terms depend on the fixed effect given $X$.
Each relaxation is strictly weaker than the baseline, and the two are not nested in each other.
For each link we generate data from the baseline~\eqref{eq:cfhn_dgp}, so the data satisfy all three specifications, and we hold the resulting distribution $\mu^*$ of $Z=(Y,X)$ fixed.
We then compute the identified set for $\beta$ at $T=2$ under each specification, with the baseline as the reference point and the comparison of interest between the two relaxations.
\footnote{Supplemental Appendix~\ref{app:computation_example1} derives the LPs and, where needed, the row bounds, and Additional Appendix~\ref{app:parametric_details} reports implementation details.}

\begin{figure}
    \centering
    \includegraphics[width=0.49\linewidth]{figures/1-beta-intervals-serial-dependence-certified}\hfill
    \includegraphics[width=0.49\linewidth]{figures/1-beta-intervals-fixed-effects-certified}
    \caption{Identified sets for $\beta$ at $T=2$.
    Solid: baseline.
    Dashed: serial dependence (left) and fixed effect--error dependence (right).
    Dotted: $\beta_0=1$.
    Open circle: point identification.}
    \label{fig:numerical_example1_beta}
\end{figure}
Under the baseline, Figure~\ref{fig:numerical_example1_beta} shows that $\beta$ is point identified at the logit link, the classical conditional logit result \parencite{cham1980,Chamberlain2010}, while the other five links leave sets of width $0.066$ (probit) to $0.27$ (uniform).
\footnote{Along a path from standardized probit to standardized logit, the coefficient sets contract toward $\beta_0$ (Additional Appendix~\ref{app:parametric_mixture}).}
Across all six links, serial dependence widens the identified set but leaves it bounded away from zero, while fixed effect--error dependence produces a set that contains zero.
In these designs, independence between the fixed effect and the error terms therefore carries more identifying content than serial independence.

In Figure~\ref{fig:numerical_example1_beta}, a coefficient value belongs to the identified set when its enclosure is the single point zero, and lies outside when the enclosure lies entirely above zero, so neither verdict rests on a membership threshold for the ADF (Theorems~\ref{thm:enclosure} and~\ref{thm:main_sharpness}).
The values left undecided occupy a band narrower than $2\times10^{-4}$ around each endpoint, well inside the width of the plotted line (Additional Appendix~\ref{app:parametric_implementation}).

\begin{figure}
    \centering
    \includegraphics[width=\linewidth]{figures/3-joint-beta-tau-by-T}
    \caption{Joint identified sets for $(\beta,\tau_{\mathrm{ASF}})$ in the probit design, at $T=2,3,4$, and $T=5$ in the right panel.
    Asterisk: true value.
    Each panel has its own scale for $\beta$.}
    \label{fig:numerical_example1_joint}
\end{figure}
For the probit link, we also compute joint identified sets for $(\beta,\tau_{\mathrm{ASF}})$ under all three specifications, where $\tau_{\mathrm{ASF}}\equiv\Pr(\beta+A-V_1\ge0)$ is the ASF at the counterfactual covariate value $\bar x=1$.
Under the baseline, the identified set for $\beta$ is narrower than $2\times10^{-3}$ by $T=4$.
Under serial dependence, the set at $T=4$ is still wider, for both $\beta$ and the ASF, than the baseline set at $T=2$ (Figure~\ref{fig:numerical_example1_joint}).
Under fixed effect--error dependence, contraction is even slower, and a fifth period is needed before the identified set for $\beta$ excludes zero, so $T=5$ periods identify the sign of the coefficient without any restriction on the dependence between the fixed effect and the error terms.
In this design, the rapid contraction of identified sets with the number of time periods that one expects in a panel model is specific to the baseline specification.

\begin{figure}
    \centering
    \includegraphics[width=\linewidth]{figures/4-beta-power-curves}
    \caption{Rejection rates for $\beta$ in the probit design, at nominal level $0.05$, for $n=10^3,10^4,10^5$ (light to dark).
    Shaded: identified set.
    Dotted: $\beta_0=1$ and the nominal level.
    Each panel has its own scale for $\beta$.
    Based on $200$ Monte Carlo samples with $499$ bootstrap draws each.}
    \label{fig:numerical_example1_power}
\end{figure}
We next illustrate inference on the coefficient using the penalized bootstrap test of Section~\ref{sec:inference}.
Because it rejects only when the bounds settle the comparison of the statistic with the critical value (Supplemental Appendix~\ref{app:inference_computation}), the test rejects no more often than the exact test on the same bootstrap draws.
We fix the penalty at $\lambda_n=2\sqrt{\log n}$, which satisfies the conditions of Proposition~\ref{P:inf_finite}.
In the probit design at $T=2$, only $14\%$ of units switch both covariates and outcomes, which suggests that the nominal sample sizes overstate the information available.
In Figure~\ref{fig:numerical_example1_power}, the Monte Carlo rejection rate never exceeds the nominal level at any evaluated coefficient inside the identified set, in any specification, at any sample size.
Power outside the set rises sharply with the sample size.
At $n=10^5$, the rejection rate reaches $1$ away from the identified set in all three specifications, and remains low only near its boundary.
Additional Appendix~\ref{app:parametric_inference_implementation} records the implementation and reports the rejection curves for $c\in\{1/2,1,2,5\}$ in $\lambda_n=c\sqrt{\log n}$.
Constants at or below $2$ hold the nominal level in all three specifications at every sample size, while $c=5$ over-rejects inside the identified set, most at $n=10^3$.
Power increases with $c$, sharply at $n=10^3$ and little at $n=10^5$.

\subsection{Sequential exogeneity}
\label{subsec:numerical_semiparametric}

We now drop the known error distribution of Section~\ref{subsec:numerical_example1} and impose only the sequential exogeneity (SE) of Example~\ref{ex:binary_sequential}, which permits feedback from past outcomes into future covariates.
Existing identification results for the coefficient under SE require a parametric error distribution or a special regressor.
\footnote{Supplemental Appendix~\ref{app:semiparametric_binary_choice} derives the program for SE.
Additional Appendix~\ref{app:se_details} carries the results behind the computed sets.}
As a benchmark we carry along conditional stationarity (CS), which equates the error distributions across periods given the whole covariate path,
\begin{equation}
\label{eq:cond_stat}
V_t\mid (X,A) \overset{d}{=} V_1\mid (X,A), \qquad t=2,\ldots,T,
\end{equation}
whose identified set for the coefficient is characterized by the rank inequalities of \textcite{Manski1987}, which are sharp \parencite{pakesPorterMomentInequalities,gaoIdentificationNonlinearDynamic2024}.
Because CS conditions on future covariates, it rules out feedback.
A CS structure becomes an SE structure once the fixed effect is redefined as the pair $(A,X)$, so the CS set is contained in the SE set.

Both are evaluated on one family of designs,
\begin{equation}
\label{eq:cs_dgp}
Y_{t} = \one\{(t-1)+\beta_0\, X_{t} + A - V_{t} \ge 0\},
\qquad
V_t \mid (A,X_1,\ldots,X_t) \sim \mathcal N(0,1) \text{ i.i.d.},
\end{equation}
with $T\in\{2,3\}$, $\beta_0=0.4$, covariates on $\{0,1,2\}$ with $X_1$ uniform, and $A=\gamma_0+\gamma_1 X_1+\sigma_a\,\xi$ where $\xi\sim\mathcal N(0,1)$ is independent of $X_1$ and the error terms, and, for $t\ge1$,
\begin{equation}
\label{eq:cs_covariate}
X_{t+1}=
\begin{cases}
\min\{\max\{X_{t}+2Y_{t}-1,\,0\},\,2\} & \text{with probability }\rho,\\
X_{t} & \text{with probability }(1-\rho)\phi,\\
\text{a uniform draw on }\{0,1,2\} & \text{with probability }(1-\rho)(1-\phi).
\end{cases}
\end{equation}
The first line is feedback, the second persistence, and the third a fresh uniform draw, so $\rho=\phi=0$ makes the covariate independent of the past.
In both specifications the coefficient on the time trend is known and equal to one.
The baseline sets $(\gamma_0,\gamma_1,\sigma_a)=(-1,0.5,1)$ and $\rho=\phi=0$.
Because each error is drawn independently of the past, SE holds at every design in the family, while CS holds exactly when $\rho=0$.
Figure~\ref{fig:numerical_beta_comparison} varies feedback $\rho$, persistence $\phi$, and the dependence $\gamma_1$ of the fixed effect on the initial covariate, one at a time.

\begin{figure}
    \centering
    \includegraphics[width=\linewidth]{figures/4-identification-beta-knobs}
    \caption{Identified sets for $\beta$ under SE (blue) and CS (grey), at $T=2$ (top) and $T=3$ (bottom), varying one design parameter per column.
    A band reaching the frame is unbounded on that side.
    The CS band ends where the CS set becomes empty.
    Dotted horizontal: $\beta_0=0.4$.}
    \label{fig:numerical_beta_comparison}
\end{figure}

Feedback breaks CS but not SE.
Under both CS and SE, the coefficient enters only through how it orders the period indices $t-1+\beta x_t$ along each covariate path, so the endpoints of both sets are values at which two indices coincide (Additional Appendix~\ref{app:se_details}).
At every design we compute, the CS set is either empty or the same interval $(-1/2,1/2)$, at both horizons.
The set is empty once $\rho$ exceeds about $0.62$ at two periods and about $0.12$ at three.
A third period leaves the nonempty CS set unchanged but lowers the feedback strength at which it becomes empty.

Because SE holds at every design, its identified set is never empty.
At the baseline the set is $(-1/2,+\infty)$ at two periods and $(-1/2,1)$ at three, so the sign of the coefficient is not identified and the third period supplies an upper bound.

Unlike the CS set, the SE set varies with each design parameter.
At two periods, feedback widens it to $(-1,+\infty)$ and, at the strongest feedback, replaces the lower bound with an upper bound at $1$, so the set is not monotone in $\rho$.
\footnote{The feedback move also shifts the marginal distribution of the later covariates away from uniform, so the $\rho$ axis varies the composition of covariate paths together with feedback.}
Greater persistence only widens the set: from $\phi=0.75$ the two-period set is all of $\R$.
Only $\gamma_1$ shrinks the set, adding an upper bound at $1$ once $\gamma_1\le-1/2$.
Adding a third period can only shrink the identified set.
Under strong feedback it recovers a finite lower bound that the two-period set has lost, and under persistence it recovers a lower bound that falls from $-1/2$ to $-2$ as $\phi$ rises.