Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
102,724 characters · 20 sections · 28 citation commands
Inference on Auctions with Weak Assumptions on Information
\baselineskip=1.5\baselineskip
A recent literature in robust mechanism design studies the following question: given a game \vsedit{(e.g. an auction)}, what are the possible outcomes - such as welfare or revenue - that arise under different information structures? This literature is motivated by {\it robustness}, i.e., characterizing outcomes that can occur in a given game under weak assumptions on information (See Bergemann2013). For example, in auction models, in addition to specifying the details of the game in terms of bidder utility function, and bidding rules, one needs to specify the information structure (what players know about the state of the world and the information possessed by other players) to be able to derive the Nash equilibrium.
In particular, \vsedit{ in the Independent Private Values (IPV)} setting, players know their \vsedit{value for the item}, which is assumed independent from other player \vsedit{values}, and receive no further information about their opponents' values. \vsedit{The latter typically yields a unique equilibrium outcome. However, different assumptions on what signals players have about opponent values prior to bidding in the auction, lead to different equilibrium outcomes.} \vsedit{Auction data} rarely contain information on what bidders knew and what their information sets included, \etedit{and given that this information leads to different outcomes}, it would be interesting to analyze what can happen when we relax the independence assumption in such auctions \etedit{by allowing bidders to know some information about their opponents' valuations}. Bergemann2016c (BBM) examine exactly this question in an auction game and provide achievable bounds on various outcomes\vsedit{, such as the revenue of the auction as a function of auction fundamentals like the distribution of the common value}.
In this paper, we address the following econometrics question, \etedit{which is the reverse of the one posed by BBM above}: given an i.i.d. sample of auction data (independent copies of bids from a set of auctions), what can we learn about auction fundamentals, such as the distribution of values, when we make weak assumptions on the information structure? We maintain throughout that players play according to a Bayesian Nash equilibrium (BNE) but allow these bidders to have different information structures \etedit{in different auctions}\vsedit{, i.e. receive different types of signals prior to bidding.} In particular, we use these observations (bids and other observables) from a set of independent auctions to construct {\it sets} of valuation distributions that are consistent with both the observed distribution of bids and the \vsedit{known auction rules} maintaining that players are Bayesian. We exploit the robust predictions in a given auction \`{a} la BBM to conduct econometrically robust predictions of auction fundamentals given a set of data; i.e., robust economic prediction leads to robust inference.
\ \
Key to our approach is the characterization of sharp sets of valuation distributions (and other functionals of interest) via computationally attractive procedures. This is a result of the equivalence between a particular class of {\it Bayes Correlated Equilbria} (or BCE) and Bayes Nash Equilibria (or BNE) for a similar game with an arbitrary information structure. It is well known that BCE can be computed efficiently since they are solutions to linear programs (as opposed to BNE which are hard to compute). \vsedit{Moreover, exploiting a result of Bergemann2016 (see also aumann1987), we show that there is an equivalence between the set of fundamentals that obey the BCE restrictions and the fundamentals that obey the BNE constraints under {\it some} information structure.} This equivalence is the key to our econometrics approach. The formal statistics program that ensues is one where the sharp set satisfies a set of linear equality and inequality constraints. If we knew the true distribution of bids, then identifying the parameters would be a simple computational problem of solving a linear program. We do not observe the true bid distribution, but this distribution can be estimated consistently from the observed data. Thus, we use the estimated bid distribution to solve for an estimate of the identified set. We are also able to characterize sampling uncertainty to obtain various notions of confidence regions \etedit{covering the identified with a prespecified probability}.
In addition to learning about auction primitives, we show how our approach using data on auctions can be used to construct identified sets for auction welfare measures, seller surplus and other objects. Information on auction primitives, along with these measures obtained using our procedures that combine data with the theory, can be used to guide future market designers to better study particular auction setups (or use our results in other markets).
\vsedit{Importantly, we address the problem of counterfactual estimation: what would the revenue or surplus have been had we changed the auction rules? We formulate notions of informationally robust counterfactual analysis and we show that such counterfactual questions can also be phrased as solutions to a single linear program, \etdelete{that} simultaneously captur\etdelete{es}\etedit{ing} equilibrium constraints in the current auction as well as the new target auction. We show that even without recovering the information structure from the data, an analyst can perform robust counterfactual analysis and answer the following question: under an arbitrary information structure in the current auction which produced the data at hand, what is the best and worst value of a given quality measure (e.g. welfare, revenue) in the new target auction under an arbitrary information structure? Thus we can get estimates of the upper and lower bounds of a given quantity in the new auction design, in a way that is robust to information structure and without the need of recovering it.} \etedit{Hence, this approach to inference requires minimal informational assumptions on the data generating process (DGP) that generated the existing data set and minimal informational assumptions on the counterfactual auction that a market designer is contemplating to run. }
\
The closest work to our paper is \citealt*{Bergemann2016c}, where the authors provide worst-case bounds on the revenue of a common value first price auction as a function of the distribution of values. Their approach does not use the bid distribution as input, unlike our approach which obtains an estimate of the bid distribution from the data. The main approach in BBM is to show that the revenue cannot be too small since at BCE, no player wants to deviate to any other action and so players do not want to deviate to a specific type of a deviation which is the following: conditional on your bid, deviate uniformly at random above your bid (upwards deviation). Hence, the bound on the mean they provide uses a {\it subset} of the set of best response deviations that are allowed \vsedit{so as to bound the bid of a player as a function of his value. \etdelete{and} \etedit{In drawing a connection between the equilibrium bid and the value,} \etedit{this bound} is by definition loose \etdelete{in drawing a connection between the equilibrium bid and the value} (and can be very loose - bound twice as large as identified set - as we show in an example in Appendix (ref)).} Given data, we are able to learn the bid distribution and hence are not constrained to look at only these bid-distribution-oblivious upwards deviations. We can instead compute an optimal \vsedit{deviating} bid for this given bid distribution and use the constraint that the player does not want to deviate to this \vsedit{distribution-tailored action.} This allows us to bound the unobserved value of the player as a function of the observed bid \etdelete{This approach} leading to a {\it sharp characterization} of \vsedit{auction fundamentals} using the data. Also, the approach taken to inference in this paper is deliberately conservative in that we try to make weak or no assumptions on information while maintaining Nash behavior. This is in the same spirit as HaileTamer who study the question of inference in English auctions under minimal assumptions and derive estimable bounds on the distribution of bidder valuations.
Another paper that uses a similar insight of studying the econometrics of games with weak information is the recent work of Magnolfi2016 on inference in entry game models. The approach used there, though similar in motivation, does not transfer easily to studying general auction mechanisms.
The paper is organized as follows. Section (ref) introduces the problem and provides formal definitions of the objects of interest. We then state our identification results given an i.i.d. set of data on bids. This identification is constructed via a linear program where we show how various constraints (such as symmetry, parametric restrictions, etc) can be incorporated. \vsedit{We also show how computing sharp sets for the expected value of moments of the fundamentals amounts to solving two linear programs and how robust counterfactual analysis of some metric function, with respect to changes in the auction, can also be handled in a computationally efficient manner.} \vsedit{We then provide two example applications of the general setup: one for common value auctions (Section (ref)) and another for \vsedit{private value} auctions (Section (ref)). Section (ref) provides our estimation approach for constructing confidence intervals on the estimated quantities from sampled datasets using sub-sampling methods and finite sample concentration inequality approaches.} Section (ref) examines the finite sample performance of the large linear program using a set of Monte Carlo simulations. These show adequate performance in IPV and Common Value (CV) setups. Section (ref) illustrates our inference approach using auction data from OCS wildcat oil auctions and show how the statistical algorithm can be used to derive bounds on valuation distributions. \vsdelete{Section ? provides further extensions and Section ? concludes.} \etedit{Finally, the Appendix contains results on the sharpness of the BBM bounds, and bounds on the mean of the valuation in common value auctions with different smoothness assumptions. }
We consider a game of incomplete information among $n$ players. There is an unknown payoff-relevant state of the world $\theta\in \Theta$. This state of the world enters directly in each player's utility. Each player $i$ can pick from among a set of actions $A_i$ and receives utility which is a function of the payoff-relevant state of the world $\theta$ and the action profile of the players $\ensuremath{{\bf a}}\in A\equiv A_1\times\ldots\times A_n$: $u_i(\ensuremath{{\bf a}};\theta)$. This along with a prior on $\theta$ (defined below) will represent the game structure that we denote by $G$, as separate from the information structure which we will define next.
Conditional on the state of the world each player receives some minimal signal $t_i\in T_i$. The state of the world $\theta\in \Theta$ and the vector of signals $\ensuremath{{\bf t}}\in T\equiv T_1\times\ldots\times T_n$ are drawn from some joint measure\footnote{The setup is general in that the set $T$ is unrestricted.} $\pi \in \Pi\subseteq \Delta(\Theta\times T)$. The signals $t_i$'s can be arbitrarily correlated with the state of the world. We denote such signal structure with $S$. This defines the game $(G,S).$\footnote{\vsedit{We will assume throughout that the set of $n$ players is fixed and known. If in the data the set of participating players varies across auctions \etedit{and players do not necessarily know the number or bidders}, then we can consider the superset of all players and simply assign a special type to each non-participating player. Conditional on this type, players always choose a default action (e.g. bidding zero in an auction). Then dependent on whether players observe the number of participants before submitting an action can be encoded by whether the players receive as part of their default signal, whether some player's type is the special non-participating type. In our auction applications we will assume that players do not necessarily observe the entrants in the auction before bidding.}}
We consider a setting where prior to picking an action each player receives some additional information in the form of an extra signal $t_i'\in T_i'$. The signal vector $t'=(t_1',\ldots,t_n')$ can be arbitrarily correlated with the true state of the world and with the original signal vector $t=(t_1,\ldots,t_n)$. We denote such augmenting signal structure with $S'$ and the set of all possible such augmenting signal structures with $\ensuremath{{\cal S}}'$. This will define a game $(G, S')$. Subsequent to observing the signals $t_i$ and $t_i'$, the player picks an action $a_i$. A Bayes-Nash equilibrium or BNE in this game $(G, S')$ is a mapping $\sigma_i:T_i\times T_i'\rightarrow \Delta(A_i)$ from the pair of signals $t_i,t_i'$ to a distribution over actions, for each player $i$, such that each player maximizes his expected utility conditional on the signals s/he received.
A fundamental result in the literature on robust predictions (see Bergemann2013,Bergemann2016), is that the set of joint distributions of outcomes $a\in A$, unknown states $\theta$ and signals $t$, that can arise as a BNE of incomplete information under an arbitrary additional information structure in $(G, S')$, is equivalent to the set of Bayes-Correlated Equilibria, or BCE in $(G, S)$. \etedit{So, every BCE in $(G,S)$ is a BNE in $(G,S')$ for some augmenting information structure $S'$.} \etdelete{there is an augmenting information structure $S'$ such that} We give the formal definition of BCE next.
\ \ \
\ \ An equivalent and simpler way of phrasing the Bayes-correlated equilibrium conditions is that:
We state the main result in Bergemann2016 next. \ \ \
The robustness property of this result is as follows. The set of BNE for $(G,S')$ (think of an auction with unknown information) is the same as the set of BCE for $(G,S)$ where $S'$ is an augmented information structure derived from $S$. So, we will not need to know what is in $S'$, but rather we could compute the set of BCE for $(G,S)$ and the Theorem shows that for each BCE, there exists a corresponding information structure $S'$ in $\ensuremath{{\cal S}}'$ and a BNE of the game $(G,\ensuremath{{\cal S}}')$ that implements the same outcome.
We consider the question of inference on auction fundamentals using data under weak assumptions on information. In particular, assume we are given sample of observations of action profiles $a^1,\ldots,a^N$ from an incomplete information game $G$. Assume that we do not know the exact augmenting signal structure $S^1,\ldots,S^N$ that occurred in each of these samples where it is implicitly maintained that $S^i$ can be different from $S^j$ for $i \neq j$, i.e., the signal structure in the population is drawn from some unknown mixture. Also, maintaining that players play Nash, or that $a^t$ was the outcome of some Bayes-Nash equilibrium or BNE under signal structure $S^t$, We study the question of inference on the distribution $\pi$ of the fundamentals of the game. Under the maintained assumption, we characterize the sharp set of possible distributions of fundamentals $\pi$ that could have generated the data. This allows for policy analysis within the model without making strong restrictions on information.
A similar question was recently analyzed in the context of entry games by Magnolfi2016, where the goal was the identification of the single parameter of interaction when both players choose to enter a market. In this work we ask this question in an auction setting and attempt to identify the distribution of the unknown valuations non-parametrically.
The key question for our approach is to allow for observations on different auctions to use different (and unobserved to the econometrician) information structures, and, given the information structures, that different markets or observations on auctions, to use a different BNE. Given this equivalence of the set of Bayes-Nash equilibria under some information structure and the set of Bayes-correlated equilibria, and given that the set of BCE is convex, allowing for this kind of heterogeneity is possible. Heuristically, given a distribution over action profiles $\phi\in \Delta(A)$ (which is constructed using the data), there exists a mixture of information structures and equilibria under which $\phi$ was the outcome. The process by which we arrived at $\phi$ is by first picking an information structure from this mixture and then selecting one of the Bayes-Nash equilibria for this information structure. This is possible if and only if there exists a distribution $\psi(\theta,t,a)$, that is a Bayes-correlated equilibrium and such that $\sum_{\theta,t}\psi(\theta,t,a)=\phi(a)$ for all $a\in A$. Again, we start with the elementary information structure $S$ and maintain that observations in the data are expansions of this information structure. So, for a given market $i$, any BNE using information structure $S^i$ is a BCE under $S$. The data distribution of action is a mixture of such BNE over various information structures and hence it would map into a mixture of BCE under the same $S$. Since the set of BCE under $S$ is convex, any mixtures of elements in the set is also a BCE. So then intuitively, the set of primitives that are consistent with the model and the data is the set of BCEs, $\psi(\theta,t,a)$ such that $\sum_{\theta,t}\psi(\theta,t,a)=\phi(a)$ for all $a\in A$. \newline To conclude, the convexity of the set of BCEs allows us to relate a distribution of bids from an iid sample to a mixture of BCEs. This is possible since the distribution of bids uses a mixture of signal structures, which essentially coincides with a mixture of BCEs, itself another BCE by convexity. We summarize this discussion with a formal result.
\ \ \
\ \
\ \ \
Finally, we provide next the main engine that allows for construction of the observationally equivalent set of primitives that obey model assumptions and result in a distribution on the observables that match that with the data. We state this as a Result.
\ \
\ \ \
\etedit{The above result is generic, in that it handles general games with generic states of the world $\theta.$ In particular, it nests both standard private and common value auction models and provides a mapping between the distribution of bids $\phi$ and the set of feasible distribtions over signals and $\theta.$} An iid assumption on bids along with a large sample assumption allow us to learn the function $\phi(.)$ (asymptotically). So, given the data, we can consistently estimate $\phi$. This is a maintained assumption that we require throughout. Given $\phi$, the above result tells us how to map the estimate of $\phi$ to the set of BNE that are consistent with the data and are robust to any information structure that is an expansion of a minimal information structure $S$. Suppose we assume that both $t$ and $\theta$ take finitely many values (an assumption we maintain throughout), then a joint distribution on $(t,\theta)$, $\pi(t,\theta)$ is consistent with the model and the data if and only if it solves the above {\it linear program}. The LP formulation is general, but in particular applications, it is possible to use parametric distributions for the $\pi$. In addition, it is possible for the above LP to allow for observed heterogeneity by using covariate information whereby this LP can be solved accordingly \vsedit{(see an example such adaptation in the common value Section (ref)).} Though the above LP holds in general (and covers both common and private values for example), we specialize in the next Sections the above LP to standard cases studied in the auction literature, mainly common values an private values models.
\vscomment{I think we need to move the corollary that we can calculate the identified set of the expected value of any function of $\theta,t$, by simply solving two linear programs, which will yield the upper and lower bounds of this “projection” of the identified set. This will make it way more general in this general setup. I also think that we might want to move the counter-factuals up here, as they can be phrased in the general setup and don't need to be a common value or private value auction. I implemented these two sections. Let me know what you think. \etedit{I AGREE!} }
{
In the case of non-parametric inference where we put no constraint on the distribution of fundamentals, i.e. $\Pi=\Delta(\Theta\times T)$, observe that the sharp set is linear in the density function of the fundamentals $\pi(\cdot,\cdot)$. Therefore, maximizing or minimizing any linear function of this density \etdelete{variables} can be performed via solving a single linear program. This implies that we can evaluate the upper and lower bounds of the expected value of any function $f(\theta, t)$ of these fundamentals, in expectation over the true underlying distribution. The latter holds, since the expectation of any function with respect to the underlying distribution is a linear function of the density. \etedit{We state this as a corollary next.}
}
\vsedit{Suppose that we wanted to understand the performance of some other auction when deployed in the same market, with respect to some \etedit{objective or} metric: $F:\Theta\times T\times A\rightarrow \ensuremath{\mathbb R}$ that is a function of the unknown fundamentals and the action vector. Examples of such metrics in single-item auctions could be social welfare: $F(\theta, t, a)=\sum_{i=1}^n \theta_i x_i(a)$ or revenue, i.e. $F(\theta, t, a)=\sum_{i=1}^n p_i(a)$, where $\theta_i$ is the value of player $i$ for the item at sale, $x_i(a)$ is the probability of allocating to player $i$ under action profile $a$ and $p(a)$ is the expected payment of player $i$ under action profile $a$.}
\vsedit{We are interested in computing an upper and lower bound on this metric under this new auction which has different utilities $\tilde{u}_i(a;\theta)$ and under any Bayes-correlated equilibrium which would map into a BNE with an augmenting information structures. This is important since it allows us to obtain bounds on welfare or other metrics in a new environment \etdelete{under general information structures}. The welfare bounds computed in this manner will inherit the robustness property in that they will be valid under all information structures. }
\vsedit{This is straightforward in our setup since computing a sharp identified set for any such counter-factual can be done in a computationally efficient manner, in both the common value setting and in the correlated private value setting. The upper bound of the counter-factual \etdelete{boils down to} \etedit{can be obtained} using the following linear program, that takes as input the observed distribution of bids in our current auction, the metric $F$, and the primitive utility form $\tilde{u}=(\tilde{u}_1,\ldots,\tilde{u}_n)$ under the alternative auction. We state this result in the next Theorem.}
\vsedit{
}
\etedit{Note that getting sharp bounds on welfare measures for example using the above Theorem does not require one to infer in a prior step the distribution over the primitives. Rather, the above procedure provides sharp bounds on this welfare measure using a linear program.}
\etedit{We specialize the above results to important classes of auctions.} We begin with the common value model in which the game of incomplete information $G$ is a single item common value auction. In this case the unknown state of the world is the unknown common value of the object $v$, which we assume to take values in some finite set $V$. Moreover, we initially assume that the minimal information structure is degenerate, i.e., players receive no minimal signal about this unknown common value\footnote{Other constraints on the initial signals are allowed and here we take the degenerate signal for simplicity.}. The signal set $T$ becomes a singleton and is irrelevant. Thus we will denote with $\pi\in \Delta(V)$ the distribution of the unknown common value, which is the parameter that we wish to identify. This is a particularly simple model to illustrate the structure of the LP approach and showcase the flexibility of our methods. So, in this particular model, we want to learn the distribution of the state of the world which is the common valuation distribution.
Prior to bidding in the auction, the players receive some signal which is drawn from some distribution; this signal can be correlated with the unknown common value and with the signals of his opponents. We wish to be ignorant about which information structure realized in each auction sample and want to identify the sharp identified set for $\pi$. Moreover, we will assume that the players' bids take values in some discrete set $B$ and players play a BNE. The characterization of the identified set for $\pi$ in this model is stated in the Theorem below.
\ \
\ \ \
Observe, that in this setting, the latter linear program is also linear in $\pi$. Thus we get that the sharp identified set is a convex set and is defined as the set of solutions to the above linear program, where $\pi$ is also a variable.
This observation also allows us to easily infer upper and lower bounds on any linear function of the unknown distribution $\pi$. This is stated next as a Corollary to the above Theorem.
\ \ \
Observe that the latter linear expressions are simply: $\mathbb{E}_{v\sim \pi}[f(v)]$ and so the above shows that we can compute in polynomial time upper and lower bounds of any moment of the unknown distribution of the common value.
Also, note that the Corollary above shows that to do set inference with respect to any moment of the unknown distribution, we do not need to discretize the space of probability distributions and enumerate over all probability vectors, checking whether they are inside the sharp identified set. Rather we can just solve the above LP.
Essentially this observation says that we can easily compute the support function $h(z; \Pi_I(\phi))$ of the identified set $\Pi_I(\phi)$ at any direction $z$, by simply solving a linear program. In Section ((ref)), we provide an upper bound on the mean valuation distribution when the latter is continuous. This upper bound is derived in terms of the observed bids distribution.
\ \ \
\ \ \
\ \ \ \
We now consider the case of a private value single item auction. In this case the (unknown) state of the world is the a vector of private values $\ensuremath{{\bf v}}=(v_1,\ldots,v_n)\in V^n$. We assume that these private values come from some unknown joint distribution $\pi\in \Delta(V^n)$. Moreover, we initially assume that players know at least their own private value. Thus the (minimal) signal set $T_i$ is equal to $V$ and moreover, we have that conditional on a value vector $v$, $t_i=v_i$, deterministically. Since the signal is a deterministic function of the unknown state of the world, we will again denote with $\pi\in \Delta(V^n)$ the distribution of the unknown valuation vector, which is the parameter that we wish to identify. Here, each player first draws a valuation (as an element of the state of the world), and then each player's own valuation is revealed to the player through a signal. After that, a signal is further revealed before players play a BNE given this signal.
In this setting the sharp identified set is again slightly simplified. The result is stated in the next Theorem. \vskip .1in
Observe, that in this setting, the latter linear program is also linear in $\pi$. Thus we get that the sharp identified set is a convex set and is defined as the set of solutions to the above linear program, where $\pi$ is also a variable. Note also here that no assumption is made on the correlation between player valuation. This result allows for the recovery of valuation distribution with arbitrary correlation (and general signaling structures). As a special case of the above, we study next the IPV model of auctions.
\paragraph{Independent Private Values.} The situation becomes more complex if we also want to impose an extra assumption that the distribution of private values is independent. In that case, we have the extra condition that $\pi$ must be a product distribution which is a non-convex constraint. For instance, if we want to assume that the value of each player is independently drawn from the same distribution $\rho$, and hence $\rho$ is what we wish to identify, then we also have the extra constraint that:
Adding this constraint into the above LP, makes the LP non-convex with respect to the variables $\rho(v)$ (even though checking whether a given $\rho$ is in the identified set, is still an LP). Thus in this case we cannot compute in polynomial time upper and lower bounds on the moments of the distribution $\rho$ using the above LP.
However, we make the following observation which simplifies the constraints of the LP: we note that conditional on a player's valuation and on a bid profile $b$, the effect of a deviation $u_i(\ensuremath{{\bf b}};v_i)-u_i(b_i',\ensuremath{{\bf b}}_{-i};v_i)$ is independent of the values of opponents, in a private value setting. Thus we can re-write the best response constraint as:
where $x_i(v_i|b) = \Pr[V_i=v_i | B=b]$, where $V_i$ is the random variable representing player $i$'s value and $B$ is the random variable representing the bid profile at a BCE.
Then we can formulate the consistency constraints, by simply imposing a constraint per player, i.e. if $\rho_i$ is the distribution of player $i$'s value, then it must be that:
These are constraints that are still linear in $\rho_i(v_i)$. We state the LP as a corollary next.
\ \
\ \ In particular if we assume that player's are symmetric, i.e. $\rho_i(v) = \rho(v)$, then we can compute upper and lower bounds on any moment $E[f(v)]$ of the common value distribution:
A by-product of this analysis is that the linear program allows us to test for symmetry in the independent private values model. In particular it is not clear that when we assume that all marginals are the same, then the LP is feasible. Thus by checking feasibility of the LP we can refute the assumption of symmetric independent private values.
\ \ \
\etedit{ Note here that given that we allow the augmenting signals to be arbitrary correlated, we are not able to infer any information about the joint distribution of valuation, such as correlation, given a bid profile. } \vscomment{The observation below is wrong. We can still infer something about correlation, due to correlation in bids. What we cannot infer at all is what is the correlation in values conditional on a bid profile. This could still be arbitrary and there constraints only on the marginal distribution of each value conditional on a bid profile.} \vsdelete{\paragraph{Non-identifiability of Correlation in Values.} The above discussion shows a stronger point: the BCE constraints for the case of private values and under the assumption that players observe their own valuation, yields no constraints on the correlation of player valuations, but only constraints the marginal distributions of each player's value. Thus assuming that the observed distribution is the outcome of a BCE, does not allow for identification of the correlation among player's valuations. The reasoning behind this result is that it in a Bayes-correlated equilibrium it is impossible to distinguish between the case where valuations are genuinely correlated and where bids are correlated through signaling. Since a BCE allows for arbitrary such signaling, no inference can be made on the underlying correlation of values. We can learn the marginal distributions but the model under general information structures contains no information on the copula without further restrictions on the model.}
\vscomment{If we make the assumption that the equilibrium is strictly monotone, i.e. that there is a one-to-one mapping between bids and values, then we should be able to extend this to the correlated values case, and claim that marginals are uniquely identified in the correlated values case and the formula which I think is the same as some prior work on correlated first price auctions is robust to informational assumptions. The paper that does exactly that without informational robustness is Li2002. So we would be stating something like the results of Li2002 for estimating private values in the affiliated private values setting is robust to }
Generally in the independent private values model, if we allow the bids and valuations to be continuous, we show here that the bidder specific valuation distributions are all identified even allowing for general information structures. This is important since the information here is allowed to be correlated but yet the valuation distributions $\rho_i$ of each player $i$ (under independence) are point identified.
Denote with $G_{-i}(\cdot|b_i)$, the CDF of the maximum other bid at a BCE, conditional on a bid $b_i$ of player $i$, and let $g_{-i}(\cdot|b_i)$ be the density. These are observable in the data. Now the utility of a player from submitting a bid $b_i'$ conditional on being recommended by BCE to player $b_i$ and observing a value of $v_i$ is:
Since, $b_i$ is maximizing the above quantity, by the best-response constraints, then we get that the derivative of the utility with respect to $b_i'$ has to be equal to $0$ at $b_i'=b_i$. The latter implies:
We summarize our result in the following Theorem.
Hence, if we know the population bid distribution and if we assume that it is continuous and admits a density, then given the bid of a player we can invert and uniquely identify his value. Thus we can write the CDF of the value distribution of each player as a function of the observables.
\vsedit{A similar result was shown for the first price IPV model (without signals) in guerre2000 and later also generalized to the affiliated private values setting by Li2002. Interestingly, the inversion formula that we arrived to in the independent private values setting, but under robustness to the information structure, is the same as the inversion formula of Li2002 under the correlated private value setting. It is interesting to note that under the assumption that the bid of a player is strictly increasing in his value at any Bayes-Correlated equilibrium, then we can re-do the analysis in this section in the more general correlated private values setting and show that we can invert the value of a player from the observed correlated bid distribution. The inversion formula is identical to Equation (ref) and identical to the inversion derived in Li2002. This essentially shows that in terms of inference, robustness to information structure is in some sense equivalent to robustness to correlation in values. }
Finally, as in the common values case, it is possible to adjust the LP to account for observing only the winning bids in the Private Values case.
In general, we do no have access to the distribution $\phi(\ensuremath{{\bf b}})$. Instead, we have access to $N$ i.i.d. observations \vsedit{$\omega_1,...,\omega_N$ from said distribution ; let $\omega_t(\ensuremath{{\bf b}}) = 1\{\omega_t = \ensuremath{{\bf b}}\}$. The sampled distribution is then given by
} In this section, we develop techniques for estimating the sharp identified set using samples from $\phi(\ensuremath{{\bf b}})$. We showcase our techniques for the CV setup for simplicity. Also, we focus on inference on the identified set for $\pi$ which is the object of inference here. Often times, we parametrize the distribution of the common values, and so inference will be on the identified set for the vector of parameters\footnote{It is possible to construct the CI for the unknown parameters -rather than the identified set by inverting test statistics. }. \etedit{We start with finite sample approaches to constructing confidence regions for sets using concentration inequalities. These seem to be the first application of such results on using such inequalities to settings with partial identification\footnote{Finite sample inference results are particularly attractive in models with partial identification since standard asymptotic approximations are not typically uniformly valid especially in models that are close to/or are point identified.}. We also show how existing set inference methods can also be used to construct confidence regions. }
We explore first the question of set inference using finite sample concentration inequalities. This allows us to obtain a confidence set for the identified set where the coverage property holds for every sample size. In addition, we highlight tools from the concentration of measure literature applied to inference on sets in partially identified models.
As a reminder, given a probability distribution over bids $\phi(\ensuremath{{\bf b}})$, the sharp set of compatible equilibria $\phi(v,\ensuremath{{\bf b}}) = \phi(\ensuremath{{\bf b}}) x(v|\ensuremath{{\bf b}})$ is the set $\Theta_I$ of joint probability distributions satisfying:
for the proper $x(.|.)$. See the statement of Theorem ((ref)) above. Let $\pi(.)$ be the vector characterizing the distribution of the common value. Let $\{F_j(x;\pi, \phi): j\in M\}$, denote the negative of the best-response and density constraints, associated with the BCE LP in Theorem ((ref)). Then in the population $\pi$ is feasible iff:
Then we have that the identified set is defined as:
Now we consider a finite sample analogue. Observe that $F_j(x;\pi, \phi) = \mathbb{E}[f_j(x;\pi,\omega)]$, where expectation is over the random vector $\omega$ and the function $f_j(\cdot;\pi,\omega)$ takes the form:
for some triplet $(i, b_i^*, b_i')$ for the case of best-response constraints and similarly for density consistency constraints. Then we consider the sample analogue:
Observe that due to the linearity of $F_j$ with respect to $\phi$, we can re-write: $F_j^N(x; \pi)=F_j(x;\pi, \phi_N)$. One can then define the estimated identified set by analogy, i.e., replacing $\phi(b)$ with $\phi_N(b):$
for some decaying tolerance constant $\sigma^N$ (which can be set to zero).
\vsedit{Because the pdf of a non-parametric distribution is a very high-dimensional object, inference on it will require many samples. Hence, we will instead focus on two more structured inference problems. In the first one we are interested in inferring the identified set for {\it a moment} of the distribution and in the second we make assume that the said pdf is known up to a finite dimensional parameter and infer the identified set of these lower dimensional parameters. One can in principle recover non-parametric inference by simply making the parameters be the values of the pdf at the discrete support points, albeit at a cost in the sample complexity.}
\vsedit{
} We begin with the non-parametric setting and show how to construct inference on the identified set of any moment function $m:V\rightarrow [-H,H]$ of the common value distribution that is valid in finite samples. \vsdelete{Even though the identified set on the whole distribution $\pi$ is very high dimensional, we can still answer probabilistic questions for any moment of the distribution via the Bayesian Bootstrap.} Observe that the identified set for any moment $m:V\rightarrow [-H,H]$ is an interval $[L,U]$ defined by:
where $x$ in both optimization problems is ranging over the convex set of conditional distributions, i.e. $x(\cdot|\ensuremath{{\bf b}})\in \Delta(V)$, which we omit for simplicity of notation. We can then define their finite sample analogues as:
Where $F_j^N(x)$ and $\phi_N$ are given in Equations (ref) and (ref) respectively.
The following result gives finite sample high probability bounds on the coverage of the interval $[L^N(\sigma^N)-\epsilon^N, U^N(\sigma^N)+\epsilon^N]$. As a reminder, we use $n$ to designate the number of bidders, $N$ is sample size, and $|B|$ is the number of support points for bids.
For the parametric case, we assume that the distribution $\pi(v)$ is parametric of the form $\pi(v,\theta)$ for some finite parameter set $\Theta$. In the parametric setting we need to augment the constraint set $M$, apart from containing the best-response constraints, to contain the parametric form consistency constraints:
These can also be written of the form $F_j(x;\phi)\leq 0$ for some function $F_j(x;\phi)=\mathbb{E}[f_j(x;\omega)]$.\footnote{Simply add one constraint of the form $\pi(v,\theta) - \sum_{v\in V} \phi(\ensuremath{{\bf b}})\cdot x(v|\ensuremath{{\bf b}})\leq 0$ and one of the form $\sum_{v\in V} \phi(\ensuremath{{\bf b}})\cdot x(v|\ensuremath{{\bf b}}) - \pi(v,\theta)\leq 0$.} We overload notation and let $M$ be this augmented set of constraints. Then the parameter of interest is $\theta$ and the identified set for $\theta$ takes the form:
and its sample equivalent
for some decaying tolerance constant $\sigma^N$ (which can be set to zero).
The next Theorem provides finite sample high probability coverage bounds for $\Theta_I$.
\etcomment{Is there a way to say something about how loose the bound in (44)? just heuristically? Either way it is fine...} \paragraph{Sample Variance Based Bounds.} We now show how to improve upon the prior analysis and getting tolerance variables $\sigma^N$ whose leading $1/\sqrt{N}$ term does not depend on $H$, but rather on the sample variance of the constraints. We will modify the finite sample identified set $\Theta_I^N(\sigma^N)$ defined in Equation (ref), as follows: instead of using a uniform tolerance $\sigma^N$ for all the constraints, we will define tolerance for each constraint differently based on its sample variance:
Where $\ensuremath{\text{Var}}_N(X)$ for a random variable $X$, denotes the sample variance: $$\ensuremath{\text{Var}}_N(X) = \frac{1}{N(N-1)}\sum_{1\leq t < t'\leq n} (X_t - X_t')^2.$$ The extra variance modification, which is reminiscent of sample variance penalization in empirical risk minimization Mauer2009 and optimism in bandit algorithms Audibert2009 , will allow us to set $\sigma^N$ to be of order $O(H/N)$ rather than $O(H/\sqrt{N})$.
The latter theorem is asymptotically an improvement over Theorem (ref), when the variances of the constraints are not as large as their worst case bound of $4H^2$, which is essentially what is assumed in Theorem (ref). However, one drawback of this theorem is that the new estimated set $\hat{\Theta}_I^N(\lambda, \sigma^N)$ requires solving more than linear programs. In particular for every parameter $\theta\in \Theta$ we need to solve an optimization problem of the form:
The negative part is a concave function of $x$, as it can be thought of as a norm of a vector that is a linear function of $x$. Thus this is a non-convex minimization problem. Thus even though it offers a statistical improvement, it is at the expense of computational efficiency.
Here, we use existing set inference approaches and adapt them to the problem at hand. The first approach exploits the mapping between the reduced form parameter (the distribution of bids) and the identified set via the linear program, and the second uses subsampling approaches to set inference.
In this section, we describe approaches to inference on the identified set via a Bayesian Bootstrap for both parametric and nonparametric models. We start with the nonparametric case. As a reminder, the identified set in this nonparametric case is
This is a case of a separable problem where knowing $\mathbf \phi^*= \{\phi^*(b_1), \ldots, \phi^*(b_B)\}$ we can solve for the identified set via the mapping:
for some decaying tolerance constant $\sigma^N$ (which can be set to zero). Inference on parameter sets in separable models is analyzed in (KlineTamer) where (posterior) probability statements related to the identified set\footnote{Examples of such statements are: the probability that the identified set belongs to a particular set (e.g, when $\Theta_I$ is an interval, the probability that this interval lies in $[a,b]$) or the probability that a particular vector belongs to the identified set, etc.} can be computed using the mapping between the reduced form parameters, $\phi(\mathbf b)$ in this case, and the structural parameters $\pi.$ Intuitively, for every draw $s$ from the posterior for $\phi$ (this is a multinomial distribution and so obtaining draws from the multinomial is simple using the Bayesian bootstrap\footnote{For example, under a limiting uninformative Dirichlet prior for $\phi(\mathbf b)$, the posterior for this $\phi(\mathbf b)$ approaches the Dirichlet posterior $Dir(n_1, \ldots, n_B)$ where $n_j = \sum_{i=1}^N 1[b_i = b_j]$. Therefore, using results that connect the Gamma and Dirichlet distributions, a draw from given posterior can be approximated by a weighted Bootstrap with Gamma weights. See also ChamberlainImbens.}) we can solve via ((ref)) for a “copy” of the implied identified set $\Pi_N^{s}$ by solving the LP above. Using this procedure, we can get a sequence $\{\Pi_N^{s}\}_{s=1}^S$ that we can use to answer probability statements about the identified set $\Pi_I.$ The computational constraint in this nonparametric problem is that the parameter of interest $\pi$ is a vector of probabilities with $|V|$ support points. So, the bigger $|V|$ is the larger the number of parameters the more difficult it is to solve for the identified set via ((ref)) above. \etedit{Hence, with the finite support condition on the bids standard (Bayesian) inference on $\phi(\mathbf b)$ can be easily mapped into inference on the set $\Pi_I$. } \etedit{The exact same bootstrap procedure can be used to provide (posterior) probability statement on identified intervals $[L,U]$ that are defined through the mapping in (ref) above. This is relevant if one is interested in linear functional of the distribution. Inference on such objects reduces the computational burden substantively.}
\ \ \
In the parametric case, we assume that the distribution function of $v$ belongs to a parametric class that is known up to a finite dimensional parameter $\theta$ and so the LP in Theorem (ref) is modified by setting $\pi(v)=\pi(v,\theta)$. Then the parameter of interest is $\theta$ and the identified set for $\theta$ takes the form:
and so the separable mapping between $\mathbf \phi^*$ and $\theta$ takes the form again of
for some decaying tolerance constant $\sigma^N$ (which can be set to zero).
Here again, we can use the approach in (KlineTamer) to answer probability statements on $\Theta_I$. The computational step here is much easier than the nonparametric case as now solving for the identified set for every draw from the posterior for $\phi$ is a lower dimensional problem. The Bayesian bootstrap is used to draw vectors of bid probabilities from the multinomial distribution for bids.
In addition to this Bayesian bootstrap approach, we also provide methods based on subsampling next.
\etedit{Subsampling methods can be used to conduct inference on the identified set in a large sample framework. In particular}, let $\theta$ be the parameter of interest. Again, our problem is isomorphic to one where the objective function $Q(\theta)$ is as follows:
with a the corresponding sample analogue
We first re-state the above random variable as a minimax problem over convex and compact sets. Let $p\in \Delta_M$ lie on the simplex on $M$ constraints. Then $Q^N(\theta)$ is equivalently defined as:
and
\etedit{where $Q^*$ is zero for all feasible $\theta's.$} In addition, following the approach in cht, we get a confidence region for the identified set $\Theta_I =\{\theta: Q^*(\theta) \leq 0\}$ by studying the asymptotic distribution\footnote{Alternatively, we can define the objective function $\tilde Q^N(\theta) = [Q^N(\theta)]_+$ which is the positive part of $Q^N(\theta)$. The identified set now is the minimizer of $\tilde Q^N(\theta)$ and then similar approaches to cht can be used to construct a CI based on subsampling.} of $$ \mathcal C_N= \sup_{\theta \in \Theta_I} \min_{x} \sup_{p} \sum_{j\in M} p_j F^N_j(x;\theta) $$
Let $\tau_N(1-\alpha)$ be the $(1-\alpha)-$quantile of $\mathcal C$ where $\mathcal C$ is the nondegenerate limit of $\sqrt N \mathcal C_N.$ Then, define the $C_N(1-\alpha)$ as follows $$ C_N(1-\alpha) = \{ \theta: Q^N(\theta) \leq \tau_N^+(1-\alpha)\}$$ where $\tau_N^+(1-\alpha)= \max(\tau_N(1-\alpha),0).$ Notice here that the event $\Theta_I \subseteq C_N(1-\alpha)$ is equivalent to the event $\mathcal C_N \leq \tau_N^+(1-\alpha)$. The next Theorem states the result.
\ \
\ \ \ The asympotic distribution $\mathcal C$ above can be characterized using for example results from Shapiro2009ch5 and sufficient conditions for nondegeneracy are given there which requires second moments to hold (these hold in our case trivially since we maintain the assumptions that bids take finitely many values).
Finally, to be able to feasibly implement the above approach, one needs to get a value for the cutoff $\tau_N(1-\alpha)$. It is clear here that the standard nonparametric bootstrap may not work. Even if the asymptotic distribution is normal\footnote{The distribution of $\mathcal C$ is likely to be the supremum of a Gaussian process since the optimal value of the LP is generally not unique.}, we may not be able to estimate its variance because it depends on the number of binding constraints. Given the nondegeneracy of the limit, one approach is to use subsampling to compute the cutoff. This can be accomplished by getting $m$ subsamples of the data such that $m/N \rightarrow 0$ as $N \rightarrow \infty.$ For every subsample, we compute $\mathcal C^m_N$ using a preliminary estimate $\widehat \Theta_I$ and using the sequence $\{\mathcal C^m_N\}_{m=1}^{M}$ to get its upper $(1-\alpha)$ quantile. This will result in an estimate of $C_N(1-\alpha)$. We can then use that as our new $\widehat \Theta_I$ and iterate one or two times. This approach was implemented in the Monte Carlo section below and in the empirical application and it provided adequate results.
In this section, we examine the techniques of Section (ref) to obtain an estimated set on simulated data. We assume in all the simulations that the values and bids are taken from discrete sets, respectively $V$ and $B$. In the whole section, we fix the following parameters: i) The maximum value/bid $H$ is given by $H = 20$. ii) The set of possible bids is given by $B = \{0, \ldots,H\}$. The set of possible common values is $V = B = \{0, \ldots, H\}$. We generate equilibrium observations as follows: First, we generate a density $f(v)$ of valuations with support $V$. We consider the four following densities: \paragraph{Normal density.} $f_n(.)$ is the density of a normal random variable with mean parameters $\mu = 4$ and standard deviation parameter $\sigma = 1$, discretized and truncated to have support $V$, i.e.
\paragraph{Poisson density.} $f_p(.)$ is the density of a Poisson distribution with parameters $\lambda = 4$, truncated inside $V$, i.e.:
\paragraph{Binomial density.} $f_b(.)$ is the density of a binomial random variable with probability $p = 0.2$ and number of draws $n = H = 20$, i.e.:
\paragraph{Geometric density.} $f_g(.)$ is the density of a geometric random variable with probability $p = 0.2$ truncated to have support $V$, i.e.:
For each of those densities, we then generate one distribution of equilibrium bids $\phi$, through solving the BCE linear program for the given distribution of values and with variables $\phi(\ensuremath{{\bf b}})$ rather than $\pi(v)$. We then generate $N$ samples of bid vectors from $\phi$, to generate the observed empirical bid distribution $\phi_N$.
In the non-parametric setting, since we cannot directly formulate the estimated set of the variance (as it cannot be written as $E[m(v)]$ for some $m$), we obtain \etedit{a superset} for the identified set as follows: First obtain upper and lower bounds $E_{min}, E_{max}, E^{2nd}_{min}, E^{2nd}_{max}$ for the first and second moments respectively. Then set the bounds on the standard deviation by the conservative ones: $ E^{2nd}_{min} - E_{max}^2 \leq Var \leq E^{2nd}_{max} - E_{min}^2$. \etedit{A computationally more tedious procedure of estimating the identified set for the variance is to first get the identified set for the (discrete) distribution of $V$ and then using that we can “solve” for a bound on the variance\footnote{For example, for every “draw” from the identified set for the distribution of $V$, we can obtain a variance. We can repeat the process to build the identified set for the variances.}. }
In the parametric case, bounds on the variance can be obtained directly from recovering bounds on the possible parameters for the distribution we consider, and tight bounds on the parameters imply tight bounds on the second and higher order moments. All simulations in the parametric case are therefore presented in terms of identified and estimated sets on the parameters of the distribution of values.
In this section, we compare the identified sets when we use and do not use parametric knowledge of the distribution of $v$. Figure (ref) uses the true distribution $\phi$ and shows:
We do so for the four different distributions of the common value (Gaussian, Poisson, binomial and geometric) mentioned above. We remark that the non-parametric linear programs seems to recover the mean of the common value accurately in all cases. However, the bounds obtained on the second moment/standard deviation of the distribution of common values are far from being tight.
Figure (ref) compares the identified set for the true distribution to the estimated set using Hoeffding with $\delta = 0.10$, as described in Section (ref) for the Gaussian density function, for a number of samples $N \in \{10^3, 10^4, 10^5, 10^6\}$. We see that as $N$ grows larger, the estimated set grows smaller and smaller and closer to the true identified set.
In this section, we characterize the identified and estimated sets in the parametric case for the four distributions described above. For the Gaussian distribution, we provide figures of the identified parameters in the $(\mu,\sigma)$ space. For the other, $1$-dimensional parameter distributions, we provide intervals for the parameter of the chosen distribution. We consider two different techniques to determine the estimated set:
In all figures, the brown region is the true identified set while the union of the brown and the green region is the estimated set. Figure (ref) plots the identified and estimated set when using a Gaussian distribution for the common value and tolerances determined through Hoeffding with $90$ percent confidence ($\delta = 0.10$), for the number of samples $N \in \{10^3, 10^4, 10^5, 10^6\}$. We remark that the number of samples needs be large (of the order of at least $10^5$) for Hoeffding to perform well.
Figure (ref) plots the identified and estimated set when using a Gaussian distribution for the common value and tolerances determined through subsampling and quantile estimation for the $90$, $95$ and $99$ percent quantiles, as seen in Section (ref); we only use $N = 100$ samples for the bid distribution in the three figures and $k = 50$ subsamples of size $s = N/4 = 25$. We note that the quantile estimation technique covers the true identified set fairly sharply even though $N$ is only equal to $100$ and hence should be preferred to Hoeffding when small amounts of data are available.
Figure (ref) gives the true identified interval for the parameters of the Poisson, binomial and geometric distributions and the estimated interval using subsampling for quantile estimation, for the $90$, $95$ and $99$ percent quantiles -- see (ref) -- for a bid distribution sampled with $N = 100$. We see that for the binomial distribution, the estimated set is very close to the true identified set. While the recovered sets for the Poisson distribution are not as sharp as the recovered sets for the binomial distribution, they still restrict the space of possible parameters in a reasonable way: the recovered interval is $[1.5,6.5]$ while the space of possible parameters is $[0,20]$. However, the recovered sets for the geometric distribution contains more than half of the possible parameters, which is unsatisfying. A reason for this comes from the fact that the geometric distribution has a much higher variance than a Poisson or binomial with comparable means, hence there is a lot of variability across different subsamples that can lead to large tolerances.
Using $N = 500$ samples and $k = 50$ subsamples of size $s = N/4 = 125$, we obtain Figure (ref). We can see that the estimated sets for the binomial and geometric distributions are now significantly tighter.
Finally, Figure (ref) plots the identified and estimated set when using a Gaussian distribution for the common value and tolerances determined through subsampling and quantile estimation for the $90$, $95$ and $99$ percent quantiles, as Figure (ref). However, we now use $N = 500$ samples for the bid distribution in the three figures and $k = 50$ subsamples of size $s = N/4 = 125$. We note that the estimated set now almost coincides with the true identified set.
In this section, we illustrate our framework for common value auctions on real data. We use the Outer Continental Shelf (OCS) Auction Dataset that was used in the seminar work of Hendricks and Porter (see hendricksporter). The dataset contains bidding information on $3036$ tract auctions in Louisiana and Texas. In particular, for each auction, our dataset contains the acreage of the tract and the total bid of each participant in the auction. We assume that the bidders participate in a first price common value auction, where the value is defined per acre; our goal is i) to show that indeed, the bidders' behavior in the data can be explained by a common value auction (via the testable restriction of whether the estimated identified set is empty), and ii) to recover the first and second order moments of the distribution of said value. This is under weak assumptions on information in that the framework allows bidders in different auctions to know more information about the environment.
\paragraph{Pre-processing of data:} The dataset contains $3036$ auction with varying number of players. We consider $2$-player common value auctions, hence we only keep the entries in the dataset that contain exactly $2$ bidders; there are $584$ such auctions. We model the two bidders as being the same over the $584$ auctions, and assign bidders' identities to be $1$ or $2$ uniformly at random in each auction. Many of the bids we have are zero and Figure (ref) plots the distribution of bids; it has mean $\$991.48$ and standard deviation $\$1825.43$.
We assume the distribution of the common value per acre has bounded finite support $V = \{0, \ldots, H\}$. We renormalize the bids per acre to be in $[0,\lceil \frac{H}{2} \rceil]$ -- we pick $H/2$ because the common value could have a distribution whose support goes beyond the observed bids --, and discretize the set of bids to be $\{0, \ldots, \lceil \frac{H}{2} \rceil \}$; we do so by rounding each renormalized bid in each auction to the closest integer. We remark that the dataset contains a few outliers whose bid per acre is significantly higher than in all other auctions; we therefore delete the auctions that contain bids over threshold $t = \$20000$. We further assume that the distribution of the common value is given by a truncated normal distribution that takes discrete values in $\{0,\ldots,H\}$, exactly as described in the simulations of Section (ref), and parametrize all optimization problems we solve accordingly.
\paragraph{Results:}
In all figures, the value of $(\mu,\sigma)$ are given as the values in dollars instead of the corresponding discretized and renormalized value, for the sake of comparison with the $\$20,000$ threshold and the corresponding maximum value of $\$40,000$. Figure (ref) plots a heat map of the estimated set as a function of the chosen tolerance in the $(\mu,\sigma)$ space for two different values of $H$. Each color on the heat map corresponds to a tolerance level, and the mapping from tolerance levels to colors is given by the colorbar on the right of each figure. The color that is assigned to any given $(\mu,\sigma)$ pair corresponds to the minimum level of tolerance that the analyst needs to add to the equilibrium constraints for $(\mu,\sigma)$ to belong to the estimated set; therefore, the heat map shows how the estimated set grows as the tolerance picked by the analyst increases.
Figure (ref) plots the estimated set when using the tolerance determined by the subsampling approach of Section (ref), in green using 95% level, and compares it to the estimated set using the minimum tolerance for which the estimated set is non-empty, in brown, for $H = 400$. No matter what method is used for picking the tolerance and determining the corresponding confidence intervals, the region that is obtained cannot be possibly smaller than the minimum tolerance set unless it is empty; in this sense, the minimum tolerance set is the best possible set that one could hope to obtain through any method for determining tolerances. We remark that i) the estimated set for $(\mu,\sigma)$ remains small relatively to the upper bound of $\$20,000$ on the bids and of $\$40,000$ on the maximum common value, and ii) the estimated set is not much bigger compared to the best estimated set we could hope to obtain, indicating that on top of covering the true identified set with high probability, our techniques cannot possibly overestimate the size of the identified by too much. We can see from the confidence regions that the mean of the common values varies from zero 0 to around \$4000 while the standard deviation varies from close to zero to 6000. It is possible to estimate these means as functions of covariates and hence to allow for observed heterogeneity.
The above plots the mean and variance of a normal density that is then truncated. Since the truncated normal has a different mean and variance then the underlying normal, we plot in Figure (ref) the minimum tolerance and the estimated sets as a function of this overall mean and standard deviation of the distribution of common values that we identify, instead of the parameters $\mu, \sigma$ of the truncated normal distribution. We obtain these plots by computing a mapping from $(\mu,\sigma)$ pairs to (mean, standard deviation) pairs, and plot the image of Figure (ref) by said mapping in the (mean, standard deviation) space. We note that the estimated set is fairly small, indicating that our approach identifies the first two moments of the true, underlying distribution of the common value in an accurate fashion, despite only having access to a limited number of samples.
We provide a framework for inference on auction fundamentals without making strong restrictions on information. Using data, we use the recent results in theory to characterize the identified set for these primitives by exploiting the linear programming structure of the set of Bayesian correlated equilibria. We have several applications of the approach mainly to common value and private value auctions and other scenarios. Our results can also be used by mechanism designers in that the data allows us to restrict the domain of signal/state of the world distribution, which would lead to sharper mechanisms. We also provide approaches to inference, and propose finite sample approaches to building confidence regions for sets in partially identified models.