Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,330 characters · 10 sections · 31 citation commands
When Surveys Become Conversations: Adaptive Matrix Validation for AI-Assisted Interviews
\noindentKeywords: AI-assisted interviewing; adaptive matrix validation; survey measurement; validation.
With a traditional, structured survey, a person with a complex day, illness episode, political view, household arrangement, or service experience is asked to express their experience through pre-specified questions and allowed responses. The respondent has to understand what the question means, remember the relevant details, decide how those details add up, and then fit that answer into one of the survey’s available response options. tourangeau2000psychology. The person is, in effect, fitting their experience into the tabular language that modern statistical software can understand. In other words, the burden of mapping from a complex, lived experience to tabular data falls largely upon the respondent. AI-assisted interviews could let respondents describe their experiences conversationally while the interview system maps those accounts into structured variables for analysis, sometimes with error. This mapping has the potential to reduce survey time, improve respondent experience, and reach a broader group of individuals who may be willing to converse but not be polled (see, for example, barari2024generative).
This paper addresses the statistical and design challenges that arise when moving to an AI-assisted interview. Here, an AI-assisted interview is any interview in which an AI system helps ask, choose, or interpret questions during measurement. To give a few representative examples, Xiao et al. build an AI-powered chatbot for conversational surveys that asks open-ended questions, interprets free-text answers, and probes when responses are incomplete or unclear xiao2020tell. More recent large-language-model (LLM) systems go further: Wuttke et al. evaluate LLMs as adaptive interviewers that conduct conversational interviews on political topics wuttke2025ai, while Barari et al. study text-based AI interviewers that dynamically probe respondents for elaboration and code open-ended answers during the survey itself barari2025ai.
In these systems, AI selects or suggests probes, interprets natural-language responses, may apply stopping rules for follow-up questions, and maps responses into structured variables. These steps determine what was asked, what was left unasked, and how the respondent's language was mapped into structured, or tabular, data. As it performs these roles, the AI system can affect multiple error sources simultaneously. It can change comprehension by rephrasing or probing; retrieval and judgment by deciding what details to ask for; response mapping by assigning narratives to categories; processing error by extracting and coding variables; nonresponse or breakoff by changing burden and perceived privacy risk; and comparability by changing prompts, models, or stopping rules over time. While all of these require careful attention, this paper focuses specifically on a statistical and design framework to address the measurement and processing errors introduced in an AI-assisted interview.
Specifically, this paper proposes Adaptive Matrix Validation (AMV) as a strategy to reduce bias and quantify uncertainty for population averages, proportions, subgroup estimates, and regression coefficients. It is called adaptive matrix validation because it validates a small randomized set of items in a respondent-by-variable matrix of structured survey responses. AMV first maps the interview to a structured answer for each item. It then asks a small random set of ordinary survey questions and uses those answers to correct the mapped values. For item means and subgroup estimates with fully observed subgroup membership, the estimate averages the mapped values after a validation-based correction. For regressions, the same idea is applied to the estimating equation. This distinction matters because the same structured response variable can be an outcome in one analysis, a predictor in another, and a control or interaction in a third. A tool that probes respondents well and maps their accounts closely to the structured variables can reduce the number of validation questions needed. When the mapped values are weak, calibration shrinks their contribution and the design relies more heavily on the validation questions themselves.
This work connects survey-statistical tools for measurement error, adaptive data collection, matrix sampling, planned missingness, computerized adaptive testing, active questionnaires, and model-assisted estimation biemer2010total,groves2010total,groves2006responsive,tourangeau2017adaptive,raghunathan1995split,graham2006planned,cochran1977sampling,sarndal1992model,zhang2020active,yoshida2023bayesian,wang2020methods with work on conversational interviewing, machine coding of free responses, LLM coding of open-ended survey data, and generative AI in survey practice schober1997conversational,conrad2000clarifying,frisbie1968computers,mellon2024issues,halterman2025codebook,kreuter2025aapor,aapor2026responsibleai. AMV differs by treating the AI-assisted interview as part of the survey design. The interview maps each item to a structured response, while randomized validation questions reveal only some of the corresponding survey answers with known item, pair, and planned-analysis probabilities. The goal is population and regression estimation tied to answers from validation questions, rather than better text coding or summarization alone. Appendix Table (ref) places AMV alongside adjacent literatures in more context.
The remainder of the paper proceeds as follows. Section (ref) presents the data, design, and methods. Section (ref) derives sample-size calculations for the number of structured validation questions required. Section (ref) gives a design-calibration simulation. Sections (ref) and (ref) apply AMV to data from the American Time Use Survey (ATUS) and the CHAMPS network for global health surveillance. Section (ref) summarizes the implications and limits of the design. Appendix (ref) expands the literature comparison, Appendix (ref) gives baseline estimator forms, Appendix (ref) gives the pairwise-moment form for linear regressions, Appendix (ref) provides supporting simulation and ATUS results, Appendix (ref) reports the local-model ATUS interview-coding check, Appendix (ref) gives CHAMPS data details, and Appendix (ref) summarizes disclosure elements.
This section defines the data, validation-question randomization, and estimators used by AMV.
Let \(i=1,\ldots,n\) index sampled respondents or cases, write \(\mathcal I_n\) for the sample, and let \(w_i\) be the survey weight. Let \[ Z_i=(Z_{i1},\ldots,Z_{ip}) \] denote the structured-response vector. This is the list of coded survey answers that a conventional questionnaire would produce, such as symptoms, activity minutes, categorical responses, time-use indicators, care-seeking variables, or derived domains. Not every item applies to every respondent, so \(E_{ij}=1\) indicates that item \(j\) applies to case \(i\). Inapplicable items are excluded from the relevant denominator rather than treated as zero. Eligibility and routing rules should be fixed before the validation answer for item \(j\) is observed. If eligibility is uncertain, the eligibility question should be asked as an anchor or included with the related validation item.
Each respondent also has fully observed background variables. Some questions are asked of everyone because they define the interview, such as the reference period, eligibility, age, sex, site, language, or diary day. Other fully observed variables are used for weights, subgroup estimates, response adjustment, or regression controls. The notation keeps these roles visible by writing core anchors as \(H_i\), analysis variables as \(W_i\), and all anchor and analysis variables available before validation assignment as \(A_i=(H_i,W_i)\). The same variable can play both roles.
The AI-assisted interview produces a natural-language record \(D_i\) and a mapped structured response \(\tilde{Z}_i\) saved before the relevant validation answer is observed. In a time-use interview, for example, the interview record may be mapped into minutes spent sleeping, working, and commuting. If the respondent is later asked a direct commuting question, the commuting value used in the estimator must be the one saved before that direct answer was seen. The saved value may include pre-specified coding rules, reconciliation of impossible combinations, or score-to-category mapping. If validation questions are interleaved with the interview, the item-specific value used for item \(j\) must be saved before the validation answer for item \(j\) is observed. Later text or later values that use that validation answer can be saved as post-validation outputs, but they cannot be used in the validation probability or in the mapped value corrected for that answer. AMV does not require a particular interviewer, prompt, model, or rule for follow-up questions. It requires the tool to save the mapped structured responses before validation answers are revealed and to record the information needed to reproduce the measurement: question wording, software version, questions asked, order of the conversation, stopping rule, and any recorded confidence or “could not be determined from the interview” labels. This paper studies how to estimate survey quantities after an interview tool has produced coded answers. Whether respondents understand, trust, or react differently to the interview must be studied separately.
Each respondent is also assigned a subset of structured validation questions. The chance of obtaining a usable validation answer may depend on what is already known, but it cannot depend on that validation answer itself. Let \(R_i=(R_{i1},\ldots,R_{ip})\), with \(R_{ij}=1\) when item \(j\)'s usable validation answer is observed and \(R_{ij}=0\) otherwise, including when the item is inapplicable, not selected, or selected but unanswered. The item probability is \(\pi_{ij}=\mathbb{P}(R_{ij}=1\mid \mathcal F_{ij},E_{ij}=1)\), where \(\mathcal F_{ij}\) is the information available when deciding whether item \(j\)'s usable validation answer will be observed. It may include anchors, paradata, earlier conversation, uncertainty scores, accumulated fieldwork and validation-coverage history, and the saved mapped value, but not the validation answer to item \(j\). If validation questions are interleaved with the interview, earlier volunteered information may enter \(\mathcal F_{ij}\) and \(\tilde{Z}_{ij}\); later text cannot justify \(\pi_{ij}\), though it can remain in the final record. If selected validation questions can go unanswered, \(\pi_{ij}\) should include that response process or be paired with an explicit response adjustment.
The validation answer defines the structured-response scale under the AMV protocol. Treating it as equivalent to a legacy survey item requires evidence on wording, order, context, and coding, because these features affect comprehension and response formation biemer2010total,groves2010total,tourangeau2000psychology,schober1997conversational,conrad2000clarifying. Selected validation questions should therefore use documented wording, response options, reference periods, and coding rules. When the interview setting changes those features, the validation item should be treated as a versioned measurement protocol.
The resulting data are $\{D_i,A_i,E_i,\tilde{Z}_i,R_i,R_iZ_i,\Pi_i,P_i\}$, where \(R_iZ_i\) denotes the observed validation answers and \(R_i\) marks which entries are observed. The object \(\Pi_i\) contains the known probabilities for individual items, item pairs, and preplanned analyses that require several answers from the same person, including response adjustments when needed. The record \(P_i\) contains the question path, timing, language, software version, coding rules, stopping rule, and uncertainty or inferability labels saved before validation answers were observed.
Define \(S_i\) as respondent \(i\)'s validation tile, which is the subset of applicable structured items selected for validation. One respondent might receive symptom and care-seeking questions, while another receives timing and location questions. Across respondents, these tiles form the second-phase validation sample for the structured-response matrix. The tile is usually much smaller than the full structured questionnaire, and the choice of validation items is separate from the order in which questions are presented. In this way, AMV uses the familiar logic of a two-phase design bose1943samplingerror,cochran1977sampling,sarndal1992model, combining proxies imperfectly predicted from a statistical model (in this case the AI-assisted interview tool) with strategically placed items for validation.
It also borrows from split questionnaires, matrix sampling, and planned missing data designs, where long instruments are divided across respondents to reduce burden raghunathan1995split,graham2006planned,gonzalez2008multiplematrix,thomas2006matrix,axenfeld2022split.
For exposition, suppose first that the background variables, narrative interview, and mapped structured responses are completed before validation tiles are selected. The information available at that point is \(\mathcal F_i\). It contains \(D_i,A_i,E_i,\tilde{Z}_i\), uncertainty scores \(U_i=(U_{i1},\ldots,U_{ip})\), contradictions or labels saying that an answer could not be determined from the interview, accumulated fieldwork and validation-coverage history, and the path and software information available before tile selection. If validation questions are interleaved with the conversation, the same logic applies using the information available just before each validation question is selected.
The tile rule, the design that assigns a validation tile to each respondent, chooses \(S_i\) before observing the selected validation answers. In the simplest version, the study fixes \(K\), the number of non-anchor validation questions per respondent. The rule removes inapplicable items, selects \(K\) applicable items when at least \(K\) are available, and otherwise selects all applicable items. \(K\) can vary only when that variation is part of the field design, such as different burden limits by mode, language, accessibility need, or interview length. The rule should reserve some probability for broad random coverage, add support for planned subgroup reports and planned regressions, and place more validation on items, groups, languages, model versions, or fieldwork waves where the mapped values appear less reliable or previous validation data are sparse. Every item should keep at least a small chance of being validated, even when it looks easy to code. When an analysis needs two or more validated answers from the same person, the design must record the chance that those answers were asked together.
Using an AI-assisted interview requires careful consideration of both the survey design and the ultimate inferential goals. Before data collection, the analyst must ensure the validation questions are assigned in a way that gives the validation answers needed for each planned analysis. Means and prevalences require each item to have a known positive probability of being asked and answered as a validation question for each eligible respondent, \(\pi_{ij}\). Subgroup estimates use the same item probability, and the design must provide enough validation observations within each reporting group. If subgroup membership is itself produced from the interview, the subgroup variable and outcome need same-respondent validation. Regressions require the variables in the estimating equation to be observed together. For example, if a care-seeking regression uses fever, transport, and care outside the home, those answers must sometimes be validated for the same respondent. This requirement has to be built into the field design before data collection. As the structured-response dictionary grows, the number of possible outcome, predictor, control, and interaction combinations grows quickly, so a sparse validation design cannot be expected to support every regression an analyst might later imagine. AMV therefore requires the study to name the main means, subgroup reports, item pairs, and regression blocks in advance, and the validation-tile rule must record the probabilities needed for those planned analyses. Let \(\mathcal J_q\) be the structured variables needed for planned analysis \(q\), let \(Q_{iq}=1\) when usable validation answers are observed for every variable in \(\mathcal J_q\), and define \(\pi_{iq}=\mathbb{P}(Q_{iq}=1\mid \mathcal F_i,E_{iq}=1)\), where \(E_{iq}=1\) indicates that respondent \(i\) is in the analytic population. Joint probabilities such as \(\mathbb{P}(R_{ij}=1,R_{ik}=1\mid \mathcal F_i,E_{ij}E_{ik}=1)\) are useful for item pairs, correlations, or later two-variable analyses. For analyses added after fieldwork, the recorded design must show that the needed items could have been selected and that enough respondents actually received the relevant validation questions.
The estimator uses two related ideas. The first is calibration. It uses the validation answers to learn how much the mapped value for item \(j\) should be trusted. If the mapped value is strongly related to the structured answer, the estimator leans on it more; if it is weak, the estimator leans on it less. This part follows an extensive literature in survey research on model-assisted estimators and double-sampling which uses information observed for everyone to improve precision, while checking it against a smaller set of higher-quality measurements cochran1977sampling,sarndal1992model. For regression, the closest ancestor is Chen and Chen's double-sampling estimator, which uses first-phase proxy information together with second-phase validated measurements to improve regression estimates without changing the target parameter chen2000unified. Second, after the mapped value has been calibrated, the estimator uses the validation answers to estimate what the calibrated mapped values still miss on average. It adds that missing piece back in, weighting each validation answer by the inverse of its known chance of being asked. This correction is the same basic idea as augmented inverse-probability estimation in missing-data and validation-sample problems robins1994estimation,bang2005doubly,tsiatis2006semiparametric. Recent work on inference with predicted or imputed outcomes uses the same prediction-plus-correction logic in modern prediction settings wang2020methods,angelopoulos2023prediction,zrnic2024cross,gronsbell2024anotherlook,chen2025unified,salerno2025modern,salerno2025moment.
This subsection begins with individual item means/prevalences, then moves to subgroups and, finally, regressions. For item \(j\), calibration starts with the mapped answer \(\tilde{Z}_{ij}\) and pulls it toward the validation-question mean estimated from other folds. The amount of pull is governed by \(\lambda\). When the mapped answers closely match validation answers in the training folds, \(\lambda\) is closer to one. When they do not, \(\lambda\) is closer to zero. This paper uses a simple calibration because sparse validation tiles can make more detailed item-specific models unstable. A respondent's validation answer can correct the final estimate, but it should not also decide how much to use that respondent's mapped answer. To keep those roles separate, partition the respondents into \(K\) groups, called folds, and let \(k(i)\) denote the fold containing respondent \(i\). Then, for fold \(k\), define the item-specific mean: \[ \widehat\mu_{j,-k} = \frac{ \sum_{\ell:k(\ell)\ne k} w_\ell E_{\ell j}R_{\ell j}Z_{\ell j}/\pi_{\ell j} }{ \sum_{\ell:k(\ell)\ne k} w_\ell E_{\ell j}R_{\ell j}/\pi_{\ell j} }. \] For respondent \(i\), define the calibrated response to item \(j\) by \[ M_{ij}^{(-k(i))} = \widehat\mu_{j,-k(i)} + \widehat\lambda_{j,-k(i)} \{\tilde{Z}_{ij}-\widehat\mu_{j,-k(i)}\}, \qquad \widehat\lambda_{j,-k(i)}\in[0,1] \]
In this simple version, \(\widehat\lambda_{j,-k(i)}\) controls how much the mapped answer can contribute to precision. A value of 1 leaves \(\tilde{Z}_{ij}\) unchanged. Choosing \(\lambda\) matters because calibration alone can reduce bias, while interval-width protection comes from using the mapped answer only when the training-fold validation answers show that it helps precision. More detailed calibration rules can use uncertainty scores, inferability labels, question-path features, or subgroup features, but they must be learned from other folds rather than from the held-out respondent's validation information.
The value of \(\lambda\) is chosen in the training folds. For each candidate value between 0 and 1, compute the calibrated mapped answer implied by that value and compare it with the observed validation answers in the training folds: \[ \widehat Q_{j,-k}(\lambda)= \sum_{\ell:k(\ell)\ne k} \frac{R_{\ell j}}{\pi_{\ell j}}w_\ell^2E_{\ell j} \left(\frac{1}{\pi_{\ell j}}-1\right) \left[ Z_{\ell j} - \{\widehat\mu_{j,-k}+\lambda(\tilde{Z}_{\ell j}-\widehat\mu_{j,-k})\} \right]^2. \] The chosen value \(\widehat\lambda_{j,-k}\in\arg\min_{\lambda\in[0,1]}\widehat Q_{j,-k}(\lambda)\) is the value that gives the smallest estimated validation-assignment variance for this item. To see where the weights come from, fix a calibrated value \(M_{ij}\) and write the remaining error as \(e_i=Z_{ij}-M_{ij}\). The validation-weighted correction contains \(R_{ij}e_i/\pi_{ij}\). Under the diagonal validation-assignment calculation, with sampled records held fixed,
\[ \operatorname{Var}_R\left\{ w_iE_{ij}\frac{R_{ij}}{\pi_{ij}}e_i \mid \mathcal F_{ij},Z_{ij},E_{ij}=1,M_{ij} \right\} = w_i^2E_{ij} \left(\frac{1}{\pi_{ij}}-1\right)e_i^2. \] Survey weights enter squared because this is a variance calculation. The factor \(R_{\ell j}/\pi_{\ell j}\) in \(\widehat Q_{j,-k}\) appears because \(e_\ell^2\) is observed only for respondents who answered item \(j\) as a validation question, and those observed residuals must represent the full eligible training-fold population. This criterion is the training-fold version of the conditional validation-assignment variance for a fixed calibrated value. Because \(\widehat\mu_{j,-k}\) is itself estimated in the training folds, an inner cross-fit version can be used when an exactly fold-external tuning criterion is desired. If the training folds contain too few usable validation answers for item \(j\), the analysis must use a pre-specified fallback, such as a previous wave, a pooled estimate across related items, or \(\widehat\lambda_{j,-k}=0\), or report that the item lacks enough validation support for calibrated AMV. When selected items are linked within fixed-size tiles or blocks, the same idea should use the relevant covariance terms, pair or block probabilities, replicate weights, or a nested resampling analogue.
After calibration, the estimator uses the validation answers to correct for remaining error. The calibrated value $M_{ij}^{(-k(i))}$ is a working approximation available for every respondent. Among respondents who answered item $j$ as a validation question, the gap between their validation answer and their calibrated value, $Z_{ij} - M_{ij}^{(-k(i))}$, is observed. These gaps reveal what the calibration still misses. Because only some respondents receive each validation question, the estimator reweights each observed gap by the inverse of its known selection probability, $1/\pi_{ij}$, so that the correction applies to the full eligible population rather than just the validated subset:
For any calibrated mapped value fixed before respondent \(i\)'s own validation assignment, response status, and validation answer for item \(j\) can affect it, the known-probability validation assignment gives
The expectation is over the validation-tile assignment, holding the sampled records and their structured responses fixed. Under those conditions, calibration can reduce noise, while the validation-weighted correction keeps the estimate tied to the structured-response target.
Let \(G_i\) denote a subgroup label observed before validation assignment, such as sex, age group, site, or language. For subgroup \(g\), the corresponding estimator is:
If a priority subgroup has few validation-question answers for item \(j\), the estimate will be noisy and any remaining error will be hard to see. The validation rule must assign enough validation questions within subgroups that will be reported. If subgroup membership is itself a mapped structured response item, then subgroup estimation requires the subgroup item and \(Z_j\) to be observed through validation questions for the same respondent. That case needs a version that validates both the subgroup label and the outcome for the same respondent.
For regressions, AMV uses the same order as the item estimator, but applies it to an estimating equation. The design requirement is stronger than for a single item. A regression score can combine outcomes, predictors, controls, interactions, and transformations, so the structured answers needed to evaluate that score must sometimes be validated together for the same respondent. Separate validation of the outcome for one respondent and a predictor for another is not enough for this score form.
For a planned regression \(q\), let \(\mathcal J_q\) be the structured responses needed to evaluate its score, and let \(\beta\) be the coefficient vector for that regression. Define the unweighted complete-data score contribution \[ \psi_i(\beta)=\psi_q(Z_i,A_i;\beta), \] so that the planned weighted estimating equation would use the sample sum of \(w_iE_{iq}\psi_i(\beta)\). Let \[ \widetilde\psi_i(\beta)=\psi_q(\tilde{Z}_i,A_i;\beta) \] be the same score evaluated from the mapped record. This setup covers estimating equations whose complete-data score can be evaluated from a same-respondent validation block, including interactions, transformations, outcomes, predictors, and controls when the needed variables are observed together robins1994estimation. Here \(E_{iq}=1\) means respondent \(i\) is in the analytic population for regression \(q\) and the variables in \(\mathcal J_q\) apply. Let \(Q_{iq}=1\) when usable validation answers are observed for every structured response in \(\mathcal J_q\), and set \(Q_{iq}=0\) otherwise, including when \(E_{iq}=0\). Let \(\mathcal F_{iq}\) be the information available before the validation block for \(\mathcal J_q\) is assigned; when validation tiles are selected after the interview, this is the corresponding block version of \(\mathcal F_i\). Among applicable analytic records, \(\pi_{iq}=\mathbb{P}(Q_{iq}=1\mid \mathcal F_{iq},E_{iq}=1)\) is the known probability that the usable same-respondent validation block is observed, with \(0<\pi_{iq}\le 1\). For inapplicable records, the factor \(E_{iq}\) makes the contribution zero.
Calibration first chooses a mapped score for each respondent. Let \(C_{iq}^{(-k(i))}(\beta)\) be that calibrated mapped score. It can use the mapped record, anchors, saved uncertainty scores, inferability labels, and other information logged before the needed validation answers are observed. It cannot use respondent \(i\)'s own validation assignment, response status, or validation answers for regression \(q\). A simple first version uses one multiplier for the mapped regression score: \[ C_{iq}^{(-k(i))}(\beta)=\widehat\lambda_{q,-k(i)}\widetilde\psi_i(\beta), \qquad \widehat\lambda_{q,-k(i)}\in[0,1]. \] This is close in spirit to Chen and Chen's double-sampling regression estimator chen2000unified. In their setting, rough or proxy measurements are observed for the full primary sample and exact measurements are observed for a simple random validation subsample. They start from the validation-sample regression estimate and improve it by adding a matrix-weighted difference between a working-model estimate from the validation subsample and the corresponding estimate from the full primary sample. That matrix is built from the large-sample covariance structure and gives the most efficient constant-matrix adjustment in their class chen2000unified. AMV uses the same broad idea: proxy information improves precision, while validation answers keep the target on the structured-response scale. The implementation here uses a simpler scalar score calibration rather than a matrix adjustment. That choice is deliberate. Sparse validation tiles often leave only a limited number of respondents with the full validation block for a planned regression, especially when the regression uses several mapped variables. A single multiplier is easier to pre-specify, easier to estimate with limited block support, and easier to report.
To choose \(\lambda\), start with a preliminary coefficient without the held-out fold. It can come from validation blocks in the training folds, a previous wave, a pilot, or a pre-specified training-fold AMV estimate. If validation assignment or sampling is clustered, the folds should be formed at that level. At a preliminary value \(\widehat\beta^{(0)}_{q,-k}\), the training-fold rule chooses \(\lambda\) by comparing the validated score with the \(\lambda\)-scaled mapped score: \[ \widehat Q_{q,-k}(\lambda)= \sum_{i:k(i)\ne k} \frac{Q_{iq}}{\pi_{iq}}w_i^2E_{iq} \left(\frac{1}{\pi_{iq}}-1\right) \left\| \psi_i(\widehat\beta^{(0)}_{q,-k}) - \lambda\widetilde\psi_i(\widehat\beta^{(0)}_{q,-k}) \right\|^2 . \] The fitted multiplier is \(\widehat\lambda_{q,-k}\in\arg\min_{\lambda\in[0,1]}\widehat Q_{q,-k}(\lambda)\). Values near one give more weight to the mapped score. Values near zero make the final equation rely mainly on same-respondent validation blocks. The weights in \(\widehat Q_{q,-k}\) have the same source as the item weights. For \(r_i(\lambda)=\psi_i(\widehat\beta^{(0)}_{q,-k})-\lambda\widetilde\psi_i(\widehat\beta^{(0)}_{q,-k})\), the validation-assignment part of the score correction has diagonal variance contribution proportional to \[ w_i^2E_{iq} \left(\frac{1}{\pi_{iq}}-1\right) \|r_i(\lambda)\|^2 . \] The factor \(Q_{iq}/\pi_{iq}\) estimates this quantity from training-fold records where the full validation block was observed for the same respondent. The norm should be chosen before analysis. Ordinary squared differences tune the score on its own scale; a pre-specified matrix or norm is needed when the goal is to put more weight on selected coefficients or contrasts. Recent reference-sample methods use related tuning ideas for fitted mapped values angelopoulos2023prediction,zrnic2024cross.
After calibration, the validation blocks estimate what the calibrated mapped score still misses. For any calibrated mapped score fixed before respondent \(i\)'s own validation block can affect it, the corrected score starts with \(C_{iq}^{(-k(i))}(\beta)\) and adds back the validation-weighted difference between the validated score and the calibrated mapped score:
The calibrated AMV regression estimate solves \(\widehat U_q(\beta;C)=0\). For a calibrated mapped score fixed without respondent \(i\)'s own score-validation block, the known validation probability gives
The weighted estimating equation therefore targets the same planned structured-response score that would be available if the needed structured responses were observed for every applicable sampled record. With the scalar calibrated mapped score above, the final equation is \[ \sum_{i\in\mathcal I_n} w_iE_{iq} \left[ \widehat\lambda_{q,-k(i)}\widetilde\psi_i(\beta) + \frac{Q_{iq}}{\pi_{iq}} \{\psi_i(\beta)-\widehat\lambda_{q,-k(i)}\widetilde\psi_i(\beta)\} \right] =0. \] If a training fold lacks enough same-respondent validation blocks, the analysis must use a pre-specified fallback, such as a previous wave, a pilot, or a pooled training estimate, or report that the planned regression lacks support for that fold. Appendix (ref) gives the corresponding pairwise-moment form for linear regressions.
The same corrected quantities are used for uncertainty estimates. For item means, define the calibrated pseudo-outcome \[ \phi_{ij}=E_{ij} \left[ M_{ij}^{(-k(i))} + R_{ij}\{Z_{ij}-M_{ij}^{(-k(i))}\}/\pi_{ij} \right]. \] If the sample is treated as independent after weighting, a first-order variance estimator is \[ \widehat V(\widehat\theta^{cal}_j)= \frac{1}{\widehat N_j^2} \sum_{i\in\mathcal I_n} w_i^2(\phi_{ij}-E_{ij}\widehat\theta^{cal}_j)^2, \] with the usual stratum and cluster modifications for complex samples. Standard errors should, of course, reflect both the original survey design and the random validation questions. When the survey has weights, strata, clusters, fixed-size validation tiles, validation blocks, or adaptive assignment, the variance calculation should use those features cochran1977sampling,sarndal1992model. Replicate weights are preferred when they are available. If the calibration parameter is estimated from the validation data and has a meaningful effect on the reported interval, it should be re-estimated inside each replicate or bootstrap sample when feasible.
This section outlines a framework for understanding how changes to AMV design impact the (expected) variance of the estimators, thereby giving a sense of when using a particular AI-assisted interview tool would increase efficiency over a traditional survey. AMV depends on three quantities. The first is how well the interview record \(D_i\) predicts the structured response. Better mapping leaves less error for validation questions to correct. The second is the number of validation questions per respondent, denoted by \(B\). The third is the effective sample size. A design can sometimes trade one for another. A stronger mapping can support fewer validation questions. More validation questions can compensate for a weaker mapping. A larger effective sample can make either design more precise. These tradeoffs are partial, not automatic, because rare items, routing, uneven weights, and planned regressions all need their own validation support.
For item \(j\), let \(n_{\mathrm{eff},j}\) be the effective number of applicable cases, let \(\sigma_j^2=\mathrm{Var}(Z_{ij}\mid E_{ij}=1)\), and let \(q_j=\mathbb{P}(R_{ij}=1\mid E_{ij}=1)\) be the planning probability that an applicable respondent receives item \(j\) as a validation question. If selected validation questions can go unanswered, \(q_j\) should include that response process or be paired with a response adjustment.
The mapping-quality input is how much item-level variation remains unexplained after calibration. For a mapped value \(M_{ij}\), either raw or calibrated, define
This ratio compares the error left after using the mapped value with the natural variation in the structured item. When \(M_{ij}=\tilde{Z}_{ij}\), it describes the raw mapped value. When \(M_{ij}\) is the fold-external calibrated mapped value, it describes the calibrated AMV input. It is often easier to read this as \(1-\rho_j(M)\): the share of item-level variation explained by the mapped value, relative to using only the item mean. A value of \(\rho_j(M)=0.10\) means the mapped value explains about 90 percent of the item-level variation and leaves 10 percent for validation questions to correct. A value near one means the mapping gives little precision gain. For a binary item, the denominator is \(p_j(1-p_j)\).
For planning, use a conservative unexplained-variation ratio \(\rho_j^{plan}\). When validation support is thin, set \(\rho_j^{plan}=1\). Otherwise, use a capped upper bound of the estimated calibrated unexplained-variation ratio, such as \[ \rho_j^{plan}=\min\{1,\widehat\rho_j^\star+c\,\widehat{SE}(\widehat\rho_j^\star)\}, \] where \(\widehat\rho_j^\star\) is the estimated unexplained-variation ratio after calibration and \(c\) is chosen before fielding. The cap at one reflects the separate benchmark that uses only validation questions: if the calibrated mapping appears worse than that benchmark for planning purposes, do not plan on a precision gain from it.
If each applicable respondent has the same validation probability \(q_j\) for item \(j\), a useful planning approximation for variance is
The first factor is ordinary sampling uncertainty with complete structured responses. The bracket is the extra uncertainty from giving item \(j\) as a validation question to only some applicable respondents. When \(q_j=1\), the bracket is one. When \(q_j<1\), sparse validation is affordable only to the extent that the calibrated mapping leaves little unexplained variation.
The same calculation can be read as a sample-size requirement. Suppose the target margin of error is \(h_j\). For a proportion, a 95 percent interval with half-width 5 percentage points uses \(h_j=0.05\) and \(z=1.96\). The required effective sample size is approximately
This is the main planning formula. It says that the complete-structured-response sample size is multiplied by an extra factor. That factor is small when the mapping is strong or the item is often included in the validation tile. It is large when the mapping is weak and the item is rarely validated.
If every item applies to everyone and validation questions are spread roughly evenly across \(p\) items, then \(q_j\approx B/p\). In that simple case,
This version shows the tradeoff. For a fixed precision target, the study can increase \(B\), improve the mapping, increase the sample size, or relax the target. None of these choices is free. More validation questions increase respondent burden. Improving the mapping requires a better interview protocol or coding rule. Increasing \(n_{\mathrm{eff},j}\) requires more usable sample. \(p\) counts structured fields after routed subparts are expanded. It is larger than the number of printed questionnaire prompts. For scale, the ACS household form has 27 numbered housing questions and a Person 1 section that runs through 44 numbered person questions. Census estimates the average household completion time at 40 minutes CensusACS2025Questionnaire. After expanding subparts into structured fields, a one-person ACS-like instrument is on the order of 120 structured items. The example below therefore uses \(p=120\).
Figure (ref) is an illustration. It focuses on settings where AMV could reduce burden because the mapped values are strong enough for sparse validation to improve precision without an unrealistically large sample. If early validation data suggest that the unexplained share \(\rho_j^{plan}\) is much larger than these values, the same formula gives a practical warning. The study should ask that item more often, improve the interview so it probes the content more directly, increase the sample size, or treat the item as one that needs ordinary structured questioning.
Sometimes the study instead starts with a fixed sample size and asks how often item \(j\) must be validated. Define \[ m_j=\frac{n_{\mathrm{eff},j}h_j^2}{z^2\sigma_j^2}. \] The number \(m_j\) says how much room the target leaves for not asking item \(j\) of everyone. If \(m_j=1\), the target is as tight as the interval would be if every applicable respondent answered that structured item. If \(m_j>1\), the target allows some extra uncertainty. Solving (ref) gives
A value of \(q_j=0.06\) means asking the item as a validation question of about 6 percent of applicable respondents. If \(m_j<1\), the requested margin of error is tighter than this sample can reach, even with a perfect mapping. Ordinary survey sampling error would still exceed the target if every applicable respondent answered item \(j\). The design should still keep a pre-specified minimum validation probability for every item, because mapping accuracy can differ across groups, languages, field periods, or versions of the interview protocol.
The item-level probabilities add up to the average number of non-anchor validation questions per respondent. If every item applies to every respondent, \[ B_{\mathrm{required}}=\sum_{j=1}^p q_j. \] With routing, each item is weighted by the share of respondents to whom it applies: \[ B_{\mathrm{required}} \approx \sum_{j=1}^p \bar E_j q_j, \] where \(\bar E_j\) is the population share for whom item \(j\) applies. This count covers only item-by-item precision. Core anchors, same-respondent validation for planned regressions, minimum validation for important subgroups, and allowances for unanswered validation questions must be added when they are part of the design.
Realized validation support should be reported, not only planned. For item \(j\), define validation weights \(v_{ij}=w_iE_{ij}R_{ij}/\pi_{ij}\). A useful effective validation sample size is \[ n_{\mathrm{eff},j}^{val}= \frac{\left(\sum_i v_{ij}\right)^2}{\sum_i v_{ij}^2}. \] For binary items, also report effective positives and negatives. Claims should be weaker when support is low, outcomes are rare, propensities are very small, or a few weights dominate, even if the calibrated estimate looks precise. Subgroup claims require the same support inside the subgroup, not only in the full sample.
Regressions need overlap, not only coverage. To estimate a mean for item \(j\), the design needs enough validation answers for item \(j\). To estimate a regression, it needs the outcome and predictors validated for the same respondents. If each respondent receives \(B\) validation items chosen at random from \(p\) items, the chance that any specified pair is validated together is \[ q_2=\frac{B(B-1)}{p(p-1)}. \] The expected effective number of respondents with both answers is \(n_{\mathrm{eff}}q_2\). For a larger item universe, say \(p=420\) and \(n_{\mathrm{eff}}=5{,}000\), random selection needs about \(B=43\) validation questions per respondent to get about 50 effective observations for a given pair. This is why planned regressions should be built into the question-selection rule instead of left to random overlap.
For a planned regression indexed by \(q\), use the effective number of applicable records and the probability that all variables needed for the score are validated together. With \(v_{iq}=w_iE_{iq}Q_{iq}/\pi_{iq}\), the realized same-respondent support is \[ n_{\mathrm{eff},q}^{val}= \frac{\left(\sum_i v_{iq}\right)^2}{\sum_i v_{iq}^2}. \] Before fielding, use \(n_{\mathrm{eff},q}\bar\pi_q\), where \(\bar\pi_q\) is the design-average probability that the validation tile contains all variables needed for the score. The residual variation that matters is the calibrated score residual \(\psi_i(\beta)-C_{iq}^{(-k(i))}(\beta)\), not only item-by-item mapping error. The design implication is simple. Broad item coverage supports means and proportions. Planned subgroup reports need validation inside the subgroup. Planned regressions need same-respondent validation for the variables in the score. Sparse overlap across unrelated topics can help later analysis, but it does not support regressions whose score variables were not validated for the same respondents.
Before the empirical examples, a small simulation checks the basic behavior of the estimators. It creates \(n=5{,}000\) respondents and repeats 800 times. In the item-mean setting, a binary structured response depends on a covariate and subgroup, and the mapped value is informative but underestimates the subgroup effect. Each respondent receives the target validation question with probability \(q\). In the regression setting, the outcome and one predictor are noisily mapped from the interview record, and with probability \(q\) the validation tile contains both variables for the same respondent. In both settings, the target is the finite-population quantity that would be computed from the complete structured responses. Appendix (ref) gives the validation-question-only baseline estimators used for comparison.
Figure (ref) shows RMSE for validation probabilities of 0.05, 0.10, and 0.25. Mapping-only estimation remains biased when the mapped values have systematic error, because it has no validation answers to correct that error. Validation questions only and AMV both improve as \(q\) increases, but AMV has lower RMSE when the mapped value is informative and stays close to the finite-population target. For \(q=0.10\), Appendix Table (ref) reports bias, RMSE, and coverage. The next sections evaluate the methods in an ATUS emulation and a CHAMPS narrative study.
The American Time Use Survey (ATUS) is a nationally representative U.S. time-diary survey that asks respondents age 15 and older to report activities over a 24-hour day, with activities coded into a structured lexicon and released with survey weights for population estimates BLS2025ATUSMicrodata,HamermeshFrazisStewart2005ATUS. This experiment uses public ATUS respondent and activity files for 2018, 2019, and 2021--2024. The analytic sample includes 52,468 respondent-days and 970,712 activity episodes.
For each respondent-day, 32 variables are derived from the complete diary and form \(Z_i\). This reference retains the usual recall, reporting, coding, and processing errors of a time-use survey. The experiment creates a synthetic \(\tilde{Z}_i\) for every respondent and all 32 variables. The error introduced by filling in the structured items from the AI-assisted interview is random and varies across three settings. Strong, moderate, and weak scenarios differ in error and systematic bias, with larger errors for domains that are plausibly hard to infer from a short narrative, like commute and other short travel, direct and secondary childcare, work-at-home distinctions, and the boundary between screen time and other leisure. While this strategy makes the experiment semi-synthetic, it also facilitates isolating the types of errors AMV is designed to correct. Running full interviews would make it harder to tell whether AMV corrects the estimates, because the results would also depend on question wording, probing, coding, and software behavior. The size of these errors was calibrated with a small LLM experiment. One LLM played the interviewer, using ATUS diary facts to ask fixed questions. A second LLM played the respondent and answered in natural language. A third LLM coded those answers back into ATUS variables. Appendix (ref) gives the details.
Validation assignment rules are defined over 250 possible validation items, including the 32 reported ATUS variables plus 218 additional diary checks. The 218 additional checks count toward respondent burden and affect the validation assignment, even though they are not used in the reported means or regressions. The validation tiles contain \(B=12,18,\) or \(25\) structured items per respondent-day over 30 simulation repetitions. The results presented here for regression use \(B=18\), or 7.2 percent of the validation universe. In the moderate-error setting, each priority item is directly validated for about 7.4 percent of respondent-days, and the two planned regression blocks are validated for about 6.0 and 6.1 percent. This gives roughly 2,100 effective item-validation observations and about 1,700 effective same-respondent regression-block observations, far below the full sample of 52,468 respondent-days. Most validation questions came from extra diary checks, but the design also reserved some questions for broad random coverage and some for the variables needed in the planned regressions. This made sure that enough respondents answered the outcome and predictor validation questions together.
Figure (ref) reports RMSE for \(B=12,18,\) and \(25\) validation items under the strong, moderate, and weak error settings. The figure shows two patterns. Mapping-only estimates (what would happen if the researcher used the output of the AI-assisted interview naively) carry the bias built into the simulated mapped values. Validation-tile H\'ajek and calibrated AMV improve as \(B\) increases, because a larger tile raises the validation probability for target items and gives more observations for estimating mapped-value errors. The calibrated AMV item estimator uses a calibrated mapped value derived from folds that don't include the respondent and then tunes how much to use the mapped term. The average item \(\lambda\) values for the seven priority variables are close to one in the \(B=18\) setting, corresponding to a highly effective AI-assisted interview. In this 32-variable recode, with the 250-item validation design, the simulated error mechanism, and the logged validation propensities, the calibrated AMV item biases are near zero for the priority variables. The validation-tile H\'ajek benchmark is also centered near the reference value, but it is noisier for these item means because it discards the mapped values observed for everyone. Across the seven items, three error settings, and three validation burdens, calibrated AMV has lower RMSE than validation-tile H\'ajek in 62 of 63 comparisons and a smaller linearized standard-error approximation in all 63 comparisons. Appendix Table (ref) lists supporting item-level numerical values and additional simulation details.
For regression, the first example is a sleep-minutes regression motivated by time-use work on sleep and waking activities basner2007sleep. The structured outcome is sleep minutes (nightly). The mapped structured predictors are paid-work minutes, commute minutes, and screen-time minutes. Age, sex, education, employment, weekend diary day, children in the household, and survey year are observed background variables for all respondents. The second regression is a linear-probability model for any direct childcare, motivated by the parental time-use literature on employment, gender, education, and household structure guryan2008parental,raley2012fathers,pepin2018marital. Its structured outcome is the direct-childcare participation indicator. The mapped structured predictor is paid-work minutes. The same background covariates are included. The childcare coefficient is children in the household, an always-observed, anchor predictor. This makes the childcare panel mainly about correction of the mapped childcare outcome. The paid-work coefficient in this same regression is not shown because its mapping-only bias is close to zero through offsetting errors in mapped childcare and mapped paid work, rather than because the mapped regression is generally accurate.
Figure (ref) reports the regression results in the moderate error setting. For the sleep regression, mapping-only estimation (uncorrected AI-assisted interview) would materially distort all three structured predictor effects, with biases of about \(-4.3\) to \(-5.7\) minutes of sleep per predictor hour. Correcting only the outcome variable still leaves the commute coefficient biased by about \(-3.2\) minutes, so the validation tiles need to include outcomes and predictors for the same people. Because these two ATUS examples are linear regressions, the implementation uses the linear-moment version in Appendix (ref). It corrects the mapped \(X'X\) and \(X'Y\) moments and then solves the resulting normal equations. The regression estimator first forms calibrated mapped moments, then tunes how much of the mapped moment term to use. In the main \(B=18\) setting the average \(\lambda\) is about 0.78 for the sleep regression and 0.80 for the childcare regression. In the sleep regression, calibrated AMV reduces the commute bias to about \(0.8\) minutes. The commute coefficient remains the least precise sleep coefficient. Commute minutes are sparse and highly skewed, and the regression needs same-respondent validation of sleep, commute, paid work, and screen time. With \(B=18\), the design has enough same-respondent validation to reduce most of this bias, although the commute interval remains wider than the paid-work and screen-time intervals. The block H\'ajek estimate based only on validation tiles uses only records whose validation tile contains every variable needed for the score. In the childcare regression, mapping-only estimation understates the children-in-household coefficient by about 14 percentage points. Calibrated AMV reduces this bias to about 0.5 percentage points and has a smaller standard-error approximation than the validation-only block estimate. In the moderate \(B=18\) setting, calibrated AMV has lower RMSE and a smaller linearized standard-error approximation than this validation-only benchmark for all four displayed coefficients. Appendix Table (ref) lists supporting coefficients, biases, RMSE values, and standard-error approximations.
Subgroup results show the same pattern. Across sex, employment, education, children-in-household, and weekday/weekend groups, mapping-only sleep bias ranges from about \(-9.5\) to \(-2.1\) minutes, while calibrated AMV subgroup sleep biases are close to zero. For secondary childcare, interview-coding errors vary much more by group, from roughly 101 minutes below to 75 minutes above the complete-diary reference means. The corresponding calibrated AMV subgroup biases are much smaller. These results matter because a design that performs well on average can still leave large subgroup bias. Appendix Table (ref) reports the supporting subgroup ranges. Table (ref) reports the supporting uncertainty and question-selection values.
The second example uses data from the Child Health and Mortality Prevention Surveillance (CHAMPS) program blau2019champs,taylor2020champs,bassat2023champs,champsdata2025, which monitors health and delivers public health interventions through sites in low-resource settings. We focus on verbal autopsies of children under 5. These are interviews with a surviving caregiver or relative, often after deaths outside hospitals, used to understand the causes and circumstances of death. Field evidence identifies structured-interview length as a practical challenge for verbal-autopsy teams surekclark2020va. That burden makes verbal autopsy a useful setting for asking whether a future conversational protocol could ask fewer structured items while still producing estimates on the structured-response scale.
The CHAMPS data contain de-identified verbal-autopsy and verbal/social-autopsy narratives, translated into English, together with structured verbal-autopsy categorical responses. The narrative prompt asks the family member for a general description of the circumstances leading to death, with no additional probing. These narratives therefore represent a limited version of an open-response interview. They provide free text that can be translated into structured variables, but they do not include follow-up questions targeted to missing items. This exercise uses the narratives directly as transcribed. The narrative is \(D_i\), the structured verbal-autopsy fields define \(Z_i\), and site, age group, sex, and location of death are fully observed case variables in \(A_i\).
The data contain 9,299 CHAMPS records. Of these, 4,693 have nonempty narratives and 4,606 are blank or missing. Appendix (ref) reports the narrative availability, word-length summaries, construct denominators, and data-handling details. This exercise relies entirely on existing narratives as the open component and on structured verbal-autopsy responses from the same respondent as validation items. The “narrative only” results use pre-specified phrase and duration rules applied to the narrative text to extract features.
The validation rule uses a 30-construct CHAMPS dictionary as the question universe. Individual items are revealed with probability 0.09. For each displayed regression, the validation-block draw is scaled so that about 9 percent of nonempty-narrative records receive the whole set of validation questions needed for that regression. Across the displayed item fractions, the realized reveal fractions average 0.09. The two regression examples reveal the needed validation questions for about 424 and 421 records on average, or 9.0 percent of the 4,693 nonempty-narrative records. The item results use 100 assignment draws, and the regression results use 400 draws.
Figure (ref) gives the structured-response fraction results, ordered by the structured VA fraction. Because the narratives were collected from a broad prompt rather than from item-specific probes, the fixed narrative rules work unevenly across items. Narrative-only errors are large for several shown items, including motorized transport, smaller than usual at birth, care sought outside home, and treatment received during illness. The largest item-level precision gains occur for symptoms that are often stated in the narrative. For fever, vomiting, and cough, calibrated AMV standard errors are lower than validation-question-only standard errors by about 22 percent, 15 percent, and 12 percent. Smaller gains appear for convulsions, traditional medicine, neurologic symptoms, gastrointestinal symptoms, and respiratory symptoms. Appendix Table (ref) gives the supporting numerical comparisons. For the selected items shown here, calibrated AMV under the sparse validation rule is close to the corresponding structured-verbal-autopsy fractions. Additional numeric values are reported in Appendix Table (ref), while Appendix Table (ref) reports direct mention rates for six indicators that are often absent from the open narrative.
Figure (ref) reports two linear-probability regressions. The first uses traditional medicine as the structured outcome and age group, site, sex, and death location as observed case variables. This example is useful because traditional medicine is sometimes stated directly in the narrative, while the predictors do not require mapping from text. Calibrated AMV is close to the structured-verbal-autopsy coefficients and reduces the average assignment spread by about 9 percent relative to validation questions only; the corresponding RMSE ratio is 0.90.
The second regression uses treatment received during illness as the structured outcome. The five mapped predictors are fever or infectious symptoms, cough, difficulty breathing, vomiting, and traditional medicine use, with age group and site included as controls. These quantities are common in verbal-autopsy and child-mortality work because they describe illness severity, symptom presentation, and care-seeking pathways who2024va,li2023openva. In this more demanding model, the narrative only coefficients differ visibly from the structured-verbal-autopsy coefficients for several predictors, especially difficulty breathing and cough. Calibrated AMV and validation questions only both fall near the structured-verbal-autopsy coefficients. Their across-assignment spreads are nearly identical, with a calibrated-to-validation spread ratio of 0.999. Together, the panels show that narratives help for some quantities and add little for others. The validation questions keep both examples on the structured-response scale. Calibrated AMV improves precision only when the narrative rules add information beyond those validation questions. Appendix Tables (ref) and (ref) give the numerical coefficient and support summaries.
AMV treats AI-assisted interviewing as a survey measurement design problem, not as ready-to-analyze conversational data. It pairs a conversational interview with sparse, known-probability validation items, then uses those items to estimate means, subgroup quantities, and regressions on the structured-response scale while showing how many structured validation questions are needed. The simulations and ATUS emulation show the burden-precision tradeoff under stated interview-coding error settings. The CHAMPS analysis shows that existing narratives omit many verbal-autopsy items, and that validation items can return selected item and regression estimates to the structured verbal-autopsy scale. Several limits remain.
Spoken dialogue must be transcribed, and often translated across languages, before it can be mapped to structured responses. Multilingual speech-recognition systems such as Whisper perform unevenly across languages; recent Bangla evaluations find that language-adapted alternatives can outperform Whisper on word and character error rates radford2022robust,ridoy2025adaptability. Work on language-weighted training, language-model adaptation, and Bangla-specific speech resources is ongoing pineiromartin2024weighted,dezuazo2025whisperlm,rakib2023oodspeech, but dialect, code-switching, recording quality, and regional vocabulary can still create group differences in what is captured and coded koenecke2020racial.
AMV also addresses only part of what AI-assisted interviews can change. It validates and corrects the structured variables produced from the interview record, but phrasing, probing, perceived privacy, trust in automated interviewers, cultural expectations about AI use, device quality, connectivity, literacy, disability, and comfort with conversational tools can affect who responds, what they disclose, and how they interpret the exchange tourangeau2000psychology,schober1997conversational,conrad2000clarifying,kreuter2025aapor,aapor2026responsibleai. These issues need direct study because they can create group differences before the validation design is applied.
Finally, there are substantial opportunities for next steps. This paper treats unstructured data as spoken dialogue, but the same logic could extend to other respondent-provided material: photos of meals, product labels or shelf prices, receipts, housing conditions, or short videos recorded during an AI-assisted interview shao2021integrated,tahir2021comprehensive,wireduk2016premise. Those settings could combine automated extraction with validation items or human review on a known-probability subset, but image and video data would also require methods for objects, locations, timestamps, repeated frames, privacy masking, and correlated errors within each uploaded file.
Overall, this paper articulates the potential and caveats associated with AI-assisted interviews. AI-assisted interviewing will not automatically reduce respondent burden. It will reduce burden only when the AI maps responses well enough that a small number of validation questions can correct what remains. The formulas presented here let practitioners check whether that threshold is met before data collection begins. When it is, the efficiency gains are real. When it is not, the same framework says so clearly and points to the remedy, which is to ask more validation questions, probe more directly, or measure the item with a structured question.