Ask, Don't Answer: Where AI Belongs in Research Interviews

Ask, Don't Answer: Where AI Belongs in Research Interviews

A decision is being made quietly inside insight teams, product organisations, policy units and university departments, and it is usually being made without deliberation. The decision is whether a language model may stand in for a human research participant.

It arrives disguised as a workflow improvement. Recruitment for a hard-to-reach segment is slow, so someone generates a "panel" of simulated buyers. A concept test needs to run before Friday, so the concept is put to synthetic personas instead of people. A qualitative study needs coding, so transcripts go to a model. Each step is individually reasonable. Together they change what the word finding means in an organisation, and they do so without anyone signing off on the change.

The topic is timely for two reasons. First, the evidence base has matured enough to support real distinctions — we now have peer-reviewed work on both sides of the question, not just enthusiasm. Second, the professional standards have moved: the 2025 revision of the ICC/Esomar International Code, its fifth edition, was written specifically to remain fit for purpose amid AI and synthetic data, with emphasis on ethics, accountability, transparency and human oversight. Practice is now measurable against something. ICC

My position, offered as argument rather than consensus: the evidence supports putting AI on the asking side of the interview far more strongly than on the answering side. An AI interviewer that elicits testimony from real humans generates new information about the world. A synthetic respondent recombines information the model already contains. Those are different epistemic acts, and treating them as interchangeable is the central methodological error of the current moment.

Who benefits: market and insight researchers; UX and product researchers; social scientists and policy analysts; healthcare and medical-education researchers; strategy and commercial leaders who consume research; and research-ops and governance functions.

What you will learn: what the strongest studies on each side actually found; a substitution ladder for deciding what a simulated answer may be used for; a disclosure and validation standard you can adopt; and where the evidence is genuinely unresolved.


Executive summary

  • Two peer-reviewed papers in the same journal bracket the debate. Argyle and colleagues introduced "algorithmic fidelity" and "silicon samples," showing that conditioning a model on sociodemographic backstories can emulate response distributions across human subgroups. Bisbee and colleagues found that although averages matched a benchmark survey closely, synthetic responses showed less variation than real ones, regression coefficients often differed significantly, results shifted with minor prompt changes, and the same prompt produced significantly different results three months later. Cambridge CoreCambridge Core
  • The strongest pro-simulation result depends on real interviews as input. Stanford-led work reported that generative agents built from two-hour interviews replicated participants' General Social Survey answers about 85% as accurately as those participants replicated their own answers two weeks later. Interview-grounded agents also showed lower performance disparities across ideology, race and gender subgroups than agents built from demographic information or persona descriptions. This is a preprint and policy brief, not a peer-reviewed journal article. MediumStudocu
  • The asking side has independent support. An AI interviewer conducted 381 interviews on stock-market non-participation, producing rich qualitative evidence at a fraction of the cost of human-led interviews, with follow-up showing the interview data predicted economic behaviour eight months later. This is a working paper. ifo InstituteSSRN
  • Analysis assistance is real but partial. A study in Academic Medicine combining experiments with ChatGPT-4 and a scoping review of 130 studies found AI useful for summarisation and keyword detection while many complex analyses remained challenging, underlining the need for human oversight. R Discovery
  • Specialised approaches are advancing: a NAACL 2025 paper fine-tuned models on first-token probabilities to minimise divergence between predicted and actual survey response distributions, using country-level results from two global cultural surveys. Progress is real and narrow. ACL Anthology
  • Variance collapse is the central technical failure. Simulated respondents tend to under-represent disagreement, which is precisely the signal most commercial and policy decisions depend on.
  • Professional standards now distinguish humans from simulations. The 2025 Code revision added definitions and responsibilities covering AI, synthetic data and synthetic personas, with new emphasis on transparency and human oversight. EIN Presswire
  • The practical answer is not prohibition but grading: match the evidentiary weight of a simulated answer to the reversibility of the decision it informs.

Two different acts that look like one

Both practices are called "AI in interviews." They are not the same operation.

Elicitation extracts information that exists only inside a person's head and was not previously recorded anywhere. It adds to the world's stock of evidence.

Simulation produces text consistent with patterns already present in a model's training data and prompt. It rearranges the existing stock.

Simulation can be extremely useful — but its usefulness is bounded by a structural fact: a model cannot tell you something no one has yet said. This is not a limitation of current systems that scaling will remove. It is what simulation is. When a chief marketing officer asks how customers will respond to a product category that does not yet exist, a simulated panel answers by analogy to categories that do. The answer will be fluent, plausible and unfalsifiable — the three properties most likely to survive a slide deck unchallenged.

That is why I find the asking side more defensible. Delegating the interviewer role to an AI system still puts a human on the other end supplying the testimony, and the resulting data can be validated against subsequent behaviour rather than only against internal plausibility. ifo Institute


What the evidence actually supports

Table 1 — The evidence map: what each strand of work established, and what it did not

StrandRepresentative findingEvidence statusWhat it does not establish
Silicon samplingConditioning on sociodemographic backstories emulates subgroup response distributionsPeer-reviewed (Political Analysis, 2023)That individual-level or novel-topic responses are recoverable
Statistical critiqueMatching averages while collapsing variance, unstable coefficients, prompt and time sensitivityPeer-reviewed (Political Analysis, 2024)That simulation is useless for all purposes
Interview-grounded agents~85% relative accuracy against participants' own two-week retest; lower subgroup disparity than persona-only agentsPreprint / policy brief (2024–25)That agents work without rich, consented human interviews as input
Distribution specialisationFine-tuning to minimise divergence from actual response distributionsPeer-reviewed conference (NAACL, 2025)That group-level fit transfers to unseen populations or topics
AI-led interviewing381 AI-conducted interviews; predictive of behaviour eight months laterWorking paper (CESifo, 2023–26)That AI interviewers match skilled human interviewers on sensitive topics
AI-assisted analysisEffective summarisation and keyword tasks; complex thematic work still difficultPeer-reviewed (Academic Medicine, 2025)That AI can replace human interpretive judgement

Three observations follow.

First, the headline numbers and the underlying statistics diverge. The 2024 critique is precise about this: mean scores tracked the American National Election Study closely while regression estimates often differed significantly, and responses clustered more tightly than real survey data. A synthetic study can therefore look right on the summary slide and be wrong in every relationship that matters — which is worse than being obviously wrong. Cambridge Core

Second, reproducibility is not assured. The same prompt yielded significantly different results across a three-month window. Any research programme built on a commercial model's behaviour is building on a moving substrate, and few teams version-pin or archive the model as they would a panel definition. Cambridge Core

Third, the best pro-simulation result is an argument for human interviews. The Stanford work simulated 1,052 individuals using interviews plus a large language model; agents informed by interview transcripts outperformed those informed by demographics or persona descriptions, and showed smaller disparities across demographic subgroups. Read plainly, this says the qualitative interview is the scarce ingredient. Simulation is a way of amplifying human testimony, not of avoiding the cost of collecting it. Stanford HAIStudocu


An original standard: the substitution ladder

Most governance failures here come from a binary framing — "are synthetic respondents valid?" — that has no correct answer. The useful question is what a given output may be used to decide.

Table 2 — The substitution ladder: evidence grade by input type and permitted use

LevelWhat produces the answerEvidence gradeAppropriate usesValidation required before decisions
S0Real humans, human interviewerPrimary evidenceAny, within sampling limitsStandard methodology
S1Real humans, AI interviewerPrimary evidence, novel instrumentExploratory and confirmatory qual; scale-up of open-ended probingBenchmark subset against human-led interviews
S2Agents grounded in that population's own interview transcriptsDerived evidenceExtending a study to more questions; internal consistency checksHold out real respondents; compare distributions, not just means
S3Models fine-tuned or calibrated on that population's survey dataDirectional estimateInstrument piloting; question triage; prioritisationTrain-synthetic / test-real on unseen items
S4Off-the-shelf model prompted with demographic personasHypothesis generator onlyIdeation, objection brainstorming, stimulus stress-testingNever sufficient alone

Two rules make the ladder operational.

Rule 1 — Grade the decision, not the method. Reversible, low-cost, easily corrected decisions (which of forty concepts to explore further) tolerate S3 or S4. Irreversible, expensive or publicly consequential ones (market sizing, pricing commitments, clinical or policy claims, regulatory filings) require S0 or S1.

Rule 2 — Never let a level be inferred. Every deliverable should state its level on the front page. In my view this single convention would prevent most of the damage now accumulating, because the practical harm is rarely that someone used simulation. It is that the reader could not tell.

The professional standards point the same way. The 2025 Code revision introduced clear definitions and responsibilities around AI, synthetic data and synthetic personas, with new emphasis on transparency and human oversight — and drafting for the revision explicitly defined "person" as a human being, to differentiate it from a synthetic, virtual or digitally created persona. The distinction only does work if it appears in deliverables, not just in codes. EIN PresswireEsomar


Where AI-led interviewing earns its place

The strongest practical case is not simulation but scale in elicitation. Human-led depth interviews are expensive, which is why most organisations run twelve and generalise from them. The AI-interviewer approach produced 381 interviews on a single research question at a fraction of the cost of human-led interviewing, and reported that AI-led interviews could generate novel hypotheses while examining the internal and predictive validity of responses against other survey methods. Its most useful validation is external: the interview data predicted economic behaviour eight months later, which addresses the "cheap talk" objection that participants say whatever sounds good in the moment. ifo Institute + 2

This should be read with care. It is a working paper, from a specific domain (household finance) with a non-sensitive topic and a motivated online sample. It does not establish that AI interviewers handle grief, trauma, workplace power dynamics, clinical disclosure or culturally loaded subjects competently. My reading of the current state: AI interviewing is promising for breadth on tractable topics, unproven for depth on difficult ones, and — importantly — subject to disclosure duties. Under the EU AI Act's Article 50, applicable from 2 August 2026, systems that interact directly with people must be designed so that individuals are informed they are dealing with an AI, from the start of the first interaction. Consent to be interviewed is not consent to be interviewed by a machine. europa

Similarly, on analysis: the Academic Medicine study found that ChatGPT-4 produced accurate brief summaries but that initial prompts failed for other tasks, with iterative prompt engineering succeeding for some (keyword counting, summarisation) while others remained problematic. That is a fair description of the frontier — genuine time savings on mechanical work, no safe delegation of interpretation. Oxford Academic


Figure specifications for the design team

Figure 1 — "Elicitation vs Simulation: Two Different Acts"
Purpose: Establish the article's core distinction visually.
Layout: Two parallel horizontal flows. Top flow (Elicitation), rendered in the primary accent colour: Research question → Human participant → Interview (human or AI interviewer) → Transcript → Analysis → New evidence. A small "+" badge sits above the arrow between Transcript and Analysis, labelled "adds to world knowledge". Bottom flow (Simulation), rendered in a muted grey: Research question → Prompt/persona → Language model → Generated response → Analysis → Recombined prior knowledge. A circular arrow loops from "Recombined prior knowledge" back to "Language model", labelled "closed loop". A vertical dashed divider separates the flows. Caption: "Simulation cannot return information that no human has yet expressed; elicitation can."

Figure 2 — "The Substitution Ladder"
Purpose: Give teams a single reference image for governance.
Layout: Five stacked horizontal bars forming a staircase descending left to right, labelled S0 to S4 from top-left (widest, darkest) to bottom-right (narrowest, lightest). Each bar carries three text fields in fixed columns: Input, Evidence grade, Permitted decisions. To the right of the staircase, a vertical arrow labelled "Decision reversibility" points upward, with "Irreversible / high-cost" at the top aligned to S0–S1 and "Reversible / exploratory" at the bottom aligned to S3–S4. A footer strip beneath all bars reads: "State the level on every deliverable." Caption: "Match the evidentiary weight of the method to the reversibility of the decision."


Myths worth retiring

Myth 1: "Synthetic respondents remove bias because they have no incentives." They import the biases of their training distribution while removing the disagreement that would reveal them. Reduced variance relative to real surveys means dissent is systematically under-sampled — the opposite of what a fair method does. Cambridge Core

Myth 2: "It matched the real data, so it works." Matching a benchmark you already possess demonstrates fit, not predictive validity. The test that matters is performance on data withheld from the process, on questions the model has not effectively seen.

Myth 3: "Simulation is cheaper." It is cheaper per response and potentially far more expensive per wrong decision. The relevant unit of cost is the decision, not the interview.

Myth 4: "Adoption figures prove maturity." Most published adoption statistics for synthetic research come from vendor surveys with undisclosed sampling. Treat them as marketing inputs, not evidence — a caution I would apply to any figure in this space that lacks a documented method, including favourable ones.


Barriers, trade-offs and open questions

  • Validation is expensive, which defeats the stated purpose. Doing simulation properly requires holding out real respondents — reintroducing the cost that motivated simulation. Teams under budget pressure will skip it precisely when it matters most.
  • Model drift breaks reproducibility. Documented shifts in the same prompt's output over three months mean a synthetic study is not straightforwardly repeatable unless models are versioned and archived. Cambridge Core
  • Consent and provenance for grounded agents. Agents built from real people's interviews raise questions about withdrawal, secondary use and downstream commercial deployment that most consent forms do not yet address. Stanford HAI
  • Genuinely unresolved: whether interview-grounded agents generalise beyond the population and question types they were built from; whether AI interviewers match humans on sensitive topics; whether variance collapse can be corrected without simply fitting to the answer you already have; and whether simulated evidence, once normalised, degrades organisations' willingness to fund real fieldwork.

Practical takeaways

For research and insight leaders

  1. Adopt the substitution ladder, or an equivalent, and require the level to appear on the cover of every deliverable.
  2. Set a bright line: no S3 or S4 output may be the sole basis for an irreversible commitment.
  3. Version-pin and log the model, prompt and date for any simulated study, and treat the prompt as a survey instrument subject to review.
  4. Validate on holdout humans, comparing distributions and relationships — not means alone.
  5. Budget the savings from simulation into more real fieldwork, not out of the research line entirely.

For social scientists and academic researchers

  1. Report simulation as a method with its own limitations section, including model version and date.
  2. Prefer designs where simulated results are checked against pre-registered human data.
  3. Where you build agents from participants' interviews, extend consent explicitly to agent creation, retention and withdrawal.

For healthcare and clinical researchers

  1. Restrict AI to mechanical stages — transcription, summarisation, first-pass indexing — consistent with evidence that complex thematic analysis remains challenging and human oversight is required. R Discovery
  2. Never simulate patient voice for evidence intended to inform care, guidelines or safety claims.

For executives who commission research

  1. Ask one question of every deck: were these answers given by people? Require the answer in writing.
  2. Fund the validation step explicitly, or you will receive validation-free work by default.

Key insights

  1. Elicitation and simulation are different epistemic acts; conflating them is the core error.
  2. Matching averages while collapsing variance makes simulated data most dangerous where it looks most convincing. Cambridge Core
  3. The strongest simulation result depends on two-hour human interviews as input, which argues for funding fieldwork, not replacing it. Medium
  4. Interview-grounded agents showed smaller subgroup performance disparities than persona-based ones — grounding improves fairness as well as accuracy. Studocu
  5. Prompt sensitivity and drift over three months make reproducibility a first-order problem. Cambridge Core
  6. Predictive validity against later behaviour is the strongest available test; demand it. SSRN
  7. AI is reliable for summarising and counting, not for interpreting. Oxford Academic
  8. Specialised fine-tuning improves distributional fit without resolving generalisation to new populations or topics. ACL Anthology
  9. Professional standards now require transparency and human oversight over synthetic personas; disclosure is a duty, not a courtesy. EIN Presswire
  10. Grade the method to the reversibility of the decision — the only rule that scales across use cases.

Frequently asked questions

What are synthetic respondents?
Language-model outputs generated to stand in for human survey or interview participants, usually by conditioning a model on demographic profiles, personas, or real participants' prior data.

Are synthetic respondents accurate?
Partially and conditionally. Group-level emulation of response distributions has peer-reviewed support, but the same literature documents reduced variance, unstable regression estimates and sensitivity to prompt wording. Accuracy on averages does not imply accuracy on relationships. Cambridge CoreCambridge Core

What is silicon sampling?
A method introduced by Argyle and colleagues in which a language model is conditioned on sociodemographic backstories from real survey participants to generate a virtual population of respondents. Cambridge Core

What is algorithmic fidelity?
The property, proposed in the same paper, whereby a model's biases are fine-grained and demographically correlated, so that proper conditioning reproduces subgroup response patterns rather than a single average. Cambridge Core

Can AI conduct qualitative interviews?
Documented work has delegated the interviewer role to an AI system across 381 interviews, producing rich data at much lower cost. Performance on sensitive or emotionally complex topics remains unestablished. ifo Institute

Do I have to tell participants they are talking to an AI?
Under EU rules applicable from 2 August 2026, providers must design directly interacting systems so people are informed they are interacting with AI from the start of the first interaction — and disclosure is sound ethics regardless of jurisdiction. europa

Can synthetic data replace a panel?
Not for population estimates or irreversible commitments. It is defensible for hypothesis generation, instrument piloting and question triage, with real-human validation before decisions.

Is simulated research cheaper?
Per response, substantially. Per correct decision, unknown — and the validation needed to make it trustworthy reintroduces much of the cost saved.

Why does variance matter more than the average?
Most business and policy questions turn on how much people disagree and which subgroups differ. Synthetic responses cluster more tightly than real ones, so they systematically understate exactly that. Cambridge Core

Does the ICC/Esomar Code address this?
The 2025 revision added definitions and responsibilities for AI, synthetic data and synthetic personas, with emphasis on transparency and human oversight. It is the fifth edition of the Code. EIN PresswireICC

Can AI code and analyse interview transcripts?
Partly. Summarisation and keyword tasks performed acceptably; thematic and interpretive tasks remained difficult and required human oversight. R Discovery

How should I validate a synthetic study?
Hold out real respondents, compare full distributions and key relationships rather than means, and test on questions the model was not calibrated against.

Is this improving quickly?
Yes in narrow directions — for example, fine-tuning models to match survey response distributions at country level. Gains have so far been strongest where the target distribution is known in advance, which is not the case in genuine research. ACL Anthology

What should never be simulated?
Patient experience informing clinical or safety claims; regulatory submissions; population estimates presented with confidence intervals; and any research whose value depends on discovering something no one has said before.


Glossary

  • Algorithmic fidelity: The degree to which a model's conditioned outputs reproduce the response patterns of specific human subgroups.
  • AI interviewer: A system that conducts a live or asynchronous interview with a human participant, generating follow-up probes.
  • Distributional fit: How closely simulated responses match the full shape of real responses, not merely the mean.
  • Elicitation: Obtaining information from a person that was not previously recorded.
  • Holdout validation: Reserving real human data, unseen during generation or tuning, to test simulated output.
  • ICC/Esomar Code: The international professional code for market, opinion and social research and data analytics; fifth edition published 2025.
  • Silicon sample: A model-generated virtual population of respondents conditioned on real sociodemographic profiles.
  • Simulation: Generating plausible responses from a model's existing learned distribution.
  • Synthetic persona: A digitally created character used to represent a segment or individual in research.
  • Train-synthetic, test-real: A validation pattern in which models are developed on simulated data and evaluated against real human data.
  • Variance collapse: The tendency of simulated responses to show less spread than real ones, understating disagreement.

References

Academic Papers (peer-reviewed)

Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351. https://doi.org/10.1017/pan.2023.2

Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. Political Analysis, 32(4), 401–416. https://doi.org/10.1017/pan.2024.5

Cook, D. A., Ginsburg, S., Sawatsky, A. P., Kuper, A., & D'Angelo, J. D. (2025). Artificial intelligence to support qualitative data analysis: Promises, approaches, pitfalls. Academic Medicine, 100(10), 1134–1149. https://doi.org/10.1097/ACM.0000000000006134

Conference Papers

Cao, Y., Liu, H., Arora, A., Augenstein, I., Röttger, P., & Hershcovich, D. (2025). Specializing large language models to simulate survey response distributions for global populations. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 3141–3154). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-long.162

Preprints and Working Papers (not peer-reviewed)

Chopra, F., & Haaland, I. (2023). Conducting qualitative interviews with AI (CESifo Working Paper No. 10666). CESifo. https://doi.org/10.2139/ssrn.4583756

Park, J. S., Zou, C. Q., Bernstein, M. S., et al. (2024). Generative agent simulations of 1,000 people (arXiv:2411.10109). arXiv. https://arxiv.org/abs/2411.10109

Standards and Professional Codes

International Chamber of Commerce & Esomar. (2025). ICC/Esomar international code on market, opinion and social research and data analytics (5th ed.). https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/

Esomar. (2025, June 25). Esomar announces major update to the ICC/Esomar international code to reflect AI and emerging technology [Press release]. https://www.einpresswire.com/article/825277849/esomar-announces-major-update-to-the-icc-esomar-international-code-to-reflect-ai-and-emerging-technology

Official Documentation

European Commission. (2026). Transparency obligations under Article 50 of the AI Act — Frequently asked questions. Shaping Europe's Digital Future. https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act

Other Authoritative Sources

Stanford Institute for Human-Centered Artificial Intelligence. (2025). AI agents simulate 1,052 individuals' personalities with impressive accuracy. https://hai.stanford.edu/news/ai-agents-simulate-1052-individuals-personalities-with-impressive-accuracy



One Tech & AI · Thursday, August 6, 2026 · 21 min read

Expert Insights – Discover exclusive interviews with industry leaders, researchers, and innovators sharing their perspectives on emerging technologies and trends.

Real-World Experience – Learn from professionals discussing challenges, success stories, practical strategies, and lessons from their careers.

Future Perspectives – Explore expert opinions on the future of technology, innovation, policy, and the opportunities shaping tomorrow's industries.

The research interview has always been an expensive way to learn something true. Its cost was never incidental — it was the price of reaching information that existed nowhere else. Simulation is attractive precisely because it removes that cost, and troubling for exactly the same reason. The evidence, read carefully, does not support a prohibition and does not support a substitution. It supports a division of labour. AI has demonstrated real value in conducting interviews at scale and in the mechanical stages of analysis. Simulation has demonstrated genuine group-level fidelity under conditioning while also demonstrating collapsed variance, unstable estimates and instability over time. And the most impressive simulation result to date was achieved only by first conducting a thousand long human interviews — a finding that reads less like a replacement for fieldwork than a tribute to it. ifo Institute + 4 Important uncertainties remain. Much of the strongest work on both sides is preprint or working-paper stage; generalisation beyond the studied populations is unproven; and the adoption statistics circulating in the industry are largely vendor-produced and should not be mistaken for evidence. But one conclusion follows from what has been established rather than what has been forecast. The failure mode of synthetic research is not that it produces obviously wrong answers. It is that it produces confident, coherent, well-formatted answers whose errors live in the second-order statistics no one checks. Organisations that grade their methods explicitly, and say so on the page, will catch that. Those that do not will find their research pipeline quietly converted from an instrument for discovering things into an instrument for confirming them — which is a change no one will have decided to make, and no one will notice having made.

TOPIC

Opinion