
AI can replace research moderators.
Maybe yes; maybe no.
When we encounter interesting claims like that, we generally take the following approach:
- Define the claim.
- Assess evidence and the quality (focusing on peer-reviewed publications).
- Conduct our own experiments.
- Refine and revise our opinion of the claim.
When it comes to AI moderators, we are now on step two in that process. We covered step one in a previous article in which we asked whether AI can reasonably be used to moderate UX interviews.
We’ve also written previously about what separates an adequate moderator from an excellent one, and asking good questions is only one part of that list. It starts before the session, with understanding why a stakeholder wanted the study at all—that’s what tells a skilled moderator when to probe and when to go off script. In the session, it means knowing when to assist a stuck participant without contaminating the task, when someone is misrepresenting who they are, when to stop probing and move on, and how to handle observers who want more than the discussion guide asks for.
Some data suggest that AI could be used at scale to interview job candidates, but the interviews in that case were straightforward questions and answers for a low-level position. This is qualitatively different from semi-structured UX sessions conducted by skilled UX moderators who build rapport and can judge when to go off script to follow up an interesting lead versus when to stick to the guide.
For this article, we reviewed the published literature to investigate evidence for and against the effectiveness of AI moderators conducting interviews. Specifically, we examined five key questions about the advantages (or lack thereof) of using generative AI as research moderators:
- Do participants say less to AI than to human moderators?
- Do participants provide the same depth of insights to AI and human moderators?
- Does AI moderate as competently as a human?
- Do participants want to talk to AI moderators?
- Do findings generalize across models or prompts?
How we picked these studies
We searched the peer-reviewed literature for work on AI-led interviewing published between 2024 and mid-2026. To ensure a reasonable approximation of the capabilities of current LLM agents, all papers in this review used ChatGPT-3.5 or a later model.
We prioritized peer-reviewed papers but included two exceptions: Chopra and Haaland’s CESifo working paper (a large, empirical AI-interview dataset with n = 766) and Rosala’s Nielsen Norman Group report (the only source describing real-world practitioner tool use), for a total of ten sources. For each question in the following sections, findings are tagged as pro AI (apparent advantage for AI moderation), neutral (no apparent advantage or disadvantage for AI), or anti AI (apparent disadvantage for AI). All studies are summarized in the appendix.
Do participants say less to AI than to human moderators (number of words, engagement, elaboration)?
There are some claims that people talk more to AI moderators about sensitive topics, but for UX research, which doesn’t usually involve sensitive topics, this is less likely. Consequently, we don’t expect an AI moderator to generate more talking from a participant, but we would be concerned if people spoke less to AI moderators.
There were two studies with relevant findings:
The results of Wuttke et al. (2025) were mixed. They ran a small classroom pilot in which student pairs each completed one AI-led (GPT-4 Turbo) and one human-led interview on politics and democracy, analyzing six human-led and five AI-led sessions.
In their Table 1, they showed participants’ raw response lengths were longer when the moderator was AI (52 vs. 33 words/answer; pro AI), but ratings of engagement and elaboration were better for human moderators (anti AI). Note that the human moderators in this study were students, not seasoned UX professionals.
Zhu and colleagues (2026) ran a randomized controlled trial comparing an agentic voice moderator (using Chinese ByteDance technologies like Doubao-1.5-pro-32k-250115) against a human moderator in think-aloud testing of a note-taking app (n = 60). They found no significant difference in measures of total words in responses to their AI and human moderators (Table 6, neutral).
Other findings of interest
Cuevas and colleagues (2025) ran a large study (n = 399) pitting two LLM-based interview chatbots using GPT-3.5-turbo against a hard-coded-question baseline. They found no significant difference in the number of words per session.
Chopra and Haaland (2026) ran two large-sample open-ended qualitative interviews (one on stock market nonparticipation in 2023, n = 381; one on attitudes toward U.S. tariffs in 2025, n = 385) with a multi-agent AI chatbot, comparing conditions with different levels of probing. They reported 29 words per minute (wpm) in response to their AI interviewer in their stock market study and compared that to a published benchmark for human-led chat-based interviews (20 wpm; Namey et al., 2020). However, there appears to be an error in the wpm calculations for the AI interviewer because the mean number of words was 654 and the mean number of minutes was 33, so the actual wpm was 19.8 (654 / 33). We classified this finding as not relevant due to the apparent calculation error and the weakness of this type of uncontrolled cross-study comparison.
Bottom line: The evidence is mixed regarding whether participants say more to AI than human moderators, but there is no compelling evidence that they say less.
Do participants provide as much depth of insights to AI and human moderators (richness, novel insight)?
Quantity of words does not equate to quality of insights. A good moderator should be able to get participants to go deep on topics to help generate insights for stakeholders. Do AI moderators go as deep as humans?
The relevant studies were Wuttke et al. (2025) and Zhu et al. (2026).
The classroom study by Wuttke et al. (their Table 1) reported lower human-coded ratings of specificity (level of detail in the response) for their AI moderator (anti AI).
Zhu et al. reported “The analysis indicates that the depth of participants’ analytical thinking and problem-solving articulation remained similar across conditions” (Section 4.3, neutral).
Other findings of interest
- Chopra and Haaland (2026) reported faster discovery of themes for their AI moderators than single or multiple open-ended survey questions (their Figure 8), roughly equivalent to discovery rates with human participants.
- Cuevas et al. (2025) reported their AI moderator had better follow-up than a hard-coded bot but no practical difference in richness of the responses.
- Kuric et al. (2025) reported poorer results for their AI moderator compared to static pre-written follow-up questions.
Bottom line: The evidence is mixed regarding whether AI moderators can achieve human levels of response depth from participants, but there is no compelling evidence that responses to AI moderators are shallower or deeper.
Does AI moderate as competently as a human (protocol fidelity, probing, task outcomes)?
Focusing on comparison with human moderators, once again the relevant studies were Wuttke et al. (2025) and Zhu et al. (2026).
The findings in Wuttke et al. were about evenly distributed among pro AI, neutral, and anti AI. AI had the advantage in listening/paraphrasing participant responses and in being less likely to lead participants (pro AI). On the other hand, AI often failed to follow up on unclear/surprising answers and had a potentially biasing tendency to praise respondents (anti AI). Overall, both AI and humans “faithfully followed the provided questionnaire” and operated at similar levels of competence (neutral). When interpreting these findings, keep in mind that this finding is complicated, as the respondents in each condition were students who also played other roles (interviewer, observer) across the two conditions.
Zhu et al. reported no practical difference for procedural adherence or think-aloud guidance quality (neutral). AI moderation was affected by occasional task fixation and flow disruption and had significantly poorer context-aware follow-up and trust building (Table 8, anti AI).
Other findings of interest
- Kuric et al. (2025) reported poorer performance of their AI moderator relative to static pre-written follow-up for surfacing new usability issues, leading participants, and perceived reasonability of questions.
- Cuevas et al. (2025) reported better performance of their AI moderator relative to a hard-coded baseline chatbot for follow-up quality and continuity, but no practical difference for relevance, clarity, or specificity.
- Panfilova et al. (2026) measured absolute protocol compliance and probing judgment scores across six different LLMs, excluding DeepSeek early in the study for compliance failures and noting that Grok over-probed already complete answers.
- Chopra and Haaland (2026) reported high protocol fidelity for their AI moderator.
- Jacobsen and colleagues (2025) tested four different probe types (descriptive, idiographic, clarifying, and explanatory) embedded directly in an online survey (n = 64, 16 participants per probe type condition). They found that the idiographic strategy was generally the most competent and explanatory the least (based on human-coded metrics of relevance, specificity, and clarity).
- Rosala (2026) reported Nielsen Norman Group’s hands-on test of AI-moderated interview tools, Marvin and UserFlix, with ten research leaders and ResearchOps professionals across eight countries evaluating the tools. They described rigid script-following, timing problems, and sycophancy with the AI interviewing products they evaluated.
Bottom line: Evidence relative to the competence of AI versus human moderation currently rests on two studies (Wuttke et al. with five AI-led and six student-led interviews; Zhu et al. with 30 each AI- and human-led interviews). The findings indicate that AI matches (but does not exceed) human skill for mechanical, procedural aspects of moderation, might do a little better at properly paraphrasing without leading, but is inferior to humans on judgment-dependent competence.
Do participants want to talk to AI moderators (rapport, trust, willingness)?
Rapport is more than just breaking the ice; it’s a tool to help establish trust and elicit deeper insights. How well does AI do this compared to humans?
There were five relevant studies for this question:
- Chopra and Haaland (2026) reported, “a majority of participants would prefer an AI interviewer over a human interviewer” (pro AI).
- Zhu et al. (2026) found, for most subgroups, a strong preference for human moderation (anti AI) but a preference for AI from their subgroup of introverts (pro AI).
- Wuttke et al. (2025) reported no difference in overall satisfaction between AI- and human-led moderation (neutral) but better Interestingness and Repeatability (willingness to do it again) ratings for human moderation (anti AI).
- Jacobsen et al. (2025) found a slight preference for disclosing to an AI chatbot relative to a human (14 AI, 21%; 6 human, 9%; pro AI), but most participants had no preference (44, 67%; neutral).
- Cuevas et al. (2025) reported “no significant differences … [in] preference for a human versus an AI interviewer” (neutral).
- Danó and colleagues (2025) studied Hungarian attitudes toward AI interviewers, finding 49% of respondents were reluctant to engage at all with a virtual interviewer (anti AI).
Bottom line: Results were mixed regarding participant preference for AI or human moderation. For this question, however, even if only a substantial minority of people are reluctant to engage with AI moderators, for many types of research this could be a problem for recruiting and increased likelihood of study abandonment.
Do findings generalize across models or prompts?
LLMs are changing weekly and are also probabilistic. If changing prompts and models changes results (for better or worse), it’s difficult to generalize the findings to practitioners.
For this question, we classify evidence of generalizability as pro AI and evidence against it as anti AI. There were four relevant studies:
- Chopra and Haaland (2026) reported using the same prompt across multiple studies and models with “no application-specific instructions, making it portable across domains” (pro AI).
- Panfilova et al. (2026) ran six different LLMs on an identical task, finding meaningfully different rankings and failure modes (anti AI).
- Jacobsen et al. (2025) reported different probing behaviors as a function of manipulating prompts for probing strategies (anti AI).
- Rosala (2026) reported that two commercial AI moderation tools (Marvin, UserFlix) had different strengths and weaknesses (anti AI).
Bottom line: Many authors noted the impermanence of rapidly changing models and different results for different prompts as a characteristic of generative AI. The moving target of models and prompts makes it difficult to conduct research on these topics. The model often credited with kicking off the AI craze in 2020, GPT-3, ran on 175 billion parameters with a context window of 2,048 tokens. Kimi K3, released in July 2026, carries 2.8 trillion parameters and a one-million-token context window. When the underlying technology changes that fast, any result you attribute to the model of the moment necessarily has a short shelf life.
Summary and Discussion
Is AI a suitable replacement for human moderators? A review of the (mostly) published literature found mixed results. Table 1 summarizes the results for the five questions, pro AI (6), neutral (6), and anti AI (9). Across the sources, pro- and anti-AI findings were about balanced.
| The Five Questions | Pro AI | Neutral | Anti AI |
|---|---|---|---|
| Do participants say less to AI than to human moderators (number of words, engagement, elaboration)? | 1 (Wu) | 1 (Zh) | 1 (Wu) |
| Do participants provide as much depth of insights to AI and human moderators (richness, novel insight)? | 0 | 1 (Zh) | 1 (Wu) |
| Does AI moderate as competently as a human (protocol fidelity, probing, task outcomes)? | 1 (Wu) | 1 (Wu, Zh) | 1 (Wu, Zh) |
| Do participants want to talk to AI moderators (rapport, trust, willingness)? | 3 (Ch, Zh, Ja) | 3 (Cu, Ja, Wu) | 3 (Da, Wu, Zh) |
| Do findings generalize across models or prompts? | 1 (Ch) | 0 | 3 (Ja, Pa, Ro) |
Table 1: Findings across the relevant sources, grouped by question. The two-letter codes are the first two letters of the lead author’s last name (e.g., Ch for Chopra, Cu for Cuevas). The same author appears in multiple columns for a question when the study findings were mixed.
AI moderators achieved surface-level adherence to practices like asking open-ended, non-leading questions and staying on the topic guide. Their most consistent weakness was judgment about when and how much to probe. Thus, on the more desirable outcomes of depth and richness, our review finds them falling short.
A potentially problematic issue for researchers is the possibility of low willingness of many participants to engage with AI moderators, leading to reduced recruitment and increased survey abandonment. So, when a client asks what we think about AI moderators, our answer is, “It depends on the job.”
For structured, high-volume, consistency-driven data collection, an AI moderator seems like a reasonable tool today. For discovery, emotional nuance, and studies where the whole point is to find what you didn’t know to ask about, the human moderator isn’t going anywhere yet. AI seems proficient with the mechanical aspects of this work but lacks in the key area of judgment.
Little of the work above is UX research, and none of it is ours. So, we decided to run our own study: a UX-focused evaluation of AI moderators against the work we do every day as human UX researchers. Stay tuned for that article.
Appendix: References and Summaries
Chopra and Haaland ran two large-sample open-ended qualitative interviews (one on stock market nonparticipation in 2023, n = 381; one on attitudes toward U.S. tariffs in 2025, n = 385) with a multi-agent AI chatbot at a scale and marginal cost no human team could match. The comparisons were between conditions with different levels of probing (a single open-ended question, multiple open-ended questions, AI-led interview with no probing, and AI-led interview with probing), claiming five times as many unique themes discovered for AI-led probing relative to a single open-ended survey item. Adding more open-ended questions or AI without probing closed about half that gap. Human coders judged the transcripts to hold substantive qualitative content, participants rated the experience favorably, and most participants approved of the AI-generated summaries of their statements presented after the interview. AI and human review of transcripts found the AI interviewers consistently adhered to core methodological guidelines for interviewing (open-ended, relevant, non-leading) as guided by the prompts (95% in the first survey, near 100% in the second). Most participants indicated a preference for AI over human moderation and a willingness to take future AI-led interviews.
Cuevas and colleagues ran a large user study (n = 399) pitting two LLM-based interview chatbots using GPT-3.5-turbo against a hard-coded-question baseline. They scored the transcripts in two ways: on established communication-quality metrics, and on a purpose-built “richness” scale: how well a response captured the complexity and specificity of the respondent’s actual situation. The chatbots scored well on the standard metrics but no better than the hard-coded baseline except for follow-up quality. Furthermore, the responses rarely surfaced a participant’s specific motives or personal examples, so they scored poorly on richness. Surface fluency does not necessarily lead to genuine qualitative depth.
Danó and colleagues studied Hungarian attitudes toward AI interviewers. They conducted a large-scale survey (n = 1077) in June 2024 that asked people how they would feel about being interviewed by AI with a human voice, finding that 49% of respondents were reluctant to engage at all with a virtual interviewer. An AI interviewer that half of your sample doesn’t want to use could create a data-quality problem.
Kuric and colleagues ran a between-subjects experiment (n = 60) comparing unmoderated usability debriefs either with or without real-time GPT-4 follow-ups, then compared four ways of asking follow-up questions within that data (none, researcher-authored static, GPT-4-generated, and a blend of the last two). The follow-ups added depth to known issues but did not surface any new ones. Participants also rated the questions as significantly less reasonable when the AI was probing, complaining they felt repetitive, and their answers to the primary seed question got worse when they knew AI follow-ups were coming.
Jacobsen and colleagues tested four different probe types (descriptive, idiographic, clarifying, and explanatory) embedded directly in an online survey (n = 64). In response to a participant’s statement, a descriptive probe asks what you were doing, feeling, and thinking; an idiographic probe asks for a specific example; a clarifying probe asks what it means to you; and an explanatory probe asks why you believe what you said. The idiographic version won overall and was the only one still working well by the evaluation stage. The explanatory “why” questions never came out on top.
Panfilova et al. ran a controlled protocol across six frontier models at the time of data collection. AI interviewers created with five of the models completed the protocol (Claude Sonnet 4, Gemini 2.5 Pro, GPT-5, Grok 4, Qwen3), and one (DeepSeek) was removed due to serious performance issues. The researchers started by collecting ten baseline interviews in which the moderator and respondents were human, getting responses to 54 questions with no follow-up. The AI interviewers reviewed the transcripts question by question, decided whether follow-up was warranted and, if so, asked the additional question. Questions were answered by an AI agent conditioned with the human respondent’s Big Five personality profile, with the possibility of additional probing. The appropriateness of following up and quality of the additional questions were assessed by human evaluators on five binary criteria (benevolence, necessity, context-awareness, openness, and justified skip). Interviewing competence varied substantially by model but with no clear winner. For example, Gemini was rated as the most empathetic model, while Grok produced the most follow-ups but tended to over-probe already complete responses. The best-performing models on quality were not the cheapest or fastest.
The Nielsen Norman Group’s AI Interviewers article reports a hands-on test of two AI-moderated interview tools, Marvin and UserFlix, with ten research leaders and ResearchOps professionals across eight countries (tool evaluators rather than recruited respondents). Their verdict: AI interviewers suit a bounded set of uses (product-feedback collection, recruitment screening, translated interviews, and teams without a dedicated researcher), but not exploratory research, high-stakes decisions, or work that demands domain expertise and real-time judgment.
Wei and colleagues built chatbots that took on four roles: a sleep expert, a dietitian, a life coach, and a fitness coach. They manipulated prompts to create four AI interviewer types by crossing two formats (structured list and descriptive narrative) and two personality modifiers (with and without “who always shows empathy and engages my customer in conversations” in the prompt), abbreviated SP, SN, DP, and DN, with 12 participants assigned to each version (n = 48). Participants spent less than 20 minutes completing eight conversations with their assigned interviewer type (four roles by two scenario paths, one positive and one negative). A key dependent measure was the percentage of 18 predefined information slots (facts) the AIs were directed to gather during the conversations. Using nothing but the LLM itself (no training data and no fine-tuning), on average the chatbots captured about 79% of the targeted facts (SP: 83%, SN: 77%, DP: 72%, DN: 83%). How the prompt was written mattered. The researchers reported that the personality modifier improved performance for the structured format but hurt the descriptive format, even though both formats contained the same factual information.
Wuttke and colleagues ran a small classroom pilot study in which student pairs each completed one AI-led (using GPT-4 Turbo) and one human-led interview in randomized order, rotating roles. The researchers analyzed six human-led and five AI-led sessions, using an identical questionnaire on politics and democracy. The AI matched the student interviewers’ overall rate of guideline violations, but the errors differed in kind. The AI consistently missed chances to ask a follow-up on unexpected or unclear answers (88% of the violations of this rule). The students, on the other hand, often missed engaging in active listening, which includes restating what the participants had just said (94% of the violations of this rule). Both students and AI interviewers behaved in ways that potentially biased respondents’ answers. Participants rated the AI-led interviews as less interesting and were less willing to repeat them, even though task-level measures like response quality and understanding came out similarly. This finding is complicated by the respondents in each condition having also played other roles (interviewer, observer) across the two conditions.
Zhu and colleagues ran a randomized controlled trial of an agentic voice moderator (using Chinese ByteDance technologies like Doubao-1.5-pro-32k-250115) against a human in think-aloud testing of a note-taking app (n = 60). They found no significant difference in participants’ task performance or verbalization behavior, but significantly lower social-perception ratings such as anthropomorphism, animacy, likeability, intelligence, and social presence for the AI. The verbalization behavior of AI and human moderators differed, with the AI moderator speaking more, faster, and more frequently than the humans during the think-aloud phase, behaviors that likely lowered the AI’s rapport score.