Can We Trust AI to Moderate UX Interviews? – MeasuringU


Feature image showing an AI moderatorCan AI Replace UX Researchers?

Some criticized us for using that headline in 2023, calling it clickbait. Fair criticism, given the inflammatory nature of AI (for the record, we did have a subtitle). But here we are three years later, and we’re not talking about just the tedious task of having AI code comments—a task that most researchers are happy to offload. We’re talking about AI being able to do some of the core jobs of the UX researcher, in particular, moderating participant sessions.

When we wrote our article three years ago about AI doing comment coding, the idea of replacing a UX moderator seemed about as probable as replacing a waiter in a restaurant. Both jobs seemed immune to AI. Yet here we are. AI isn’t just crunching numbers; it’s being tasked with doing qualitative work. It’s a serious enough threat that 419 professionals have signed letters against the use of AI for qual research.

And just like we did with synthetic users, we need to separate the hype from the data. That means asking what the claim actually is, who it comes from, and how strong the evidence behind it is. Is it anecdotal? Is it peer-reviewed? Is it from a company selling AI moderators?

Before we dig into the claims about AI and moderation and start letter-writing campaigns, it’s important to understand what moderating is, what an AI moderator is, and what one can do.

What Is Moderating in UX Research?

Moderating is a general term that’s not to be confused with content moderation on social networks. In UX research, moderating is a broad term that spans several UX research methods: usability testing, interviews, contextual inquiry, and occasionally focus groups. The moderator’s job changes depending on the method and goals.

An unstructured in-depth interview where you need to uncover problems and opportunities for a new product requires a different approach than a large-scale summative evaluation where you need minimal interaction to collect metrics (yes, you can moderate and collect metrics).

What Makes a Moderator Good?

Beyond knowing the study type, the practical skills that separate strong moderators from weak ones are things like building enough rapport that participants give honest reactions, probing appropriately (asking a follow-up once or twice to get past a surface answer, without dragging out a dead end), catching participants who are misrepresenting themselves or reading from a script, using silence to let someone think rather than rushing to fill it, and knowing when to go off-script versus when to stick to the guide.

None of this is really a checklist. Instead, it’s judgment applied in real time to what a specific person just said. Even with that list, there is no agreed-upon way to score whether a moderator did a good job. What makes a good moderator isn’t even clearly defined, let alone measured objectively.

What Is an AI Moderator?

An AI moderator, as the name suggests, is software that uses AI to interactively interview participants using voice and even video to replicate the interactivity of a human. But it’s more than software reading a script and asking questions. An AI can be prompted to probe and follow up based on what the participants say. The technology has advanced considerably since Dragon NaturallySpeaking in the 1990s and Google’s cutting-edge transcription failures of the 2010s. AI conversations now flow with high accuracy and little latency.

AI Moderators (Not Quite HAL 9000)

While the technology has come a long way, AI moderators are not like the diabolical AI in 2001: A Space Odyssey. They are limited by the script you feed them and in their interactive ability.

An AI moderator is software (prompts) built on top of the same LLM models that make headlines (Claude, ChatGPT, Gemini). The moderators can be voice-only or have a range of appearances, from simple visualizations to more life-like avatars.

ChatGPT, Gemini, and Claude all ship voice modes that hold unstructured back-and-forth conversations, interruptions included, with no command list behind them. Figure 1 shows the appearance of ChatGPT’s voice interface compared with HAL 9000 from 2001: A Space Odyssey.

ChatGPT’s visual placeholder for its voice.HAL 9000 

Figure 1: ChatGPT’s visual placeholder for its voice (left) and HAL 9000 (right).

An example of a simplified prompt for the LLM could be something like the following for doing research on a new sleep-aid product:

You are a UX researcher running semi-structured interviews with participants in a research study. The goal of the study is to find out more about their sleeping behavior and to find potential market opportunities for wake-up devices. Follow the discussion guide you are given directly, covering all required questions, but follow up on interesting threads.

AI moderators can also have avatars to go with the voice, so respondents aren’t talking to a circle or blank screen. Figure 2 shows an example of one used for job interviews.

Example of a visual AI moderator used for job interviews from Humanly.io.

Figure 2: Example of a visual AI moderator used for job interviews from Humanly.io.

AI Moderation in Job Interviews

We can talk about AI moderators as a concept, but are they really being used at scale? One clear application: interviewing for a job, where services like CodeSignal, Humanly, or Eightfold offer AI interviewers.

One estimate suggests over 60% of job seekers have had exposure to an AI interview. Note that this estimate comes from a company selling AI moderating technology, so it’s not necessarily objective.

However, a large pre-registered experiment from 2025 with over 70,000 job applicants in the Philippines provides a more objective data point. Participants were randomly assigned to either a human interviewer (20% of sample), an AI job interviewer (60% of sample), or allowed to choose between the two (20% of sample) after a brief intro to the AI interviewer.

Candidates interviewed by the AI were 12% more likely to get an offer, 18% more likely to start, and 17% more likely to still be employed after a month. Of the subset that had a choice, a surprisingly high 78% picked the AI.

These numbers are hard to ignore. Consider a few caveats, though, before we fire all the job interviewers and UX moderators out there. First, this was for entry-level customer service jobs (call center) in the Philippines. Second, humans made the hiring decision, not AI. Probably most importantly, the value that was derived was in the perception and reality of a consistently delivered interviewer. The type of questions being asked were mostly closed, factual, verification-style questions, things like commute time, salary fit, availability, and contact info.

Job interviews are notorious for both explicit and implicit bias. There does seem to be a benefit if those biases can be significantly reduced (or even eliminated).

Interviewing a candidate for a job isn’t the same as interviewing a participant to understand usage patterns. While we are concerned about identifying and reducing bias in UX research, UX moderation is often best deployed for unstructured and unknown problems.

Job interviews, especially screening interviews, are usually very structured with a clear set of questions. They are designed to be consistent. A structured UX interview is really more akin to a verbalized survey (something we’ll revisit). Where interviews really matter is when you don’t know what you don’t know.

AI in Qual Research

Job interviewing is qualitatively different from the sort of inquiry in moderated research. In one of the best-known AI interviewer studies, in December 2025, Anthropic conducted a large study of 80k Claude users across 159 countries using their new Anthropic Interviewer tool. Their interview questions were a bit closer to the types of topics investigated in UX moderated interviews:

  1. What’s the last thing you used an AI chatbot for?
  2. If you could wave a magic wand, what would AI do for you?
  3. Has AI ever taken a step towards that vision for you?
  4. Are there ways AI might be developed that would be contrary to your vision or what you value?

Anthropic Interviewer followed up on each, probing for the underlying values and experiences behind people’s answers. Claude was then used to synthesize the transcripts for themes (much to the chagrin of the 419 professionals who rejected such usage).

Claude’s analysis determined that 88% to 98% of the open-ended comments were “substantive.” There aren’t many details to assess the quality relative to a human, but the authors reported that a subset of comments were validated as having at least 90% agreement with human coders on 25 labels. The authors were surprised by how candid some people were. Respondents shared things like grief, mental health crises, financial precarity, and relationship failures, all of which our human user researchers rarely encounter in traditional interviews. It’s unclear if those same comments would have been shared with a human, but it’s certainly a possible benefit worth investigating more when a topic may elicit more sensitive comments from participants.

What is clear is that this is a huge sample that would almost certainly never happen with human moderation. It’s less clear how well the insights from a smaller sample conducted by humans would have performed.

Anthropic isn’t the only one in the AI interviewer game, and others have noticed the self-disclosure. Marvin’s AI Interviewer “leads human-like conversations, searches for the why behind responses, and gathers insights faster than ever before.” Or, from Perspective, “customers share things in these conversations they’d never put in a form and would rarely say on a Zoom call with a stranger. Not a chatbot, not a survey, not a junior researcher reading from a script—the best interviewer you’ve ever seen, available at any hour, in any language, for every single customer.”

Bold vendor claims like these are worth investigating. NN/Group’s analysis of ten experienced researchers using two AI moderators (including Marvin from above and UserFlix) found that AI interviewers can handle structured, scripted interviews (at scale). However, they didn’t think AI interviewers were adequate yet for semi-structured interviews.

Are We Ready to Deploy the AI Moderators?

So, we have some data suggesting that AI can be used at scale to interview job candidates. We have at least a proof of concept that AI can moderate a massive number of sessions for a handful of more open-ended research questions. And there’s some evidence that people may be more willing to disclose more to an AI moderator than to a human.

But if you recruit a dozen IT decision makers, or Chief Product Officers, and want to build them a better product by conducting a semi-structured interview that requires a moderator to go off script, would you use an AI moderator? NN/Group certainly suggests we aren’t there.

But this raises the question of how well an AI moderator would do compared to a human. That’s a question we’ll take up in our upcoming articles. First, we’ll review the published literature, and then we’ll report on the results of our own controlled experiment, where we compare an AI moderator with a human moderator in a UX research context.

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Register New Account
Compare items
  • Total (0)
Compare
0
Shopping cart