Do Participants Say Less to AI Than to Human Moderators? – MeasuringU


Feature image“Are AI moderators as good as human moderators?”

That’s a good question. It’d be good to have data to help answer that. Instead, we have claims from vendors selling AI moderator technology. That’s not nothing, but it’s hardly the objective data you’d want before you retire your research team from moderating.

When we encounter interesting claims like that, we generally take the following approach:

  1. Define the claim.
  2. Assess evidence and the quality (focusing on peer-reviewed publications).
  3. Conduct our own experiment.
  4. Refine and revise our opinion of the merits of the claim.

We completed step two in our previous article, where we reviewed the published literature. There wasn’t much literature to review, and some of the findings may be a bit dated. But even newer work is lacking because most studies compared AI moderators to surveys or static follow-up questions, not to a live human moderator. Only two studies actually put an AI moderator head-to-head against a human one.

One was a small classroom-style study (Wuttke et al., 2025). Student pairs each ran one AI-led and one human-led interview on politics and democracy, netting just five AI-led and six human-led sessions, with students standing in for trained moderators.

The other (Zhu et al., 2026) was a properly powered randomized control trial (RCT) with 30 sessions per condition, pitting an agentic voice moderator against a human moderator in think-aloud usability testing of a note-taking app. It’s a stronger study, but think-aloud moderation is a lower-demand job than in-depth-interview moderation. In a think-aloud session, the moderator’s main task is prompting someone to keep narrating what they’re already doing. It’s not that far removed from unmoderated think-aloud, where there’s no moderator at all, just a periodic “keep talking” nudge or “tell me more.” A semi-structured or in-depth interview asks much more of a moderator: building rapport from a cold start, deciding in real time whether an offhand comment is worth five extra minutes, and steering a conversation that has no task to anchor it.

In the Zhu et al. study, they found no difference in the number of words participants produced, but the AI moderator talked more, faster, and more often during the session itself. The AI moderator also scored significantly worse than the human moderator on context-aware follow-up and on building trust, even though the two were indistinguishable on procedural adherence.

That review left us with five questions that the literature couldn’t settle:

  • Do participants say less to AI than to human moderators?
  • Do participants provide the same depth of insights to AI and human moderators?
  • Does AI moderate as competently as a human?
  • Do participants want to talk to AI moderators?
  • Do findings generalize across models or prompts?

Each of those questions needs its own analysis, and all require more data. To help us (and the industry) answer these questions, we built our own AI moderator and ran a controlled experiment. In this article, we’ll cover the setup of that experiment and address the first question of whether participants say less, more, or about the same to AI moderators as human moderators. Of course, speaking more isn’t necessarily better, but if humans systematically say LESS to an AI moderator, that would create fewer opportunities for uncovering insights.

Experimental Study Setup

To test an AI moderator, you need a few things: an AI moderator, human moderators, a study goal and discussion guide, and human participants. The last three are staples of UX research; AI moderators are the new additions.

AI Moderator

There are AI moderators on the market, but we wanted to build our own for two reasons. First, we wanted to understand how prompting and programming decisions affect the output around probing decisions, especially as models change. Second, we wanted the AI moderator to look as realistic as possible (and uncannily similar to the real human moderators in the study). We built our AI moderator using an AI agent and created photorealistic avatars with voices of two of our researchers (see Video 1).

Video 1: A quick comparison between our AI and human moderators.

The AI moderator used Gemini 3 Flash. To minimize latency and the risk of hallucinations due to the context window overflowing, the AI moderator was made up of several agents working sequentially. Each agent covered a broad topic, summarized its findings, and then handed those findings off to the next agent. This happened behind the scenes, so to participants, it seemed like one fluid conversation with latency times about 400ms longer than human moderator latency.

Research Topic and Discussion Guide: Sleep Tech

In the study, instead of asking a simple set of questions (like asking people about their AI usage), we wanted to simulate a realistic exploratory study that would benefit from an in-depth interview. We generated a study for a fictitious company looking to build a product to help with waking up and sleep management. We selected sleep tech as the research topic because it was both realistic (there are many sleep tech products) and relatively easy to find recruits because, well, everyone sleeps.

We built discussion and moderator probing guides (Appendix A) that generated questions to help understand people’s sleep patterns, current pain points, waking habits, and any techniques and tech they may use.  The goal of the research was to find current frustrations with people’s wake-up methods and to identify potential new product opportunities for waking up. For details, see the appendix.

Recruitment and Screener

Twenty participants were recruited using a convenience sample from the Denver metro area to participate in a 30-minute in-person interview. Before the interview, participants completed a screener covering topics like caffeine use, melatonin intake, sleep accessories, and living arrangements (e.g., having kids or pets) that may affect sleep. We also captured AI sentiment (favorable or unfavorable) so we could assess and evenly balance AI attitudes across conditions. If the AI moderator underperformed, it wouldn’t be because we stacked the AI-moderated group with AI skeptics. Both human and AI moderators were instructed to probe beyond the discussion guide whenever it served the research goals.

Unlike a typical study, neither human nor AI moderators had access to the screener data beforehand. We withheld that information to see how well the moderators would be able to uncover prior established facts. This would become a key quantitative measure of probing effectiveness that we’ll cover in a future article.

Participant Assignment: 2 × 2 Factorial Design

We randomly assigned the 20 participants to a human moderator or AI moderator condition, crossed with moderator gender. This was a full 2 × 2 between-subjects design with five participants per cell. Two MeasuringU researchers, one female and one male, ran the human sessions. To keep the conditions as similar as possible, the AI moderators were built using the voices and likenesses of the human moderators, resulting in ten AI moderated sessions and ten human moderated sessions.

Participants were greeted upon arrival at our Denver labs and escorted by a different MeasuringU researcher who did not participate in the sessions. Participants didn’t know ahead of time if they were to be paired with a human or AI moderator. Participants sat in our lab and the moderators connected from the other side of the one-way mirror, avoiding direct human contact with the participants. Participants were connected to a video call with either the AI or the human moderator.

The 20 participants split evenly across the two conditions (ten AI-moderated, ten human-moderated). Gender skewed slightly male in both groups (7/3 in the AI condition, 6/4 in the human condition) and ages spanned the same 25–65+ range in each. Attitude toward AI (Enthusiast, Neutral, or Skeptic) was reasonably balanced across conditions, so a weaker AI-moderator showing couldn’t be chalked up to one group being stacked with AI skeptics.

Attitude toward AI
AI moderator (n = 10)
Human moderator (n = 10)
Enthusiast
4
6
Neutral
3
2
Skeptic
3
2

Table 1: Attitude toward AI by moderator condition (n = 20).

Study Results

So, do participants say less to AI than to human moderators?

In short, yes. And not just a little less, a lot less. We used several methods to assess how much participants spoke to AI versus human moderators (Table 2).

AI Human
Session Duration (Min) 14 22
Moderator Words Per Minute (WPM) 99 55
Participant Words Per Minute (WPM) 54 98
Participant % Share of Talking 34% 64%
Participant to Moderator Word Ratio 0.5 1.8
Delay in Responding (seconds) 1.9 1.5

Table 2: Comparison of AI and human results.

AI Moderated Sessions Were 37% Shorter

Participants were scheduled for 30-minute slots, but as is typical with research sessions, that 30 minutes included set-up time and some buffer time. Once the interview started with the AI or human moderators, participants were told the interview would last about 20 minutes.

Human sessions ran on average 22 minutes, whereas AI-moderated sessions lasted only 14 minutes (37% shorter). Or put another way, participants spent 57% more time with human moderators than AI moderators despite having the same script and planned session duration.

AI Moderators Spoke 80% More Than Human Moderators

Because average session duration differed a lot between human and AI moderators, we normalized the words spoken by both moderators and participants by dividing the total number of words spoken per session by the total session minutes to derive words spoken per minute (WPM). Human moderators spoke on average 55 WPM compared to AI moderators’ 99 WPM. That is, on average, AI moderators were speaking 80% more than human moderators.

Participants Spoke 45% Less to AI Moderators

Moderators speaking more isn’t necessarily bad as long as participants are also speaking. The data told another story. Participants spoke 98 WPM to human moderators compared to 54 WPM to AI moderators (45% less—see Video 2 for examples).

Video 2: Examples of different participant responses to the same question with an AI and a human moderator.

AI Moderators Dominated the Talking

Another way to look at the amount of time participants speak to AI moderators versus humans is to look at the ratio of words spoken per session. The ratio tells a similar story, with participants speaking 64% of the words in human-moderated sessions compared to 34% in AI-moderated sessions. In nine of the ten human-moderated sessions, participants out-talked the moderator; in ten AI-moderated sessions, participants spoke more than the moderator only three times.

We also looked at the ratio of words spoken by participant to moderator. Like the other metrics, AI moderators just spoke more. For every human moderator word, we got 1.8 words out of participants. With the AI moderators, every word generated only 0.5 words (less than 1/3 than in human sessions).

This result differs from the Zhu et al. study, where they found that the total participant word count didn’t change much between conditions. That could be a function of their research context being think-aloud during a usability test versus in-depth interviewing (which requires the moderator to draw more out of participants).

AI Moderators Spoke a Bit Faster and Responded Slightly Slower

We measured the time actually spent talking and found the AI moderator was about 11% faster on average than the human moderators (consistent with Zhu et al.). Participants spoke at the same rate no matter who they were talking to, so what changed was who had the floor, not the speed of talking.

When we looked at the delay between the end of participant speech and the beginning of moderator speech, we saw a slight difference in response time. AI moderators took 1.9 seconds to respond versus 1.5 seconds for human moderators. Human moderators were slightly faster in responding on average but it’s not clear whether participants even perceived differences this small (400ms on average).

Summary and Discussion

A controlled experiment with 20 participants randomly assigned to human or AI moderator in-depth interview conditions found that participants spoke less to AI moderators. Specifically, AI moderators’ sessions were 37% shorter, participants spoke 45% less to AI moderators, and AI moderators dominated the conversations, speaking 64% of the words. Human moderators spoke only about 34% of the words, a ratio that matches the typical expectation of moderators: let the participant do the talking. This is just one study, but the results match those of Zhu et al., who also found that AI moderators talked more, faster, and more often than human moderators.

Of course, the number of words and amount of speaking doesn’t mean AI moderators are bad or useless. In fact, future prompt modifications could improve results. And even if participants still talk less to AI moderators than humans, if the same or more information is extracted from participants and the same conclusions are reached, then it might not matter who’s talking too much. We’ll examine the content of the sessions and what’s extracted in future articles where we dig more into the experimental data.

Appendix

Discussion Guide

Sleep schedule: Tell me a bit about your sleep schedule. When do you usually go to sleep and wake up? Probe if needed: Do you wake up at the same time every day, or do you wake up at different times depending on the day?

Wake-up method: How do you make sure you’ll wake up at the time you intend to? (Most likely they will mention some kind of alarm clock, app, or device.) What is it, how does it work, and does it ever fail? (If they say someone else wakes them up) How do they wake you up? Are there ever any days when they aren’t available to wake you? What do you do to wake up on time on those days? (If they previously mentioned they wake up at a different time on different days) Since you wake up at a different time depending on the day, what do you do to make sure you’ll wake up when you want to each day?

Sleeping arrangement: Do you sleep alone, or does anyone else sleep in the same room with you? (If they share a room) Do they wake up at the same time you do? (If no) How do you make sure each of you wakes up on time without disturbing the other?

Current experience: How has your experience been so far using your current wake-up method? What do you like about using it to wake up? What problems or difficulties have you experienced? What workarounds or backup solutions have you used after running into those problems?

Ideal wake-up: What would be your ideal way to be woken up in the morning?

Moderator Probing Guide

1. GOAL: Cover ALL required probe areas AND uncover deeper insights related to RESEARCH GOALS.

2. REQUIRED PROBE AREAS must be addressed (non-negotiable, regardless of time) – adapt phrasing based on context.

    • If already mentioned: Confirm or dig deeper rather than re-asking.
    • If not yet covered: Ask directly but naturally

3. WHEN TO MOVE ON FROM A THREAD:

    • You’ve exhausted it (participant repeating themselves or has nothing more to add).
    • The participant explicitly says they don’t know or has nothing more to add.
    • It’s a minor detail that doesn’t advance research understanding.

4. BALANCE: Cover all required probes + follow valuable tangents related to research goals. Don’t chase irrelevant details or go too in depth on minor details.



We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Compare items
  • Total (0)
Compare
0
Shopping cart