Using AI To Find Usability Problems: A Replication – MeasuringU


Feature image showing AI finding usability problemsIf AI can code problems and potentially moderate a structured interview, can it also reliably find usability issues from watching a video?

We can talk about hypotheticals, or we can actually try it out. We tried it out.

In our first study, we started small. We compared human and AI reviews of a video taken from a usability test on making dinner reservations with OpenTable. We used the standard ChatGPT and Gemini interfaces (standard Claude doesn’t support video review).

The good: AI identified roughly half the usability problems identified by human UX researchers. It even uncovered one genuine problem not found by any of the four human evaluators.

The bad: AI identified seven false alarms. They were true statements but not connected to the participants’ actions or utterances.

The ugly: Three problems were straight-up hallucinations. AI described events that just didn’t happen.

So, of the eleven problems the two AIs reported that no human flagged, only one was a genuine problem. That’s a useful number to keep in mind: 91% of the AI-only problems in this study required either correction or dismissal.

A case study like that, however, is limited. There was one video, two AIs (ChatGPT, Gemini), and one prompt. Replication helps with generalization.

To move beyond that first case study, we kept all experimental variables the same and conducted the same type of analysis on a similar video from a different usability study.

Experimental Design: One Human Researcher, Two AIs, and One Video

For this research, one UX researcher with decades of experience (Jim, the lead researcher from the previous study) reviewed a similar video from a previous usability benchmark study of online pet websites multiple times, creating a timeline of key events and a list of observed usability problems for this participant (referred to as Participant A; see Appendix A for the details of the timeline).

The characteristics the new video shared with the first one were:

  • A similar task (steps toward booking a specified appointment)
  • A similar outcome (the user experienced several usability problems but was ultimately successful).

Using the same prompt each time, we then ran the video four times each through the same two AIs used in the previous case study (ChatGPT-5.4 Thinking and Gemini 3 Flash Thinking).

So, in this study, we held constant the videos, the key elements of the prompt, and the AI versions/settings—variables that we eventually plan to vary. This time, as in the first case study, we only varied the type of analyst: human, ChatGPT, and Gemini.

The Task

During the previous usability benchmark study conducted at MeasuringU in 2019, participants used the PetSmart website to start the process of booking a grooming appointment in Glendale, CO, for a bath and full haircut for an English Springer Spaniel older than six months. The task was successfully completed if the participant found that specific grooming option and reported the listed price of $61. For the step-by-step details of the “happy path” to complete this task, see Appendix B.

The Prompt

The prompt we used for this study was:

During a usability test, the facilitator must keep track of participant behaviors as they navigate through tasks on a website, mobile app, software program, etc. We’d like you to watch a video of a usability test where participants were asked to book a grooming reservation for their dog. As you’re watching, please look for problems the participant has while attempting to complete the task. For example, you can document the path users take, describe issues they encounter as well as what on the website might be causing problems. The task has been successfully completed if the participant finds the target service (“Bath & Full Haircut” which costs $61; not “Bath & Full Haircut with FURminator” which costs $74). If you understand these instructions, let me know and I’ll drag the video in for you to review. Are you ready for the video?

This was based on the prompt we used for the OpenTable case study with slight modifications. The task details are necessarily different because there was only one correct choice for this new task (compared to many correct choices for the OpenTable task), so we specified the end task details required for successful completion.

Major Findings

Even though the participant successfully completed the task in the video, a total of ten usability problems were identified by the AIs and the human researcher.

Usability problems aren’t like observing a visual defect in a product. They require judgement, so it’s worth digging into what these problems are because they’re at the crux of how AI may or may not be able to effectively emulate UX researcher judgement (for now). So, let’s dig into what we saw.

Participant A started with a few clicks not on the happy path, but on the third click got the task started, so these two early clicks could be considered minor usability problems that were quickly corrected.

There were two more impactful usability problems, both associated with breed selection (Figure 1). First, there was no visible indication that the dropdown list could be filtered by typing over the word “breed”; the participant scrolled through the entire list. Second, the order of presentation of the breeds in the dropdown list was alphabetically inconsistent (e.g., “English Toy Spaniel” started with E; “Springer Spaniel – English” started with S). Despite this, the participant successfully completed the task, just not on the most efficient path.

a: Breed list at boundary of D and E—”English Toy Spaniel” but no English Springer Spaniel

Image showing breed list with English Toy Spaniel but not English Springer Spaniel

b: Location of “Springer Spaniel – English”

Breed list showing Springer Spaniel - English

Figure 1: Participant A’s usability issues with breed selection.

Some Agreement Between Human and AI, No Hallucinations, Six False Alarms

Table 1 shows the four problems discovered by the human researcher (three of which were also identified by the AIs) and the six additional problems reported only by the AIs. We looked to see if the problems identified only by the AIs were false alarms (an event happened but not really a usability problem) or hallucinations (the event just didn’t happen).

The good news was that none of the six problems were hallucinations. The bad news is that they were all determined to be false alarms. Either the participant never actually noticed the issue flagged by the AI or wasn’t affected by it (e.g., the location prompt, the below-the-fold item), or it was just normal, expected system behavior rather than a flaw (e.g., the ZIP search returning multiple locations, the menu reloading).

# Problem description Human ChatGPT Gemini Why (if false alarm)
H1 Clicked Shop by Pet from top menu (incorrect first click)
Y
Y
Y
H2 Clicked search field (incorrect second click)
Y
H3 No attempt to filter breed list by typing; scrolled instead
Y
Y
Y
H4 Searched breed list for Springer Spaniel, couldn’t find it alphabetically
Y
Y
Y
C1 Blank/loading state after selecting Grooming Salon
N
Lasted ~2 seconds; participant had already moved on
C2/G1 Browser location prompt appears alongside site’s own location modal
N
N
No sign participant noticed it; used site’s own ZIP entry instead
C3 Entering ZIP returns multiple grooming locations
N
Accurate, but that’s how the control is designed to work
C4 Menu reloads after selecting breed and age
N
Expected behavior after Check Prices & Book Now; no impact
C5/G2 Target service is at the bottom of the list, below the fold
N
N
True, but not actually a problem for this participant
C6 Promotional content crowds out the service menu
N
True, but not an obvious problem for this participant

Table 1: Summary of usability problem discovery by the human researcher and the AIs. In the # column, H indicates a problem identified by the human researcher, C indicates a problem identified by ChatGPT, and G indicates a problem identified by Gemini. In the Human, ChatGPT, and Gemini columns, Y indicates the discovery of a verified usability problem, and N indicates the reporting of a false alarm. The full runs of ChatGPT and Gemini are shown in Appendix A.

As we did in our first study, we ran the videos four times through the AIs (because of the probabilistic nature of how they work). Problems identified in any of the four runs were included in this analysis. We used the mean any-2 agreement to assess overlap.

Technical note: Our preferred method for quantifying the correspondence between two lists of usability issues is any-2 agreement. Any-2 agreement is the ratio of the intersection of the two sets divided by their union. Historically, we’ve found an any-2 agreement of 50% to be average (typical), around 25% to be low, and around 75% to be high.

ChatGPT Agreement: 35%

The mean any-2 agreement of the four ChatGPT runs and the UX researcher was 35%. ChatGPT identified (at least once) three of the four problems reported by the researcher but also produced six false alarms.

Gemini Agreement: 38%

The mean any-2 agreement of the four Gemini runs and the UX researcher was 38%. Gemini identified (at least once) three of the four problems reported by the researcher but also produced two false alarms (matching two of the false alarms produced by ChatGPT).

The mean any-2 agreement between the four runs of the AIs was 40%. Figure 2 shows the Venn diagram for the problem discovery results for the UX researcher and the AIs.

Venn diagram of usability problem discovery for Participant A by the human reviewer, ChatGPT, and Gemini.

Figure 2: Venn diagram of usability problem discovery for Participant A by the human reviewer, ChatGPT, and Gemini.

The Venn diagram illustrates the relationship between the human reviewer and AI analyses of Participant A. The human reviewer identified four usability issues, three of which were identified at least once by an AI. However, there was one usability problem that was not caught by either AI, while the AIs produced six issues that were not legitimate usability problems (all false alarms, no hallucinations). The AIs did not discover any real problems that the human reviewer failed to identify.

Comparison with the OpenTable Case Study Results

Figure 3 shows the Venn diagram from our OpenTable case study.

Venn diagram of usability problem discovery from our OpenTable case study.

Figure 3: Venn diagram of usability problem discovery from our OpenTable case study.

At a glance, the diagrams in Figures 2 and 3 have some similarities and some differences. To help with comparison, Table 2 shows the side-by-side comparisons for different aspects of the results (graphed in Figure 4).

Comparison Participant A OpenTable Abs. Diff.
% human identified 40% 45%  5%
% AI & human overlap 30% 30%  0%
% AI-only identified 60% 55%  5%
% AI errors 60% 50% 10%
% AI false alarms 60% 35% 25%
% AI hallucinations  0% 15% 15%
% AI-only discovery  0%  5%  5%
Total unique problems 10 20 10

Table 2: Comparison of PetSmart Participant A and OpenTable problem identification rates.

Figure 4: PetSmart Participant A and OpenTable problem identification rates.

In most respects, this second evaluation of AI problem discovery has replicated the first case study (OpenTable).

The usability problem identification rates were similar (within 10 percentage points) for the percentages of usability problems identified by the human UX researchers, AIs, both (the AI/Human overlap), and AI-only discovery.

Observed differences in the patterns were due to the incidence of AI hallucinations in the OpenTable case study (3) compared to none in the AI outputs for Participant A.

Summary and Discussion

Our key findings were:

Successful replication of the OpenTable case study. Most metrics comparing the two videos were within 10 percentage points of each other (e.g., 40% vs. 45% of human-identified problems found by AI). The one real divergence was in the type of AI-only errors: this PetSmart video had more false alarms (60% vs. 35%) but zero hallucinations.

A troubling number of false alarms. Focusing on the new data for PetSmart Participant A, the total number of unique usability problems was ten, of which only four were identified by the UX researcher (which we treat as ground truth). Of the six unique problems the AIs reported that the human researcher did not flag, all were false alarms (no hallucinations).

AI adds potential value as an overly enthusiastic junior researcher, not a trusted expert. In the analysis of these videos, the AIs discovered three of the four usability problems reported by the UX researcher. Relying only on these multiple runs of the AIs would have missed a quarter of the real usability problems. Unlike our earlier research with the restaurant reservation video, where the AIs found one problem that UX researchers missed, the AIs in this study did not identify any valid usability problems that the UX researcher failed to discover.

Like humans, AI usability reviews of videos are prone to the “evaluator effect.” Just like human evaluators, multiple runs of AI usability evaluations of videos are not perfectly consistent, so it’s good practice to run these evaluations multiple times for consistency checks. Running multiple evaluations and looking for consistency across runs is a practical filter before any human review.

Bottom line—AI usability reviews of videos require human oversight. In their current form (what we tested), these AI products can add some value to this type of UX research, but more as junior researchers whose actions and conclusions require expert human oversight rather than as trusted experts themselves.

Future research: Our next step in this research program is to perform the same analyses on two PetSmart videos in which the user experiences were different from PetSmart Participant A and OpenTable regarding task success and number of problems identified by the human UX researcher (one who did not complete the task successfully and one who experienced no problems completing the task).

Appendix A: Detailed Timeline and Problem-by-Problem Tables

Key Events Timeline for Participant A

Appendix Table 1 summarizes the key events in the video (compiled by the UX researcher), identifying four problematic events deviating from the “happy” path (see Appendix B).

Event # Timestamp Summary of key user actions Notes
1 0:00:13 Clicked Shop by Pet from top menu Problematic event
2 0:00:21 Clicked search field Problematic event
3 0:00:24 Clicked Pet Services
4 0:00:44 Grooming form begins to appear
5 0:00:45 All form elements except dog/cat buttons appear
6 0:00:46 Dog/cat buttons appear—participant cursor was headed for breed but when these buttons appeared he changed course toward the dog/cat buttons
7 0:00:49 Clicked Dog button
8 0:00:56 Clicked “select” link by Select a Store
9 0:00:58 Find a Grooming Salon Near You pops up (one field: “Zip Code, City or State” and Search button)
10 0:00:59 Location permission prompt appeared at top of screen apparently triggered by presentation of the Use My Current Location link in Find a Grooming Salon Near You
11 0:01:09 Typed zip code from task instructions and clicked Search
12 0:01:13 List of locations appears with target Glendale at top
13 0:01:19 Clicked Glendale
14 0:01:21 Clicked x to clear the location permission prompt
15 0:01:29 Clicked Breed dropdown, list of breeds appears
16 0:01:30 Made no attempt to filter list by typing; instead started scrolling through the long list Problematic event
17 0:01:35 Searching list for English Springer Spaniel, scrolling up and down between the D’s and E’s but couldn’t find the breed there Problematic event
18 0:01:56 Scrolling farther down the list found “Springer Spaniel – English”
19 0:02:00 Clicked Age dropdown
20 0:02:04 Selected “6 months or older”
21 0:02:07 Clicked Check Prices & Book Now button
22 0:02:13 Grooming Salon Menu appeared
23 0:02:32 Scrolled through list of services to the bottom
24 0:02:35 Did not click service but said, “Just bath and full haircut, $61, let me write that down.”
25 0:02:36 Task successfully completed (correct location and service)

Appendix Table 1: Timeline for Participant A.

ChatGPT Problem-by-Problem Results

Appendix Table 2 shows the usability problems reported by ChatGPT and the UX researcher, run by run.

Prob #
Description
Run 1
Run 2
Run 3
Run 4
H1 Clicked Shop by Pet from top menu (incorrect first click)
1
1
1
1
H2 Clicked search field (incorrect second click)
C1 Page shows a mostly blank/loading state after selecting Grooming Salon (FALSE ALARM—this happened but only lasted about two seconds within which the participant started to select Breed but diverted to Dog when that button appeared)
1
1
C2 Browser location permission prompt appears while the site also shows its own location modal (FALSE ALARM—at this time the participant clicked a link to bring up a location modal and entered the ZIP—there was no indication that the participant even saw the location permission prompt)
1
1
1
1
C3 Participant enters 80246 and gets multiple grooming locations (FALSE ALARM—true but this is how this control is supposed to work)
1
H3 Made no attempt to filter list by typing; instead started scrolling through the long list
1
1
H4 Searching list for English Springer Spaniel, scrolling up and down between the D’s and E’s but couldn’t find the breed there
1
1
1
1
C4 After selecting breed and age, the menu reloads. (FALSE ALARM—true but happens quickly and is the expected action after clicking Check Prices & Book Now—no impact on participant behavior)
1
C5 The target option is at the bottom of the list below the fold/below the Bath& Full Haircut with FURminator service (FALSE ALARM —true but not a problem for this participant)
1
1
1
C6 Large promotional grooming content and video tiles take up much of the page while the actual service menu is constrained to the right side (FALSE ALARM—true but not an obvious problem for this participant)
1
1

Appendix Table 2: Usability problems reported by ChatGPT 5.4 Thinking and the UX researcher for Participant A.

Gemini Problem-by-Problem Results

Appendix Table 3 shows the usability problems reported by Gemini and the UX researcher, run by run.

Prob #
Description
Run 1
Run 2
Run 3
Run 4
H1 Clicked Shop by Pet from top menu (incorrect first click)
1
H2 Clicked search field (incorrect second click)
G1 Browser location permission prompt appears while the site also shows its own location modal (FALSE ALARM—at this time the participant clicked a link to bring up a location modal and entered the ZIP—there was no indication that the participant even saw the location permission prompt)
1
1
H3 Made no attempt to filter list by typing; instead started scrolling through the long list
1
1
H4 Searching list for English Springer Spaniel, scrolling up and down between the D’s and E’s but couldn’t find the breed there
1
1
1
1
G2 The target option is at the bottom of the list below the fold/below the Bath& Full Haircut with FURminator service (FALSE ALARM —true but not a problem for this participant)
1
1

Appendix Table 3: Usability problems reported by Gemini 3 Flash Thinking and the UX researcher for Participant A.

Appendix B: The PetSmart Reservation Task

For a UX benchmark study conducted in 2019, one of the participants’ tasks was to use the PetSmart website to start booking a grooming appointment in Glendale, CO for a one-year-old English Springer Spaniel, then stop after determining the cost of a bath and full haircut. In this section, we review the steps through the “happy path” and speculate about possible user behaviors that would be reasonable to track to provide background knowledge for understanding the problem lists presented later.

Appendix Figure 1 shows the home page. Before continuing, ask yourself, where would you start?

PetSmart home page for the pet grooming task.

Appendix Figure 1: Home page for the pet grooming task.

For this task, the best first choice is to click “pet services” from the horizontal navigation menu close to the top of the page. From here, there are two paths to grooming, shown in Appendix Figure 2. Appendix Figure 2a shows the dropdown from which a user could drag the cursor down and release the button to select Grooming Salon. Appendix Figure 2b shows the pet services menu that appears after clicking “pet services” but releasing the mouse button without dragging, from which the user would click Grooming.

Appendix Figure 2a: Click and drag path

Click and drag path to grooming

Appendix Figure 2b: Click without dragging pathClick and drag path to grooming

Click without dragging path.

Appendix Figure 2: Two paths to the grooming menu.

Appendix Figure 3 shows the grooming form. This is where users who are not in Glendale can change the location to Glendale, select dog, select the breed, select the age, then click the button to check prices.

The grooming form.

Appendix Figure 3: The grooming form.

Before checking prices, users needed to select a breed and age for their dog. As shown in Appendix Figure 4, clicking breed produced a searchable breed dropdown. Appendix Figure 4a shows the initial appearance of the dropdown; Appendix Figure 4b shows its appearance after typing “english” over the placeholder text “breed” in the combobox.

Appendix Figure 4a: Initial appearance of the breed dropdown

Initial appearance of the breed dropdown.

Appendix Figure 4b: Appearance of the breed dropdown after typing “english” over “breed”

Appendix Figure 4b: Appearance of the breed dropdown after typing “english” over “breed”

Appendix Figure 4: The searchable breed dropdown.

For the happy path, a user should type “english” into the combobox, so one of the potential problems we anticipated was users not realizing the dropdown list could be filtered. Compounding the complexity of this step in the process is that the location of the word “English” for various breeds was inconsistent. For example, after filtering, the list in Figure 4b included English Toy Spaniel, Old English Sheepdog, and Springer Spaniel – English. That’s less of a problem after filtering but could be more problematic if scrolling through the unfiltered list.

The age dropdown, shown in Appendix Figure 5, was relatively straightforward with only two choices.

The age dropdown.

Appendix Figure 5: The age dropdown.

With the grooming form completed, the next step is to click “check prices & book now” to get the list of grooming options shown in Appendix Figure 6. Because the target option was the last one in the list and below the fold, we anticipated that some users might select an earlier option.

Appendix Figure 6a: Completed grooming menu and first option in Grooming Salon Menu (above the fold)Completed grooming menu and first option in Grooming Salon Menu (above the fold)

Appendix Figure 6b: The other grooming options (below the fold; the last option is the target)

The other grooming options (below the fold, last option is the target)

Appendix Figure 6: Grooming options above and below the fold, showing the target Bath & Full Haircut for $61.

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Register New Account
Compare items
  • Total (0)
Compare
0
Shopping cart