
Can ChatGPT Replace UX Researchers?
We wrote that title into our article three years ago, and some people thought it was click-bait. However, the subtitle spoke of what we thought was a potentially good job for AI—coding open-ended comments from a survey. These comments often get ignored, with only a quick summary or relegation to a word cloud, though they can reveal valuable insights.
Our initial results showed promise. And since then, the AI landscape has exploded. What started as a one-horse race with OpenAI is now a battle among industry heavyweights, with Google’s Gemini and Anthropic’s Claude emerging as serious contenders in their own right. Models have gotten faster, smarter, and significantly better at reasoning. With the tech maturing so rapidly, we decided it was time to revisit our original experiment and see where things stand today.
Back in 2023, our study found substantial agreement among human coders, among repeated runs of ChatGPT-4, and between the two (mean kappas from .632 to .704), suggesting that ChatGPT could be useful for helping to automate this analysis. But how much better has ChatGPT gotten, and how do the new industry leaders stack up against one another?
To find out, we’re kicking off a series of benchmarking tests, starting right where we began with re-running our original dataset and prompt through version 5.6 of ChatGPT. In upcoming posts, we’ll put Gemini and Claude through the exact same paces.
In short, we found that ChatGPT-5.6 not only matched human levels of interrater agreement and achieved near-perfect self-consistency, but has also evolved into a far more granular analyst, subdividing core themes into small subcategories.
Study Details
To maintain direct comparability, we used the SUPR-Q® benchmarking data from our original 2023 study. The dataset comprises participant-reported problems and frustrations from their most recent visit to one of three websites:
All datasets and prompts remained identical to the 2023 evaluation (see the appendix for the full prompt). On August 11, 2026, we evaluated each dataset across three separate runs using ChatGPT-5.6 Sol set to “high” thinking. Each run was executed in a fresh chat session with the “Memory” feature disabled to ensure independence between trials. For our comparison, we retained the baseline coding from the three human evaluators (1, 2, 3) and the three original ChatGPT-4 runs (A, B, C), and designated the three new ChatGPT-5.6 Sol runs as D, E, and F.
How we aligned themes this time (and how the approach differs from 2023)
In our initial 2023 study, ChatGPT-4 consistently grouped survey responses into six to nine broad categories. In contrast, ChatGPT-5.6 configured with “high” thinking frequently split responses into finer subcategories. While these granular divisions preserve the same core conceptual insights, they introduce practical analytical trade-offs. For instance, across all new runs on the Office Supplies data, the “Out of Stock” category was subdivided into two to four narrower groups. This added granularity may offer researchers deeper contextual nuance, but it carries the risk of fragmenting macro-level themes if an LLM over-subdivides this data.
This structural shift required us to adapt our alignment methodology. In 2023, we aligned themes using a bottom-up approach based strictly on the highest number of overlapping statements across runs. However, applying the overlap-based approach to the new outputs would have artificially depressed agreement metrics. Instead, we evaluated the new runs against the original framework, manually combining and mapping subcategories that logically aligned with the baseline structure. See Table 1 for an example of how the new ChatGPT-5.6 categories were combined to align with our existing themes.
By shifting from a purely data-driven overlap model to a top-down frame, we preserved structural continuity across iterations that allowed for a more direct, apples-to-apples comparison between ChatGPT-5.6 and our original baseline codes.
Office Supplies Data: Theme Alignment
| Category | GPT D | GPT E | GPT F |
|---|---|---|---|
| No Issues | No problems or frustrations | No problems or frustrations reported | No problems or frustrations reported |
| Appearance | Unappealing, bland, or outdated visual design | Unappealing, bland, or outdated visual design; Product image quality |
Visual design is bland, dated, or unappealing |
| Navigation | Navigation and browsing difficulties; Too many or overly complex product categories; Search, filtering, and product-findability problems |
Navigation and information architecture; Search, filtering, and product findability |
Navigation, browsing, and category organization; Search, filtering, and product-finding difficulties |
| Cluttered | Cluttered, busy, or overwhelming interface; Distracting ads, promotions, pop-ups, or graphics |
Cluttered, overwhelming, or distracting interface | Cluttered, busy, or overwhelming interface; Promotions, ads, recommendations, and pop-ups are distracting |
| Out of Stock | Inventory information is inaccurate; Website does not fully reflect local-store inventory; Out-of-stock, unavailable, or limited product selection; In-store pickup availability limitations |
Store inventory accuracy and online/in-store synchronization; Product availability, out-of-stock items, and pickup limitations; Limited or incomplete product assortment |
Store inventory and website availability mismatch; Product availability, stock, and assortment limitations |
Table 1: Alignment of themes for ChatGPT-5.6 runs in the Office Supplies dataset.
Study Results
We used the baseline metrics from our 2023 study as a starting point to evaluate interrater agreement and assess how qualitative analysis has evolved with the newest generation of ChatGPT.
Comparing the number of themes
As we’ve previously discussed, qualitative coding often reflects a fundamental tension between lumpers (where coders group smaller concepts into broad themes) and splitters (where coders use more granular categories). In our initial study, we identified two distinct baseline behaviors: substantive theme granularity (the number of core categories containing more than one statement) and fringe-theme sensitivity (the willingness to create standalone categories for single, isolated comments).
In 2023, ChatGPT-4 generated almost half as many total themes as human coders, averaging seven total themes per dataset compared to thirteen for humans. However, this gap was driven almost entirely by fringe sensitivity. Human coders frequently created dedicated categories for one-off comments, accounting for nearly half of their total generated themes. Once those single-statement themes were removed, human coders and ChatGPT-4 converged at a similar macro baseline of six or seven substantive themes. Both operated as macro-lumpers for primary patterns.
Human coders were lumpers at the macro level, splitters at the margins. Human researchers built high overall category counts by isolating single-statement fringe themes and maintaining lower resolution on core patterns. In our 2026 re-run, ChatGPT-5.6’s approach to qualitative coding evolved, demonstrating a distinct shift towards aggressive splitting of substantive themes.
ChatGPT-5.6 closed the theme gap with human coders, generating a similar volume (15.6 for ChatGPT-5.6 versus 13.3 for humans) by mainly subdividing core themes into smaller buckets of two to four statements. On average, the model generated almost 70% more substantive themes than humans and nearly double that of ChatGPT-4. For example, where 2023 runs grouped feedback under a single “Cluttered” category, ChatGPT-5.6 split the node into two distinct subthemes: “Cluttered, busy, or overwhelming interface” and “Distracting ads, promotions, pop-ups, or graphics.” Table 2 shows the full counts of themes generated by each coder and AI; humans and ChatGPT-5.6 often had many more themes than the ChatGPT-4 runs. Table 3 shows that removing the one-off categories made the human coders’ major themes more comparable to ChatGPT-4, while ChatGPT-5.6 retained about twice as many themes (mainly due to categories with two to four statements).
| Grocery | Wireless | Office Supplies | |||||
|---|---|---|---|---|---|---|---|
| Number of Themes | Number of Themes with 1 statement | Number of Themes | Number of Themes with 1 statement | Number of Themes | Number of Themes with 1 statement | ||
| Coder 1 | 12 | 6 | 14 | 6 | 13 | 5 | |
| Coder 2 | 25 | 16 | 16 | 8 | 17 | 10 | |
| Coder 3 | 7 | 0 | 10 | 4 | 6 | 1 | |
| Coder Average | 14.7 | 7.3 | 13.3 | 6 | 12 | 5.3 | |
| ChatGPT 4 2023 |
ChatGPT A | 8 | 1 | 8 | 2 | 6 | 0 |
| ChatGPT B | 6 | 0 | 7 | 0 | 6 | 0 | |
| ChatGPT C | 7 | 0 | 8 | 1 | 8 | 2 | |
| ChatGPT Average | 7.0 | 0.3 | 7.7 | 1 | 6.3 | 0.7 | |
| ChatGPT 5.6 2026 | ChatGPT D | 22 | 9 | 14 | 2 | 14 | 1 |
| ChatGPT E | 24 | 12 | 15 | 2 | 13 | 3 | |
| ChatGPT F | 14 | 1 | 13 | 2 | 11 | 0 | |
| ChatGPT Average | 20 | 7.3 | 14 | 2 | 12.7 | 1.3 |
Table 2: Number of category themes generated by coder and ChatGPT runs.
| Grocery Themes | Wireless Themes | Office Supplies Themes | Average | ||
|---|---|---|---|---|---|
| Coder 1 | 6 | 8 | 8 | 7.3 | |
| Coder 2 | 9 | 8 | 7 | 8 | |
| Coder 3 | 7 | 6 | 5 | 6 | |
| Coder Average | 7.3 | 7.3 | 6.7 | 7.1 | |
| ChatGPT 4 2023 | ChatGPT A | 7 | 6 | 6 | 6.3 |
| ChatGPT B | 6 | 7 | 6 | 6.3 | |
| ChatGPT C | 7 | 7 | 6 | 6.7 | |
| ChatGPT Average | 6.7 | 6.7 | 6 | 6.4 | |
| ChatGPT 5.6 2026 | ChatGPT D | 13 | 12 | 13 | 12.7 |
| ChatGPT E | 12 | 13 | 10 | 11.7 | |
| ChatGPT F | 13 | 11 | 11 | 11.7 | |
| ChatGPT Average | 12.7 | 12 | 11.3 | 12 |
Table 3: Number of category themes by coder and ChatGPT containing more than one statement.
Interrater agreement between humans and ChatGPT runs
To evaluate how accurately ChatGPT-5.6 aligned with human judgment, and how reliable it was across repeated runs, we computed kappa for each pairing of raters, as we did in 2023. We averaged kappas across themes within products and then across products to get the overall results shown in Table 4.
A brief reminder about kappa. There are different methods for assessing the magnitude of interrater agreement. One of the best-known is the kappa statistic (Fleiss, 1971). Kappa measures the extent of agreement among raters that exceeds estimates of chance agreement. Kappa values can be between −1 (perfect disagreement) and 1 (perfect agreement) and are often interpreted with the Landis and Koch guidelines (poor agreement: ≤ 0, slight: 0.01–0.20, fair: 0.21–0.40, moderate: 0.41–0.60, substantial: 0.61–0.80, almost perfect agreement: 0.81–1.00).
|
2023 |
2026 |
|||||||
| Rater | Coder 1 | Coder 2 | Coder 3 | ChatGPT A | ChatGPT B | ChatGPT C | ChatGPT D | ChatGPT E |
| Coder 2 | 0.703 | |||||||
| Coder 3 | 0.726 | 0.683 | ||||||
| ChatGPT A | 0.699 | 0.608 | 0.662 | |||||
| ChatGPT B | 0.643 | 0.548 | 0.615 | 0.684 | ||||
| ChatGPT C | 0.681 | 0.607 | 0.624 | 0.784 | 0.584 | |||
| ChatGPT D | 0.756 | 0.754 | 0.734 | 0.656 | 0.610 | 0.612 | ||
| ChatGPT E | 0.767 | 0.738 | 0.732 | 0.665 | 0.627 | 0.642 | 0.952 | |
| ChatGPT F | 0.779 | 0.744 | 0.740 | 0.654 | 0.635 | 0.643 | 0.949 | 0.949 |
Table 4: Kappa matrix showing chance-corrected agreement for each pairing of human coders and ChatGPT runs, averaged over fourteen themes pulled from three datasets, expanded to include ChatGPT-5.6 runs (D, E, F).
Looking at the full pairwise matrix in Table 4, three clear trends emerge. First, ChatGPT-5.6 consistently achieved higher agreement with human coders than ChatGPT-4 did. Second, cross-generational agreement between GPT-4 and GPT-5.6 remained firmly substantial, confirming underlying thematic continuity despite the newer model’s increased granularity. Most strikingly, ChatGPT-5.6 demonstrated near-perfect self-consistency across independent runs, elevating intra-model agreement to a remarkable 0.949–0.952.
Human Coders vs. ChatGPT Runs vs. Combined
To test whether these shifts were statistically meaningful, we aggregated the pairwise kappas into group-level averages. Below are the average kappas for each pairing type, and these same values are shown visually in Figure 1.
- Humans Coders: .704 (substantial) — unchanged from 2023
- ChatGPT-4: .684 (substantial) — unchanged from 2023
- ChatGPT-5.6: .950 (almost perfect)
- Humans ↔ ChatGPT-4: .632 (substantial) — unchanged from 2023
- Humans ↔ ChatGPT-5.6: .749 (substantial)
- GPT-4 ↔ ChatGPT-5.6: .638 (substantial)
Figure 1: Average agreement between ChatGPT and human coders.
There was a statistically significant difference among these group means (F(5,30) = 40.5, p p p p = .475).
These findings suggest that ChatGPT has improved to a point where it’s highly consistent and agrees considerably well with humans. Consistency is, of course, good, but you can also be consistently wrong.
No Issues vs. Other Themes
Consistent with our 2023 findings, interrater agreement remained exceptionally high for the “No Issues” category, maintaining an almost perfect mean kappa of .937 across all three datasets. However, there were some cases where ChatGPT-5.6 called out unclear statements and pulled them out separately from the main analysis. This behavior is illustrated by how different runs of ChatGPT-5.6 evaluated two ambiguous participant responses in the Wireless dataset, Statement 16 (“Na”) and Statement 44 (“Not sure.”):
Wireless – ChatGPT D:
Here, the model created explicit reasoning notes, observing that “‘Na’ provides no usable information. It may mean ‘N/A,’ but that is not explicit enough to confidently code as ‘no problems,’” and “‘Not sure’ does not identify a specific problem or indicate clearly that no problem occurred.”
Wireless – ChatGPT E:
During this run, the model created a separate “Uncategorized Statements” bucket, designating statement 44 as an off-topic/service issue, while identifying statement 16 as a non-substantive response.
Wireless – ChatGPT F:
In this case, the model created a dedicated category titled “Statements without enough information to classify,” grouping 16 and 44 together because they lacked substantive feedback.
Humans often included these statements in the “No Issues” theme. Depending on your lens of analysis, either classification could be justified. Subtle shifts in classification like this could lead to lower agreement. Fortunately, this type of model variance can likely be mitigated with a more explicit prompt instruction regarding edge-case handling.
Summary and Discussion
Our key conclusions from these analyses were:
Agreement with human coders improved with the newer ChatGPT-5.6 model; human-to-ChatGPT-5.6 agreement was as high as human-to-human. Average interrater agreement increased, with kappa going from .632 to .749 when comparing ChatGPT to human coders. While ChatGPT-4 already fell within the established guidelines for “substantial” agreement, this improvement suggests that using ChatGPT-5.6 for coding open-ended survey responses is now virtually as reliable as comparing multiple human researchers.
ChatGPT-5.6 responses seem to be dramatically more consistent in this setting. In addition to achieving better agreement with human coders, ChatGPT-5.6 exhibited remarkable agreement across repeated runs. In contrast to our 2023 testing, running open-ended comments through ChatGPT-5.6 is now sufficiently consistent that a single analysis run will likely yield the same taxonomic output.
ChatGPT-5.6 responses may split categories into further detail. In this setting, ChatGPT-5.6 demonstrated the ability to construct more precise subcategories while maintaining structural alignment with the overarching core themes identified by humans.
Caveats
Prompt frozen at 2023 for comparability. To isolate model performance over time, we controlled prompt structure and data inputs. However, optimized prompting techniques customized for newer reasoning architectures could likely yield even stronger performance.
Alignment method changed from 2023. To align outputs with pre-existing categories, we shifted to a top-down alignment strategy. Because ChatGPT-5.6 frequently split baseline categories into narrower subthemes, those subcategories were conceptually combined for cross-model evaluation.
Small datasets, one prompt, three runs, one model. In 2023, LLMs struggled with high-input tasks due to limited context windows, forcing us to rely on compact datasets (50–52 statements). To preserve comparability, we retained these small datasets, tested a single prompt across three runs, and limited this initial phase to OpenAI models. Keep this in mind when generalizing these findings to open-ends in different contexts.
What’s Next?
In this study, we tested a fairly straightforward replication of our original analysis. Given the large improvement in how ChatGPT performed, further expansions are justified. We plan to continue to explore and test ideas like:
- How can we improve our prompt to maximize accuracy and ease of use?
- How do Claude, Gemini, Meta AI, or other models perform on qualitative coding?
- How do LLMs perform on larger datasets now that context windows have improved?
Appendix: The Verbatim Prompt
As a UX researcher, you are tasked with analyzing the dataset provided below, which contains answers to the question, “What are some problems or frustrations you’ve had with the XXXX website?”. Your goal is to classify each numbered statement according to common themes. Create as many categories as necessary to group similar statements together. If a statement fits multiple categories, include it in all relevant categories.
START DATASET:
[dataset numbered by participant]
END DATASET
After analyzing the dataset, follow these steps:
- List the categories you have created, along with a brief description for each category.
- For each category, list the numbers of the statements that belong to it.
- If there are any statements that do not fit into any of the categories you have created, list their numbers separately.
