Streamlined Measurement of the UX of AI – MeasuringU


There’s nothing quite the same as chatting with AI. But it’s just like every other interface in one important way: we experience it. And what we experience, we can measure.

Measuring the user experience means assessing both what people do and what people think (a mix of action and attitudinal metrics). For attitudes, standardized questionnaires like the UX-Lite® are a good place to start, but they’re not diagnostic on their own and won’t tell you why Trust is low or what’s driving Anxiety.

Building a standardized questionnaire involves assessing validity, reliability, and sensitivity. Practically, this involves six broad steps:

  1. Identifying the items to ask participants
  2. Collecting data on the candidate items
  3. Checking that the items group the way you expect
  4. Selecting the best items
  5. Confirming a shorter version is still reliable
  6. Demonstrating the items can tell good experiences from bad ones

We did step one in our previous article, where we identified 34 candidate items across six constructs (AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption) based on our reading and experience with these products.

In this article, we’ll cover the next five steps: collecting data from real users of real products and seeing how well those 34 items perform.

AI-Based Chat Software Benchmark Study

In May 2026, we conducted a retrospective study of four AI-based chat software products with 420 U.S.-based panel participants. This study included the metrics we typically collect in our standard UX and NPS study of consumer software.

There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several existing questionnaires, including the NPSSUSUX-Lite, and TAC-10™. In addition to these existing questionnaires, respondents completed 34 new items focused on different aspects of the UX of AI. The AI-based chat products and sample sizes were:

  • ChatGPT: 113
  • Claude: 103
  • Gemini: 101
  • Grok: 103

The sample sizes are modest but adequate to establish baselines and identify medium-sized differences relative to each other and other software products we measure (e.g., ± 5 for 0–100-point rating scales). With a combined n = 420, the sample size is also large enough to support advanced analyses (e.g., factor analysis, regression analysis, reliability analysis, ANOVA).

Factor Analysis: Items (Almost) Perfectly Aligned with Target Constructs

We used the multivariate technique of factor analysis to determine how well the candidate items map to their intended construct. The output of a factor analysis is factor loadings, numbers that range from −1 to +1, which are interpreted like a correlation. To review the factor loadings for all 34 items, see Appendix Figure 1.

All but one of the items (“I often rely on AI chatbots to perform tasks that I would otherwise do myself”) didn’t map well to its intended construct of Trust (it loaded as high on Productivity), so this was a good candidate to exclude. The remaining 33 items strongly loaded on their intended constructs—solid evidence of construct validity.

Item Analyses: Selecting the Best Items for a Streamlined Questionnaire

A common approach to item selection in the psychometric process of developing a standardized questionnaire is to retain the items with the highest loadings on their associated factors. In our current practice, we enhance that method by examining item means and beta weights from regression models with a key outcome variable.

While still paying attention to the magnitude of item loadings, we also try to select items that vary in their observed means to better differentiate low and high levels of the construct. We also created six regression models, one for each construct, to see by examining beta weights which items accounted for larger amounts of variation in a key outcome metric: the likelihood to continue using the product.

In the following sections, we present the three criteria for all the items for each target construct to guide the items we selected for inclusion in a final, streamlined questionnaire. For easier interpretation, the means in the tables were converted from their original five-point scale to a scale ranging from 0 to 100. Our typical target for measuring a construct at this stage of questionnaire development is to select two to three items per construct with the goal of achieving scale reliability (measured with coefficient alpha) of at least 0.70 for each construct.

The selected items are at the top of each table. Table cells are highlighted for the three largest loadings, the lowest and highest means, and the three largest beta weights. We present all the items and their scores, as you may choose to try out alternative combinations of items for your own questionnaire (call us if you need to talk it through!).

AI Productivity

Image showing man working at computer

For AI Productivity, we selected two of the highest loading items, one of which also had the highest beta weight. The third selected item balanced an acceptably high loading and beta weight plus a relatively low mean.

Item
Selected
Loading
Mean
Beta
Using this AI chatbot greatly improves my productivity.
x

.959

69.2

 .259

Using this AI chatbot makes me feel more capable in my work or studies.
x

.904

65.1
−.091
I feel comfortable being accountable for work that used this AI chatbot.
x
.568

63.3

 .183

This AI chatbot adds substantial value to my personal tasks.
.701
64.9

 .245

This AI chatbot adds substantial value to my professional tasks.
.884
63.7
 .041
Using this AI chatbot helps me achieve my goals.

.959

67.5
 .079
The amount of time it takes for this AI chatbot to respond is acceptable.
.524

78.5

 .002
This AI chatbot’s responses efficiently tell me the information I need.
.618
72.8
 .144

Table 1: AI Productivity items.

AI Trust

Image showing trust in AI

The three items selected for AI Trust all had acceptably high loadings. The top two had impressive beta weights but little difference in their means (64.6, 62.7), so the third item was included to extend the lower range of the item means to 49.4.

Item
Selected
Loading
Mean
Beta
I trust this AI chatbot to provide reliable information.
x

.654

64.6

 .344

I feel confident relying on responses from this AI chatbot when making decisions.
x

.606

62.7

 .295

It’s easy to understand what happens to the information I share with this AI chatbot.
x
.570
49.4
 .008
This AI chatbot always provides accurate responses.

.813

56.4
 .003
When this AI chatbot makes mistakes, they are usually easy to detect.
.417
58.0
−.019
I don’t worry about how my data is used when interacting with this AI chatbot.
.391

44.8

 .011
My professional value is not affected by products like this AI chatbot.
.372
61.1
 .076

Table 2: AI Trust items.

AI Dependency

Image showing a man lounging while an AI does the work

There were only three items developed for AI Dependency, and we excluded one of them, “I often rely on AI chatbots to perform tasks that I would otherwise do myself,” because it loaded on both Trust and Productivity. There were no issues warranting exclusion of the other two items, so we kept them both.

Item
Selected
Loading
Mean
Beta
I tend to accept answers from AI chatbots without verifying their accuracy.
x

.983

37.3
.034
I rarely double-check information provided by AI chatbots.
x

.852

35.1

.108

I often rely on AI chatbots to perform tasks that I would otherwise do myself.
.363

49.2

.250

Table 3: AI Dependency items.

AI Anxiety

Image showing nervous man watching computer.

The three items retained for AI Anxiety had the highest loadings of the set, reasonably impactful beta weights, and a reasonable range of means.

Item
Selected
Loading
Mean
Beta
The increasing use of AI makes me uneasy.
x

.822

55.4

−.259

I am often concerned that AI could cause serious harm to society.
x

.899

58.9
−.095
AI development feels difficult to control.
x

.836

60.4

 .126

I often worry about the environmental impact of AI.
.780
61.8
−.044
AI development feels risky.
.821
57.0

−.192

There should be more government regulation for AI development.
.794

66.5

 .009
Using AI chatbots for work or school feels unethical.
.603

47.1

−.021

Table 4: AI Anxiety items.

AI Personification

Image showing man speaking to an angelic AI incarnation.

The three items selected for AI Personification had acceptably high loadings, significant beta weights, and a reasonable range of means.

Item
Selected
Loading
Mean
Beta
Interacting with this AI chatbot feels like communicating with a human.
x
.593

45.1

 .249

I feel like AI chatbots understand me well.
x

.777

43.5

 .162

I tend to feel a sense of connection when interacting with AI chatbots.
x

.985

35.5
 .145
I tend to feel like I’m socializing when I interact with AI chatbots.

1.050

34.5

−.299

Sometimes I feel like this AI chatbot is more like a friend than a tool.
.712

33.9

 .151
I’m more likely to share personal information with AI chatbots than with other people.
.684
36.6
 .081

Table 5: AI Personification items.

Early Adoption

We did not see any issues that warranted excluding any of these items, so we kept all three.

Item
Selected
Loading
Mean
Beta
I like to experiment with new technologies before most people do.
x

.961

59.9

 .118

I am usually among the first to try new digital tools.
x

.942

55.4

 .171

I actively seek out new technologies to try.
x

.920

62.4

−.087

Table 6: Early Adoption items.

Reliability Analysis: All Streamlined Scale Reliabilities Exceeded 0.80

Table 7 shows the coefficient alpha values for each measure for all items and for the streamlined item set. For research, the typical reliability goal is to exceed 0.70. There was little reduction in reliability for the streamlined versions and for AI Dependency; eliminating its one problematic item increased its reliability even though only two items were retained. The reliabilities for all the streamlined versions of the questionnaires not only met the typical research goal but exceeded 0.80—strong evidence of reliability for these new scales and statistical justification for the selection of their constituent items.

Reliability (Coefficient Alpha)
All Items
Streamlined
AI Productivity
0.92
0.85
AI Trust
0.82
0.81
AI Dependency
0.76
0.84
AI Anxiety
0.91
0.87
AI Personification
0.92
0.86
Early Adoption
0.94
NA

Table 7: Scale reliabilities (coefficient alpha) for the six new metrics.

Profile Analysis: Claude Leads in Productivity, ChatGPT Lags in Trust

We created two different visualizations of profiles for the streamlined scales, showing how the AI assistants compare. The line graph in Figure 1 makes it easy to see at a glance which scales differentiate among the products, while the column chart in Figure 2 makes it easy to compare the confidence intervals around the means.

Line graph of AI scale scores for four generative AI chatbots.

Figure 1: Line graph of AI scale scores for four generative AI chatbots.

Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals.

Figure 2: Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals.

The results show Claude leading in AI Productivity, ChatGPT lagging in AI Trust, and little difference among the products for AI Dependency. Grok scored the lowest in AI Anxiety and the highest in AI Personification, and all four products scored different levels of Early Adoption (highest for Grok, lowest for ChatGPT).

A mixed ANOVA of the ratings indicated a significant main effect of scale (F(5, 2080) = 85.4, p F(3, 416) = 2.8, p = .039), and most importantly, a highly significant scale by product interaction (F(15, 2080) = 3.6, p strong statistical evidence of the sensitivity of these new scales.

Summary and Discussion

We collected data from 420 respondents for 34 items designed to measure six constructs related to attitudes toward four generative AI chatbots (ChatGPT, Claude, Gemini, and Grok). We then conducted analyses to complete the psychometric measurement goals of construct validity (factor analysis), measurement efficiency (item selection), scale reliability (coefficient alpha), and scale sensitivity (ANOVA).

The key points are:

We have solid evidence of construct validity for the new questionnaires. Our factor analysis of the data demonstrated almost perfect alignment of items with their intended constructs for the measurement of AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. One item loaded on two factors and was thus not retained in the streamlined versions of the questionnaires.

The items selected for streamlined versions of the questionnaires produced reliable measurement. For each of the six constructs, we retained two to three items to balance item loadings, item mean ranges, and strong beta weights for their relationships with the likelihood to continue using the product. For all the streamlined versions of the questionnaires, coefficient alpha exceeded 0.80 (ranging from 0.81 to 0.87). The common criterion for acceptable reliability is > 0.70.

The statistical evidence for scale sensitivity is strong. Profile analysis of the mean ratings of the new scales by product indicated a significant main effect of scale, a significant main effect of product, and a highly significant product-by-scale interaction. Claude led in AI Productivity, and ChatGPT lagged in AI Trust. There was little difference among the products for AI Dependency, Grok scored the lowest for AI Anxiety and the highest for AI Personification, and there were different levels of Early Adoption for all four products.

Appendix: Factor Structure and Item Key

Appendix Figure 1 shows the pattern matrix from the factor analysis for the original 34 items. The numbers in the figure are item loadings, which indicate the degree of connection of the item with the target constructs (with values ranging from -1 to +1, interpreted like a correlation). The usual criterion for a meaningfully large loading is anything more extreme than ±0.3. A maximum likelihood factor analysis with Promax rotation using SPSS 23 was used following a parallel analysis that indicated, as expected, retention of six factors. A key pattern to look for when identifying problematic items is any item with strong loading on more than one factor. This happened to OftenRelyOnChatbots (“I often rely on AI chatbots to perform tasks that I would otherwise do myself,” highlighted in yellow), leading to its exclusion from the final questionnaires.

Alignment of items with constructs (values greater than 0.3 are highlighted in green). To see the complete item text, refer to the appendix.

Appendix Figure 1: Alignment of items with constructs (values greater than 0.3 are highlighted in green). To see the complete item text, refer to Appendix Table 1.

Appendix Table 1 documents the short labels for each item. Items highlighted in green are the ones retained for the streamlined versions of these questionnaires.

Short Label Item

AI PRODUCTIVITY

ImprovedProductivity

Using this AI chatbot greatly improves my productivity.

AddsValuePersonal This AI chatbot adds substantial value to my personal tasks.
AddsValueProfessional This AI chatbot adds substantial value to my professional tasks.

AIMakesMeFeelMoreCapable

Using this AI chatbot makes me feel more capable in my work or studies.

IFeelAccountableForMyAIAssistedWork

I feel comfortable being accountable for work that used this AI chatbot.

AIHelpsMeAchieveMyGoals Using this AI chatbot helps me achieve my goals.
AIResponseTimeIsAcceptable The amount of time it takes for this AI chatbot to respond is acceptable.
AIResponsesAreEfficient This AI chatbot’s responses efficiently tell me the information I need.

AI TRUST

TrustReliableInfo

I trust this AI chatbot to provide reliable information.

SupportsConfidentDecisions

I feel confident relying on responses from this AI chatbot when making decisions.

AlwaysAccurate This AI chatbot always provides accurate responses.
EasyToDetectMistakes When this AI chatbot makes mistakes, they are usually easy to detect.
NotWorriedAboutDataUse I don’t worry about how my data is used when interacting with this AI chatbot.

EasyToUnderstandHowInfoIsShared

It’s easy to understand what happens to the information I share with this AI chatbot.

ProfValueNotAffected My professional value is not affected by products like this AI chatbot.

AI DEPENDENCE

OftenRelyOnChatbots I often rely on AI chatbots to perform tasks that I would otherwise do myself.

AcceptAnswersWithoutVerification

I tend to accept answers from AI chatbots without verifying their accuracy.

RarelyDoubleCheck

I rarely double-check information provided by AI chatbots.

AI ANXIETY

UneasyWithIncreasingUse

The increasing use of AI makes me uneasy.

SeriousSocietalHarm

I am often concerned that AI could cause serious harm to society.

EnvironmentalImpact I often worry about the environmental impact of AI.
AIDevelopmentRisky AI development feels risky.

AIDevelopmentHardToControl

AI development feels difficult to control.

NeedMoreGovRegulation There should be more government regulation for AI development.
UseForWorkOrSchoolUnethical Using AI chatbots for work or school feels unethical.

AI PERSONIFICATION

AIWorkFeelsLikeHumanCommunication

Interacting with this AI chatbot feels like communicating with a human.

SometimesAIFeelsLikeAFriend Sometimes I feel like this AI chatbot is more like a friend than a tool.

ChatbotsUnderstandMeWell

I feel like AI chatbots understand me well.

FeelSenseOfConnection I tend to feel a sense of connection when interacting with AI chatbots.

FeelsLikeSocializing

I tend to feel like I’m socializing when I interact with AI chatbots.

MoreLikelyToSharePersonalInfo I’m more likely to share personal information with AI chatbots than with other people.

EARLY ADOPTION

IExperimentBeforeOthers

I like to experiment with new technologies before most people do.

FirstToTryNewDigitalTools

I am usually among the first to try new digital tools.

SeekNewTechnologiesToTry

I actively seek out new technologies to try.

Appendix Table 1: Short labels and full text for each item.

 

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Register New Account
Compare items
  • Total (0)
Compare
0
Shopping cart