
There’s nothing quite the same as chatting with AI. But it’s just like every other interface in one important way: we experience it. And what we experience, we can measure.
Measuring the user experience means assessing both what people do and what people think (a mix of action and attitudinal metrics). For attitudes, standardized questionnaires like the UX-Lite® are a good place to start, but they’re not diagnostic on their own and won’t tell you why Trust is low or what’s driving Anxiety.
Building a standardized questionnaire involves assessing validity, reliability, and sensitivity. Practically, this involves six broad steps:
- Identifying the items to ask participants
- Collecting data on the candidate items
- Checking that the items group the way you expect
- Selecting the best items
- Confirming a shorter version is still reliable
- Demonstrating the items can tell good experiences from bad ones
We did step one in our previous article, where we identified 34 candidate items across six constructs (AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption) based on our reading and experience with these products.
In this article, we’ll cover the next five steps: collecting data from real users of real products and seeing how well those 34 items perform.
AI-Based Chat Software Benchmark Study
In May 2026, we conducted a retrospective study of four AI-based chat software products with 420 U.S.-based panel participants. This study included the metrics we typically collect in our standard UX and NPS study of consumer software.
There was a roughly equal gender split (52% female, 47% male). Respondents tended to be younger, with 65% under the age of 40. Participants were asked to reflect on their most recent experiences with the software and complete several existing questionnaires, including the NPS, SUS, UX-Lite, and TAC-10™. In addition to these existing questionnaires, respondents completed 34 new items focused on different aspects of the UX of AI. The AI-based chat products and sample sizes were:
- ChatGPT: 113
- Claude: 103
- Gemini: 101
- Grok: 103
The sample sizes are modest but adequate to establish baselines and identify medium-sized differences relative to each other and other software products we measure (e.g., ± 5 for 0–100-point rating scales). With a combined n = 420, the sample size is also large enough to support advanced analyses (e.g., factor analysis, regression analysis, reliability analysis, ANOVA).
Factor Analysis: Items (Almost) Perfectly Aligned with Target Constructs
We used the multivariate technique of factor analysis to determine how well the candidate items map to their intended construct. The output of a factor analysis is factor loadings, numbers that range from −1 to +1, which are interpreted like a correlation. To review the factor loadings for all 34 items, see Appendix Figure 1.
All but one of the items (“I often rely on AI chatbots to perform tasks that I would otherwise do myself”) didn’t map well to its intended construct of Trust (it loaded as high on Productivity), so this was a good candidate to exclude. The remaining 33 items strongly loaded on their intended constructs—solid evidence of construct validity.
Item Analyses: Selecting the Best Items for a Streamlined Questionnaire
A common approach to item selection in the psychometric process of developing a standardized questionnaire is to retain the items with the highest loadings on their associated factors. In our current practice, we enhance that method by examining item means and beta weights from regression models with a key outcome variable.
While still paying attention to the magnitude of item loadings, we also try to select items that vary in their observed means to better differentiate low and high levels of the construct. We also created six regression models, one for each construct, to see by examining beta weights which items accounted for larger amounts of variation in a key outcome metric: the likelihood to continue using the product.
In the following sections, we present the three criteria for all the items for each target construct to guide the items we selected for inclusion in a final, streamlined questionnaire. For easier interpretation, the means in the tables were converted from their original five-point scale to a scale ranging from 0 to 100. Our typical target for measuring a construct at this stage of questionnaire development is to select two to three items per construct with the goal of achieving scale reliability (measured with coefficient alpha) of at least 0.70 for each construct.
The selected items are at the top of each table. Table cells are highlighted for the three largest loadings, the lowest and highest means, and the three largest beta weights. We present all the items and their scores, as you may choose to try out alternative combinations of items for your own questionnaire (call us if you need to talk it through!).
AI Productivity
For AI Productivity, we selected two of the highest loading items, one of which also had the highest beta weight. The third selected item balanced an acceptably high loading and beta weight plus a relatively low mean.
| Item | ||||
|---|---|---|---|---|
| Using this AI chatbot greatly improves my productivity. | .959 |
.259 |
||
| Using this AI chatbot makes me feel more capable in my work or studies. | .904 |
|||
| I feel comfortable being accountable for work that used this AI chatbot. | 63.3 |
.183 |
||
| This AI chatbot adds substantial value to my personal tasks. | .245 |
|||
| This AI chatbot adds substantial value to my professional tasks. | ||||
| Using this AI chatbot helps me achieve my goals. | .959 |
|||
| The amount of time it takes for this AI chatbot to respond is acceptable. | 78.5 |
|||
| This AI chatbot’s responses efficiently tell me the information I need. |
Table 1: AI Productivity items.
AI Trust
The three items selected for AI Trust all had acceptably high loadings. The top two had impressive beta weights but little difference in their means (64.6, 62.7), so the third item was included to extend the lower range of the item means to 49.4.
| Item | ||||
|---|---|---|---|---|
| I trust this AI chatbot to provide reliable information. | .654 |
64.6 |
.344 |
|
| I feel confident relying on responses from this AI chatbot when making decisions. | .606 |
.295 |
||
| It’s easy to understand what happens to the information I share with this AI chatbot. | ||||
| This AI chatbot always provides accurate responses. | .813 |
|||
| When this AI chatbot makes mistakes, they are usually easy to detect. | ||||
| I don’t worry about how my data is used when interacting with this AI chatbot. | 44.8 |
|||
| My professional value is not affected by products like this AI chatbot. |
Table 2: AI Trust items.
AI Dependency
There were only three items developed for AI Dependency, and we excluded one of them, “I often rely on AI chatbots to perform tasks that I would otherwise do myself,” because it loaded on both Trust and Productivity. There were no issues warranting exclusion of the other two items, so we kept them both.
| Item | ||||
|---|---|---|---|---|
| I tend to accept answers from AI chatbots without verifying their accuracy. | .983 |
|||
| I rarely double-check information provided by AI chatbots. | .852 |
35.1 |
.108 |
|
| I often rely on AI chatbots to perform tasks that I would otherwise do myself. | 49.2 |
.250 |
Table 3: AI Dependency items.
AI Anxiety
The three items retained for AI Anxiety had the highest loadings of the set, reasonably impactful beta weights, and a reasonable range of means.
| Item | ||||
|---|---|---|---|---|
| The increasing use of AI makes me uneasy. | .822 |
−.259 |
||
| I am often concerned that AI could cause serious harm to society. | .899 |
|||
| AI development feels difficult to control. | .836 |
.126 |
||
| I often worry about the environmental impact of AI. | ||||
| AI development feels risky. | −.192 |
|||
| There should be more government regulation for AI development. | 66.5 |
|||
| Using AI chatbots for work or school feels unethical. | 47.1 |
Table 4: AI Anxiety items.
AI Personification
The three items selected for AI Personification had acceptably high loadings, significant beta weights, and a reasonable range of means.
| Item | ||||
|---|---|---|---|---|
| Interacting with this AI chatbot feels like communicating with a human. | 45.1 |
.249 |
||
| I feel like AI chatbots understand me well. | .777 |
.162 |
||
| I tend to feel a sense of connection when interacting with AI chatbots. | .985 |
|||
| I tend to feel like I’m socializing when I interact with AI chatbots. | 1.050 |
−.299 |
||
| Sometimes I feel like this AI chatbot is more like a friend than a tool. | 33.9 |
|||
| I’m more likely to share personal information with AI chatbots than with other people. |
Table 5: AI Personification items.
Early Adoption
We did not see any issues that warranted excluding any of these items, so we kept all three.
| Item | ||||
|---|---|---|---|---|
| I like to experiment with new technologies before most people do. | .961 |
.118 |
||
| I am usually among the first to try new digital tools. | .942 |
55.4 |
.171 |
|
| I actively seek out new technologies to try. | .920 |
62.4 |
Table 6: Early Adoption items.
Reliability Analysis: All Streamlined Scale Reliabilities Exceeded 0.80
Table 7 shows the coefficient alpha values for each measure for all items and for the streamlined item set. For research, the typical reliability goal is to exceed 0.70. There was little reduction in reliability for the streamlined versions and for AI Dependency; eliminating its one problematic item increased its reliability even though only two items were retained. The reliabilities for all the streamlined versions of the questionnaires not only met the typical research goal but exceeded 0.80—strong evidence of reliability for these new scales and statistical justification for the selection of their constituent items.
| Reliability (Coefficient Alpha) | ||
|---|---|---|
| AI Productivity | ||
| AI Trust | ||
| AI Dependency | ||
| AI Anxiety | ||
| AI Personification | ||
| Early Adoption |
Table 7: Scale reliabilities (coefficient alpha) for the six new metrics.
Profile Analysis: Claude Leads in Productivity, ChatGPT Lags in Trust
We created two different visualizations of profiles for the streamlined scales, showing how the AI assistants compare. The line graph in Figure 1 makes it easy to see at a glance which scales differentiate among the products, while the column chart in Figure 2 makes it easy to compare the confidence intervals around the means.
Figure 1: Line graph of AI scale scores for four generative AI chatbots.
Figure 2: Column chart of AI scale scores for four generative AI chatbots with 95% confidence intervals.
The results show Claude leading in AI Productivity, ChatGPT lagging in AI Trust, and little difference among the products for AI Dependency. Grok scored the lowest in AI Anxiety and the highest in AI Personification, and all four products scored different levels of Early Adoption (highest for Grok, lowest for ChatGPT).
A mixed ANOVA of the ratings indicated a significant main effect of scale (F(5, 2080) = 85.4, p F(3, 416) = 2.8, p = .039), and most importantly, a highly significant scale by product interaction (F(15, 2080) = 3.6, p strong statistical evidence of the sensitivity of these new scales.
Summary and Discussion
We collected data from 420 respondents for 34 items designed to measure six constructs related to attitudes toward four generative AI chatbots (ChatGPT, Claude, Gemini, and Grok). We then conducted analyses to complete the psychometric measurement goals of construct validity (factor analysis), measurement efficiency (item selection), scale reliability (coefficient alpha), and scale sensitivity (ANOVA).
The key points are:
We have solid evidence of construct validity for the new questionnaires. Our factor analysis of the data demonstrated almost perfect alignment of items with their intended constructs for the measurement of AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. One item loaded on two factors and was thus not retained in the streamlined versions of the questionnaires.
The items selected for streamlined versions of the questionnaires produced reliable measurement. For each of the six constructs, we retained two to three items to balance item loadings, item mean ranges, and strong beta weights for their relationships with the likelihood to continue using the product. For all the streamlined versions of the questionnaires, coefficient alpha exceeded 0.80 (ranging from 0.81 to 0.87). The common criterion for acceptable reliability is > 0.70.
The statistical evidence for scale sensitivity is strong. Profile analysis of the mean ratings of the new scales by product indicated a significant main effect of scale, a significant main effect of product, and a highly significant product-by-scale interaction. Claude led in AI Productivity, and ChatGPT lagged in AI Trust. There was little difference among the products for AI Dependency, Grok scored the lowest for AI Anxiety and the highest for AI Personification, and there were different levels of Early Adoption for all four products.
Appendix: Factor Structure and Item Key
Appendix Figure 1 shows the pattern matrix from the factor analysis for the original 34 items. The numbers in the figure are item loadings, which indicate the degree of connection of the item with the target constructs (with values ranging from -1 to +1, interpreted like a correlation). The usual criterion for a meaningfully large loading is anything more extreme than ±0.3. A maximum likelihood factor analysis with Promax rotation using SPSS 23 was used following a parallel analysis that indicated, as expected, retention of six factors. A key pattern to look for when identifying problematic items is any item with strong loading on more than one factor. This happened to OftenRelyOnChatbots (“I often rely on AI chatbots to perform tasks that I would otherwise do myself,” highlighted in yellow), leading to its exclusion from the final questionnaires.
Appendix Figure 1: Alignment of items with constructs (values greater than 0.3 are highlighted in green). To see the complete item text, refer to Appendix Table 1.
Appendix Table 1 documents the short labels for each item. Items highlighted in green are the ones retained for the streamlined versions of these questionnaires.
| Short Label | Item |
|---|---|
|
AI PRODUCTIVITY |
|
|
ImprovedProductivity |
Using this AI chatbot greatly improves my productivity. |
| AddsValuePersonal | This AI chatbot adds substantial value to my personal tasks. |
| AddsValueProfessional | This AI chatbot adds substantial value to my professional tasks. |
|
AIMakesMeFeelMoreCapable |
Using this AI chatbot makes me feel more capable in my work or studies. |
|
IFeelAccountableForMyAIAssistedWork |
I feel comfortable being accountable for work that used this AI chatbot. |
| AIHelpsMeAchieveMyGoals | Using this AI chatbot helps me achieve my goals. |
| AIResponseTimeIsAcceptable | The amount of time it takes for this AI chatbot to respond is acceptable. |
| AIResponsesAreEfficient | This AI chatbot’s responses efficiently tell me the information I need. |
|
AI TRUST |
|
|
TrustReliableInfo |
I trust this AI chatbot to provide reliable information. |
|
SupportsConfidentDecisions |
I feel confident relying on responses from this AI chatbot when making decisions. |
| AlwaysAccurate | This AI chatbot always provides accurate responses. |
| EasyToDetectMistakes | When this AI chatbot makes mistakes, they are usually easy to detect. |
| NotWorriedAboutDataUse | I don’t worry about how my data is used when interacting with this AI chatbot. |
|
EasyToUnderstandHowInfoIsShared |
It’s easy to understand what happens to the information I share with this AI chatbot. |
| ProfValueNotAffected | My professional value is not affected by products like this AI chatbot. |
|
AI DEPENDENCE |
|
| OftenRelyOnChatbots | I often rely on AI chatbots to perform tasks that I would otherwise do myself. |
|
AcceptAnswersWithoutVerification |
I tend to accept answers from AI chatbots without verifying their accuracy. |
|
RarelyDoubleCheck |
I rarely double-check information provided by AI chatbots. |
|
AI ANXIETY |
|
|
UneasyWithIncreasingUse |
The increasing use of AI makes me uneasy. |
|
SeriousSocietalHarm |
I am often concerned that AI could cause serious harm to society. |
| EnvironmentalImpact | I often worry about the environmental impact of AI. |
| AIDevelopmentRisky | AI development feels risky. |
|
AIDevelopmentHardToControl |
AI development feels difficult to control. |
| NeedMoreGovRegulation | There should be more government regulation for AI development. |
| UseForWorkOrSchoolUnethical | Using AI chatbots for work or school feels unethical. |
|
AI PERSONIFICATION |
|
|
AIWorkFeelsLikeHumanCommunication |
Interacting with this AI chatbot feels like communicating with a human. |
| SometimesAIFeelsLikeAFriend | Sometimes I feel like this AI chatbot is more like a friend than a tool. |
|
ChatbotsUnderstandMeWell |
I feel like AI chatbots understand me well. |
| FeelSenseOfConnection | I tend to feel a sense of connection when interacting with AI chatbots. |
|
FeelsLikeSocializing |
I tend to feel like I’m socializing when I interact with AI chatbots. |
| MoreLikelyToSharePersonalInfo | I’m more likely to share personal information with AI chatbots than with other people. |
|
EARLY ADOPTION |
|
|
IExperimentBeforeOthers |
I like to experiment with new technologies before most people do. |
|
FirstToTryNewDigitalTools |
I am usually among the first to try new digital tools. |
|
SeekNewTechnologiesToTry |
I actively seek out new technologies to try. |
Appendix Table 1: Short labels and full text for each item.







