Measuring the UX of AI – MeasuringU


Feature image showing a researcher measuring the UX of AIAI is everywhere and getting embedded in all of our products.

If you ask the typical person in 2026 what AI is, they’ll probably say it’s a generative chat product like ChatGPT, Copilot, Claude, or Gemini. Of course, these are only the frontier Large Language Models (LLMs). That’s not all AI is.

For example, the algorithm recommending your next Netflix show, the AI drafting a formula in Excel, or the model flagging fraud on your credit card are all types of AI (none of which have chat interfaces).

But when most people say they’re “using AI,” they mean typing into a chat box, and that’s a good place to start when thinking about how to measure the UX of AI as it’s popularly understood.

How should we measure the quality of these experiences?

Measuring UX in General

Measuring the user experience in general involves assessing what people think and feel, and what people do. That means using a mix of attitudinal and action measures.

Action (behavioral) measures are more straightforward to interpret. A typical suite of action metrics includes a combination of effectiveness (completion rates, errors) and efficiency (time on task). But they’re relatively hard to collect because you have to set up task scenarios and record or observe behaviors.

Attitudinal measures are easier to collect, but you need to be sure you’re measuring the right thing.

For attitudes, we’ve found that standardized metrics like the UX-Lite® provide a good measure of overall attitudes toward a product’s usefulness (capabilities/features) and usability (ease of use). For example, see our 2026 retrospective benchmarks for ChatGPT, Claude, Gemini, and Grok.

Having an assessment of usefulness and usability from the UX-Lite provides good high-level measures that can be compared to historical benchmarks. But even though it discriminates at a high level between usefulness and usability, it doesn’t provide very granular or diagnostic measures. Consequently, we’ll also want to explore more specific attitudes to see how they affect the way people think about their use of AI chatbots.

When measuring specific interfaces, it can be helpful to identify additional constructs, features, or interactions that participants can rate so you have a better idea of which aspects of the UX are perceived as good or bad. That gives you a more diagnostic set of items.

How do you do that for AI chat interfaces like ChatGPT, Claude, Gemini, and Grok? You follow the process for creating standardized measures (for example, see our IJHCI paper on measuring the perceived clutter of websites). You need items, data, and validation.

Picking The Items: What Matters When Interacting with Generative AI Chat Software?

When building a standardized measure, the first step is to pick a set of items. We developed an initial set of 34 items based on input from the MeasuringU research team, drawing upon the existing literature and their experiences using these products to measure constructs like AI Productivity, AI Trust, AI Dependency, AI Anxiety, AI Personification, and Early Adoption. Consistent with psychometric practice, we created at least three items per construct.

AI Productivity

A researcher using an AI for productivity

One of the most attractive capabilities of generative AI chatbots is the potential for enhanced productivity, making this an important construct to measure. A survey conducted by Microsoft and LinkedIn in 2024 found that 90% of respondents said using AI helped them save time.

In 2026, results from an Anthropic survey indicated 86% of respondents reported improvements in the speed of their work. Table 1 shows the initial set of items we developed for this construct of increased productivity.

AI Productivity (Initial Set)
Using this AI chatbot greatly improves my productivity.
This AI chatbot adds substantial value to my personal tasks.
This AI chatbot adds substantial value to my professional tasks.
Using this AI chatbot makes me feel more capable in my work or studies.
I feel comfortable being accountable for work that used this AI chatbot.
Using this AI chatbot helps me achieve my goals.
The amount of time it takes for this AI chatbot to respond is acceptable.
This AI chatbot’s responses efficiently tell me the information I need.

Table 1: Initial item set for AI Productivity.

AI Trust

A researcher with computer screen showing "AI trust"

The flip side of excitement about increased productivity is distrust in AI output and data security. Even recent models hallucinate usability problems after reviewing videos of usability test sessions. A 2025 Melbourne-KPMG survey of 48,000 people across 47 countries found that less than half of the people regularly using AI were willing to trust it. Table 2 shows our initial set of AI Trust items.

AI Trust (Initial Set)
I trust this AI chatbot to provide reliable information.
I feel confident relying on responses from this AI chatbot when making decisions.
This AI chatbot always provides accurate responses.
When this AI chatbot makes mistakes, they are usually easy to detect.
I don’t worry about how my data is used when interacting with this AI chatbot.
It’s easy to understand what happens to the information I share with this AI chatbot.
My professional value is not affected by products like this AI chatbot.

Table 2: Initial item set for AI Trust.

AI Dependency

A relaxed researcher watching an AI assistant doing the work

As with Trust, there are particular concerns about AI users being overly dependent and uncritical of AI outputs. In the Melbourne-KPMG survey, 66% of respondents reported relying on AI output without evaluating its accuracy, and 56% reported making mistakes in their work due to uncritical acceptance of an AI output. See Table 3 for the AI Dependency items.

AI Dependency (Initial Set)
I often rely on AI chatbots to perform tasks that I would otherwise do myself.
I tend to accept answers from AI chatbots without verifying their accuracy.
I rarely double-check information provided by AI chatbots.

Table 3: Initial item set for AI Dependency.

AI Anxiety

Stressed researcher looking at a line chart

Conversations about AI often turn to anxiety about the potential negative effects of generative AI chatbots on society, the environment, and ethics. A 2025 University of Chicago AP-NORC survey found 44% believed AI would do more to hurt than help society, compared with 22% who expected it to do more good, and 41% were extremely or very concerned about AI’s environmental impact. Table 4 shows our initial item set for AI Anxiety.

AI Anxiety (Initial Set)
The increasing use of AI makes me uneasy.
I am often concerned that AI could cause serious harm to society.
I often worry about the environmental impact of AI.
AI development feels risky.
AI development feels difficult to control.
There should be more government regulation for AI development.
Using AI chatbots for work or school feels unethical.

Table 4: Initial item set for AI Anxiety.

AI Personification

A researcher interacting with an AI in the form of an angelic woman

Another aspect of interaction with generative AI chatbots that interests us is the extent to which users feel a personal relationship with the AI. This could range from feeling like you’re communicating with a human-like entity to feelings of friendship. Not everyone has the same emotional reaction to generative AI chatbot products, but we are interested in how products may differ in the extent to which they lead to social connection with their users. Our initial set of AI Personification items is listed in Table 5.

AI Personification (Initial Set)
Interacting with this AI chatbot feels like communicating with a human.
Sometimes I feel like this AI chatbot is more like a friend than a tool.
I feel like AI chatbots understand me well.
I tend to feel a sense of connection when interacting with AI chatbots.
I tend to feel like I’m socializing when I interact with AI chatbots.
I’m more likely to share personal information with AI chatbots than with other people.

Table 5: Initial item set for AI Personification.

Early Adoption

Although not related solely to AI, we included three items to assess respondents’ tendencies to be early adopters of new technologies (Table 6).

Early Adoption (Initial Set)
I like to experiment with new technologies before most people do.
I am usually among the first to try new digital tools.
I actively seek out new technologies to try.

Table 6: Initial item set for Early Adoption.

Summary and Discussion

In this article, we discussed how to measure the UX of AI in general and, specifically, generative AI chatbots like ChatGPT, Claude, Gemini, or Grok. The key points in the article were:

Measuring AI UX Starts with Measuring UX. Measuring UX in general means using a mix of attitudinal and action metrics. Action metrics are more rigorous but harder to collect—you need task scenarios and observation. Attitudinal metrics are easier to collect and often take the form of standardized questionnaires. Standardized metrics like the UX-Lite can discriminate among products based on their perceived usefulness and usability and are a good place to start. But if you want more diagnostic insight—to know more than just that something is off—you need a deeper, more specialized set of measures.

What Constructs Comprise the UX of AI? A deeper dive into the UX of generative AI chatbots requires investigation of specialized constructs. Based on our reading and experience with these types of products, we’ve proposed items for measuring AI Productivity, AI Trust, AI Dependency, AI Anxiety, and AI Personification.

To Validate Items, You Need Data from Real People. Creating an initial set of items is an important first step to develop standardized metrics, but it is just a first step. Items that look sensible on paper don’t always hold up once real people respond to them. In future articles, we’ll report the results of psychometric evaluation to (1) determine if the initial items, as we expect, group into statistical factors, (2) examine item quality to determine which items to retain for a final streamlined instrument, and (3) explore the connection between the new constructs and higher-level constructs like brand attitude, intention to continue use, and intention to recommend. Stay tuned!

We will be happy to hear your thoughts

Leave a reply

Som2ny Network
Logo
Compare items
  • Total (0)
Compare
0
Shopping cart