Back to Articles
When a UX Metric Becomes a Verdict

When a UX Metric Becomes a Verdict

Explained how a broad usability score can raise leadership awareness while failing to diagnose causes, then outlined a stronger UX measurement program using lightweight metrics, cohort context, open-text feedback, and ownership across teams.

What SUS gave us

Before we introduced a shared metric, evidence of customer frustration lived across research sessions, support tickets, product discussions and customer calls. Each source contained useful information, though the evidence was easy to treat as anecdotal when considered separately.

SUS made the concern visible at a leadership level. A standardized score gave people a common reference point and confirmed that usability deserved attention. It also helped move UX quality beyond personal opinion. Those outcomes mattered, and I do not regret using it.

The score also created a false sense of precision. Standardized measures look authoritative, especially when they come with established benchmarks. That authority can encourage people to treat an aggregate result as a complete assessment of the product. Ours was far from complete.

A low score can reflect confusing interactions, missing capabilities or poor performance. Users may also be responding to weak onboarding, unreliable data, unfamiliar terminology, permission problems or an experience that changed during implementation. In a B2B platform, several of those conditions can affect the same task.

The score told us that customers perceived the product as difficult to use. It offered limited evidence about the source of that difficulty.

Our product made the aggregate score difficult to interpret

Our respondents had different roles, expectations and levels of experience. Some used mature parts of the product every day. Others had recently started using workflows that were still developing. Their experiences were also affected by implementation choices, data quality and gaps between the intended design and the shipped product.

Combining those responses produced a technically valid aggregate. Its practical value was limited because we could not reliably identify which users were struggling in which workflows. We also lacked enough context to determine how recent product changes had affected their responses.

Benchmarking introduced another problem. A comparison only helps when the underlying products and respondent groups are meaningfully comparable. A single score can conceal substantial variation among users and across the work they perform in different parts of the product. The benchmark remained useful as a directional reference. It provided too little evidence for the diagnosis that leadership wanted from it.

The weakness sat in the wider measurement system. We needed stronger segmentation, more deliberate timing and qualitative evidence connected to each response.

Customer experience has shared causes

The product experience resulted from decisions made across design, product, engineering, data, implementation and support. Design had an important role in its quality and had to be accountable for that work. Assigning the full score to design would have ignored much of the evidence.

A confusing workflow may contain interaction problems. It may also reflect reduced scope, missing bulk actions or technical constraints introduced during delivery. Inaccurate data can damage confidence in an otherwise understandable interface. Repeated support complaints may indicate a product problem that has never reached roadmap planning.

UX can manage the measurement program, conduct the research and explain the findings. The work required to improve the score may belong elsewhere in the organization, or it may require several teams to coordinate their decisions.

Shared accountability needs named owners. Vague collective ownership usually leaves the underlying problem untouched. Once research identifies the causes, each contributing issue needs a person or team with the authority to address it.

The metric needs context

I would use a shorter measure such as UMUX-Lite if I were setting up this program again. Its length makes repeated use easier and provides a directional view of perceived usefulness and ease of use. Changing the instrument alone would solve very little. Most of the improvement would come from the information collected around it.

Each response should be connected to a known user role and a recent workflow. Product tenure and exposure to relevant changes also matter. Implementation or support conditions may explain part of the rating and should be available during analysis.

Timing deserves equal attention. A request to rate an entire platform in the abstract tends to produce broad dissatisfaction with limited diagnostic value. Collecting feedback after onboarding, a recurring task or issue resolution ties the response to something the team can examine.

Repeated measurement can show whether perception changes after the product changes. A customer’s first score establishes a baseline. A later response can indicate whether a revised workflow improved the experience, whether heavier use exposed new problems, or whether confidence remained low after the original issue was fixed.

That progression is usually more useful than comparing one isolated number against an industry benchmark.

Qualitative evidence explains the rating

Scores compress a large amount of experience into a number. The compression makes reporting easier and removes much of the meaning required for action.

A low rating could follow a confusing interface. The same rating could come from missing functionality, unreliable data or an unsuccessful support interaction. Without comments, interviews or behavioural evidence, internal teams supply their own explanations. Those explanations often reflect the part of the product each team already understands.

Open-text responses should accompany the rating. Interviews, support themes, product analytics and observed behaviour can then be used to test the initial interpretation. The research team can connect the rating to a specific point in the customer experience and determine whether several sources describe the same issue.

This evidence also improves prioritization. Leadership can act more confidently when a low score is tied to a particular user group, workflow and documented cause. A broad usability concern creates attention. A defined problem gives a team something it can change.

How I would structure the program now

I would begin by defining the decisions the metric needs to support. A product-wide health indicator serves a different purpose from evaluating a revised workflow. Trying to make one measure perform both jobs weakens the analysis.

The program would pair a lightweight standardized measure with segmentation and open-text feedback. Collection would happen after selected tasks and at planned intervals, with enough respondent information to compare meaningful cohorts. Research and operational evidence would be reviewed alongside the score before findings reached leadership.

Reporting would identify the affected user group, product area and likely causes. It would also separate confirmed evidence from hypotheses that needed further investigation. The relevant teams would receive ownership of the contributing issues, and the same cohort or workflow would be measured again after changes shipped.

The score would remain visible, though it would appear as one part of the evidence. Leadership would see what had changed, what the team believed caused the movement and where uncertainty remained.

What leadership does with a poor score

The response to an uncomfortable metric affects future research. When teams expect a number to be used as a performance judgment, they become cautious about surfacing bad news. Risk gets softened in presentations, and discussion shifts toward defending the result. That behaviour reduces the value of the measurement program.

A useful leadership review examines the affected customers and the evidence behind their ratings. It traces the product decisions that contributed to the experience, assigns the work and agrees on how improvement will be measured. Accountability remains clear without pretending that one function produced every part of the outcome.

Our SUS score identified a genuine usability problem and gave it organizational visibility. We lost precision when the score was treated as a complete explanation of that problem. A stronger program would have retained the signal, added the missing context and directed each finding to the people able to act on it.