Data Analysis for Design: Metrics, KPIs and Evidence-Based UX Decisions

Long parallel rows of native grass test plots glowing in low evening light, bare soil between each row
Published

2026-09-15

Author

Nural Choudhury

Data analysis for design is the practice of testing decisions against measured evidence, usage data, and controlled experiments, rather than defending them by opinion alone.

What this unblocks:

Design arguments where both sides have a number but no agreed metric, and dashboards nobody can act on because no one settled what they were meant to answer before anyone opened them.

What the output lets you do:

Take a design decision into a room with a documented baseline, a controlled test result, and a written recommendation behind it, rather than a single chart nobody has to defend.

What you have at the end:

A metric defined before the data existed, a baseline, a segmented result from a controlled experiment, and a written finding recording what changed, how confident it is, and what to do next.

Where the method comes from

Data analysis for design has no single inventor or founding paper, unlike the Eisenhower Matrix or the MoSCoW method. It is an amalgam of three separate traditions that converged on product and design teams over the course of a century.

The statistical foundation is Ronald Fisher’s 1935 book The Design of Experiments, which introduced randomised allocation and the significance test, the machinery still underneath every A/B test run today. Google Analytics, launched in November 2005 after Google acquired Urchin Software Corporation, gave ordinary teams the first free, general-purpose tool for measuring site behaviour, moving analysis out of specialist statistics departments. Fred Reichheld’s Net Promoter Score, published in Harvard Business Review in December 2003, added a third strand: reducing a business outcome to one trackable, comparable number.

That third strand has since become contested. In Net Promoter 3.0 (Harvard Business Review, November 2021), Reichheld and co-authors Darnell and Burns acknowledged that Net Promoter Score had become gamed and misused. It introduced earned growth rate, a complementary metric built from accounting data rather than survey scores. CMSWire’s State of the Digital Customer Experience research has since recorded NPS falling from the second to the eighth most used customer-experience metric, and a 2022 paper in the Journal of the Academy of Marketing Science found methodological flaws in the score itself.

The folk version of this history, that data-driven design began with the growth-hacking movement of the 2010s, understates it by seventy-five years. What changed in the 2010s was not the method. It was the cost of running it, once free analytics tools and cheap experimentation platforms put statistical testing within reach of a small design team rather than a dedicated research function.

Three men at an outdoor lectern, one in white robes speaking into a microphone, a bearded man seated at right
R. A. Fisher (right) with Satyendra Nath Bose and P. C. Mahalanobis, 1955. Public domain

The six steps

Run these in order. Skipping an earlier step to save time is the most common way a data-informed decision goes wrong.

  1. Define the metric before you look at any data. State what you are trying to move, as a single measurable number rather than a vague goal such as “better experience”, and why it matters to the business as well as the design. I insist on agreeing on this before anyone touches a dashboard, because it saves me days of arguing about which number counts once the data already exist.
  2. Record a baseline before you change anything. Take it over enough time to average out daily and weekly noise, typically two to four weeks for a metric with regular traffic, because a baseline measured on one unusual day proves nothing. I do not shorten this step for a deadline, since there is no honest way to compress an averaging period whatever the schedule demands.
  3. Verify the tracking before you trust the number. Check that the events firing match the actions they claim to measure, and that the tool’s definition matches the metric you defined in step one. I run this check myself because I have seen a team ship confidently on numbers that were wrong from the start, and when the only number available turns out to be wrong, I say so and go back to step one rather than report against it.
  4. Segment before you conclude anything from an average. Break the number down by device, channel, and user type before deciding what it means, because an aggregate can average two opposite stories into one flat line. I treat a headline number as a question rather than an answer until I have looked underneath it, and this is the step I see skipped most often under deadline pressure.
  5. Test the change with a controlled experiment. Run an A/B split, or an equivalent controlled comparison, for as long as the sample size the expected effect requires, rather than reading a graph after launch and calling it evidence. I would rather wait the extra week than call a launch-week wobble a result, and that wait is measured in days to weeks depending on traffic and the size of the effect I need to detect.
  6. Report the finding with its context, not the number alone. Describe what changed, what it is compared against, how confident the result is, and what you recommend doing next. I write this up myself rather than hand a stakeholder a bare chart, because a chart nobody has to defend is a chart nobody acts on, and that hour of writing is what turns an analysis into a decision.
Two panels of wheat sheaves from Rothamsted plots, each plot's stalk labelled by treatment, heights varying with the manure used
Wheat from the Rothamsted plots, 1878 and 1899: one variable changed per plot, the rest held still. A. D. Hall, 1909, public domain

A worked example

A checkout redesign, walked through the six steps.

Baseline: over the four weeks before any change, checkout converts 840 of 40,000 sessions that reach it, 2.1 per cent once averaged across weekdays and weekends.

Hypothesis: usability testing flagged an account-creation step before payment as the point most users abandon, so removing it should lift conversion.

Test: an A/B split runs for three weeks, half of checkout sessions see the new flow without the account-creation step and half see the old one, chosen to reach the sample size a two-percentage-point lift needs at 95 per cent confidence.

Result: the variant converts 912 of 39,400 sessions (2.3 per cent), a 0.2 percentage point lift and a 10 per cent relative improvement, reaching significance in the third week.

SegmentControlVariant
Desktop2.6 per cent2.9 per cent
Mobile1.4 per cent1.4 per cent
Overall2.1 per cent2.3 per cent

The lift holds on desktop but is flat on mobile, so the account-creation step was never the mobile problem. That finding points the mobile investigation elsewhere.

Bar chart of checkout conversion: overall 2.1 to 2.3 per cent, desktop 2.6 to 2.9, mobile flat at 1.4
The worked example, segmented: the lift is all desktop

Where data analysis for design fails

Optimising the metric while the experience degrades

I have seen a metric chosen because it was easy to move, not because it reflected the experience, get gamed the moment a team was measured on it. I have watched a button made impossible to miss inflate clicks while it buried content the user needed. I treat a metric as a proxy for the experience, never as the experience itself, and I keep a second metric, usually satisfaction or task completion, in view alongside whichever number is the headline.

Mistaking correlation for cause

Two metrics moving together have never told me that one caused the other. I have seen a checkout redesign ship the same week as a seasonal traffic spike and take credit for a conversion lift it may have contributed nothing to. I trust only a controlled experiment, where the variable under test is the sole difference between groups, to separate a cause from a coincidence.

A sample too small to support the conclusion

I do not trust a result that reaches significance on a handful of conversions; that is noise dressed as a finding. Every test carries a minimum sample size for the effect it is trying to detect, and I have stopped a test early because the early numbers looked good, a practice researchers call peeking, only to watch the false positive rate climb well beyond the stated confidence level. I calculate the sample size the effect requires before the test starts, not after the graph looks convincing.

That doesn’t mean a test can’t stop early. Group Sequential Testing and mixture SPRT permit principled early stopping while controlling the false positive rate by design, and platforms including Optimizely, Statsig, and Eppo now implement them. If your test is running on one, my rule for peeking stands: fix the sample size before you start and don’t look early.

Survivorship bias: seeing only the users who stayed

Web analytics, by construction, measure the users still using the product. A user who abandoned it entirely after one bad experience leaves no trail for me to analyse, so a dashboard built only from active users has, every time, underweighted the failure it should be catching most. I cross-reference usage data with churn, exit surveys, or research among people who left, not only the users the dashboard can show me.

A wartime press print of a B-24 bomber with its nose shattered by flak, resting on a mound of earth after a crash landing
A B-24 home from Ploesti, 1944: the planes that return only show the hits a plane can survive. US Air Force, public domain

Common questions

What counts as data analysis for design:

combining measured evidence, usage metrics, experiment results, and funnel data with design judgement to validate a decision instead of asserting it from opinion alone.

Do I need a dedicated analyst to run this:

No. The six steps need an analytics tool that is already tracking the product, a documented baseline, and the discipline to define the metric before making a change, not a dedicated data team.

How long does a test need to run:

Long enough to reach the sample size the effect requires, given current traffic. A small effect on low traffic can take weeks; a large effect on high traffic can resolve in days.

What is the single biggest mistake in this process:

Skipping the baseline. Without a documented figure from before the change, no after-the-fact comparison is credible, no matter what the dashboard shows afterwards.

Does this replace user research? No, they answer different questions. Quantitative data show what changed and by how much; research explains why. Run both together rather than choosing one instead of the other.

Key facts, current as of September 2026

FactDetail
Statistical foundationRonald Fisher formalised randomised allocation and the significance test in The Design of Experiments, 1935
Customer-outcome metricFred Reichheld’s Net Promoter Score, published as “The One Number You Need to Grow”, Harvard Business Review, December 2003
General-purpose web analyticsGoogle Analytics launched November 2005, after Google’s acquisition of Urchin Software Corporation
Data protection standardThe EU General Data Protection Regulation, enforceable from 25 May 2018
Named inventorNone. The practice combines separate statistical, analytics, and customer-research traditions rather than one method with one creator