User Research Methods: Qualitative, Quantitative and Evidence-Based UX Guide

A bearded man in glasses gestures as he speaks during a small focus group session around a table
Published

2026-09-15

Author

Nural Choudhury

User research is the systematic study of what users do, need and struggle with, combining qualitative depth with quantitative scale to ground design in evidence.

What this unblocks:

A design decision stuck on competing opinions, or a team about to build something on an assumption nobody has tested. Naming the actual question and matching it to a method that can answer it turns an argument nobody can settle into something you can check.

What the output lets you do:

You put evidence in front of the decision instead of another opinion, and hand findings to the person making the call before they decide. Findings that need turning into a recommendation feed data analysis for design. The test I carry over is whether the research could plausibly have said no.

What you have at the end:

A written research question tied to a decision, the method you chose and why, the session records or dataset you collected, and synthesised findings connected back to that decision.

Where the method comes from

Software teams often assume user research began with them. It didn’t. The discipline descends from human factors engineering, which studied how people operate complex machinery long before computers reached a general audience.

Two lineages gave it a shared vocabulary. Jakob Nielsen and Rolf Molich published usability heuristics in 1990. The heuristics gave researchers a common checklist for interface problems. Karen Holtzblatt and Hugh Beyer developed contextual inquiry at Digital Equipment Corporation: watching people work in their own setting rather than a lab. They published the method as Contextual Design in 1998.

Every method below sits on one of these two lineages, or borrows from an older social-science tradition, ethnography, ergonomics, market research, that predates computing altogether. None of them was built for software specifically. Each was adapted.

Two cards naming 1990 usability heuristics and 1998 contextual inquiry as source methods

Choosing and running a method

I answer one question at a time, with the method that question needs, not the method my team already knows how to run.

  1. Define the question, and name the decision it serves. Write down what you need to know and who needs the answer to make a decision. A question without an attached decision isn’t worth researching. Cost: under a day.
  2. Match the method to the question. Choose a qualitative method- interview, contextual inquiry, ethnography- for a why or how question, and choose a quantitative method- survey, analytics, A/B test- for a how many or how much question. Most real questions need both: qualitative work to generate a hypothesis, quantitative work to test it at scale. I reach for an interview first on a why or how question, and I will not run a survey until I know what to ask. A focus group is the exception: I keep it for early discovery and to gauge interest in a concept, and I no longer run one to settle a product decision. Nielsen Norman Group’s current guidance explains why: groupthink and social desirability bias mean a group tells you what might happen, whereas in-context testing with individuals shows you what happens. Cost: an hour of planning.
  3. Recruit participants who match your target users, not whoever answers first. A sample of colleagues or the easiest customers to reach produces confident, misleading findings. Screen for the behaviour and context your question depends on. When I don’t have a budget for a recruiting agency, I pull from my own customer list rather than a generic panel: the people on it still need to have lived the problem I am studying. Cost: two days to two weeks, depending on how narrow the target group is.
  4. Run the sessions from a guide, not a script. Prepare open questions and a rough order for interviews and usability tests, then follow what the participant says. Pilot survey questions on five people before sending them to five hundred. No-code and AI-assisted tools have lowered the barrier to running a session without a dedicated researcher, and AI-moderated interview platforms are improving quickly. I still don’t hand a study over to an AI moderator: Nielsen Norman Group’s December 2024 assessment found the technology is not there yet, while noting real potential as it develops further. Cost: 30 to 60 minutes per participant for a qualitative session; days to weeks to collect a quantitative sample.
  5. Analyse before you interpret. Code qualitative data into themes, and run descriptive statistics before inferential ones on quantitative data. Do this before you form an opinion about what the data means, or you will find the pattern you expected rather than the one that is there. Cost: at least as long as collecting the data.
  6. Connect every finding to a decision, in writing. State what the finding means for a specific choice, and hand it to the person making that choice before they make it. A finding with no decision attached gets filed and forgotten. Cost: half a day, but only if it happens before the decision.

These six steps describe how to run one study well, and they still hold. What has changed is the cadence: continuous discovery, the model Teresa Torres has argued for, keeps customer contact running throughout the product lifecycle rather than confined to a discrete project. I still work through the same six steps, but I now treat them as a loop against a standing question rather than a one-off assignment that ends when the readout is delivered.

Eight methods cover most research questions.

MethodBest forTypical cost
InterviewThe motivation and context behind a behaviour30 to 60 minutes per participant
Contextual inquiryWatching a task performed in its real setting, not described afterwards1 to 2 hours per participant, on site
EthnographyCultural norms and group dynamics that shape individual behaviourDays to weeks of immersion
Focus groupGauging interest in a concept during early discovery, not settling a design decision60 to 90 minutes per group
SurveyHow many people hold a view or report a behaviourDays to collect, once questions are piloted
A/B testWhich of two variants performs better on a defined metricDays to weeks, until the sample reaches significance
Analytics reviewWhat users already do at scale, without being askedOngoing, once instrumented
Usability testWhere a design breaks down under a real task30 to 60 minutes per participant; five for a formative round
Two axis chart mapping interview, focus group, survey and five more research methods
Diagram plotting the eight methods on attitudinal versus behavioural and qualitative versus quantitative axes.

Diary studies

Diary studies extend self-report over days or weeks instead of one session, useful when behaviour varies by day or context in ways a single visit misses. Entry design, participant fatigue and analysis load all need their own treatment, so the method has a page of its own: Diary studies.

Four day cards each with a short entry, showing a diary study spread across days

A worked example

Analytics show two out of five checkout attempts abandon between the shipping and payment steps, but the dashboard cannot say why.

You choose two methods: five moderated usability tests to see where people get stuck, then an A/B test on the leading fix. You recruit five participants who abandoned a similar checkout in the last month, screened from your own customer list rather than a generic panel, because they need to have hit the problem themselves.

Each session runs 45 minutes: ten minutes of background questions, then a task- complete a purchase on the prototype- with the participant thinking aloud. Four of the five stall at the same field: a shipping-address format the system rejects without explaining why.

Nielsen’s argument for five participants in formative usability testing holds here. The same problem showed up in four of five sessions, which is the signal the method is built to catch. A sixth or seventh participant would likely hit the same field. Five isn’t enough to decide whether shoppers prefer a blue or a green button. That is a preference question, and it needs the larger sample a statistical claim requires.

You fix the field validation and the error message, then run the new flow against the old one for two weeks to confirm the fix moves the abandonment number before rolling it out to everyone.

Checkout funnel bar chart narrowing sharply between the delivery and payment steps

Where user research fails

User research fails in the same few ways, and I have watched all four happen.

I have watched a team ask instead of observe: what people say they would do and what they do diverge, especially where money, willpower or social approval are involved. A survey question, “Would you pay for this?”, gets a different answer than watching what someone does with their own card in hand. I treat stated intent as a hypothesis, not a finding.

I have watched research recruit whoever was easiest to reach: colleagues, the loudest customers, the first names on a list. A convenience sample produces confident findings about the wrong population. When the recruitment criteria describe who was available rather than who I need to understand, I know the research will answer a different question than the one I asked.

I have watched research arrive after the decision. Commissioning a study to validate a choice the team already made turns research into theatre: the budget bought confirmation, not evidence.

I have watched findings confirm what the team wanted to build all along: confirmation bias runs through what gets asked, what gets coded as a theme and which findings make the readout. I test this by asking whether the research could plausibly have said no. A team whose every study has vindicated its existing roadmap is not running research that can surprise it.

Four dark cards naming ask instead of observe, convenience sample, after the decision and confirmation bias
Diagram testing the four common research failures against whether the study could have said no.

Common questions

What is the difference between qualitative and quantitative user research:

Qualitative methods, interviews, contextual inquiry, ethnography, answer why and how, using small numbers of participants studied in depth. Quantitative methods: surveys, analytics, A/B tests; answer how many and how much; use larger samples to support a statistical claim. Most programmes need both.

Are five participants really enough to test a design? Only for formative usability testing, the practice of finding major problems early in a design’s life. Nielsen’s 2000 argument for five participants is scoped to that case. It does not apply to research generally, and never to a question that needs a statistical claim, such as which of two designs converts better.

How do you recruit participants who represent your users:

Define the behaviour and context your question depends on before you recruit, then screen against that definition rather than convenience. A sample chosen because it was easy to reach usually answers a different question than the one you asked.

When should research happen in a project:

Before the decision it informs, not after. Research commissioned to validate a choice the team has already made produces confirmation, not evidence.

What is the biggest risk when analysing findings:

Confirmation bias, in what gets asked, what gets coded as a theme and which findings reach the readout. Ask whether the research could plausibly have said no to what the team wanted.

Key facts, current as of September 2026

ItemWhere it stands
Also known asUser experience research, UX research or design research: an umbrella term for a family of methods rather than one named technique
Formalised as a disciplineGrew out of human factors engineering and human-computer interaction in the 1980s, as computing moved from specialists to a mass audience
Founding evaluation methodJakob Nielsen and Rolf Molich published usability heuristics for spotting interface problems in 1990
Founding contextual methodKaren Holtzblatt and Hugh Beyer developed contextual inquiry at Digital Equipment Corporation, published as Contextual Design in 1998
Formative usability sample sizeNielsen argued in 2000 that five participants surface most usability problems in one round of testing. The figure is specific to formative usability testing and is often misapplied to research generally
Current statusA mixed-methods practice, standard across product and design teams, pairing qualitative work that generates hypotheses with quantitative work that tests them at scale