
2026-09-15
Nural Choudhury
User research is the systematic study of what users do, need and struggle with, combining qualitative depth with quantitative scale to ground design in evidence.
A design decision stuck on competing opinions, or a team about to build something on an assumption nobody has tested. Naming the actual question and matching it to a method that can answer it turns an argument nobody can settle into something you can check.
You put evidence in front of the decision instead of another opinion, and hand findings to the person making the call before they decide. Findings that need turning into a recommendation feed data analysis for design. The test I carry over is whether the research could plausibly have said no.
A written research question tied to a decision, the method you chose and why, the session records or dataset you collected, and synthesised findings connected back to that decision.
Software teams often assume user research began with them. It didn’t. The discipline descends from human factors engineering, which studied how people operate complex machinery long before computers reached a general audience.
Two lineages gave it a shared vocabulary. Jakob Nielsen and Rolf Molich published usability heuristics in 1990. The heuristics gave researchers a common checklist for interface problems. Karen Holtzblatt and Hugh Beyer developed contextual inquiry at Digital Equipment Corporation: watching people work in their own setting rather than a lab. They published the method as Contextual Design in 1998.
Every method below sits on one of these two lineages, or borrows from an older social-science tradition, ethnography, ergonomics, market research, that predates computing altogether. None of them was built for software specifically. Each was adapted.

I answer one question at a time, with the method that question needs, not the method my team already knows how to run.
These six steps describe how to run one study well, and they still hold. What has changed is the cadence: continuous discovery, the model Teresa Torres has argued for, keeps customer contact running throughout the product lifecycle rather than confined to a discrete project. I still work through the same six steps, but I now treat them as a loop against a standing question rather than a one-off assignment that ends when the readout is delivered.
Eight methods cover most research questions.
| Method | Best for | Typical cost |
|---|---|---|
| Interview | The motivation and context behind a behaviour | 30 to 60 minutes per participant |
| Contextual inquiry | Watching a task performed in its real setting, not described afterwards | 1 to 2 hours per participant, on site |
| Ethnography | Cultural norms and group dynamics that shape individual behaviour | Days to weeks of immersion |
| Focus group | Gauging interest in a concept during early discovery, not settling a design decision | 60 to 90 minutes per group |
| Survey | How many people hold a view or report a behaviour | Days to collect, once questions are piloted |
| A/B test | Which of two variants performs better on a defined metric | Days to weeks, until the sample reaches significance |
| Analytics review | What users already do at scale, without being asked | Ongoing, once instrumented |
| Usability test | Where a design breaks down under a real task | 30 to 60 minutes per participant; five for a formative round |

Diary studies extend self-report over days or weeks instead of one session, useful when behaviour varies by day or context in ways a single visit misses. Entry design, participant fatigue and analysis load all need their own treatment, so the method has a page of its own: Diary studies.

Analytics show two out of five checkout attempts abandon between the shipping and payment steps, but the dashboard cannot say why.
You choose two methods: five moderated usability tests to see where people get stuck, then an A/B test on the leading fix. You recruit five participants who abandoned a similar checkout in the last month, screened from your own customer list rather than a generic panel, because they need to have hit the problem themselves.
Each session runs 45 minutes: ten minutes of background questions, then a task- complete a purchase on the prototype- with the participant thinking aloud. Four of the five stall at the same field: a shipping-address format the system rejects without explaining why.
Nielsen’s argument for five participants in formative usability testing holds here. The same problem showed up in four of five sessions, which is the signal the method is built to catch. A sixth or seventh participant would likely hit the same field. Five isn’t enough to decide whether shoppers prefer a blue or a green button. That is a preference question, and it needs the larger sample a statistical claim requires.
You fix the field validation and the error message, then run the new flow against the old one for two weeks to confirm the fix moves the abandonment number before rolling it out to everyone.

User research fails in the same few ways, and I have watched all four happen.
I have watched a team ask instead of observe: what people say they would do and what they do diverge, especially where money, willpower or social approval are involved. A survey question, “Would you pay for this?”, gets a different answer than watching what someone does with their own card in hand. I treat stated intent as a hypothesis, not a finding.
I have watched research recruit whoever was easiest to reach: colleagues, the loudest customers, the first names on a list. A convenience sample produces confident findings about the wrong population. When the recruitment criteria describe who was available rather than who I need to understand, I know the research will answer a different question than the one I asked.
I have watched research arrive after the decision. Commissioning a study to validate a choice the team already made turns research into theatre: the budget bought confirmation, not evidence.
I have watched findings confirm what the team wanted to build all along: confirmation bias runs through what gets asked, what gets coded as a theme and which findings make the readout. I test this by asking whether the research could plausibly have said no. A team whose every study has vindicated its existing roadmap is not running research that can surprise it.

Qualitative methods, interviews, contextual inquiry, ethnography, answer why and how, using small numbers of participants studied in depth. Quantitative methods: surveys, analytics, A/B tests; answer how many and how much; use larger samples to support a statistical claim. Most programmes need both.
Are five participants really enough to test a design? Only for formative usability testing, the practice of finding major problems early in a design’s life. Nielsen’s 2000 argument for five participants is scoped to that case. It does not apply to research generally, and never to a question that needs a statistical claim, such as which of two designs converts better.
Define the behaviour and context your question depends on before you recruit, then screen against that definition rather than convenience. A sample chosen because it was easy to reach usually answers a different question than the one you asked.
Before the decision it informs, not after. Research commissioned to validate a choice the team has already made produces confirmation, not evidence.
Confirmation bias, in what gets asked, what gets coded as a theme and which findings reach the readout. Ask whether the research could plausibly have said no to what the team wanted.
| Item | Where it stands |
|---|---|
| Also known as | User experience research, UX research or design research: an umbrella term for a family of methods rather than one named technique |
| Formalised as a discipline | Grew out of human factors engineering and human-computer interaction in the 1980s, as computing moved from specialists to a mass audience |
| Founding evaluation method | Jakob Nielsen and Rolf Molich published usability heuristics for spotting interface problems in 1990 |
| Founding contextual method | Karen Holtzblatt and Hugh Beyer developed contextual inquiry at Digital Equipment Corporation, published as Contextual Design in 1998 |
| Formative usability sample size | Nielsen argued in 2000 that five participants surface most usability problems in one round of testing. The figure is specific to formative usability testing and is often misapplied to research generally |
| Current status | A mixed-methods practice, standard across product and design teams, pairing qualitative work that generates hypotheses with quantitative work that tests them at scale |

Task analysis is the systematic study of how people currently perform a task, breaking it down into the steps, decisions and knowledge each one requires.
Read it
A diary study is a longitudinal research method in which participants record their behaviour and experience in their environment over days or weeks, revealing patterns that a single session cannot.
Read it
A content audit is a catalogue of every piece of content you own, scored against what it is supposed to do, so you can decide what to keep, fix, merge or kill.
Read it
A wireframe shows a screen’s structure without visual design, and a prototype simulates its behaviour, so a team can test both before committing to development.
Read it
Data analysis for design is the practice of testing decisions against measured evidence, usage data, and controlled experiments, rather than defending them by opinion alone.
Read it
An interface pattern is a solution to a recurring design problem, refined through enough use elsewhere that the person encountering it doesn’t have to figure out how to use it.
Read it