Idea Evaluation Matrix: Scoring Ideas on Value, Effort and Risk

Whiteboard divided into columns with yellow sticky notes grouped under suggestions and obstacles
Published

2026-09-15

Author

Nural Choudhury

An idea evaluation matrix is a structured scoring tool that ranks competing ideas against a shared set of weighted criteria, typically value, effort, and risk, so a team can compare options that all look plausible consistently rather than by whoever argues loudest.

Each idea gets an independent score on every criterion; the scores are multiplied by an agreed weight, and the weighted totals are ranked. The matrix does not replace judgement. It puts the judgement on the table, so it can be checked and argued with rather than assumed.

What this unblocks:

The shortlist argument where several plausible ideas are championed by different people, and without a shared structure the loudest or most senior voice settles it instead of the evidence.

What the output lets you do:

You can defend a prioritisation decision to people who were not in the room, because a reason sits next to every score. It feeds straight into the MoSCoW Method once the shortlisted ideas become a backlog to sequence.

What you have at the end:

The agreed criteria and weights fixed before anyone scored, a score sheet with one recorded reason per score, and a weighted ranking that flags which gaps are real and which are too narrow to trust.

Understanding the idea evaluation matrix

The problem it solves: several plausible ideas, one shortlist

Most prioritisation problems do not arrive as one good idea against nine bad ones. They arrive as four or five ideas that all sound reasonable, championed by people who all sound confident. Without a shared structure, the idea that wins is usually the one argued for by the most senior person in the room, or the one raised last, rather than the one that would deliver the most.

An idea evaluation matrix forces every idea through the same set of questions, scored by the same people, before anyone commits resources. It does not generate ideas: it filters ideas that already exist, produced by methods such as crazy eights or mind mapping, down to the ones worth building.

Where the method comes from

The lineage is genuine, not decorative. Stuart Pugh, Professor of Engineering Design at the University of Strathclyde, published “controlled convergence” in Total Design: Integrated Methods for Successful Product Engineering in 1990: a matrix that scores each concept as better, the same, or worse than a baseline against an agreed set of criteria, then tallies the result.

Software product teams inherited the same logic through lighter, faster variants. Sean McBride’s RICE score (reach, impact, confidence, effort), built at Intercom and published in January 2018, and Sean Ellis and Morgan Brown’s ICE score (impact, confidence, ease), popularised in Hacking Growth in 2017, both do what Pugh’s matrix does: combine weighted factors into one comparable number.

WSJF and Kano scoring apply the same logic to different inputs, cost of delay and customer delight rather than reach or ease. None of these frameworks replaces Pugh’s matrix: they implement it, and their spread through product teams has made practitioners more alert to the method’s limits, not less.

The six steps, in order

  1. Agree the criteria as a group, before anyone scores a single idea.
  2. Set the weights, and write them down, before anyone sees a score.
  3. Score every idea against every criterion independently, without discussion.
  4. Reveal all scores at once, and discuss only the ones that diverge.
  5. Total the weighted scores and rank the ideas.
  6. Treat a narrow gap as a tie that needs a different decision rule, not a verdict.

When to reach for it, and when not to

The matrix earns its cost when several options are genuinely close, when the decision is expensive to reverse, or when a group needs to see its own disagreement made visible before it can resolve it. It is the wrong tool for a single obvious choice, and it is the wrong tool when nobody in the room disagrees: scoring a decision that is already settled produces theatre, not insight.

Blank index cards and coloured sticky notes arranged in a grid on a wooden desk
Several plausible options laid out side by side for comparison

Choosing the right criteria

Three axes that cover most decisions

Most idea evaluation matrices reduce cleanly to three questions, and adding more rarely changes the ranking enough to justify the extra argument:

  • Value: what does this idea move, for the customer or for the business, if it works?
  • Effort: what does it cost in time, budget, and people to find out?
  • Risk: what has to go right, and how much do you know about whether it will?

Some teams add a fourth or fifth axis, strategic alignment or novelty among the most common, and the original Pugh matrix scores against as many criteria as the team’s context demands. Every axis beyond three raises the argument’s cost faster than it raises its accuracy.

Keep the list short enough to argue about in one sitting

A criterion earns its place by changing at least one idea’s rank. A criterion that would not move the shortlist if dropped is doing no useful work: it only decorates the matrix with more rigour than the decision underneath it carries. Five or six criteria is the practical ceiling for a group that still wants to finish the discussion in one sitting.

Criteria that quietly repeat each other

“Impact” and “strategic alignment” often measure the same underlying belief twice, once directly and once through a proxy, and an idea that scores well on one usually scores well on the other for the same reason. Counting related criteria as independent inflates their combined weight without anyone deciding that on purpose. Check for this before scoring starts, not after the totals look wrong.

Three cards naming the matrix axes: Value, Effort and Risk, each with the question it asks
Three axes that cover most decisions: value, effort and risk

Weighting without rigging the outcome

Set the weights before anyone sees a score.

Weights decided after scores are visible get pulled toward whichever weighting produces the answer the room already wanted. Tversky and Kahneman’s 1974 research on anchoring showed that an early value systematically distorts the estimates that follow it, and a weight adjusted once it can be seen decides the ranking the same way. I fix the weights first, in a separate conversation from scoring, and hold them once ideas are on the table.

A worked example

Three ideas, three criteria, weights fixed in advance at value 0.5, effort 0.3, and risk 0.2. Effort and risk are scored so a higher number is always better (low effort scores high, low risk scores high), matching value’s direction, or the total quietly rewards the costliest, riskiest option.

IdeaValue, weight 0.5Effort, weight 0.3Risk, weight 0.2Weighted total
Self-serve onboarding flow4353.9
In-app referral programme5243.9
Predictive search3443.5

Predictive search is out. The other two are tied on the arithmetic, and the arithmetic alone cannot say which to build.

Weights are a claim, and claims need owners.

A weight of 0.5 on value asserts that value matters two and a half times as much as risk, stated as a number rather than as a sentence somebody has to defend. Name who set each weight and why, in one line next to the matrix. A weight nobody can explain is a weight nobody holds.

A weighted scoring table for three ideas, with two tied at 3.9 and predictive search out at 3.5
The worked example, weights fixed in advance: the arithmetic narrows the field but cannot break the tie

Scoring so the loudest voice does not win

Score independently, then compare

I ask every scorer to submit numbers before the group sees anyone else’s. A senior voice scoring first, out loud, anchors everyone who scores after them toward the same number, whether or not it is right. Independent scoring is the single change that stops the matrix from formalising whoever spoke first.

Talk about the outliers, not the average.

Once scores are revealed, I discuss the criteria where scorers disagreed by two points or more, not the ones where everyone already agreed. Disagreement is information: it usually means people are scoring against different assumptions about what the idea involves, and that gap is worth closing before it gets buried inside an average.

Write down the reason, not just the number

A score without a reason cannot be checked later, when the idea has either delivered or not. I record one sentence per score explaining what the number is based on: a customer research finding, an engineering estimate, a guess. The guesses are fine to keep, as long as everyone can see which rows are guesses.

Itamar Gilad makes the same case for ICE scoring specifically: a confidence number only earns trust once it is grounded in the type of evidence behind it, not treated as a figure in isolation.

Reading the output

What a wide gap tells you

A clear leader, well ahead of the field, usually means the group already agreed before it started scoring and the matrix confirmed a shared read of the situation. That is a legitimate use of the tool: making a real consensus visible and defensible to stakeholders who were not in the room is worth doing even when the outcome was never seriously in doubt.

What a narrow gap tells you

A close field means the group is genuinely uncertain, and the total score is reporting that uncertainty back with more decimal places than the inputs deserve. I read a narrow gap as the matrix did not decide this, not as a ranking to defer to. I treat scores within a few tenths of a point as equivalent, and look for the criterion the group cares most about as the tie-breaker instead of the compound total.

When the top two tie

I do not re-weight the matrix until the tie breaks: that is exactly the anchoring problem the weights were set up to prevent. Instead, I check whether the two ideas compete for the same resource, ask which one produces more information faster if it turns out to be wrong, or run both as small, time-boxed experiments and let real evidence separate them where the matrix could not.

Where the matrix fails

The arithmetic trap: a weighted score laundering a judgement

A score of 3.9 looks more objective than the sentence “I think this one is roughly as good as that one,” but it is built from the same subjective inputs, just with a decimal point. That is not my only caution. Teresa Torres argues against scoring ideas and stack-ranking the result, and pushes teams toward evidence-based debate and opportunity solution trees instead.

Silicon Valley Product Group makes a related case against sophisticated scoring spreadsheets and algorithms for prioritisation, arguing they push through features users do not value, and recommends product scorecards tied to business outcomes instead. I still use the matrix, because a shared, checkable structure for stating an opinion beats an unstated one, but I treat that critique as a live argument to answer. I present the matrix as a structuring tool, and say so, rather than as an oracle that settled the question.

Criteria nobody disagrees about

If every idea scores the same on a criterion, that criterion isn’t discriminating, and its weight is dead work carried through every calculation for no return. Drop it from the matrix, or replace it with something the ideas genuinely differ on, before the next round.

The matrix is an input, not the decision.

Dependencies, sequencing, and one idea being a prerequisite for another rarely show up as a matrix criterion, and a lower-scoring idea sometimes has to go first anyway. Use the ranking to structure the conversation with stakeholders, not to end it: a matrix that overrides a fact the group already knows about the real world has been given more authority than the arithmetic earns. Current guidance across product and engineering practice agrees on this much: a score is an input to a decision, not decision authority, and the field has grown more explicit about the false confidence and false precision a clean decimal point invites.

Common questions

What is an idea evaluation matrix used for:

Comparing several plausible ideas on the same weighted criteria, usually value, effort, and risk, so the shortlist reflects evidence rather than whoever argued hardest in the room.

How many criteria should the matrix use:

Three covers most decisions: value, effort, and risk. Stop adding criteria once a new one would not change any idea’s rank.

Who should score the ideas:

Everyone with a stake in the outcome, scoring independently before sharing scores, so no single voice anchors the rest of the group.

What happens when two ideas tie:

Do not re-weight to break it. Check for a real dependency between them, or run both as small experiments and let evidence decide.

Is a weighted score enough on its own to decide:

No. It structures the argument and makes disagreement visible; it does not remove the judgement behind every score and weight.

Where does the idea evaluation matrix come from? Stuart Pugh’s 1990 “controlled convergence” method for engineering design, echoed in lighter product variants such as RICE (2018) and ICE (2017).

Key facts, current as of September 2026

ItemWhere it stands
Historical rootStuart Pugh’s “controlled convergence” method, published in “Total Design: Integrated Methods for Successful Product Engineering”, 1990, while Pugh was Professor of Engineering Design at the University of Strathclyde
RICE (Reach, Impact, Confidence, Effort)Built by Sean McBride’s team at Intercom and published publicly on 5 January 2018 as a scoring formula for roadmap ideas
ICE (Impact, Confidence, Ease)The lighter, three-factor variant popularised by Sean Ellis and Morgan Brown in their 2017 book “Hacking Growth”, built for speed over precision
Anchoring riskTversky and Kahneman’s 1974 study on judgement under uncertainty found an early estimate systematically distorts the ones that follow it, the reason weights get set before any idea is scored
Current statusWeighted decision matrices, and the datum-relative form Pugh popularised, remain standard practice across engineering design and product management today