
2026-09-15
Nural Choudhury
An idea evaluation matrix is a structured scoring tool that ranks competing ideas against a shared set of weighted criteria, typically value, effort, and risk, so a team can compare options that all look plausible consistently rather than by whoever argues loudest.
Each idea gets an independent score on every criterion; the scores are multiplied by an agreed weight, and the weighted totals are ranked. The matrix does not replace judgement. It puts the judgement on the table, so it can be checked and argued with rather than assumed.
The shortlist argument where several plausible ideas are championed by different people, and without a shared structure the loudest or most senior voice settles it instead of the evidence.
You can defend a prioritisation decision to people who were not in the room, because a reason sits next to every score. It feeds straight into the MoSCoW Method once the shortlisted ideas become a backlog to sequence.
The agreed criteria and weights fixed before anyone scored, a score sheet with one recorded reason per score, and a weighted ranking that flags which gaps are real and which are too narrow to trust.
Most prioritisation problems do not arrive as one good idea against nine bad ones. They arrive as four or five ideas that all sound reasonable, championed by people who all sound confident. Without a shared structure, the idea that wins is usually the one argued for by the most senior person in the room, or the one raised last, rather than the one that would deliver the most.
An idea evaluation matrix forces every idea through the same set of questions, scored by the same people, before anyone commits resources. It does not generate ideas: it filters ideas that already exist, produced by methods such as crazy eights or mind mapping, down to the ones worth building.
The lineage is genuine, not decorative. Stuart Pugh, Professor of Engineering Design at the University of Strathclyde, published “controlled convergence” in Total Design: Integrated Methods for Successful Product Engineering in 1990: a matrix that scores each concept as better, the same, or worse than a baseline against an agreed set of criteria, then tallies the result.
Software product teams inherited the same logic through lighter, faster variants. Sean McBride’s RICE score (reach, impact, confidence, effort), built at Intercom and published in January 2018, and Sean Ellis and Morgan Brown’s ICE score (impact, confidence, ease), popularised in Hacking Growth in 2017, both do what Pugh’s matrix does: combine weighted factors into one comparable number.
WSJF and Kano scoring apply the same logic to different inputs, cost of delay and customer delight rather than reach or ease. None of these frameworks replaces Pugh’s matrix: they implement it, and their spread through product teams has made practitioners more alert to the method’s limits, not less.
The matrix earns its cost when several options are genuinely close, when the decision is expensive to reverse, or when a group needs to see its own disagreement made visible before it can resolve it. It is the wrong tool for a single obvious choice, and it is the wrong tool when nobody in the room disagrees: scoring a decision that is already settled produces theatre, not insight.

Most idea evaluation matrices reduce cleanly to three questions, and adding more rarely changes the ranking enough to justify the extra argument:
Some teams add a fourth or fifth axis, strategic alignment or novelty among the most common, and the original Pugh matrix scores against as many criteria as the team’s context demands. Every axis beyond three raises the argument’s cost faster than it raises its accuracy.
A criterion earns its place by changing at least one idea’s rank. A criterion that would not move the shortlist if dropped is doing no useful work: it only decorates the matrix with more rigour than the decision underneath it carries. Five or six criteria is the practical ceiling for a group that still wants to finish the discussion in one sitting.
“Impact” and “strategic alignment” often measure the same underlying belief twice, once directly and once through a proxy, and an idea that scores well on one usually scores well on the other for the same reason. Counting related criteria as independent inflates their combined weight without anyone deciding that on purpose. Check for this before scoring starts, not after the totals look wrong.

Weights decided after scores are visible get pulled toward whichever weighting produces the answer the room already wanted. Tversky and Kahneman’s 1974 research on anchoring showed that an early value systematically distorts the estimates that follow it, and a weight adjusted once it can be seen decides the ranking the same way. I fix the weights first, in a separate conversation from scoring, and hold them once ideas are on the table.
Three ideas, three criteria, weights fixed in advance at value 0.5, effort 0.3, and risk 0.2. Effort and risk are scored so a higher number is always better (low effort scores high, low risk scores high), matching value’s direction, or the total quietly rewards the costliest, riskiest option.
| Idea | Value, weight 0.5 | Effort, weight 0.3 | Risk, weight 0.2 | Weighted total |
|---|---|---|---|---|
| Self-serve onboarding flow | 4 | 3 | 5 | 3.9 |
| In-app referral programme | 5 | 2 | 4 | 3.9 |
| Predictive search | 3 | 4 | 4 | 3.5 |
Predictive search is out. The other two are tied on the arithmetic, and the arithmetic alone cannot say which to build.
A weight of 0.5 on value asserts that value matters two and a half times as much as risk, stated as a number rather than as a sentence somebody has to defend. Name who set each weight and why, in one line next to the matrix. A weight nobody can explain is a weight nobody holds.

I ask every scorer to submit numbers before the group sees anyone else’s. A senior voice scoring first, out loud, anchors everyone who scores after them toward the same number, whether or not it is right. Independent scoring is the single change that stops the matrix from formalising whoever spoke first.
Once scores are revealed, I discuss the criteria where scorers disagreed by two points or more, not the ones where everyone already agreed. Disagreement is information: it usually means people are scoring against different assumptions about what the idea involves, and that gap is worth closing before it gets buried inside an average.
A score without a reason cannot be checked later, when the idea has either delivered or not. I record one sentence per score explaining what the number is based on: a customer research finding, an engineering estimate, a guess. The guesses are fine to keep, as long as everyone can see which rows are guesses.
Itamar Gilad makes the same case for ICE scoring specifically: a confidence number only earns trust once it is grounded in the type of evidence behind it, not treated as a figure in isolation.
A clear leader, well ahead of the field, usually means the group already agreed before it started scoring and the matrix confirmed a shared read of the situation. That is a legitimate use of the tool: making a real consensus visible and defensible to stakeholders who were not in the room is worth doing even when the outcome was never seriously in doubt.
A close field means the group is genuinely uncertain, and the total score is reporting that uncertainty back with more decimal places than the inputs deserve. I read a narrow gap as the matrix did not decide this, not as a ranking to defer to. I treat scores within a few tenths of a point as equivalent, and look for the criterion the group cares most about as the tie-breaker instead of the compound total.
I do not re-weight the matrix until the tie breaks: that is exactly the anchoring problem the weights were set up to prevent. Instead, I check whether the two ideas compete for the same resource, ask which one produces more information faster if it turns out to be wrong, or run both as small, time-boxed experiments and let real evidence separate them where the matrix could not.
A score of 3.9 looks more objective than the sentence “I think this one is roughly as good as that one,” but it is built from the same subjective inputs, just with a decimal point. That is not my only caution. Teresa Torres argues against scoring ideas and stack-ranking the result, and pushes teams toward evidence-based debate and opportunity solution trees instead.
Silicon Valley Product Group makes a related case against sophisticated scoring spreadsheets and algorithms for prioritisation, arguing they push through features users do not value, and recommends product scorecards tied to business outcomes instead. I still use the matrix, because a shared, checkable structure for stating an opinion beats an unstated one, but I treat that critique as a live argument to answer. I present the matrix as a structuring tool, and say so, rather than as an oracle that settled the question.
If every idea scores the same on a criterion, that criterion isn’t discriminating, and its weight is dead work carried through every calculation for no return. Drop it from the matrix, or replace it with something the ideas genuinely differ on, before the next round.
Dependencies, sequencing, and one idea being a prerequisite for another rarely show up as a matrix criterion, and a lower-scoring idea sometimes has to go first anyway. Use the ranking to structure the conversation with stakeholders, not to end it: a matrix that overrides a fact the group already knows about the real world has been given more authority than the arithmetic earns. Current guidance across product and engineering practice agrees on this much: a score is an input to a decision, not decision authority, and the field has grown more explicit about the false confidence and false precision a clean decimal point invites.
Comparing several plausible ideas on the same weighted criteria, usually value, effort, and risk, so the shortlist reflects evidence rather than whoever argued hardest in the room.
Three covers most decisions: value, effort, and risk. Stop adding criteria once a new one would not change any idea’s rank.
Everyone with a stake in the outcome, scoring independently before sharing scores, so no single voice anchors the rest of the group.
Do not re-weight to break it. Check for a real dependency between them, or run both as small experiments and let evidence decide.
No. It structures the argument and makes disagreement visible; it does not remove the judgement behind every score and weight.
Where does the idea evaluation matrix come from? Stuart Pugh’s 1990 “controlled convergence” method for engineering design, echoed in lighter product variants such as RICE (2018) and ICE (2017).
| Item | Where it stands |
|---|---|
| Historical root | Stuart Pugh’s “controlled convergence” method, published in “Total Design: Integrated Methods for Successful Product Engineering”, 1990, while Pugh was Professor of Engineering Design at the University of Strathclyde |
| RICE (Reach, Impact, Confidence, Effort) | Built by Sean McBride’s team at Intercom and published publicly on 5 January 2018 as a scoring formula for roadmap ideas |
| ICE (Impact, Confidence, Ease) | The lighter, three-factor variant popularised by Sean Ellis and Morgan Brown in their 2017 book “Hacking Growth”, built for speed over precision |
| Anchoring risk | Tversky and Kahneman’s 1974 study on judgement under uncertainty found an early estimate systematically distorts the ones that follow it, the reason weights get set before any idea is scored |
| Current status | Weighted decision matrices, and the datum-relative form Pugh popularised, remain standard practice across engineering design and product management today |

Mind mapping is a technique for organising ideas around a central topic using branching, radial diagrams instead of linear lists.
Read it
A design sprint is a five-day process that takes a team from a business problem to a tested prototype, revealing whether an idea works before it is built.
Read it
The MoSCoW method is a requirements-prioritisation technique that sorts features into Must Have, Should Have, Could Have, and Won’t Have, protecting a fixed delivery date without cutting corners nobody chose to cut.
Read it
The Eisenhower Matrix is a prioritisation tool that sorts work into four quadrants by urgency and importance, so a leader spends deliberate time on what matters instead of only reacting to what shouts loudest.
Read it