How international product companies — from startups to enterprise — decide what to build: frameworks, the strategy→prioritization link, decision processes, who decides, post-launch effectiveness and experimentation, ties to metrics and plans, feature deprecation, and anti-patterns. An industry reference grounded in primary sources.
Method note: a deep-research workflow's adversarial-verify stage failed on a rate limit (all claims returned "0-0 abstain", which the pipeline mislabeled "refuted" — a false refutation, not a real one). Load-bearing claims were therefore re-verified by hand via direct WebFetch; the two SVPG (Marty Cagan) pages that returned HTTP 403 were verified through a Chrome browser session (verbatim).
Scoring frameworks compute a number for ranking. Every practitioner agrees: it's a tool for discussion and discipline, not truth. Processes and theory are in §3–§5.
| Framework | How it works / formula | Strengths | Weaknesses |
|---|---|---|---|
| RICE ✓ | (Reach × Impact × Confidence) ÷ Effort. Impact 3/2/1/0.5/0.25; Confidence 100/80/50%; Effort in person-months. Measures "total impact per time worked."1 | Balances value vs cost; Confidence curbs "exciting but ill-defined ideas" | Author: "shouldn't be used as a hard and fast rule"; estimates are subjective |
| ICE ✓ | Impact × Confidence × Ease (1–10).15 | "Faster to run, works well for early-stage teams" | Even more subjective than RICE |
| Kano ✓ | Must-be / One-dimensional / Attractive / Indifferent / Reverse, by impact on satisfaction.10 | Separates "basic" from "delight" features | "Broad and subject to interpretation" |
| MoSCoW ✓ | Must / Should / Could / Won't have.10 | "Easy to understand and to communicate" | "Rely on opinions, consensus, and rank" |
| WSJF (SAFe) ✓ | Cost of Delay ÷ Job Size; CoD = Value + Time Criticality + Risk Reduction. "Sequence work for maximum economic benefit."16 | Enterprise standard (SAFe); economic logic | "Vague and hard to estimate"; lagging indicators |
| Cost of Delay (CD3) ✓ | Value × Urgency ÷ Duration.10 | Focus on business metrics | "Optimizing just for money"; lagging |
| Value vs Effort (2×2) ✓ | Plot value against effort; take the quick wins.15 | "Fast initial triage" | "Consistency breaks at scale" |
| Opportunity Scoring (Ulwick) ✓ | Survey importance × satisfaction → underserved outcomes.1519 | Finds underserved customer outcomes | "Survey bias can skew the graph" |
| Story Mapping ✓ | User journey (horizontal) × importance (vertical).15 | "Identifying MVP and release slices" | "Doesn't account for business value or technical complexity" |
| Buy-a-Feature ✓ | Stakeholders "buy" features with a budget.14 | Builds consensus; reveals real priorities | Slow with a large backlog |
| Product Tree ✓ | Roots / trunk / branches / leaves.14 | Structural product hierarchy | Qualitative, not quantitative |
| Eisenhower ✓ | 2×2 important × urgent.10 | Simple discussion aid | Subjective |
| North Star Framework ✓ | A leading metric + input metrics; an alignment system.9 | "Brings coherence to the work of everyone" | Doesn't rank individual features |
| DHM (Biddle) ✓ | Prioritize what Delights customers in Hard-to-copy, Margin-enhancing ways.21 | Ties delight to moat and margin | Qualitative; needs strategic judgment |
| Stack ranking | Forced 1…N ordering, no formula. ❓ | Maximally fast; founder-led startups | Doesn't scale |
Scoring only weights signals. The most important theoretical lens for inputs is Jobs-to-be-Done: look not at features but at the "job" a customer hires the product to do.
| Signal | How it's used | Source |
|---|---|---|
| Customer "job" (JTBD) | Defines what is even worth prioritizing — functional/social/emotional progress | Christensen18, Ulwick19 ✓ |
| Product analytics | Reach in RICE = "real measurements... instead of pulling numbers from a hat"; usage also drives deprecation | Intercom1 ✓ |
| Data → insights → beliefs | Spotify DIBB: behavior+market → insights → hypotheses → bets | Spotify11 ✓ |
| Customer feedback / sales / CS | Valuable, but don't take requests literally — dig to the root pain; top-down requests to a feature team is an anti-pattern | Lenny13, Cagan5 ✓ |
| Revenue / economics | Cost of Delay and WSJF quantify the cost of waiting | SAFe16 ✓ |
| Maintenance cost / tech debt | Cross-reference usage × support tickets; "every line of code... is a liability" | DevProJournal12 ✓ |
The core thesis of modern PM theory: if prioritizing is hard, the strategy is broken — not the execution. Prioritization should flow down from strategy.
Five layers, each a prerequisite for the next: Mission → Company Strategy → Product Strategy → Roadmap → Goals. "Product Strategy serves a critical role—it is the connective tissue between the objectives of the company and the product delivery work of the product team." And bluntly: "It is impossible to make rigorous prioritization decisions when the guidance on how to do so is missing, unclear or disconnected from what you are trying to do."20 ✓
Discovery first (what to build), delivery second (how). Patton: "The most expensive way to test your idea is to build production quality software"; discovery and delivery are "two kinds of work, and two kinds of thinking."24 Torres: "product discovery... decisions about what to build, while product delivery is the work we do to build, ship, and maintain."25 ✓
Product reviews (Figma): decisions made at reviews "about making decisions and debating directions," where teams present an "option space" of possible solutions.7 ✓ Counter-example (Linear): "We don't do A/B tests. We validate ideas and assumptions that are driven by taste and opinions"; no product metric goals, only a company North Star; then "it's just a matter of sequencing and scoping."8 ✓
Two axes: the team model (Cagan) and the allocation of decision rights (DACI/RAPID). The common denominator of mature practice — one explicit decision owner.
| Model / framework | Essence | Source |
|---|---|---|
| Empowered team (Cagan) | "Cross-functional; focused on and measured by outcomes (rather than output); and empowered to figure out the best way to solve the problems they've been asked to solve." Litmus test: "the team is able to decide the best way to solve the problems they have been assigned." Purpose: "to serve the customers, in ways that meet the needs of the business." | Cagan / SVPG56 ✓ |
| Feature team (Cagan) | "All about output. Features... provided to the team in the form of a prioritized list that is called the roadmap." Value and viability are "the responsibility of the stakeholder or executive that requested the feature" — prioritization power sits with executives, not the team. | Cagan / SVPG5 ✓ |
| Delivery team (Cagan) | Output-focused ("dev/scrum teams... if your company is running something like SAFe"); product owner = "backlog administrator." "Really just re-packaged waterfall." | Cagan / SVPG5 ✓ |
| DACI (Atlassian) | Driver runs the process; Approver is "the one person (yes: one!) who makes the decision"; Contributors "have a voice, but not a vote"; Informed get "no vote, no voice." | Atlassian28 ✓ |
| RAPID (Bain) | Recommend / Agree / Perform / Input / Decide. "It comes down to one person who must decide—the single point of accountability who commits the organization to action." | Bain29 ✓ |
| RACI | Accountable = single owner of the outcome; Responsible = executes. In product: PM Responsible for requirements, engineering Accountable for delivery. | LaunchNotes30 ◐ |
| "Informed captain" (Netflix) | "For every significant decision, we identify an informed captain who's responsible for making a judgment call"; "context not control"; after a decision, "disagree then commit." | Netflix50 ✓ |
| Betting table (Basecamp) | CEO ("the last word on product") + CTO + a senior programmer + a product strategist; the call "rarely goes longer than an hour or two." | Basecamp2 ✓ |
The central fact driving the modern approach: most ideas don't improve the metric they were designed for. So mature companies don't "believe in a feature" — they measure.
Ronny Kohavi (Microsoft/Bing): for typical teams, "⅓ of the experiments have a positive significant result, ⅓ have no effect, and ⅓ have a significant negative effect"; in an optimized domain like Bing "the winner ratio is about 10% to 20%."38 At Airbnb, "out of 250 ideas tested... only 20... had a positive impact," producing "a 6% improvement in booking conversion, worth hundreds of millions."39 Netflix: "Every product change... goes through a rigorous A/B testing process before becoming the default," ensuring decisions are "not driven by the most opinionated and vocal Netflix employees, but instead by actual data."40 Amazon (Bezos, HBR 2007): "maximize the number of experiments you can do per given unit of time... The key, really, is reducing the cost of the experiments."45 ◐
Kerry Rodden (Google, CHI 2010): "Happiness, Engagement, Adoption, Retention, Task success" plus a Goals→Signals→Metrics process that "should lead to a natural prioritization of the various metrics."4243 ✓
Pendo: "Feature adoption measures the usage for a software product's specific features"; formula "Monthly Feature Adoption Rate (%) = [feature MAU / monthly logins] * 100"; dimensions are breadth (how widely adopted) and depth (how often key user types use it).44 ✓
Three levels: North Star (where we're going), OKR (this quarter), roadmap horizons (when). All connect daily prioritization to strategy.
Evolution in a real company (Figma): first "I actually deprecated OKRs," replacing them with "headlines"; later, after hiring a data-science leader, OKRs returned as "commitments." Cadence: annual company priorities; team goals/roadmaps revisited twice a year; mid-half adjustments.7 ✓
Mature PM removes as systematically as it adds — a separate process with its own criteria.
| Step | What teams do | Why |
|---|---|---|
| 1. Data audit | Irrefutable telemetry: how many unique users touched the feature in 6 months | "Prevents internal politics from derailing the decision" |
| 2. Decision matrix | Maintenance × value; beware the "noisy minority trap"; high-effort/low-value "must die immediately" | Compute: cost to keep vs alternatives |
| 3. Execute + communicate | 6–12 mo runway: warn support/sales → EOL announcement with dates → migration path. "Don't use corporate speak" | Honesty preserves trust |
Remove vs iterate: high-effort/low-value — remove; high-value — optimize.12 Intercom: "your product needs to maintain focus in order to maintain value."17 Google pays a reputational price for aggressive sunsets — a public "graveyard" of 306+ products.46 ✓
A strength of PM theory is its self-critique. The main traps:
The pattern: as a company grows, process replaces intuition and power shifts from founder to teams and committees. But not linearly — Linear keeps "founder-taste" at scale.
| Dimension | Startup | Scale-up | Enterprise |
|---|---|---|---|
| Who decides | Founder/CEO "last word"; small betting table2 | PM/CPO + teams; product reviews7 | DACI/RAPID, portfolio committees, exec go/no-go29 |
| Method | Intuition/taste; stack rank; "bets" | RICE/ICE, continuous discovery (OST), DIBB | WSJF/SAFe, portfolio bets, governance |
| Data | Little: "you simply don't have enough data"13 | Usage + feedback; first A/B tests | A/B at scale (thousands/mo)38 |
| Goal | "Make 10 customers very happy"13 | Grow input metrics → North Star | Portfolio economics, strategic bets |
| Plans | Often no formal metric goals; short cycles | OKR + North Star appear; half-year cycles7 | Cascaded OKR/KPI; annual + quarterly |
| Feature removal | Easy — little legacy | Deprecation becomes a process (usage audit) | Formal EOL policy + migration |
| Examples | Basecamp, Linear, early Airbnb49 | Intercom, Figma, Spotify, Duolingo, Shopify | Amazon, Google, Microsoft, Netflix, Booking |
| Company | Method | Essence | Stage/type |
|---|---|---|---|
| Intercom | RICE | Built their own scoring system "from first principles"1 | Scale-up |
| Basecamp | Shape Up "bets" | 6-week cycles, no backlog, 4-person betting table2 | Founder-led |
| Amazon | Working Backwards + Weblab | PR/FAQ gate; "maximize the number of experiments per unit of time"345 | Enterprise |
| Spotify | DIBB + bet board | Data→Insights→Beliefs→Bets, 3 bet levels11 | Enterprise |
| Figma | OKR→commitments; product reviews | Goal-system evolution; "option space"; CEO in the room7 | Scale-up→ent. |
| Linear | Taste-driven (anti-framework) | No A/B, no RICE; only a company North Star8 | Scale-up |
| OKR + HEART; aggressive sunset | Doerr brought OKRs (1999)32; HEART for UX43; 306 "killed"46 | Enterprise | |
| Microsoft | Controlled experiments | Bing ~1,200 experiments/mo; "⅓/⅓/⅓"38 | Enterprise |
| Netflix | A/B + "informed captain" | "Every product change... goes through A/B"40; "context not control"50 | Enterprise |
| Booking.com | Mass experimentation | ~25,000 tests/year41 | Enterprise |
| Duolingo | Experiment-driven | "A few hundred experiments running simultaneously... over 2,000 in total"48 | Growth/scaled |
| Shopify | GSD (Get Shit Done) | 6-week cycles; project leads free to choose approach47 | Scaled |
| Airbnb | "Snow White" storyboarding | Storyboard the customer journey; gaps become "number one priority"49 | Early→growth |
| Atlassian | DACI | One Approver; a Driver runs the process28 | Scaled/ent. |
An honest map of residual blind spots:
| Residual gap | Reason / status |
|---|---|
| HBR primaries (JTBD milkshake, Booking culture, Bezos 2007) | Paywalled — taken via secondary/intro fragments; some marked ◐ |
| Rumelt (Good/Bad Strategy) | Primaries unreachable (404/timeout); quotes from two independent summaries, ~85% ◐ |
| Amazon Weblab exact volumes (12k/yr, 546→1976) | Secondary only; not confirmed against an Amazon primary ❓ |
| Lenny's Newsletter (Figma/Linear/Shopify detail) | Partial paywall — only the accessible portion before the cut-off was used |
| Framework depth (Story Mapping, ODI mechanics) | Covered at overview level; implementation needs the primary books (Patton; Ulwick "Jobs to be Done") |
| B2B vs B2C specificity | Most large-scale A/B examples are B2C (Netflix/Booking/Duolingo); for B2B — where Linear explicitly warns against A/B — this warrants separate research ❓ |
All fetched this session (WebFetch, or Chrome where noted):