· PM Editorial · Product Sense  · 6 min read

Improve Uber Eats: Metrics and North Star

How to define a north star metric and supporting KPI tree for improving Uber Eats, covering orders per user, delivery time p50/p95, restaurant NPS, and gross merchandise value.

How to define a north star metric and supporting KPI tree for improving Uber Eats, covering orders per user, delivery time p50/p95, restaurant NPS, and gross merchandise value.

Why Metrics Questions on Marketplace Products Are Different

Choosing a north star metric for a three-sided marketplace like Uber Eats is harder than for a single-sided consumer app because any metric you pick can be gamed at the expense of another side of the marketplace. A north star that only measures eater behavior (like orders per user) can be juiced by aggressive promotions that destroy restaurant margins or overwork drivers. Interviewers in mid-2026 loops are specifically probing whether candidates understand this three-sided tension, not just whether they can name a metric.

Step 1: Rule Out Weak North Star Candidates

  • Gross Merchandise Value (GMV) alone: measures total transaction volume but can be inflated by heavy discounting that destroys unit economics — a company can grow GMV while losing more money per order.
  • Total orders: doesn’t account for order value or frequency patterns; ten $5 orders look identical to two $25 orders in this metric despite very different margin profiles.
  • App downloads or MAU: measures top-of-funnel reach, not actual marketplace health or repeat engagement.

Naming why these fail before proposing your answer demonstrates the same metric maturity interviewers look for in any product sense question.

Step 2: Propose the North Star Metric

North Star: Orders Per Active User Per Month, weighted by a minimum order-completion-quality threshold (i.e., only counting orders that were delivered without a refund or major complaint).

This is a stronger choice than GMV because:

  • It’s a frequency metric, directly tied to whether users are forming a repeat habit of using the platform, which is the actual business goal (retention-driven growth, not one-time acquisition).
  • Weighting by completion quality prevents the metric from being gamed by volume growth achieved through poor service (fast but wrong orders).
  • It naturally correlates with Uber One membership value, which is Uber’s actual strategic lever for eater retention in 2026’s competitive landscape against DoorDash.

Step 3: Supporting Metrics Tree

MetricDefinitionMarketplace SideWhy It Matters
Orders Per Active User/Month (North Star)Quality-weighted order frequencyEatersCore repeat-usage habit indicator
Delivery Time P50Median delivery timeEatersTypical experience quality
Delivery Time P9595th percentile delivery timeEatersWorst-case experience, most correlated with churn
Restaurant NPSNet Promoter Score surveyed from merchant partnersRestaurantsLong-term merchant supply health
Gross Merchandise Value (GMV)Total transaction valueAll sidesRevenue proxy, guardrail not north star
Driver Utilization Rate% of online driver-hours spent on active deliveriesDriversMarketplace liquidity/efficiency

Step 4: Why P50 and P95 Delivery Time Matter More Than Average

A common mistake is reporting only average delivery time, which hides the tail experience that actually drives churn. If P50 delivery time is 28 minutes (perfectly fine) but P95 is 75 minutes, the average might read as an acceptable 32 minutes while a meaningful fraction of customers are having a genuinely bad experience — and those are the customers most likely to churn to a competitor or abandon the platform for good.

Track both explicitly:

  • P50 (median): represents the typical, expected experience for most users. Improvements here move the “this app is reliable” perception broadly.
  • P95 (tail): represents the worst-case experience. Improvements here specifically reduce churn from the most dissatisfied segment of users, who disproportionately leave negative reviews and drive support costs.

A mature product answer proposes tracking P95 as the primary operational target, because tail latency reduction has outsized impact on trust relative to its cost, while shaving a few minutes off an already-good median has diminishing returns.

Step 5: Restaurant NPS as a Leading Indicator of Supply Health

Restaurant NPS is often overlooked in eater-focused interview answers, but it’s a critical leading indicator: if restaurant satisfaction declines, the best restaurants (with the most order flow, highest ratings, and best food quality) are the ones most likely to reduce their commitment to the platform or negotiate down their exclusivity — degrading the eater-side selection over time. This is the marketplace flywheel effect interviewers want candidates to articulate.

Track restaurant NPS with a segmented breakdown:

Restaurant SegmentSample NPS (2026 benchmark)Primary Driver
High-volume chains+35Reliable order flow, predictable commission
Independent restaurants+8Commission rate sensitivity, packaging cost concerns
New restaurant partners (<90 days)-12Onboarding friction, unclear payout timelines

This table illustrates why a single blended restaurant NPS number would hide a serious new-partner onboarding problem — a pattern analogous to the segmented retention tracking recommended for consumer products.

Step 6: GMV as a Guardrail, Not a Goal

GMV should absolutely be tracked and reported to the business, but it functions best as a guardrail metric rather than the optimization target for a product team focused on marketplace health. The distinction matters in an interview: optimizing directly for GMV incentivizes discounting and promotional order inflation that can hurt long-term unit economics, whereas optimizing for quality-weighted order frequency (the proposed north star) tends to grow GMV as a healthy byproduct rather than a forced outcome.

Propose monitoring the ratio of GMV growth to promotional spend growth — if GMV is growing 10% quarter over quarter but promotional spend is growing 25%, that’s a red flag that growth is being bought rather than earned through genuine repeat usage.

Step 7: Driver Utilization Rate as the Efficiency Counterweight

Because driver experience was named explicitly in the original prompt, include driver utilization rate (percentage of online driver-hours actually spent on deliveries) as a marketplace efficiency metric. Low utilization means drivers are online but idle, which either means there are too many drivers relative to order volume (bad driver economics) or dispatch algorithms are inefficient. High utilization is good up to a point, but utilization pushed too high (drivers constantly loaded with back-to-back deliveries with no rest) correlates with higher driver churn and burnout, which eventually degrades delivery reliability.

Step 8: How the Metrics Tree Changes With Company Priorities

If the interviewer signals the company is in a growth-investment phase (pre-profitability, prioritizing market share), weight the metrics tree toward orders per user and GMV growth, with NPS and utilization as guardrails. If the interviewer signals a profitability-focused phase (which better reflects Uber’s actual 2026 posture as a mature public company under margin pressure), weight the metrics tree toward quality-weighted order frequency and restaurant NPS, treating GMV growth achieved through unsustainable discounting as an active red flag rather than a win.

Book Reference

For more marketplace-specific metrics trees and north star frameworks applied to multi-sided platforms, The 100x Product Manager Interview Playbook (Amazon: https://www.amazon.com/dp/B0DBC1FQWH?tag=sirjohnnymai-20) walks through worked metrics answers for ride-sharing, food delivery, and other gig-economy marketplace prompts commonly asked in 2026 PM interview loops.

Summary

For a marketplace product like Uber Eats, the north star should be a quality-weighted repeat-usage metric (orders per active user per month) rather than raw GMV or order count. Support it with P50/P95 delivery time to capture both typical and tail experience, restaurant NPS to monitor supply-side health, and driver utilization as an efficiency counterweight, while treating GMV as a guardrail rather than the primary optimization target. This structure demonstrates the marketplace-equilibrium thinking that separates a strong answer from a generic one.

Back to Blog

Related Posts

View All Posts »