Steps Premier League
A competitive step league built from nothing, moving session time from 3.5 to 7.8 minutes, on a north star I would not choose again, and I explain why.
- Role
- Product Analyst, HCL Healthcare
- Team
- Cross-functional, with engineering and design · 8 teams of 7 in the shipped season
- Scope
- Cohort analysis · strategy evaluation · mechanic design · launch
HCL Healthcare’s public page for the league. Eight teams of seven, a five-day season, and the scoring ladder that converts steps into cricket runs. Published by HCL Healthcare; I am listed on it as the admin contact.
The declared north star for this launch. It is the wrong one for a retention feature, which I take up below.
- 0→1
- Built from nothing
- 3
- Strategies evaluated
I owned
- The cohort analysis that put a drop-off on the table instead of a feature request
- The evaluation of three strategies (content, incentives, gamification) and the argument for competition
- Prioritising competitive mechanics as the core of what shipped, and naming what that choice costs
We shipped
- Steps Premier League, launched 0→1
- A competitive mechanic built on step data the phone already collects
I did not own
- The build. I made the case for the mechanic; I did not build the league.
- The analytics. I did not build the instrument that produced 3.5 → 7.8, and its limits bound what I can claim from it.
A step counter competes with the phone in your pocket
Cohort analysis put a drop-off on the table. Not a feature request from anyone, but a shape in the data that said people were arriving and then not coming back.
The underlying problem is structural, and it is not specific to this app. Your phone already counts your steps. It tells you the number for free, on the lock screen, without anyone opening anything. A health app that shows you the same figure has to answer a question the operating system has already answered better.
A season ends, standings reset, and everyone gets a fresh reason to start
Three strategies were live options. Only one had no recurring bill.
Content, incentives, gamification. All three were real, all three had advocates, and there was appetite for roughly one.
Content and incentives buy attention you have to keep paying for. Another article, another voucher, next week, forever. Competition manufactures the reason to come back out of other users, which makes it the only one of the three whose cost is the mechanic rather than the fuel.
Content
Recurring
Every week needs new material, and the bill never stops.
Incentives
Recurring
Withdraw the reward and the behaviour it bought goes with it.
Competition
ShippedOne-time, then self-sustaining
Other users supply the reason to return. The cost is the mechanic, not the fuel.
Two design choices in what shipped answer part of that, and they are visible on the live league page rather than only in my account of it.
The first is the scoring ladder. Steps do not go onto a leaderboard as steps. They convert into cricket runs on four bands, one over per day, across a five-day season:
- Dot ball
- 0runs
- 0 – 5,000 steps
- Double
- 2runs
- 5,001 – 10,000
- Boundary
- 4runs
- 10,001 – 15,000
- Sixer
- 6runs
- Above 15,000
- Power playDay 3, every run doubled
- 3-day streak+10 bonus runs
- 5-day streak+20 bonus runs
Four bands, one over per day, a five-day season. The bands are the whole design argument: the first one pays nothing, which is what makes the second one worth walking for.
The bands are the mitigation. A raw step leaderboard ranks people by a continuous number, so everybody has a distinct position and somebody is always last by a visible margin. Bands collapse that: everyone between 10,001 and 15,000 steps scores the same four runs. Two people 4,000 steps apart are level on the table. It is a deliberately coarse instrument, and coarseness is what makes it survivable for the person at the bottom of a band.
The second is that the unit of competition is a team, not a person. Eight teams of seven, drawn along department lines. Being last in a squad of seven is a different experience from being eighth of eight on a public ladder: your bad week is absorbed rather than displayed.
Neither choice removes the cost I named above. A standings table still tells the last-placed team where it stands, every day, and there is an individual ranking alongside the team one. But both push in the right direction, and the honest version of this case study is that I can point at the shipped design rather than describe the design I would have wanted.
What shipped
Three retention strategies were evaluated. One shipped, and it shipped from nothing.
A recurring, time-boxed league built on step data the phone already collects: a five-day season, eight teams of seven, standings that reset when the season does. The design insight is in the restart: a league with no season end has no urgency, and one with no restart has no second chance for anybody who fell behind.
What happened, and the number that should have been here instead
Session time moved from 3.5 to 7.8 minutes. That was the declared north star for the launch, and it is the number I have.
It is the wrong north star for this feature. This was built to fix a drop-off, a drop-off is a retention problem, and the metric that answers a retention problem is whether people came back: a week-two-to-week-four return rate for the people who joined a league, measured against the same weeks for the people who did not. That is the headline this case study should have.
I deferred that instrument. The decision picked the metric, and I did not see the connection at the time.
Session time is a defensible supporting metric for a loop whose whole action is checking where you stand. It is a poor primary one for a health product, because a health product that works can shorten a session. Open it, log the thing, leave. Both are true at once, which is why the number needs an argument beside it rather than a bigger typeface.
The guardrails I would have watched
I did not have these. A retention mechanic that buys its numbers by burning notification permission is a loss wearing a win’s clothes, and none of the counter-metrics that would catch it were in the launch plan. They belong there rather than in a post-mortem, so here is the plan I would run it with now.
| Counter-metric | The false win it catches | What would have stopped the rollout |
|---|---|---|
| Notification opt-out rate | A league that manufactures its daily open by spending permission it can never get back. Opt-out is one-way. Every point you lose here is a point every future feature is also missing. | Any rise among people who joined a league, against the same weeks for people who did not. |
| Uninstalls | The person in last place leaving. Session time goes up when they go, because the people who stayed are the people who were winning. The metric improves precisely because the mechanic failed someone. | An uninstall rate among league joiners above the same weeks for non-joiners, at all. |
| Silent season non-entry | Quitting that does not look like quitting. Somebody who stops entering a new season but keeps the app installed is a churned user the install count will never show you. | Season-two re-entry below season-one enrolment by a margin I would set before launch, not after seeing it. |
| A fairness or complaint signal | Leagues that are not contests. A cohort where one person is unreachably ahead is demotivating for everyone else in it, and step data is trivially gamed by putting a phone on a treadmill. | Any support-ticket theme about rank fairness, or cohorts where the leader is unreachable by week one. |
- Notification opt-out rate
- A league that manufactures its daily open by spending permission it can never get back. Opt-out is one-way. Every point you lose here is a point every future feature is also missing.
- What would have stopped the rollout: Any rise among people who joined a league, against the same weeks for people who did not.
- Uninstalls
- The person in last place leaving. Session time goes up when they go, because the people who stayed are the people who were winning. The metric improves precisely because the mechanic failed someone.
- What would have stopped the rollout: An uninstall rate among league joiners above the same weeks for non-joiners, at all.
- Silent season non-entry
- Quitting that does not look like quitting. Somebody who stops entering a new season but keeps the app installed is a churned user the install count will never show you.
- What would have stopped the rollout: Season-two re-entry below season-one enrolment by a margin I would set before launch, not after seeing it.
- A fairness or complaint signal
- Leagues that are not contests. A cohort where one person is unreachably ahead is demotivating for everyone else in it, and step data is trivially gamed by putting a phone on a treadmill.
- What would have stopped the rollout: Any support-ticket theme about rank fairness, or cohorts where the leader is unreachable by week one.
Not measured. These are the counter-metrics I would put in the launch plan if I ran this again, and the thresholds at which I would have stopped the rollout rather than argued with it.
What I would do differently
I would instrument the cohort before picking the metric, not after. Everything above follows from that one sequencing error: the north star, the missing guardrails, and the fact that the strongest claim I could make about a retention feature is a number about session length.
How I worked this out
Why the phone already having the data is the whole opportunity
Almost every engagement mechanic in a consumer health app eventually competes with a person’s willingness to type. Steps are the exception, and not because step challenges are clever: the phone counts them whether or not anyone opens the app. The moment a loop needs a number only the user can supply, such as a meal, a reading or a home test, it acquires a daily tax, and the people who pay that tax longest tend to be the people who needed the product least. A league built on step data starts with the input problem already solved, which is why it was the cheapest of the three options to get right.
What a standings table costs the person in last place
A competition sorts a user base into people who enjoy being ranked and people who feel worse for having been ranked, and in a health product the second group is not a rounding error. The honest version of this mechanic needs a way out that is not deletion: a smaller cohort, an opt-out that does not read as quitting, or a second axis to be good at. That is the part of the design I would push hardest on with more time, and it is the reason I would not ship this mechanic into a clinical population without changing it.
Results
Measured outcomes
Declared north star. Supporting metric at best. The retention number this should have led with was never instrumented.
0→1
Built from nothing
No prior mechanic to extend.
3
Strategies evaluated
Content, incentives, gamification. One shipped.
Next case study
AI Smart Health ReportA generated health report people could actually read, that became a key USP in five enterprise deals.