Skip to content
01Performance & adoption

Step Syncing

A fifteen-second launch cut to under two seconds, and step-sync completion up 35%.

Role
Product Analyst, HCL Healthcare
Timeline
8 weeks
Team
Cross-functional initiative with the engineering team
Scope
Diagnosis · prioritisation · scope arbitration · launch gate
App launch time
15sunder 2s

At least 7.5× faster

+35%
Step-sync completion

I owned

  • Diagnosing launch time as the adoption blocker, and putting it ahead of the engagement features on the roadmap I had defined
  • The decision to hold that roadmap for eight weeks, and the argument for why a latency number belonged to product rather than to an engineering backlog
  • Treating launch time and bundle size as product requirements with a stated ceiling, not as housekeeping

We shipped

  • Launch time from 15s to under 2s
  • App bundle from 25MB to 6MB
  • Step-sync completion up 35%

I did not own

  • The engineering. I did not choose the technical approach, profile the startup path, or write any of the code. The tech team did.
  • The instrumentation. Whatever measurement existed was not built by me, and its limits bound what I can claim here.

The step count was behind a fifteen-second wall

HCL Healthcare’s consumer health app serves more than a million registered users. I had defined the product roadmap, and what was on it was engagement: reasons to open a health app in a week when nothing about your health has changed.

Then I timed the launch. Fifteen seconds, against a two-second benchmark.

That is not a performance statistic. It is the price of admission to everything else on the roadmap. Someone who taps the icon and waits fifteen seconds is not forming an opinion about your streaks feature. They are forming one about whether the app works.

The whole case study lives inside one arrow. Steps are counted by the phone whether or not the app is open. The product’s only job is to be openable.

Two frames of the same app. The one on the left is what fifteen seconds bought you:

Launch · before and afterReconstructed · all data synthetic
Before15s
Loading…
Afterunder 2s
Today7,412steps
Synced just now

The screen on the left has nothing to read and nothing to do. It is not a bad screen; it is the correct screen for an app that is not ready yet, which is the point. Fifteen seconds of it stands between a person and the one number they opened the app to see.

Layout and hierarchy accurate; every name, number and date is invented. No real user appears here.

Latency sat in the engineering column. Adoption sat on my roadmap.

The fifteen seconds were already known. They sat where slow things sit, next to refactors and platform upgrades, competing for engineering time against other engineering work. Classified that way it is a quality problem, and quality problems lose to deadlines for years at a stretch.

I reclassified it, and not because I could fix it. I could not have written any of that code. What a problem is filed under decides what it gets compared against, and a latency number owned by product competes against features. That is the only comparison in which its real value is visible.

Three ways to spend eight weeks

  • Ship the engagement roadmap as written

    The default. It needed no argument from anyone.

    What it costs
    Nothing new. It was already the plan.
    Reversible?
    Per feature, yes. The eight weeks, no.
    How you’d learn you were wrong
    You wouldn’t. Every read is taken on the people who got in.
  • Instrument the pre-launch window first

    The cheapest option to be wrong about.

    What it costs
    Engineering time on measurement, and the roadmap still waits.
    Reversible?
    Most reversible. The instrument outlives the decision either way.
    How you’d learn you were wrong
    Directly. That is the entire point of it.
  • Fix launch time first, hold the roadmap, chosen

    What I chose. It scored worst on reversibility.

    What it costs
    Eight weeks of engagement work, not shipped.
    Reversible?
    Least reversible in calendar terms. Durable in outcome, until it regresses.
    How you’d learn you were wrong
    The number moves and adoption doesn’t follow. Then the diagnosis was wrong.

The reasoning laid out. Reversibility is in the table because it is the dimension that separates a strategist from a prioritiser, and it is the one this decision scored worst on.

What shipped

Launch time under two seconds, and the app bundle down from 25MB to 6MB, inside eight weeks. The bundle is the number I trust most in this case study: two stated totals, one exact percentage, and no population to argue about.

Area is proportional to size. A 76% reduction, the one figure here that computes exactly from two stated totals.

What happened

The app bundle went from 25MB to 6MB, a 76% reduction. Launch time went from fifteen seconds to under two.

Step-sync completion rose 35%. More people reached the number they had opened the app for, because the thing standing between them and it was gone, though this shipped inside a period with other work in it, so latency was not necessarily the only thing in the way.

What “step-sync completion” actually counts

A completion rate is a fraction, and the argument is always about the denominator. Three readings of this one are defensible and they are not the same number.

  • Per attempt

    Syncs that succeeded ÷ syncs that started.

    What it measures
    Pipeline reliability.
    What inflates it
    Nothing much, but it also moves when nobody tries, because failed attempts leave the denominator too.
    Verdict
    An engineering metric wearing a product label.
  • Per session

    Sessions where the step count rendered ÷ sessions opened.

    What it measures
    How often the app showed the number.
    What inflates it
    A faster launch. Mechanically. Whether or not one extra person got value.
    Verdict
    Rejected. It is the reading my own intervention inflates.
  • Per user, per period, chosen

    Users who saw a current step count that week ÷ users who opened the app that week.

    What it measures
    Whether more people reached the number they came for.
    What inflates it
    Little. A user who opens twice and sees it once counts once either way.
    Verdict
    The one I would defend. It answers the question the work was about.

The definition work behind the 35%. Reversibility is not the axis here. Inflation is: which of these moves for reasons that have nothing to do with whether the product got better.

A metric that improves as a side effect of the thing you did is not evidence about the thing you did. That is the whole reason per-session is out: it would have given me a bigger number and a worse argument.

My record does not fix which of the three the 35% was measured against (the single thing I would most want back), which is why that figure travels with a caveat and the bundle number does not.

What I would do differently

I picked a problem whose success criterion I could not measure cleanly. Device capability is not randomisable, so the clean instrument was never available. I traded a measurable bet for an unmeasurable one because I thought the unmeasurable one was bigger.

The other thing I would change is durability. A performance win is not a state, it is a position you hold, and the next three sprints of feature work are where it goes back. If I ran this again the size ceiling would go into the build itself on week one, so the win stops depending on anyone remembering it.

How I worked this out

Why you cannot A/B test a latency fix, and what you compare instead

Randomisation is what makes an A/B test work: you assign the treatment, and because assignment is random the two arms differ only in what you assigned. Device capability breaks that. A phone’s processor, its free storage and what the other apps on it are holding are the variables that decide what a cold start costs, and none of them is assignable. You can randomise which build a device receives. You cannot randomise the device. So the instrument degrades to a staged rollout with a before-and-after comparison held within a device tier: weaker, and worth saying out loud rather than dressing up as an experiment.

Why a mean launch time is the wrong statistic here

A mean is pulled toward the fast devices, and the fast devices belong disproportionately to the people who build the app and the people least likely to leave. Abandonment lives in the tail. A mean that improves while the tail does not is a number that got better for a population that was never going anywhere, and it will look exactly like a win.

Results

Measured outcomes

App launch time
15sunder 2s

At least 7.5× faster

Against a two-second benchmark. The record states a bound, so the bar is drawn to it.

App bundle
25MB6MB

76% smaller

The one exact percentage in this case study.

+35%

Step-sync completion

More people reached the number they opened the app for.

Next case study

Steps Premier League

A competitive step league built from nothing, moving session time from 3.5 to 7.8 minutes, on a north star I would not choose again, and I explain why.

Ask me about this