The call we almost got wrong
Early this year we had two bets on the table and enough people for one.
The first was latency: making the pages where patients find, book, and connect with a provider measurably faster. The second was deprecating a legacy internal analytics platform — a system a whole operations function depended on daily, that had accumulated years of critical workflows, and that no single team owned.
The obvious answer is latency. It’s visible, it’s measurable, everyone in the building already wants it, and it maps cleanly to a customer outcome. It’s the kind of work that makes a good quarterly readout.
We didn’t pick it.
What a TPM actually is here — and what it isn’t
Most engineers’ mental model of a technical program manager is “the person who runs standups.” Our own internal definition says something different: a TPM is an end-to-end leader who drives planning, execution, and delivery of complex, technical, cross-functional initiatives that are critical to the business. The same page carries an explicit list of what TPMs don’t primarily do — schedule meetings, take notes, assign tasks and chase status. Those are means to an end, not the job.
The part that surprises people is the ratio. We run three TPMs for several hundred people across R&D.
That’s a choice, not a hiring backlog. A team of three cannot spread itself across every worthwhile program, so it doesn’t try. Scarcity is the strategy — it forces the team to pick the few problems where ownership is genuinely missing and the stakes are highest, and to hand everything else back.
Which is why the latency-versus-deprecation call was the hardest one we made, and the most useful one to explain.
Latency is not a vanity metric
It would be easy to read the decision as “we chose infrastructure hygiene over customer experience.” That’s not what happened, and it’s worth being precise about why latency mattered.
Headway is a mental healthcare company. When the pages a patient uses to find and book a therapist are slow, some of those patients don’t get to a first session — they drop off somewhere in the middle of a search, on a day when they’d worked up the will to look. On the other side, providers wait longer to fill their schedules. Milliseconds on a search page are, eventually, someone not starting care.
So both bets were real. That’s the whole point. Focus is easy when one option is obviously worse; it’s only a skill when both options deserve to win.
The question we actually asked
The tiebreaker wasn’t “which is more valuable.” Ranked by raw value, latency probably wins, and we’d have made a worse decision.
The question a small team has to ask is narrower:
- If we don’t own this, does it still happen?
- If we do own it, can we hand it back with the ownership intact?
Run latency through that. Engineering already owned the surfaces. There were existing performance patterns, real instrumentation, engineers who knew the code paths cold, and — increasingly — AI tooling that made profiling and refactor work dramatically cheaper to attempt. Latency was going to keep improving whether or not a TPM stood next to it. A TPM would have added a tracker to work that already had momentum.
Now run the analytics deprecation through it. Dozens of business-critical workflows, users spread across several functions, no owning team, no catalogue of what “critical” even meant, and a cost line that grew quietly every month. Nobody was going to wake up and volunteer for it. It had survived years of good intentions precisely because it was everyone’s dependency and nobody’s problem.
One bet needed us. One didn’t. That’s the call.
| The work | The call | Why |
|---|---|---|
| Broad latency across patient and provider critical paths | Hand back to Eng | Clear owners, existing patterns, strong instrumentation, and AI leverage already compounding. TPM presence adds overhead, not throughput. |
| Legacy analytics platform deprecation | TPM owns end-to-end | Business-critical, cross-functional, high ambiguity, and structurally unowned. Will not happen otherwise. |
| Status tracking and reporting for healthy, single-team projects | Decline | The team already has this. Adding a TPM converts a working process into a reporting tax. |
| Data safety, technical risk, AI governance | Take on, sequenced later | Same profile as the deprecation — cross-functional, high-stakes, no natural owner — but they had to wait their turn. |
What “owning it end to end” actually looked like
Here’s the part that separates ownership from coordination, because a deprecation is not a Gantt chart.
We catalogued before we migrated. You cannot deprecate what you haven’t enumerated, and you cannot enumerate it from the platform’s own usage logs — logs tell you what runs, not what matters. So the catalogue was built with the operations teams who depend on the system, and “critical” was signed off by them, not by us. That step is boring and it is the entire program.
Definition of done was continuity, not migration count. The success metric wasn’t “assets moved.” It was zero migration-attributable disruptions to the business, with migrated assets set read-only in the old system before anything was decommissioned. Read-only first means rollback is a flag flip instead of an archaeology project.
Training was a workstream with dates, not a launch email. A migrated workflow that its users can’t operate hasn’t migrated. So the training catalogue had target dates and per-item sign-off, the same as the technical work.
One cadence, one escalation point. A biweekly checkpoint and a single named person to escalate to, so that no operator anywhere in the company had to work out who to ask. Small thing. It’s most of what “cross-functional” costs.
And a note on the system itself: it wasn’t a bad platform. It was load-bearing. That’s why it lasted, and why unwinding it took a dedicated owner rather than a wiki page of good intentions.
Handing work back is the multiplier
Meanwhile, latency kept getting better — because the engineers who owned those pages kept making it better, with more AI leverage each month than the month before. We didn’t ask them to run a program. We got out of the way of one.
This is the part I’d most want another engineering leader to take: a TPM’s win condition is frequently to become unnecessary in a space. Establish the ownership, leave behind the cadence and the rollback pattern and the escalation path, then step out. If the program still needs you in a year, something went wrong.
Coordination overhead is what you get when a TPM is added to work that already has an owner. Acceleration is what you get when a TPM is pointed at work that has none.
Building the muscle for bigger bets
The compounding effect is the real return. An organization that can make one uncomfortable focus call — and be seen to make it deliberately, in public, with the reasoning written down — can make the next one faster.
That’s what let the same three-person team pick up harder, scarier programs after: data safety, technical risk management, AI governance, and our largest internal AI effort. None of those had a natural home either. Each one was easier to start because the org had already watched a small team say no to something popular and be right about it.
What we’d tell another startup
If you’re weighing whether to add a TPM function: yes, and keep it smaller than feels comfortable.
Point it at problems that are cross-functional, business-critical, ambiguous, and — the filter that matters most — structurally unowned. If a team already owns the work and is executing, a TPM will make it slower. If nobody owns it, a TPM is the only thing that will make it happen at all.
The honest risks, since we’re still living them: a three-person team has a real bus factor, and depth on a few programs means we are genuinely absent from things that deserve attention. The signal that it’s time to hire the next TPM isn’t backlog size — every team has a backlog. It’s a second unowned, high-stakes problem sitting untouched while the first one still needs its owner. We watch for that explicitly, and we keep a list of efforts that aren’t TPM-ready yet so the choice stays visible instead of accidental.
Saying no is uncomfortable in exactly the way it’s supposed to be. You are choosing the work that won’t happen otherwise over the work that looks impressive. But the point of the focus isn’t the focus — it’s that more people get care, faster, because a few of the right things actually got finished.
If you’re building a TPM function at a growing company and want to compare notes on where you drew the line, I’d genuinely like to hear it.



