Note · Method · 2026
The progress metric that punished the work I was asked for
I built MedStar’s desktop toolkit myself. Then the owner asked me to direct the developers migrating its capabilities into the company platform, and to own the feature backlog — with no framework, no formal backlog and no KPI. The first number I reached for would have made every good idea look like a setback.
- 10+Features delivered into the platform
- 0KPIs or frameworks handed down
- 4Kinds of work in one backlog
- 2Numbers that replaced the ratio
Context
For the first stretch of my time at MedStar I built every workflow in the desktop toolkit myself — one person, one codebase, no handoffs. Mid-2026 that changed. The company was standing up a platform of its own, built by a development team, and the owner asked me to direct those developers as the toolkit’s capabilities moved across into it, and to own the feature backlog: what went in, in what order, and what stayed out.
No title came with it and none was conferred. What came with it was a sentence: give the developers any ideas you come up with to upgrade the platform. No framework, no KPI, no formal backlog — and a weekly-to-biweekly sync in a group chat, where we talked and then I posted a summary.
That is a generous brief and a difficult one to work under for long, because an open brief with no scoreboard quietly becomes unaccountable. Nobody was going to tell me I was behind. I also couldn’t have told them whether I was.
The constraint
Some weeks in, I went looking for the state of the work and found there wasn’t one. The whole effort existed as a trail of chat summaries — each accurate on its own, useless in aggregate. Nobody could say what had been delivered, what was still open, or what was moving. Not the developers, not the owner, and not me.
So the first job wasn’t a feature. It was reconstructing the backlog out of the conversation record: reading back through every summary, pulling out each functionality that had been discussed, and sorting it into delivered or open. That reconstruction was the first time anyone had seen the whole picture at once — which tells you how much of this work had been real but invisible.
What I did
- Rebuilt the backlog from the conversation record. Every functionality that had been raised, lifted out of the chat history onto one list, each marked delivered or open. A backlog is a conversation record before it is a list, and nobody had turned the one into the other yet.
- Rejected the ratio I first reached for. My instinct was the obvious metric: delivered divided by total, as a percentage. I threw it out in the same sitting I designed it, because of what it does. Every good new idea I added to the backlog made the number fall — and new ideas were precisely what the owner had asked me for. A metric that punishes the behaviour you were asked for will either stop the behaviour or stop being reported.
- Replaced it with two numbers that can’t fight each other. Delivered, cumulative and monotonic — it only ever goes up, and only when something is actually live. Open backlog, itemized rather than merely counted, so the list reads as work rather than as a quantity. Adding an idea grows the second number and leaves the first alone; shipping moves one to the other. Nothing I was asked to do makes either number look like failure.
- Tagged the delivered work by its nature, not just its count. Reading back over what had shipped, the backlog wasn’t pure migration. Four kinds of work were in it: capabilities migrated from the toolkit as they were, capabilities migrated and improved on the way across, fixes to things that were wrong, and friction removers — small changes that make a daily task less annoying and never appear on a roadmap. One number hides that mix. The tags make the shape of the contribution legible.
- Scoped what went in by what had stopped moving. The pipelines I was still calibrating stayed with me: a bridge between my own ongoing patches and a development team uploading them into a platform wasn’t practical, and anything depending on those stayed out with them. What had settled went across. The rule underneath it is the one I’d reuse anywhere — what is still moving stays with the person moving it; what has settled goes to the platform.
- Handed over in layers, not in a spec. A workshop walking the developers through the entire toolkit, backed by videos I recorded myself, the codebase and its user manual, and infographics explaining how the thing works. A single specification document would have been faster to write and slower to use.
- Filtered my own ideas before they reached the developers. Everything I proposed was still subject to their judgement, but I screened out what I already knew was infeasible or outside what the platform was for. That filter is borrowed from the other side of the table: at TOTVS I was the consultant receiving requests from users who couldn’t see the project’s scope, or who were tailoring a system to taste. Same skill, reversed.
- Ran the sync to decide, not to report. For one meeting with the developers’ coordinator I posted the backlog in delivered/open form beforehand and used the time itself to choose between two proposed integrations — one accounting, one mail — instead of reading out status. It was quick and direct, and going through the list with him point by point surfaced an item that was built but not live — a state my two columns had nowhere to put, and one that only became visible because there was a list to ask against.
Pick a metric that rewards the behaviour you were asked for. A completion percentage is the default choice, and it is the wrong one for any brief that asks you to keep generating work — it turns your best contribution into a drop in the number. Before you publish a measure, ask what it would punish, and whether that thing is the job.
Results
- 10+ features delivered into the platform, counted only once they were live. Built-but-not-live doesn’t count, which is the whole point of a number that can only go up.
- The work became visible. One list, in one place, that the developers, the owner and I could all read the same way — where there had been a trail of chat summaries and no answer to “where are we?”
- The measure survived the backlog growing. A new item surfaced in a meeting and went straight onto the open list. The delivered figure didn’t move backwards, and discovering more work didn’t read as losing ground. The ratio would have dropped.
- The meeting changed shape. Prioritizing two integrations against each other was worth more than a status read-out, and it only became possible once the status was something I could post in advance.
- No efficiency claim attached to any of it. There is no before-and-after measurement for the platform work, and I’m not going to manufacture one. The time-reduction figure elsewhere on this site belongs to the toolkit I built myself, not to the platform the developers are building.
What I’d do differently
- I’d have started the backlog on day one, not months in. Rebuilding it from the summaries worked, and I’ll take some credit for having kept records disciplined enough to make that possible at all. But it was still archaeology: what I recovered was whatever the summaries happened to preserve, and until I did it nobody could say what was done and what was pending. The same list, started in the first week, costs almost nothing.
- I’d have defined the status gate before I counted anything. I had delivered and open, and nothing in between. It took mapping the backlog and then asking the coordinator through it point by point for one item to turn out to be built but not live — a state my two columns had no room for. TOTVS had this solved years ago: backlog, in development, in test, deployed. I’d set those states up front, with the rule for crossing each line agreed before the first count rather than discovered by hitting an item that didn’t fit one.
- I’d have run it on a real cadence, and written the scheme down. I worked task to task and changed the meeting method once it was obvious I had to — with no record of when I changed it or why. I’d fix a cadence with a stated objective per cycle, the way a sprint does, and write down the scheme itself: how the backlog gets discussed with the developers, and how the summary reaches the owner and the other stakeholders. The way the work was managed deserved the same deliberateness I gave the measure.
Looking back
I expected the hard part here to be technical — explaining a codebase I knew intimately to people who had never seen it. It wasn’t. The hard part was that nobody had given me a way to tell whether the work was going well, so I had to build one, and my first attempt at it would have quietly argued against the thing I was there to do. A measure is a design decision, with the same failure modes as any other. It deserves the same care.