Image: Sasoriza · CC BY-SA 2.0 · Wikimedia Commons
Technology·🏆 Entry of the Week

We Shipped a Feature Nobody Used for Eleven Months Because the Metric Loved It

AR
May 9, 2022 · 4 min read · Edited

The feature was a smart summary that sat at the top of the feed. It went out in March. By May, the dashboard showed engagement with it climbing steadily, week over week, in a curve that made everyone in the room happy. We put the chart in the board deck. We hired against it.

In February the following year an intern doing a routine audit found that roughly ninety-four percent of the recorded interactions were the summary being scrolled past. The event fired on the element entering the viewport. Every single user of the app entered that viewport every single session, because the summary was at the top.

I want to be careful how I tell this, because the easy version is a story about a bug, and it was not a bug. The event was implemented exactly as specified. The specification said "summary viewed". The summary was, in the ordinary meaning of the word, viewed.

The failure was that nobody in the chain had to state what the feature was supposed to make happen, and so the measurement became the statement, and the measurement was true.

Here is what makes this worth writing about rather than just embarrassing. In eleven months, at least forty people looked at that chart. Some of them were extremely good. The head of design asked twice whether the numbers felt high, and was told the summary was at the top of the feed, which is a complete and satisfying explanation of why the numbers would be high, and which is also the reason the numbers were meaningless. The explanation and the flaw were the same sentence, and it passed as reassurance.

I have been in the Bengaluru product world for nine years and I have watched some version of this happen at four companies, twice on my own teams, and the pattern is always the same three parts.

First, a metric is chosen because it is available rather than because it is right. Nobody decides "impressions are what matter"; somebody decides "impressions are what we can instrument this week", and the qualifier falls off within a month.

Second, the metric goes up, because metrics chosen for availability are usually chosen for something correlated with usage in general, and usage in general goes up when you put a thing at the top of a screen.

Third — and this is the part I have never seen written down — the metric acquires defenders. Once a number is in a board deck it is attached to someone's credibility. Auditing it is no longer a technical act. The intern who found ours was not being brave; she did not know yet that the number had an owner, which is exactly why she found it.

The fix people reach for is better metrics, and better metrics do help at the margin. Instrument dwell time, instrument the action the summary was meant to cause, instrument the counterfactual with a holdout group. All correct, all worth doing, and all of it still measurement, which means all of it still open to the same failure one layer up.

What I actually believe now, having tried the sophisticated version and watched it fail differently, is smaller and less satisfying: write down, before you build, the specific user behaviour that would change if the feature worked, and the number that would move if the feature did nothing. The second half is the part everyone skips. If you cannot describe what the chart looks like when the feature is useless, you cannot read the chart.

For our summary, the answer would have been: if it does nothing, impressions will be roughly equal to daily sessions, because it is at the top of the feed. Anyone could have written that sentence in March. It takes about forty seconds. Nobody wrote it because in March everyone was busy shipping, and because it is a sentence that pre-commits you to a way of being wrong, and the honest reason we skip it is that nobody wants to write down in advance what their own failure will look like.

We killed the feature in March of the following year. Sessions did not move. The person who built it took it better than I did, and said something I have repeated since: that the worst part was not the eleven months, it was that the chart had been going up the entire time and it had felt, genuinely, like progress. It looked exactly like the real thing. That is the whole problem — a good chart and a useless feature produce the same picture, and the picture is what gets shown.

Since publishing this I have been asked several times for the exact wording, so: if this feature does nothing, the number will be roughly equal to daily sessions, because the element is at the top of the feed. Write that sentence, about your own feature, before you build it. It takes forty seconds and nobody wants to.

◉ 78 views

Comments

3
AN
andrewcollins· May 11

"Write down what the chart looks like when the feature is useless." I have now made three teams do this and the interesting part is how much resistance it gets, which I think is the essay's real finding.

RJ
rjh_writes· May 13

This is the same structure as a study without a pre-registered analysis plan. You can always find a chart that went up.

AR
arjunmehtaAUTHOR· May 16

It is, and the analogy is closer than I would like, because we also had a version of a file drawer — the two experiments that did not move anything were never written up.