How businesses fool themselves with their own data

Christopher Ross

10 min read

AI and learning, kept human · Niagara, Ontario

Title card: how businesses fool themselves with their own data. An article on survivorship bias, vanity metrics and proxy drift by Christopher Ross.

Part 3 of Trust Your Own Data, a four-part series on whether you can believe the numbers your business collects about itself.

I have spent a fair chunk of my working life reading reports that a business wrote about itself, and the ones that worry me most are never the reports full of bad news. Bad news at least gets someone in a room. The cheerful ones are the ones that keep me up. A satisfaction score parked at ninety-something. A wall of five-star reviews by the front desk. A dashboard where every number is green and pointing up, in a quarter where the same business quietly lost customers it cannot explain.

That gap is the whole subject of this series. In Part 1 I called it the steak-dinner problem, the way people hand you a kind answer instead of a true one when you ask at the wrong moment. In Part 2 I made the case that real consent, freely given, is the safeguard most feedback skips right past. This piece is about what happens when those small errors stop being one-off moments and get wired into the way a business measures itself, on a schedule, year after year.

Here is the uncomfortable part. None of what follows feels like a mistake while it is happening. Each one hands you a real number in a real spreadsheet, and the number flatters the person holding it. Picture the mirrors in a department store changing room, the ones tilted back a few degrees so you stand a little taller and a little slimmer. Nothing in the glass is fake. It is honestly you. The angle was just chosen to help you reach for your wallet. A business can build that same flattering angle into its own research and never catch it, because the reflection keeps coming back looking great.

Survivorship bias: the customers who already left

The first trap hides in plain sight, and it hides by absence. When you survey your customers about why they love you, everyone who answers is, by definition, still a customer. The people who cancelled in January and never looked back are not on the list. They cannot be. They left.

Think of a gym running its annual member survey. It goes out to current members, and the results are lovely. People praise the classes, the hours, the clean showers. What the survey physically cannot tell you is why the roughly forty percent who joined last winter stopped showing up by spring, because those people were gone before the survey was written. So the gym learns, in great detail, why the stayers stayed. The reason people leave, which is the number that actually decides whether the gym makes rent, never enters the data at all. You end up studying the survivors and calling it the whole population. If you only ever ask the people who are still in the room, the room will always tell you it is a wonderful room.

Selection bias: asking only the winners

Close cousin, slightly different mechanism. Here the data is not missing by accident. Someone chose the sample based on the result they already got.

The everyday version is the review request. A support team sends the “how did we do” email after a five-star chat and somehow never quite gets around to sending it after a refund or a two-hour hold. The reviews that come back are glowing, and they are real, and they describe an experience a shrinking slice of customers actually had. The consulting version is the case-study shelf. I write up the project that went beautifully and stay quiet about the one that stalled, so my portfolio becomes a record of my wins with the losses edited out, which is exactly the material I would learn the most from. When you only measure the outcomes you are proud of, your data turns into a highlight reel, and nobody plans a business off a highlight reel.

Leading questions that already know the answer

Sometimes the sample is fine and the question does the fooling. A leading question smuggles the preferred answer into the wording, so the customer is really just agreeing with you.

“How much did you enjoy our fast, friendly service today?” has already decided that the service was fast and friendly. All the customer gets to pick is a volume knob. A loaded question does worse, it hides an assumption the person may not even hold, the classic being anything shaped like “do you agree that our new pricing is fair?” People agree with statements far more readily than they disagree, a real and well-documented tendency called acquiescence bias, so a survey built entirely from “do you agree” prompts will tilt positive no matter what you sell. The tell is simple. Read your own question out loud and ask whether a genuinely unhappy customer has an easy, natural place to say so. If the honest negative answer feels awkward to give, you did not write a question. You wrote a compliment with a checkbox.

Vanity metrics: counting what is easy to count

This one is less about a bad answer and more about measuring the wrong thing on purpose, because the wrong thing is easier and it looks better on a slide.

I ran a webinar once with a registration list I was proud of. Hundreds of names. Then a couple of hundred people actually showed up, which still felt like a good day, and I put the attendance number in a report and moved on. The number I did not chase was how many of those attendees did anything differently afterward, or booked a call, or became a customer, because that number was harder to get and I already suspected it was small. Impressions, sign-ups, followers, attendance, these are vanity metrics, and they are seductive precisely because they are easy to collect and they only ever go up and to the right. The hard numbers sit underneath them. Did behaviour change. Did people come back. Did revenue move. A metric that only ever climbs is applause dressed up as measurement.

Proxy drift: when the stand-in forgets it was a stand-in

The last one is the quietest, and in my experience the most expensive. You pick a stand-in for the thing you actually care about, because the real thing is hard to measure directly, and then over months everyone forgets it was ever a stand-in.

A training course is the cleanest example, and it happens to be my trade. What I care about is whether someone can do the work after the course. That is hard to measure, so I lean on a proxy like the completion rate, the percentage of people who click through to the end. Reasonable enough, at first. But completion is easy to move. Shorten the modules, drop the quiz that made people stop, add a big “next” button, and completion climbs beautifully while the thing I actually cared about, competence on the shop floor, does not budge. The economist Charles Goodhart gave us the warning for this decades ago, usually summed up as: when a measure becomes a target, it stops being a good measure. The proxy drifts loose from the real goal and starts leading a life of its own, and because the proxy is the number on the dashboard, the whole team quietly optimises for the stand-in instead of the thing.

Line these up and the same shape runs through every one. Survivorship, sampling on the outcome, leading questions, vanity metrics, a proxy that wandered off. Each produces an honest number. Each flatters the person who commissioned it. And each one is Part 1’s steak-dinner problem again, just promoted from a single awkward moment at a dinner table into standing business process, running on a schedule, quarter after quarter. The one-question test still works on every one of them. Was this collected in a moment designed to get a real answer, or a kind one?

The good news, and this is where the series is heading, is that a business that can name these traps can also design its way out of them. Part 4 is about building data you can actually stand on, the measurement that survives contact with a bad quarter. For now, the useful move is smaller and honestly a bit humbling. Take your favourite number, the one on the slide you are proudest of, and go looking for the flattering angle you built into it. It is almost always there.

If you would like a second set of eyes on where your own numbers are quietly telling you what you hoped to hear, that is more or less what a measurement audit is, and it is a good chunk of what I do.

Common questions

What is survivorship bias in customer research?

It is the error of studying only the people who stayed. Survey your customers about why they love you and everyone who answers is, by definition, still a customer. The ones who cancelled in January are not on the list and cannot be, which is awkward, because they are holding the answer you actually needed. The same trap has a long history in survivorship bias outside business too.

What is the difference between a vanity metric and a useful one?

A vanity metric moves reliably and changes nothing. The test is to ask what you would do differently on Monday if the number doubled overnight. If the honest answer is nothing, it is decoration, and it should not be sitting at the top of the report. Move it to an appendix if somebody is attached to it. The top of a report is expensive space, and every number sitting there that nobody acts on is teaching the room to skim the ones that matter.

How do I check my own research for bias?

One question catches most of it: who is missing from this data, and would they have answered differently? Then ask who benefits from the number looking good. If that person is also the one who designed the study, get somebody else to read it. An outside reader does not need to know your industry to be useful here. They only need to be free to say that the question looked leading, or that a whole group of people appears to be missing, without it costing them anything at work to say so.

Trust Your Own Data: the full series

A four-part series on whether you can believe the numbers your business collects about itself. Each piece stands on its own; together they build one test.

  1. The steak-dinner problem. Why your warmest feedback may be your least honest.
  2. Consent you can trust. The yes most business feedback quietly skips.
  3. How businesses fool themselves. The everyday ways your own data flatters you. You are reading this one.
  4. Data worth trusting. The toolkit for collecting honest answers.

Working through something on your own site? Get in touch →

Leave a reply

Your email address will not be published. Required fields are marked *

Your rating (optional)

Your name and email are stored with your comment; only your display name is shown publicly. See our privacy policy.