Five Stars?
What Uber can teach us about measurement

I was off on my travels again this week, and took a couple of Ubers along the way—and I was reminded of what Uber can teach us about measurement.
As you probably know, when you complete your Uber ride, the app asks you to assign a star rating to the driver on a one-to-five scale—the intent of which, presumably, is to measure the quality of the driver, or the experience overall. But there’s a tweak. If you click on five stars, all is fine and you are free to go about the rest of your day. But if you click on four stars or fewer, more screens of questions come up. What was it about the experience that was less than stellar?, the app wants to know, and it gives you some options. If the thing that’s leading to your score doesn’t fit in one of its pre-determined categories, it wants you to type some text in. And then it wants a bit more information beyond that. Until you complete the questions, it won’t let you complete the process of rating your ride—presumably to ensure that the app captures some feedback for every sub-five rating.
Someone at Uber’s corporate HQ is presumably smiling as they imagine the wonderful data this is generating for the company.
But actually, I suspect it’s generating a lot of fake fives.
It’s a pain to give less than five stars, so you give five stars every time, and Uber gets bad data.
(Of course, I don’t know this for sure, because I don’t know that anyone who’s not me reacts to the difficulty of assigning a low rating in the way that I do. But I do know that the incentive to tweak the rating is there, and that in itself is enough to taint the scores that result.)
There is a parallel problem in performance management, which is imagined, in many organizations, as a way to have people give star ratings to other people (and yes this is generally a bad idea and doesn’t help with very much of anything, but that’s a topic for another day). Leaders and HR folks worry about how to get “good” data from ratings, particularly as the tendency of managers to want to give higher ratings collides with the conviction of leaders that people can’t all be that good, and that even if they are that telling them might provoke complacency. So they impose a forced distribution of ratings—only a certain number of the high ones can be given out.
Uber makes it hard to give a low rating; organizations make it hard to give a high one; and neither of them gets good data as a result.
But in both cases, the good data is right there, hiding in plain sight.
It can be found by looking for a decision that is made, instead of an opinion that is sought, and a decision that is made, furthermore, under a real constraint, instead of under an artificial one.
In the Uber app, before you even reach the star-rating bit, you are asked if you wish to tip the driver, and you’re presented with a few suggested tip amounts to choose between. It’s just as easy to choose one as it is to choose another; and the constraint is not the artificial difficulty of choosing a lower option, but rather the real value you place on your own money.
The data is the tip.
And it would be quite simple for Uber to figure out, for each rider, what their typical tipping pattern is—whether they’re typically generous or not—and then to treat a divergence from this pattern as a useful piece of data. And then they could go further, and determine if there are certain drivers who generally cause riders to be lower or higher than the rider’s usual tipping pattern—and there is the rating.
(And again, I don’t know for sure that they don’t do this, and if anyone knows it would be great to hear from them, but I suspect they don’t because of all the rigmarole about star ratings.)
At the end of the performance year at work, in many companies, managers assign a bonus to each person on their team, and they do this with a real constraint—there is only so much money in the bonus pool—instead of an artificial one of a required ratings distribution. And again, similar possibilities present themselves: we can consider the relative size of any individual’s bonus, when compared to the pool available, as performance data on that person. We can go further, and look for patterns over time in how a given team leader chooses to distribute the funds, and then treat variance from that pattern as additional data, and we can look for patterns over time in how much money a given team member is given, and then treat variance from that pattern as data, too.
If we want data more than once a year, we can go further still. At Deloitte, we asked a simple question of each team leader about each of their team members once a quarter or when a project ended: If it were my money, I would give this person the highest possible compensation and bonus. Now, this is not quite as informative as the act of giving money itself, intent and action being different more often than we want to acknowledge, but it’s the next best thing—and note the constraint: if it were my money.
This question was a key part of a new approach to performance measurement across the entire organization, and it allowed us to dispense with rating scales, consensus meetings, and, for that matter, frustrated managers who were prevented by the system and its encoded policies from doing what they felt was the right thing for their people.
A good way to measure, in our world, is to measure action with consequence, preferably when that consequence represents a real trade-off of some sort. Distributing money, via tips or bonuses, fits that formula. Handing out ratings doesn’t—until, that is, you write some rules to create an artificial constraint. And then your users chafe against the constraints.
I have always found it telling that employees intuitively understand that some people will get more money than others; and at the same time resent that some people will get higher ratings than others. They understand, in other words, that ratings and their required distributions are make-believe, and that money is real.
So if you find yourself trying to improve measurement by manipulating the act of measurement itself, pause and think again. Improve it instead by finding a real thing in the real world that, unforced, captures something of what you are trying to understand.
In addition to writing about work, I advise businesses around the world on leadership, performance, and people. If you’d like to explore how I can help your organization, please check out my website here.

