Rendered at 17:33:51 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
makeitdouble 5 hours ago [-]
A less catchy but more pragmatic approach is to understand the reliability and/or value of what you measure and hold back from making them targets.
If you can't trust your org to properly set its targets and need to blind yourself from the metrics you really want, you're in a pretty dire place.
gregw2 6 hours ago [-]
The author proposes that the corollary to Goodheart's law is:
"Only measure that which you are comfortable turning into a target."
This is an interesting thought but I disagree and certainly logically wouldn't call it a corollary.
As someone managing a system, one approach is to have a balanced set of measurements and to hold some of them back so they can't be gamed by those being measured. I.e. have more measurememts than targets so you can detect alignment drift.
simianwords 11 hours ago [-]
What’s the alternative to this? Without having a fallback to human driven judgement that itself can be gamed.
Maybe you’d suggest that the VP must ask all their direct reports to verify it manually. Then you rely on each person below you to have good judgement and also act in good faith. The director asks the managers who asks the leads who may or may not give accurate reports.
It’s not clear that’s any better?
aunderscored 10 hours ago [-]
This isn't "don't measure things" it's, "understand that when you measure something, it often turns into a goal. This can have unexpected consequences".
Even with human judgement this happens. Yhrtr are jokes about payment by line of code that have existed for decades.
The main thing is we need to measure secondary things. User satisfaction, defect rate, bugfix rate, etc (and this isn't to say that those are good measures that work everywhere. They may or may not.)
At the end of the day, the challenge is to think, and not assume they a number means what you think it means, or more or fewer of something will always be good.
simianwords 10 hours ago [-]
No yes, I agree with you but frankly I'm tired of the cliche that you can't measure PRs or LOC when in reality it is very much correlated with whatever you want to optimise. I do think it can be gamed, but relying on vibes and human judgement can also be gamed (which is what I tried to point out).
You are completely right that the VP can measure user satisfaction but here's the thing: that feedback loop has a much longer time period. You could also measure your company by revenue or its stock valuation. But the point is to have metrics that have a shorter latency. How can you achieve it? Its a hard problem to solve.
aunderscored 10 hours ago [-]
Yes, whatever you want to optimise is a good way to phrase this. However, it'd not just what you want to optimise, you will get side on effects. Even when things seem reasonable, longer term things come out. I don't have a good solution for this. But training people to target one (or even a few) things is incredibly difficult.
I also question the trust aspect. Part of this (which can also be gamed of course) is the level of trust you have that a team is doing their best to work towards some goal. More metrics indicate lower trust in some ways.
I think specifically this (metrics vs judgement), gaming happens in both, but metrics lead to much more wild problems if gamed, because they have a much harder "you said do this more so I did" backstop than "well this seemed to be what you wanted".
Thinking about it a bit more, if you make the argument that human judgement here is itself a compound metric of many different inputs, this semi resolves itself. But we still have the problem that the metric has no backing that can be handed around other than "seems good to me", which changes for many people.
simianwords 9 hours ago [-]
Yeah the fork in our disagreement is on trusting human judgement vs metrics.
My position is that if you don’t trust a human to use metrics in good spirit, then you also may not trust them without metrics. But agree to disagree I guess. I do buy your point that metrics give the excuse of “hey that’s what you asked!”
aunderscored 5 hours ago [-]
Agreed, trusting a single human is fine, but there are weird extra tertiary effects as soon as you mention your looking at some metric. Especially if you're above someone in a hierarchy. They may subconsciously try to game it as well.
sans_souse 22 hours ago [-]
Excellent post.
atoav 10 hours ago [-]
The main problem is that there exists a type of management, that does not want to understand the meat of the process they are managing, but just wants to compare numbers.
Certain aspects of human culture, certain aspects of engineering quality are hard to put into numbers, as it turns out.
When we talk about (for example) writing a book, a simple productivity metric would be pages-written-per-day. Not that this isn't a useless metric, but the actual metric you want would be more something like good-pages-written-per-day, because throwing some garbage text out quickly is easy, the hard bit is doing good writing.
But then the obvious question is, what makes a good page? How do you judge the quality of a process' output aside from simple to create metrics? And many managers don't have the understanding to make that judgement. Quite frankly, if you're a manager and you cannot reliably judge the quality of the product you're working on, everything becomes hit and miss like throwing pudding onto the wall and seeing what sticks. In that case relying on a simple metric is worse than e.g. trusting the judgement of experienced engineers or users of your product.
Paul_Clayton 5 hours ago [-]
For a book, the local value of a page may not accurately represent the value added to the completed book. In fiction, a relatively boring page may be intentional as a kind of palate cleanser. In non-fiction, an early hard to understand section may make later sections easier to understand and avoid unlearning.
In programming, it may be challenging to distinguish between You Aren't Going To Need It (wasted work that adds complexity) and brillant foresight that avoids trying to retrofit an abstraction over independently developed code.
I am not sure that a manager needs to reliably judge the quality of the product to reliably deliver quality products. Judging skill and integrity (and promoting the development of such) does not seem to require having the same skill.
illusive4080 20 hours ago [-]
Title should be changed to the post title. “Lambda Land” means nothing to me.
compil3d 17 hours ago [-]
I was hoping to read about serverless functions :/
If you can't trust your org to properly set its targets and need to blind yourself from the metrics you really want, you're in a pretty dire place.
"Only measure that which you are comfortable turning into a target."
This is an interesting thought but I disagree and certainly logically wouldn't call it a corollary.
As someone managing a system, one approach is to have a balanced set of measurements and to hold some of them back so they can't be gamed by those being measured. I.e. have more measurememts than targets so you can detect alignment drift.
Maybe you’d suggest that the VP must ask all their direct reports to verify it manually. Then you rely on each person below you to have good judgement and also act in good faith. The director asks the managers who asks the leads who may or may not give accurate reports.
It’s not clear that’s any better?
Even with human judgement this happens. Yhrtr are jokes about payment by line of code that have existed for decades.
The main thing is we need to measure secondary things. User satisfaction, defect rate, bugfix rate, etc (and this isn't to say that those are good measures that work everywhere. They may or may not.)
At the end of the day, the challenge is to think, and not assume they a number means what you think it means, or more or fewer of something will always be good.
You are completely right that the VP can measure user satisfaction but here's the thing: that feedback loop has a much longer time period. You could also measure your company by revenue or its stock valuation. But the point is to have metrics that have a shorter latency. How can you achieve it? Its a hard problem to solve.
I also question the trust aspect. Part of this (which can also be gamed of course) is the level of trust you have that a team is doing their best to work towards some goal. More metrics indicate lower trust in some ways.
I think specifically this (metrics vs judgement), gaming happens in both, but metrics lead to much more wild problems if gamed, because they have a much harder "you said do this more so I did" backstop than "well this seemed to be what you wanted".
Thinking about it a bit more, if you make the argument that human judgement here is itself a compound metric of many different inputs, this semi resolves itself. But we still have the problem that the metric has no backing that can be handed around other than "seems good to me", which changes for many people.
My position is that if you don’t trust a human to use metrics in good spirit, then you also may not trust them without metrics. But agree to disagree I guess. I do buy your point that metrics give the excuse of “hey that’s what you asked!”
Certain aspects of human culture, certain aspects of engineering quality are hard to put into numbers, as it turns out.
When we talk about (for example) writing a book, a simple productivity metric would be pages-written-per-day. Not that this isn't a useless metric, but the actual metric you want would be more something like good-pages-written-per-day, because throwing some garbage text out quickly is easy, the hard bit is doing good writing.
But then the obvious question is, what makes a good page? How do you judge the quality of a process' output aside from simple to create metrics? And many managers don't have the understanding to make that judgement. Quite frankly, if you're a manager and you cannot reliably judge the quality of the product you're working on, everything becomes hit and miss like throwing pudding onto the wall and seeing what sticks. In that case relying on a simple metric is worse than e.g. trusting the judgement of experienced engineers or users of your product.
In programming, it may be challenging to distinguish between You Aren't Going To Need It (wasted work that adds complexity) and brillant foresight that avoids trying to retrofit an abstraction over independently developed code.
I am not sure that a manager needs to reliably judge the quality of the product to reliably deliver quality products. Judging skill and integrity (and promoting the development of such) does not seem to require having the same skill.