Be careful what you measure

(lambdaland.org)

28 points | by speckx 17 hours ago

6 comments

  • makeitdouble 0 minutes ago
    A less catchy but more pragmatic approach is to understand the reliability and/or value of what you measure and hold back from making them targets.

    If you can't trust your org to properly set its targets and need to blind yourself from the metrics you really want, you're in a pretty dire place.

  • gregw2 31 minutes ago
    The author proposes that the corollary to Goodheart's law is:

    "Only measure that which you are comfortable turning into a target."

    This is an interesting thought but I disagree and certainly logically wouldn't call it a corollary.

    As someone managing a system, one approach is to have a balanced set of measurements and to hold some of them back so they can't be gamed by those being measured. I.e. have more measurememts than targets so you can detect alignment drift.

  • simianwords 5 hours ago
    What’s the alternative to this? Without having a fallback to human driven judgement that itself can be gamed.

    Maybe you’d suggest that the VP must ask all their direct reports to verify it manually. Then you rely on each person below you to have good judgement and also act in good faith. The director asks the managers who asks the leads who may or may not give accurate reports.

    It’s not clear that’s any better?

    • aunderscored 4 hours ago
      This isn't "don't measure things" it's, "understand that when you measure something, it often turns into a goal. This can have unexpected consequences".

      Even with human judgement this happens. Yhrtr are jokes about payment by line of code that have existed for decades.

      The main thing is we need to measure secondary things. User satisfaction, defect rate, bugfix rate, etc (and this isn't to say that those are good measures that work everywhere. They may or may not.)

      At the end of the day, the challenge is to think, and not assume they a number means what you think it means, or more or fewer of something will always be good.

      • simianwords 4 hours ago
        No yes, I agree with you but frankly I'm tired of the cliche that you can't measure PRs or LOC when in reality it is very much correlated with whatever you want to optimise. I do think it can be gamed, but relying on vibes and human judgement can also be gamed (which is what I tried to point out).

        You are completely right that the VP can measure user satisfaction but here's the thing: that feedback loop has a much longer time period. You could also measure your company by revenue or its stock valuation. But the point is to have metrics that have a shorter latency. How can you achieve it? Its a hard problem to solve.

        • aunderscored 4 hours ago
          Yes, whatever you want to optimise is a good way to phrase this. However, it'd not just what you want to optimise, you will get side on effects. Even when things seem reasonable, longer term things come out. I don't have a good solution for this. But training people to target one (or even a few) things is incredibly difficult.

          I also question the trust aspect. Part of this (which can also be gamed of course) is the level of trust you have that a team is doing their best to work towards some goal. More metrics indicate lower trust in some ways.

          I think specifically this (metrics vs judgement), gaming happens in both, but metrics lead to much more wild problems if gamed, because they have a much harder "you said do this more so I did" backstop than "well this seemed to be what you wanted".

          Thinking about it a bit more, if you make the argument that human judgement here is itself a compound metric of many different inputs, this semi resolves itself. But we still have the problem that the metric has no backing that can be handed around other than "seems good to me", which changes for many people.

          • simianwords 4 hours ago
            Yeah the fork in our disagreement is on trusting human judgement vs metrics.

            My position is that if you don’t trust a human to use metrics in good spirit, then you also may not trust them without metrics. But agree to disagree I guess. I do buy your point that metrics give the excuse of “hey that’s what you asked!”

  • sans_souse 16 hours ago
    Excellent post.
  • atoav 4 hours ago
    The main problem is that there exists a type of management, that does not want to understand the meat of the process they are managing, but just wants to compare numbers.

    Certain aspects of human culture, certain aspects of engineering quality are hard to put into numbers, as it turns out.

    When we talk about (for example) writing a book, a simple productivity metric would be pages-written-per-day. Not that this isn't a useless metric, but the actual metric you want would be more something like good-pages-written-per-day, because throwing some garbage text out quickly is easy, the hard bit is doing good writing.

    But then the obvious question is, what makes a good page? How do you judge the quality of a process' output aside from simple to create metrics? And many managers don't have the understanding to make that judgement. Quite frankly, if you're a manager and you cannot reliably judge the quality of the product you're working on, everything becomes hit and miss like throwing pudding onto the wall and seeing what sticks. In that case relying on a simple metric is worse than e.g. trusting the judgement of experienced engineers or users of your product.

  • illusive4080 14 hours ago
    Title should be changed to the post title. “Lambda Land” means nothing to me.
    • compil3d 11 hours ago
      I was hoping to read about serverless functions :/
    • dang 6 hours ago
      Fixed now.