
Metrics Design Checklist
Metrics design checklist: Is it aligned with goals? Actionable and influenceable? Gameable by cheating? Leading or lagging indicator?
Contributions
Every accepted correction to this page is recorded with the exact change, so readers can see how the page improved over time.
-
Corrected a mislabeled condition in Steven Kerr's famous paper, removed an entirely unverified study, and corrected unverified statistics and a nonexistent publication across three real-world case studies.
What the page claimedArticle listed 'hysteresis' as one of Steven Kerr's four real conditions for metric dysfunction in his famous 1975 paper, when his real fourth condition is 'hypocrisy' - an unverified term substitution. It cited an entirely unverified 2010 Journal of Applied Psychology study by 'Nhung Nguyen' and 'Michael Levy' with unverified figures (1,247 sales reps, 23% performance difference) - no such paper, authors, or affiliation pairing could be located anywhere. It attributed an unverified '31 percent higher user adoption' statistic to a nonexistent '2017 re-Wired investigation' of Google's real OKR system (no publication called 're-Wired' exists). It included an unverified specific '3,200 employees surveyed' figure for Vanity Fair's real 2012 Microsoft investigative article (which was based on interviews, not a formal survey of that size) and included an unverified '27 percentage points' internal trust-score statistic with no real source. And it misdated the NHS's real four-hour A&E target (stated 2003; the real target was introduced in the mid-2000s with the 95% threshold formalized later) and included an unverified statistics ('22% to 6%,' a nonexistent 2015 BMJ Quality and Safety study with a '14% more likely to be readmitted' figure).
What was correctedKerr's fourth condition corrected to 'hypocrisy,' his real term. The unverified Nguyen/Levy study was removed and replaced with an accurate general statement about the real, well-documented attentional-cost principle. The Google OKR passage corrected to remove the unverified investigation and statistic while keeping the real, accurately-described calibration principle. The Microsoft passage corrected to accurately describe the real Vanity Fair investigative journalism without the unverified sample size, with the unverified trust-score statistic replaced by an accurate general statement. The NHS passage corrected to remove the wrong year, unverified percentages, and nonexistent study, while keeping the real underlying Goodhart's Law dynamic that has been documented in the real NHS target's history.
Why: Independent verification found Steven Kerr's paper (aside from the one mislabeled condition) and Bent Flyvbjerg's real megaprojects research (258 projects, 86% overrun rate, 28% average overrun) were both fully accurate as cited - continuing the pattern where some citations in a given article hold up completely while others in the same article are wholesale fabrications.
View the full record →