Metrics Design Checklist
Correction Performance Metrics

Metrics Design Checklist

Corrected by Emir Baycan · on When Notes Fly · 30 July 2026 · View published page ↗

Metrics design checklist: Is it aligned with goals? Actionable and influenceable? Gameable by cheating? Leading or lagging indicator?

Needs stronger evidenceFactually incorrect

What was corrected

What the page claimed

Article listed 'hysteresis' as one of Steven Kerr's four real conditions for metric dysfunction in his famous 1975 paper, when his real fourth condition is 'hypocrisy' - an unverified term substitution. It cited an entirely unverified 2010 Journal of Applied Psychology study by 'Nhung Nguyen' and 'Michael Levy' with unverified figures (1,247 sales reps, 23% performance difference) - no such paper, authors, or affiliation pairing could be located anywhere. It attributed an unverified '31 percent higher user adoption' statistic to a nonexistent '2017 re-Wired investigation' of Google's real OKR system (no publication called 're-Wired' exists). It included an unverified specific '3,200 employees surveyed' figure for Vanity Fair's real 2012 Microsoft investigative article (which was based on interviews, not a formal survey of that size) and included an unverified '27 percentage points' internal trust-score statistic with no real source. And it misdated the NHS's real four-hour A&E target (stated 2003; the real target was introduced in the mid-2000s with the 95% threshold formalized later) and included an unverified statistics ('22% to 6%,' a nonexistent 2015 BMJ Quality and Safety study with a '14% more likely to be readmitted' figure).

What was corrected

Kerr's fourth condition corrected to 'hypocrisy,' his real term. The unverified Nguyen/Levy study was removed and replaced with an accurate general statement about the real, well-documented attentional-cost principle. The Google OKR passage corrected to remove the unverified investigation and statistic while keeping the real, accurately-described calibration principle. The Microsoft passage corrected to accurately describe the real Vanity Fair investigative journalism without the unverified sample size, with the unverified trust-score statistic replaced by an accurate general statement. The NHS passage corrected to remove the wrong year, unverified percentages, and nonexistent study, while keeping the real underlying Goodhart's Law dynamic that has been documented in the real NHS target's history.

Why this is better

Independent verification found Steven Kerr's paper (aside from the one mislabeled condition) and Bent Flyvbjerg's real megaprojects research (258 projects, 86% overrun rate, 28% average overrun) were both fully accurate as cited - continuing the pattern where some citations in a given article hold up completely while others in the same article are wholesale fabrications.

All of Emir Baycan's contributions →