Part of our Statistical paradoxes series · Goodhart's law
For most of the 2000s, Atlanta Public Schools was the urban district other urban districts were told to study. Scores on Georgia's state test, the CRCT, climbed year after year, including in schools serving some of the poorest neighborhoods in the city. In 2009 the American Association of School Administrators named Atlanta's superintendent, Beverly Hall, the national Superintendent of the Year. The district had set annual score targets for every school, and school after school was hitting them.
Then people started looking at the answer sheets instead of the scores. Reporters at the Atlanta Journal-Constitution had already flagged gains so large that the odds of them happening by chance were, for some schools, worse than a billion to one. The state then ran an erasure analysis on every spring 2009 CRCT answer sheet in Georgia, counting how often a bubble had been erased and changed from a wrong answer to a right one. Dozens of Atlanta schools were flagged. A state investigation followed, with more than 2,100 interviews, and its 2011 report found cheating in 44 of the 56 schools it examined. It named 178 educators, 38 of them principals; 82 confessed. Investigators described a district run on a "culture of fear, intimidation and retaliation," where test results and public praise mattered more than integrity. In 2015, eleven educators were convicted, several under Georgia's racketeering law.
It's tempting to file this under "some people cheated." But the more useful reading is structural. The CRCT was designed to tell the public how much Atlanta's children knew. Then it became the thing that decided who got praised, who got a bonus, and who got pushed out. Once a number carries that weight, people have two ways to move it: improve what it measures, or move the number directly. The second is almost always faster.
This is Goodhart's law, named for the economist Charles Goodhart and usually paraphrased as "when a measure becomes a target, it ceases to be a good measure." The social scientist Donald Campbell put the same idea in terms evaluators should have taped to their monitors: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures, and the more apt it will be to distort and corrupt the social processes it was meant to monitor. Note the two halves. The indicator gets corrupted, and so does the work. Atlanta got both: scores that no longer meant anything, and classrooms where time and trust went into hitting the target instead of teaching.
Two details make this a statistics story and not just an ethics one. First, nothing in the score data looked wrong on its own. Rising averages are exactly what a successful reform looks like. The tell was in a second, unglamorous measurement nobody was being rewarded on: the eraser marks. Second, the gains were too consistent. Real improvement is noisy. Schools have off years, cohorts differ, and small schools bounce around. A district where nearly everyone hits a stretch target, every year, is a pattern that should invite a closer look rather than a press release.
Most Goodhart failures involve nobody breaking the law. A hospital meets a wait-time target by redefining when the clock starts. A job program reports placement rates and quietly enrolls the people most likely to be placed anyway. An AI team tunes its model on the benchmark until the benchmark stops predicting anything about real users. Atlanta is just the version where the gaming left physical evidence.
Whenever a number is about to start deciding things (funding, bonuses, renewals, launches), ask one question: what second measurement would catch this number being gamed, and who collects it? It should be something the people being judged can't easily move, gathered independently, and checked even when the headline number looks great. Especially then. Atlanta had that measurement in the answer sheets the whole time. It just wasn't the number anyone was looking at.
If you're designing an evaluation, a performance dashboard, or an AI eval suite and want the metrics to keep meaning something after people start chasing them, get in touch. We'll help you pick the target and the check that goes with it.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}