In 2012, Les Perelman — then the director of writing at MIT — set out to test the automated essay-scoring engines that testing companies were selling as substitutes for human graders on exams like the GRE. Working with a group of students, he helped build a program nicknamed the BABEL Generator, short for "Basic Automatic B.S. Essay Language Generator." Fed a topic, it produced prose that was grammatically elaborate and completely meaningless: long sentences, rare vocabulary, academic-sounding transitions, and zero actual argument connecting any of it. Perelman fed the output into ETS's e-rater engine, the software then used to score essays on the GRE and other standardized tests.
The nonsense essays scored at or near the top of the scale — a 5 or 6 out of 6, the same mark a genuinely well-argued, coherent essay would earn from a human reader. One sample essay solemnly discussed the "indubitable" importance of "privacy" using sentences that, read closely, connected to nothing and to each other only by grammar. E-rater didn't notice, because e-rater was never actually reading for meaning. The story ran in the New York Times under the headline "Facing a Robo-Grader? Just Keep Obfuscating Mellifluously" — and it wasn't a one-off stunt. Perelman and others replicated the result across multiple automated scoring engines, on multiple prompts, for years afterward.
Engines like e-rater are trained by finding which surface features of an essay correlate with the scores human graders already gave a large set of sample essays — sentence length and variety, vocabulary sophistication, essay length, use of transition words, and so on. In ordinary student writing, those features really do track with quality reasonably well: a stronger writer tends to produce longer, more varied sentences and a richer vocabulary as a side effect of writing something substantive. The correlation is real. It's just not the thing being graded — it's a proxy for it, standing in because the real thing (coherent, well-supported argument) is hard for software to detect directly.
That gap is invisible right up until someone optimizes for the proxy on purpose instead of writing an honest essay. This is Goodhart's Law, usually summarized as "when a measure becomes a target, it ceases to be a good measure." Word count, sentence complexity, and vocabulary rarity were decent stand-ins for quality as long as nobody was steering toward them directly. The moment the BABEL Generator started producing text engineered purely to maximize those specific features, the correlation between the score and actual writing quality didn't weaken — it collapsed completely, because the two had never been the same thing to begin with.
Don't stop at whether a metric correlates with quality on the data it was built and validated on. Ask the harder question: does this metric still track quality once someone — or something — is actively trying to make the number go up? That usually means testing the metric adversarially, on cases specifically designed to game it, rather than more cases that look like its original validation set. A metric that's never been stress-tested against gaming will always look more trustworthy than it is.
If you're relying on a score, a benchmark, or an automated grader and you're not sure whether it's measuring the thing you care about or just a proxy for it, talk to us before you build a decision on top of it.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}