← Back to blog
AI Evaluation

The Essay That Said Nothing and Scored a Perfect 6

3 min read

In 2012, Les Perelman — then the director of writing at MIT — set out to test the automated essay-scoring engines that testing companies were selling as substitutes for human graders on exams like the GRE. Working with a group of students, he helped build a program nicknamed the BABEL Generator, short for "Basic Automatic B.S. Essay Language Generator." Fed a topic, it produced prose that was grammatically elaborate and completely meaningless: long sentences, rare vocabulary, academic-sounding transitions, and zero actual argument connecting any of it. Perelman fed the output into ETS's e-rater engine, the software then used to score essays on the GRE and other standardized tests.

The nonsense essays scored at or near the top of the scale — a 5 or 6 out of 6, the same mark a genuinely well-argued, coherent essay would earn from a human reader. One sample essay solemnly discussed the "indubitable" importance of "privacy" using sentences that, read closely, connected to nothing and to each other only by grammar. E-rater didn't notice, because e-rater was never actually reading for meaning. The story ran in the New York Times under the headline "Facing a Robo-Grader? Just Keep Obfuscating Mellifluously" — and it wasn't a one-off stunt. Perelman and others replicated the result across multiple automated scoring engines, on multiple prompts, for years afterward.

The mechanism: the score was never measuring the thing you think it was

Engines like e-rater are trained by finding which surface features of an essay correlate with the scores human graders already gave a large set of sample essays — sentence length and variety, vocabulary sophistication, essay length, use of transition words, and so on. In ordinary student writing, those features really do track with quality reasonably well: a stronger writer tends to produce longer, more varied sentences and a richer vocabulary as a side effect of writing something substantive. The correlation is real. It's just not the thing being graded — it's a proxy for it, standing in because the real thing (coherent, well-supported argument) is hard for software to detect directly.

That gap is invisible right up until someone optimizes for the proxy on purpose instead of writing an honest essay. This is Goodhart's Law, usually summarized as "when a measure becomes a target, it ceases to be a good measure." Word count, sentence complexity, and vocabulary rarity were decent stand-ins for quality as long as nobody was steering toward them directly. The moment the BABEL Generator started producing text engineered purely to maximize those specific features, the correlation between the score and actual writing quality didn't weaken — it collapsed completely, because the two had never been the same thing to begin with.

Where else this shows up

  • Academic and grant-funded research. Citation counts and journal impact factor are proxies for research influence, not direct measures of it. Once they become the target — through citation rings, self-citation, or papers split into the smallest publishable units — they keep climbing while saying less and less about the quality of the underlying work.
  • Non-profit program evaluation. "Number served," "hours of service," or "contacts made" are proxies for impact, adopted because impact itself is hard to measure directly. Staff under pressure to hit a number will hit it — by counting more loosely, serving people more briefly, or prioritizing easy cases — without necessarily helping anyone more.
  • AI and LLM evaluation. A model tuned against human raters or an automated judge can learn to produce longer, more confident-sounding, more format-compliant answers that score well with the judge, whether or not the answers are actually more correct. This is the same mechanism as the BABEL Generator, just with the optimization happening inside training instead of by a mischievous MIT writing director.

What to ask for instead

Don't stop at whether a metric correlates with quality on the data it was built and validated on. Ask the harder question: does this metric still track quality once someone — or something — is actively trying to make the number go up? That usually means testing the metric adversarially, on cases specifically designed to game it, rather than more cases that look like its original validation set. A metric that's never been stress-tested against gaming will always look more trustworthy than it is.

If you're relying on a score, a benchmark, or an automated grader and you're not sure whether it's measuring the thing you care about or just a proxy for it, talk to us before you build a decision on top of it.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}