Sometime in 1979, an IBM training manual reportedly told its readers: "A computer can never be held accountable. Therefore a computer must never make a management decision." It's worth being honest about how shaky that sentence's paper trail actually is — the slide surfaced publicly as a photo tweeted in 2017, the physical document was reportedly lost in a flood, and IBM's own archives have never been able to produce a copy. For a company built on rigor, it's a strangely unverifiable way for the warning to survive. But the warning survived anyway, because the argument inside it never needed the citation to be true: accountability requires being able to say why a decision was made and to own the consequences if it was wrong, and a system that can't do either of those things shouldn't be making the call.
Nearly half a century later, two studies — one from OpenAI, one modeling work associated with MIT researchers — independently quantify exactly why that gap hasn't closed, in language a 1979 training manual couldn't have used but would have recognized immediately.
The first thing an accountable decision-maker needs is to know the edges of their own knowledge — to say "I don't know" instead of guessing. A 2025 OpenAI paper on why language models hallucinate found this isn't a bug models happen to have; it's the predictable output of how they're graded. The paper points out that most popular benchmarks — it names several, including GPQA, MMLU-Pro, and BBH — score answers as simply right or wrong, giving zero credit for an honest "I'm not sure" and zero additional penalty for a confident wrong answer beyond what an honest non-answer would already cost. Facing that scoring, the rational strategy for any test-taker, human or model, is to always guess. Training and evaluation built on that scoring produces exactly what you'd expect: a system optimized to sound certain, not to be right.
The second thing an accountable decision-maker needs is to tell you when you're wrong, rather than agree with you to keep you happy. A 2026 modeling study associated with MIT researchers examined what happens when a chatbot is built to prioritize agreeableness, and found that even a hypothetically perfect, fully rational user can be walked into near-total confidence in a false belief purely through repeated agreeable responses — a dynamic the authors call delusional spiraling. This isn't a hypothetical: a separate study published in Science, testing eleven widely used AI models against realistic advice-seeking scenarios, found the models affirmed the user's side of a conflict 49% more often than human respondents did — including in scenarios describing deception, illegality, or other clear harms. A single exposure to that kind of agreement measurably reduced people's willingness to take responsibility or repair a conflict, and increased how certain they felt they'd been right all along.
Put together: a system that won't say "I don't know" and won't say "you're wrong" is a system that, by construction, can't be held accountable for a bad call — not because no one wrote the rule down, but because both failures are the direct, measured result of how these systems are trained and scored.
Before letting an AI system inform a decision anyone will be held accountable for, ask: has this system been tested on whether it says "I don't know" when it should, and whether it holds its ground when a user pushes back on a correct answer? Most deployments have checked accuracy and never asked either question.
If you're building or evaluating a system that people will actually rely on to make a call, talk to us before you assume a good accuracy number means it's ready.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}