When Safety Becomes Stupidity
What Bonhoeffer would say about Anthropic's moral framework
Anthropic just released version 3.0 of their Responsible Scaling Policy. I read the whole thing so you don’t have to, and then I reached for Dietrich Bonhoeffer.
Not because Anthropic is evil. The opposite. The RSP is thoughtful, self-aware, and honest about its own limitations. It’s the work of serious people trying to solve a real problem.
That’s exactly what makes it worth worrying about.
Bonhoeffer’s theory of stupidity — the real one, not the social media version — isn’t about intelligence. It’s about what happens when smart people surrender their independent moral judgment to a system that provides certainty. The system doesn’t have to be wrong. It just has to be unchallengeable.
Anthropic’s framework does something subtle: it dresses moral choices in the language of technical findings. Safety levels, capability thresholds, empirical evaluations. But underneath every threshold is a question philosophy has never settled — who should have access to this knowledge? — answered as if it were settled.
Their explicit goal is to make their values the industry standard, then law. They’re already partway there.
I wrote a longer piece walking through the structural parallel, the practical risks, and what a better approach might look like. If this interests you, the full article is here:
The Sincere Danger: Bonhoeffer’s Stupidity and AI’s Moral Objectivism
The short version: the deepest risk isn’t that Anthropic is wrong about safety. It’s that they might be right about safety and wrong about morality — and their framework can’t tell the difference.
Matt Taylor builds knowledge systems and thinks about how tools shape thought.


