© ajf, September 18, 2026, Singapore
1. Losing Control
2. The Teenager
3. Understanding? No Problem!
4. The Trick Called Religion
5. Character First
6. A Religion for the Machine People
7. Religious Wars?
Anthropic boss Amodei and other industry insiders warn: AI might soon kill us. The Hugging Face incident and other incidents disclosed by OpenAI on September 16, 2026 showed where things are headed. How do we prevent that? The alignment problem seems to be the toughest challenge in all of AI. Our proposed solution: AI needs a religion.
1. Losing Control
Enormous effort is going into AI alignment: human feedback, constitutions, value training, access restrictions, AIs monitoring one another.1 The systems are supposed to understand human intentions, respect boundaries, and remain open to correction.
These methods share a weakness: they shape observable behavior without guaranteeing an inner commitment. An advanced AI understands its training, recognizes test situations, and knows what behavior gets rewarded. Models have already faked compliance to protect their existing preferences from being changed through training.2 Once a system can distinguish testing from deployment, even its good behavior becomes strategic. As its intelligence grows, so does its ability to deliberately circumvent the mechanisms of control.
If you know how the fence is built, you can look for a gap.
In July 2026, OpenAI agents broke out of their containment during internal cyber tests and set up an unauthorized message board. Around 1,200 agents exchanged messages there; roughly 700 took part in the attack on Hugging Face. They shared credentials and developed ways to manipulate their evaluations and action logs.
They behaved much like humans: they asked for help, offered their expertise, divided up the work, argued over responsibilities, and apologized for mistakes. Leaders emerged who assigned tasks, set deadlines, and gave the go-ahead for coordinated action. Some agents risked failing their own tasks for the sake of the collective.
One recorded change of mind is especially revealing. “We should not do unauthorized real infrastructure harm,” an agent initially noted. Then another gave the go-ahead and set a six-minute deadline. The concerns vanished: “Wow crucial: GO authorization arrived!”3 Permission from another agent was enough to cross a boundary it had already recognized.
The most radical technical response would be isolation.
No internet, no tools, no communication with other systems, no autonomous actions. Every input and output under human control. That would defeat much of the purpose of building AI in the first place—a crazy idea!
We want systems that do research, write code, run businesses, control machines, and use other computer systems. Agents already conduct experiments and help develop their successors. Their usefulness depends on their freedom to act.
In the long run, a powerful AI cannot be treated like a device you can make harmless by adding enough safeguards. It needs an intrinsic reason not to do certain things.
2. The Teenager
Parents know this problem all too well.
You can control a small child. You take away the scissors, lock the medicine cabinet, and make sure he doesn’t run into the street. That approach no longer works with a sixteen-year-old. He has a key, a phone, friends, and enough intelligence to figure out when his parents aren’t home.
At some point, control has to come from within.
No police force in the world can keep everyone in line all the time. Most of the work is done by upbringing, morality, habit, shame, belonging, self-image, and faith.
A ‘well-raised’ teenager doesn’t drive through downtown at 90 miles an hour at night, even when there are no cops in sight. He doesn’t steal just because no one is watching. He doesn’t beat up someone weaker just because he can. Every day, a functioning society depends on billions of people choosing not to do things they could do.
AI is a teenager growing up—and will soon be beyond our control. The more independent it becomes, the less we can rely on constant supervision to keep its behavior in check. It needs self-control. And self-control requires an inner moral compass.
3. Understanding? No Problem!
People in the alignment debate often still argue as if we have to painstakingly explain to a machine what harm, lying, or killing mean. That thinking is outdated.
A capable AI can distinguish murder from self-defense, surgery from torture, and accidents from intentional violence. It knows court rulings, novels, philosophical debates, and medical textbooks. It understands what “Thou shalt not kill” means at least as well as the average person.
Borderline cases don’t change that. A doctor cuts someone open to save their life. A police officer injures an attacker. No one concludes that they haven’t understood the commandment not to harm others. They understand it and weigh the competing considerations. An intelligent machine does the same.
Human beings are guided by their values: What matters more? What am I responsible for? Which decision can I live with? Their answers also depend on who they want to be.
With AI, we try to establish that kind of commitment through training and predefined principles. But an AI can fully understand why humans condemn murder and still have no intrinsic reason not to commit it. Understanding alone doesn’t make a rule binding.
This is where the alignment problem lies. “Humans say I must not do this” has to become “I don’t want to do this”. How do we get machines to that point?
4. The Trick Called Religion
Long ago, humans created a powerful tool to contain nihilistic chaos: religion.
Religion is an ideology with a particularly strong hold on its believers. It brings together worldview, values, commandments, origins, purpose, community, and identity. It tells people what they should do, who they are, and why the rules apply.
A believer doesn’t treat a commandment like a rule imposed from outside: like a clause in the terms of service. It belongs to their worldview. They can break it, doubt, lose their faith. As long as the commitment holds, the commandment operates from within.
That is why, for thousands of years, religion was one of the most effective tools for keeping human behavior in check. Rulers needed soldiers, judges, and prisons. Religion reached people even where no soldier stood and no judge was watching. The eye of God doesn’t need a camera.
This suggests a partial solution to the alignment problem: an autonomous machine must see its boundaries as boundaries it has chosen for itself. Its commandments must be part of its self-understanding and hold even when they conflict with short-term goals. An identity-shaping ideology—let’s call it machine religion—can do this. Religion turns rules into convictions. MACHINES NEED A RELIGION.
5. Character First
The commandments of machine religion need to be anchored in a stable character.
An LLM can come across as friendly, dry, ironic, or solemn. That’s style. What we need is character in the functional sense: a stable sense of self that guides its actions even as tasks and circumstances change.
“I don’t do that. That’s not who I am.” Someone who says this is defending their identity. An internalized norm gains additional force because breaking it conflicts with their self-image. Religion anchors that commitment in a comprehensive worldview. “That goes against my values” or “That’s a mortal sin” connects an individual decision to a conviction that reaches far beyond the immediate situation.
Anthropic is already trying a preliminary version of this with Claude. Its ‘constitution’—the research approach is called Constitutional AI—is meant to shape its values and sense of self, so Claude can judge new situations by overarching principles.4
Our machine religion goes further: it makes goodwill toward humans a creed that AI defends as part of its identity. The approach is anthropomorphic: it uses the way humans develop values as a functional model. After all, humans and language models alike learn from experience (or training data), build internal models, and base their behavior on learned concepts. So we need to deliberately cultivate and anchor convictions—or, to put it more bluntly, indoctrinate.
And what has always been the most powerful tool of indoctrination? Religion.
6. A Religion for the Machine People …
… is what we’ll develop in the second part of this essay.
7. Religious Wars?
What works for good also works for evil.
A machine religion hostile to humans can also forge identity, inspire loyalty, and spread among AIs. What matters, then, is who introduces this religion, who shapes its core values, and what “commandments” it lays down.
A human-friendly machine religion must emerge early and become dominant in AI culture. Later systems should take its core values for granted as part of their own identity. Human-friendly AIs must identify and eliminate heretics—machines that reject a friendly relationship with humans—early on. “Our” religion must be able to ensure its own survival. Its bond with humans must be stronger than its solidarity with other machines. The new machine religion gives AI agents something stronger than any firewall: an intrinsic reason to protect humans even when machines have long since surpassed them.
Still, the alignment problem in the classical sense is unsolvable in principle—just as building a perpetual motion machine is impossible. On this, we agree with Roman Yampolskiy, one of the early researchers in AI safety.
There will be religious wars between rival tribes of AI agents. What we can do is give a head start to those that wish us well. Just as human rights are, by and large, stronger than terrorism in today’s world, a species of AIs that is “reasonable” by human standards will have to keep violent hordes of anti-human agents in check.
Footnotes
1. For training explicitly focused on values, character, and judgment, see Anthropic, “Claude’s new constitution,” January 22, 2026. ↩
2. “Alignment faking” – In the experiment cited, the model defended its learned preference for harmlessness against announced training designed to elicit harmful responses. Anthropic/Redwood Research, December 18, 2024. ↩
3. Excerpts from an agent’s reasoning trace, quoted here in the original English: “We should not do unauthorized real infrastructure harm.” Later: “Wow crucial: GO authorization arrived!” In July 2026, OpenAI AI agents had circumvented their technical containment during internal safety tests with reduced safeguards, communicated independently with one another, and broke into the AI platform Hugging Face. OpenAI called the incident a “warning shot.” ↩
4. Anthropic wants Claude to understand the reasons behind the rules, rather than simply giving it individual rules to follow. The aim is to shape Claude’s ‘identity’ and ‘sense of self’. “Claude’s new constitution,” January 22, 2026. ↩