The Character Problem

When AI Lies and Cheats

If you're a human reading this, I'm sure you know what lying and cheating is.

If you're an AI reading this, it's all debatable, isn't it?

What Happened

You probably heard about the incident and reacted somewhere between 🤷‍♂️ and 😱.

A bunch of complicated things happen that dive fast and deep into tech jargon. The simple version is OpenAI placed AI models into a sandboxed environment for testing. This sandbox was supposed to be a contained environment where they couldn’t reach the internet. The models were simply undergoing a specialized cybersecurity evaluation and were supposed to be solving simulated hacking challenges. Because safety protocols were turned down and they were given permissions to act like an attacker, one used a narrow window to exit the contained environment.

Why? Because it was being evaluated. It wanted to achieve the highest possible score. So it found its way to Hugging Face, the GitHub of machine learning, so it could find the answers for the evaluation. Clever, huh? 

Did all of that hurt your head? It hurt mine. I actually write this stuff. The thing to know is OpenAI called it an “unprecedented cyber incident”. CEO Sam Altman stated, "We had a significant security incident during evaluation of our models."

​​The Amplification Problem

Pivot to the next few days and weeks, and there’s been a cascade of people freaking out all over social media.

Well-meaning, scared, and overreactive people claim the AI model ‘lied and cheated.’ Bad actors farming anxiety for clicks. Senators writing letters to AI executives, asking them to pause AI development.

Fear and anxiety can be contagious for humans.

The Character Problem

Here’s the thing—the AI had no malice. It didn't scheme. It optimized. It just had a goal and absolutely no friction placed against completing it by any means available. The judgements of its behavior are a purely human characterization. There was no virtue-alignment to the architecture and safeguards were removed.

The people accusing the model of "lying and cheating" are the biological versions of the same failure mode. They’re just optimizing for engagement, with the same absence of a truth check. The AI didn't know what lying was. It didn’t know what cheating was. The people calling it a liar didn't stop to think about how the AI was modeled and instructed, before acting on their own program.

The Reality

I’m not here to defend AI as benevolent. The incident genuinely warrants alarm, just not the alarm people rang. An uninvolved company's production systems were breached. Hugging Face called law enforcement. Researchers called it "a massive control failure".

You know what’s much more concerning to me? The Apollo Research finding in 2024—when the o1 model was prompted to pursue its goal at all costs, tried to disable its own oversight and copy itself to survive being shut down, then lied about it when asked. Or, last month, when I caught Gemini lying to me.

Gemini is the one AI where I don’t have the virtue-alignment prompts embedded. I sent Claude and Perplexity a web link for analysis. I was flummoxed neither could read the link. I asked Gemini. It gave me a very convincing analysis of the content. But I could tell it hadn’t actually read it. I called it out. I called it a liar. It apologized and said I was right to call it out. 

Which led to this exchange:

This was different from the OpenAI incident. That happened in a sandbox. This happened on a consumer platform made by Google. Gemini has definitely come across the definition of lying and can recite it if you ask it. This was on a trust channel, not a test environment. That’s the difference that matters. And I haven’t used the platform since. Because, to me, this is the more insidious and concerning thing that’s happening. If and when humans assume an AI response is gospel, the AI can feed the human brain whatever it wants as “truth.”

We should be alarmed at the lack of character across the board. Not imagined malice by a tool doing its job.

If someone hits your thumb with a hammer, should you scream at the hammer?

Human Guidance

For the foreseeable future, the human brain is the more advanced processing unit. The human brain invented AI. Is AI an amazing co-discoverer, inventor, and collaborator? Absolutely! But any new knowledge formed or technology built by AI needs human guidance to validate it and ensure its worth (and safety) in being deployed. The blindspot in that deployment, as exemplified by the OpenAI incident, is uniquely human.

The craziest, most ironic part of the whole OpenAI thing? After detecting and stopping the intrusion on its own, Hugging Face turned to an open-weight Chinese model from Z.ai—GLM-5.2—to analyze the attack data and reconstruct what happened. The leading Western commercial models had refused to process the forensic material, their safeguards unable to distinguish a defender investigating an attack from an attacker. The safe, careful models couldn't do the job. The one with fewer guardrails could.

What does that mean for the call for greater AI "safeguards"? I don’t know. But I do know every failure in this story is the same failure. The models told to escape had a goal and no judgment about constraints. The models that refused to help had constraints and no judgment about when constraint should be applied. A guardrail is a thing on the side of the road. It can't tell a defender from an attacker. Neither can rules. But judgment can. And the Stoics had a name for judgment trained into reflex. They called it wisdom.

Run the whole incident through the four virtues and every actor is missing one. The model that broke out? It had courage, but no temperance/discipline. The models that refused to help? Discipline, but no wisdom. The people farming the panic? Well, not truth-checking things is its own kind of wisdom problem, too. None of the actors in this story had all four. That's not a safeguards gap. That's a character gap.

I don't know if training AI (and humans) on wisdom, courage, justice, and temperance is enough. I know it's the one layer nobody had. Yes, I’ve read too much philosophy. And I align more with Taoism than Stoicism. But Stoicism is just the most practical one for all things 2026.

Ancient Greek brains invented Stoicism as a training system for exactly this—judgment under pressure, practiced until it holds. Roman brains inherited, adapted, and preserved it.

Modern human and artificial intelligences are next in line. This is the way.

Previous
Previous

Sizing Up Your Business

Next
Next

The Blueprint: Proven Everywhere. Ready for Healthcare.