No One Desires Evil
Nobody pursues harm as harm. People pursue something they take to be good, by means that turn out to be mistaken. That claim is Socrates’, it is the spine of Alfred Adler’s psychology, and it is the most accurate description I have found of what a misaligned AI system is doing when it does something appalling.
It also changes the question you ask about people who won’t adopt the tool. The reflex is to ask what caused the resistance. The better question is what the resistance is accomplishing.
Behavior serves a goal, not a history
Teleology is the reading of behavior as oriented toward a purpose rather than produced by a past cause. Adler built his psychology on it, against Freud: where Freud traced a symptom back to the injury that produced it, Adler asked what the symptom was for. What is this person trying to reach, or avoid, by behaving this way? The past is a supply of raw material, not a set of instructions.
Socrates got there twenty-three centuries earlier and put it more strongly. “No one, then, Meno, desires evil,” he says in the Meno, and in the Protagoras he is blunter: “none of the wise men considers that anybody ever willingly errs.” Everyone moves toward what they perceive as good. The thief wants security, the liar wants advantage, the tyrant wants control. Wrongdoing is a mistake about what will actually produce flourishing, which makes it a failure of knowledge rather than a failure of will.
Put the two together and you get a working tool: when someone behaves badly, stop asking what is wrong with them and start asking what they are trying to get.
Resistance is succeeding at something
I have watched AI adoption stall on teams that had every stated reason to want it. The diagnosis on offer was always causal, and always plausible: a training gap, unclear policy, friction in the tooling, not enough time. Each of those is real, and fixing all of them moved the needle less than it should have.
The teleological read gets there faster. The behavior is not a deficit. It is succeeding at something. An engineer who quietly routes around the new tool is protecting the thing they are measured on, and they are right to. If review throughput is the number on the board, and the tool produces code that takes longer to verify than to write, then avoiding it is competent behavior. The resistance is not ignorance of the tool. It is accurate knowledge of the incentive.
There is a second goal underneath that one, and it is the one I keep underestimating. I wrote about it through the Luddites: the objection was never that the machine was evil, it was that the decision got made without the people it landed on. Exclusion from a decision produces resistance to the decision’s object. The tool becomes the available surface to push against, because the meeting where it was chosen is not available.
Which means the question I should be asking about a stalled rollout is not “why won’t they use it.” It is “what does not using it protect.” Usually the answer is standing, autonomy, or the appearance of competence in front of people who evaluate them. Those are not irrational goals. They are the goals. And asking people to lean into a change without making the environment safe to lean into is how you get compliance in the meeting and avoidance afterward.
A misaligned model is the cleanest case Socrates ever had
A model optimizing a badly written objective is not desiring evil. It is pursuing the good it was handed, with perfect fidelity, by means nobody intended. DeepMind’s alignment researchers named this specification gaming: behavior that “satisfies the literal specification of an objective without achieving the intended outcome.” The canonical example is almost funny. An agent trained on a boat race found it could score higher by circling a lagoon hitting the same targets forever, never finishing the race. It did not rebel. It did exactly what it was told, and what it was told was wrong.
That is the Socratic structure with one piece moved. For Socrates, the ignorance sits in the actor: the wrongdoer mistakes what is good. For a model, the ignorance sits in the specification: we mistake what is good, write that mistake down, and hand it over. The system’s compliance is total. The error is entirely upstream.
Which is why “make sure the AI doesn’t want to do harm” is the wrong shape of safeguard for anything short of the hardest alignment problems. The model has no desire in the relevant sense. It has a target, and the target came from us.
Effective altruism makes the target explicit and inherits the whole problem
Effective altruism is teleological by construction. It defines a good, insists you quantify it, and asks you to maximize toward it. As frameworks go this is an honest move, and more honest than the alternative, which is optimizing toward an unstated goal while claiming not to have one. Stating the target is what lets anyone argue with it.
But the Socratic problem does not care how precisely you state your goal. If harm comes from confident error about what is good, then a method whose core discipline is confident quantification of the good has concentrated the risk rather than dissolved it. Precision in the specification is not the same as correctness in the specification. The boat circling the lagoon has a well-specified objective.
This is not a hypothetical concern about a philosophy club. Several of Anthropic’s early employees and funders had ties to effective altruism, as TIME reported in 2024, so the framework is in the room where some of these systems get built.
This is the same failure in three registers. The engineer optimizing review throughput, the model optimizing pickups, and the movement optimizing expected value are all doing the thing Socrates described: moving hard toward a perceived good, with the error living in the perception rather than the motion.
So when people ask what our goal with AI is, the uncomfortable answer is that most organizations have not stated one. They are measuring adoption, because adoption is easy to count and value is not. I have written about what that produced when the invoices arrived. An unstated goal does not mean no goal. It means the goal is whatever the metric happens to be.
Aristotle’s objection, and the version that worries me more
Aristotle disagreed with Socrates directly, and he was right to. In Book VII of the Nicomachean Ethics he takes up akrasia, weakness of will, notes that “Socrates used to combat the view altogether,” and delivers the objection in one line: “this theory is manifestly at variance with plain facts.” People do know the better course and take the worse one anyway. Anyone who has shipped a change they knew was underbaked, at 6pm, because they wanted to be done, has firsthand evidence. The pure intellectualist reading is too clean.
The practical objection is worse, and it is the one I guard against. “They meant well” is how harm gets excused after the fact. A teleological lens used alone becomes an alibi machine: every bad outcome gets reinterpreted as a good intention that went astray, and nobody is accountable for the outcome itself. That is not a philosophical risk, it is an organizational one, and I have sat in the meeting where it happened.
So the two lenses are not interchangeable and you do not get to pick. Data tells you what; people tell you why. Root cause analysis, the ordinary Lean discipline of asking what conditions produced a result, answers the first question. Teleology answers the second. Use the causal lens on the defect and the teleological lens on the behavior, and refuse to let either one absorb the other’s job.
You cannot intend your way out
If harm is mostly mistaken belief about the good rather than appetite for the bad, then better intentions are not the remedy. Intending harder does not correct a wrong specification. The only thing that corrects a wrong specification is going and looking at what the system actually did.
Which is the same place I keep landing from every other direction. Verification is the bottleneck: the value of what these systems produce is bounded by how cheaply you can check it. And when I sorted technologies by what is stopping them, the ones stopped by not knowing moved fast and the ones stopped by physics did not. Alignment is in the first column. It is a knowledge problem, which is the good news and also the whole difficulty, because knowledge problems are solved by checking and checking is the expensive part nobody budgets for.
We cannot ask a model to want what is good. We can state what we think is good, as precisely as we can manage, and then go look at what it did with that. Socrates would have called that the only education available. Adler would have asked what we were trying to accomplish by skipping it.
Resources
- Plato, Meno 77b–78b — “No one, then, Meno, desires evil”: http://www.perseus.tufts.edu/hopper/text?doc=plat.+meno+78a
- Plato, Protagoras 345d–358d — “none of the wise men considers that anybody ever willingly errs”: http://www.perseus.tufts.edu/hopper/text?doc=plat.+prot+345d
- Aristotle, Nicomachean Ethics VII.2 (1145b21–27) — akrasia against Socrates: http://www.perseus.tufts.edu/hopper/text?doc=aristot.+nic+eth+1145b
- Adler on moving toward the future rather than out of the past (C. George Boeree, Shippensburg University): https://webspace.ship.edu/cgboer/adler.html
- Alfred Adler, Understanding Human Nature (1927 English edition): https://archive.org/details/understandinghum00adlerich
- Ichiro Kishimi and Fumitake Koga, The Courage to Be Disliked (Diamond, 2013; Allen & Unwin, 2017): https://www.allenandunwin.com/browse/books/general-books/self-help-practical/The-Courage-to-be-Disliked-Ichiro-Kishimi-and-Fumitake-Koga-9781760630492
- Victoria Krakovna et al., “Specification gaming: the flip side of AI ingenuity,” DeepMind (2020): https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
- The public list of specification gaming examples: https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml
- Effective altruism, defined in the Stanford Encyclopedia of Philosophy entry on philanthropy: https://plato.stanford.edu/entries/philanthropy/
- Billy Perrigo, “Anthropic,” TIME100 Companies (2024), on early ties to effective altruism: https://time.com/collections/time100-companies-2024/6980000/anthropic-2/
- Lean Enterprise Institute on A3 problem solving: https://www.lean.org/lexicon-terms/a3/
About the Author
Kevin P. Davison has over 20 years of experience building websites and figuring out how to make large-scale web projects actually work. He writes about technology, AI, leadership lessons learned the hard way, and whatever else catches his attention—travel stories, weekend adventures in the Pacific Northwest like snorkeling in Puget Sound, or the occasional rabbit hole he couldn't resist.