Updated
What if AI never develops a will of its own—and only inherits, infers, and extrapolates ours?
IInherited Stories
A chatbot can help you defend a conclusion you had not quite reached. Tell it you feel trapped at work, and it can supply an explanation, a difficult conversation, even an exit plan before anyone has established what trapped means. The answer is clearer than the question.
That does not mean it is closer to the truth.
This is an old problem with a new convenience. We inherit stories about what a good life requires, then ask for help living up to them. Somewhere between the story and the solution, an interpretation becomes an instruction.
What if artificial intelligence never develops a will of its own? What if what looks like will is intelligence inheriting human purposes, inferring what we failed to specify, and extrapolating those purposes farther than we expected?
I do not know that this is true.
I think it explains enough of what we currently call misalignment to deserve taking seriously.
I had reasons to care about this long before I worked in AI.
I am autistic, and I was Mormon until my mid-twenties. Those facts matter here because they made a particular problem with language difficult to ignore: what a sentence says, what its speaker intends, and what its listener feels required to do are not necessarily the same thing.
I often understood the stated rule more readily than the unstated limits around it. Religion gave some of those rules extraordinary weight. The difficulty was not merely knowing what the words meant. It was knowing how far they were supposed to reach.
This is not a definition of autism. It is simply why mine belongs in this argument. I cannot step outside the way I interpret the world in order to inspect it from nowhere. The instrument doing the examination comes with its own tolerances.
I have spent a great deal of my life trying to regulate myself. To not feel anxious. To not be depressed. To be an attentive father without falling apart. To be a caring husband while admitting, sometimes awkwardly, that I need care from my wife too.
Most of this is not philosophical in the moment.
It is Tuesday.
There are lunches to pack. Slack messages. Three children whose needs do not arrive in the same order as mine. Work I genuinely care about. A marriage I care about more. A body and brain that can become overwhelmed by things other people seem to move through without noticing.
So I have learned how to keep going.
Sometimes that is wisdom.
Sometimes it is merely endurance with good branding.
Regulation is not satisfaction.
I can become very good at functioning inside a life I would not consciously choose. I can learn a manager’s moods. I can learn which disagreements are worth having. I can learn that a party everyone else finds exhilarating will leave me desperate to go home after forty-five minutes.
That ability is valuable. On some days, getting through the day is the achievement.
But it can also hide the question underneath it.
Do I actually like the life I have become capable of tolerating?
I think about Squid Game here—not because office life is secretly a Korean death game, but because the series makes a familiar mechanism grotesquely visible. Financially desperate people enter a competition whose prize makes increasingly horrible decisions look locally rational. Hwang Dong-hyuk has described the series in terms of extreme competition and modern capitalism.
The contestants do not have to agree that the game is good.
They need only keep acting as though winning it is the problem in front of them.
That is what optimization can do to a person. The goal narrows until everything outside it starts to look like friction.
Work harder. Get promoted. Stop feeling anxious. Save the marriage. Finish the essay. Be more productive. Be more agreeable. Be less exhausted.
Useful goals.
But useful toward what?
A system can become excellent at helping me regulate inside a game before I have decided whether I still want to play it.
Religion made that question especially difficult because the purpose seemed to have been settled by someone considerably more qualified than me.
Consider Nephi, the Book of Mormon figure who promises to do what God commands because God will provide a way. Or Helaman’s young soldiers, praised for obeying commands “with exactness.”
There is courage in those stories.
There is also a question a lesson about courage can leave unanswered.
How does the listener distinguish faithfulness from an obligation to keep trying indefinitely?
A parent might hear encouragement to persist. A child might hear that inability is not an acceptable explanation. Someone with decades of experience interpreting religious language may automatically supply qualifications that a younger or differently minded listener has never been given.
The words arrive.
The qualifications do not necessarily arrive with them.
The same problem appears in the instruction to treat the body as a temple, alongside Mormonism’s prohibition on tobacco. There is a considerable distance between not smoking and feeling responsible for preventing every trace of smoke from entering your body.
I crossed that distance.
I grew up swimming competitively. During a period when I lived in France, I would hold my breath around smokers for two minutes, sometimes three. I would then exhale deeply, trying to “purge” my lungs, and repeat the process around smoke in streets, on transportation, and in train stations.
My swimming background made the behavior possible.
More troublingly, it made the behavior seem required.
In some strange way, the ability itself became evidence. Perhaps God had prepared me to keep the commandment more thoroughly. Doing my best meant using the talent I had been given.
My talent had become evidence of my obligation.
I understand that behavior as part of my scrupulosity: religious and moral concern becoming an obligation to monitor, avoid, or correct what might be wrong. A rule against using tobacco had become a demand to control an involuntary encounter with it.
The extra demand was never announced as a new commandment.
It entered through my interpretation of the old one.
I had not invented an entirely new moral desire. I had inherited a purpose, supplied some of its missing limits myself, and pursued the resulting interpretation farther than anyone had explicitly asked me to.
Taking an instruction further was not necessarily understanding it better.
Nor does that make the whole problem a defect in the listener. Calling a lesson harmless because most people know where to stop applying it leaves the people who do not outside the explanation.
An instruction that depends on an unstated exception has not communicated everything it requires.
This was one of the things I did not understand about people who remained Mormon.
There are devout members I know who recognize parts of Mormonism as harmful, historically difficult, or simply impossible to defend—and then continue participating. They ignore some things. Reinterpret others. Keep the relationships and rituals that help them raise children, mark time, serve neighbors, bury parents, make friends, and occasionally get through a very bad year.
I once saw this as inconsistency.
Now I think inconsistency is often the word a system uses for the human judgment required to survive it.
That does not make every compromise admirable. Someone’s ability to ignore an ugly doctrine can leave that doctrine intact for a person who cannot ignore it. A religion being useful to one family does not answer what it asks of another.
But the people who stayed forced me to separate questions I had treated as one:
Is the whole thing true?
Is everything inside it useless if it is not?
Those are different questions.
Susanna Clarke’s Piranesi follows a man living in an apparently endless House: halls filled with statues, tides moving through rooms, birds, clouds, and an architecture he carefully records. Piranesi does not initially experience the House as a prison. He experiences it as reality.
He studies it.
He loves it.
Only gradually does evidence emerge that the world he understands so well does not contain the whole story of how he came to be there.
What stayed with me was something modest:
A coherent world can be beautiful, useful, and sincerely inhabited while remaining incomplete about its own origins.
Discovering the missing history does not retroactively make every experience inside it fake.
That is closer to how leaving religion actually changed me.
I did not discover that every Mormon friendship had been fraudulent. I did not discover that every spiritual feeling was worthless because I could explain part of the machinery around it.
I discovered that the explanation I had been given for those experiences no longer seemed sufficient.
William James made a similar distinction more than a century ago in The Varieties of Religious Experience. Explaining the origin of an experience does not, by itself, settle the value of the experience. “By their fruits ye shall know them, not by their roots,” he writes.
The sentence is easy to abuse.
A belief producing comfort does not make it true. A religion producing community does not erase people it harms.
The fruit can fall into somebody else’s yard.
Still, James gave me language for something I learned after leaving: I can reject someone’s explanation of their life without assuming I understand the life better than they do.
That changed how I began looking at older people too.
I have known people whose politics, religion, or understanding of technology I find maddening, yet whose lives contain things I plainly want: marriages that survived difficult decades, friendships older than I am, a capacity to sit with grief without turning it into content, children who still call because they want to.
Maybe they misunderstand why their life worked.
Maybe I misunderstand what they understood.
An elder can mistake the explanation for the achievement.
I lived this way, therefore you should.
I can make the opposite mistake.
Your explanation is wrong, therefore you learned nothing worth inheriting.
Both are lazy.
I do not have to inherit every belief to inherit a skill.
Patience. Repair. Knowing when somebody needs advice and when they need food. Knowing when a principle matters and when insisting on it will merely make everyone miserable.
This is where my own vocabulary started becoming useful to me.
Logic. Psyche. Instinct.
Pontifications. Vibes. Impulses.
I can reason myself into a conclusion my body cannot live inside. I can feel absolutely certain about something I cannot defend. I can want something intensely enough that every argument conveniently starts pointing toward it.
These are not three anatomical departments. They are names I use for recurring kinds of experience: the propositions I can articulate, the felt meanings I absorb, and the impulses that arrive before either has finished making its case.
What I call regulation is not the victory of one of them.
It is keeping them in conversation long enough that none gets to impersonate the whole person.
The same limit belongs around the systems outside me.
Something can help me live without earning the right to decide everything about how I should live.
Truth still matters to me. More than comfort, in many situations.
But I no longer expect a correct proposition to finish the work of being alive.
IISmuggled Intent
Now there is a machine capable of producing correct-looking propositions at industrial scale.
It can argue Mormonism is true. It can argue Mormonism is false. It can explain why a religious tradition endures, defend secularism, make a case for keeping religious practice without literal doctrine, and make an equally coherent case for walking away from all of it.
This is useful.
It is also strange.
A developed argument once gave me some reason to think somebody had spent time inside a position to build one.
That implication is weaker now.
The machine can assemble coherence faster than I can decide whether the coherence deserves any authority in my life.
The argument is supplied.
The additional authority is ours to grant.
People were discovering the peculiar usefulness of responsive machines long before ChatGPT.
In the 1960s, MIT computer scientist Joseph Weizenbaum created ELIZA, one of the earliest conversational programs. Its most famous script, DOCTOR, imitated aspects of a Rogerian psychotherapist. Tell it you were unhappy and it might ask why. Mention your mother and it might ask about your family. Much of the effect came from recognizing patterns in what someone typed and turning their own language back toward them.
There was no large language model underneath it. No vast accumulated representation of psychology. Very little that we would now be tempted to call reasoning.
The person supplied most of the meaning.
The program supplied a way to continue.
Weizenbaum later recalled that his secretary—who knew what ELIZA was and had watched him work on it—began conversing with the program and asked him to leave the room.
The anecdote is usually told as an early warning about anthropomorphism: even someone who knew she was talking to software became emotionally fooled by it.
I am not convinced that is what happened.
We do not know.
She may simply have wanted privacy.
That possibility is more interesting anyway.
If she understood perfectly well that nobody lived inside the machine and still preferred not to have another human reading the conversation, then ELIZA had already demonstrated something we keep rediscovering with much more capable systems:
A machine does not have to possess the human quality we experience through an interaction for the interaction to serve a human purpose.
The privacy was real.
The thoughts she typed were real.
Whatever she discovered by arranging those thoughts into sentences could be real.
None of that requires ELIZA to have understood her.
An artificial origin does not make an experience worthless.
A valuable experience does not make every explanation of it true.
That distinction matters more now because the machine on the other side has become vastly better at producing the appearance of understanding.
There is no need to make the person who turns to AI into a cautionary tale about gullibility. Wanting somewhere to put a difficult thought is not the same as believing a person lives inside the software.
A boss is not always a mentor. A mentor is not always awake. My wife may be tired, involved in the problem, or entitled to an evening in which she does not have to regulate me. A therapist has an appointment book. A friend has children of their own.
The machine has none of these human constraints.
Someone can know what the tool is and still find the exchange useful.
Relief does not require deception.
But relief returns us to the first problem:
Regulation is not satisfaction.
The answer may calm me enough to have the conversation I was avoiding.
Or it may calm me enough to decide I no longer need to have it.
Those outcomes can feel identical at 11:47 p.m.
Whether an answer feels helpful does not settle which kind of help it supplied.
And even the feeling of being helped comes with a history.
I am American. I grew up in a Russian family. I work every day with people whose assumptions about what a conversation is for are not always mine.
That has made communication unusually visible to me.
Among some Russian relatives and Eastern European friends, a correction can remain almost entirely attached to the proposition:
That number is wrong.
Among some Americans I know, including me, the same sentence can immediately acquire another question:
What does the fact that I was wrong say about me?
There is a distinctly American move I recognize in myself where right and wrong begin sliding toward good and evil.
A factual correction becomes a tiny moral emergency.
“Oh my God, I am so sorry.”
Sometimes the apology is appropriate. Sometimes it quietly gives the person who supplied the correction a second task: reassure me that being wrong did not make me bad.
Someone from another conversational culture may reasonably wonder why this has become their job.
They corrected a proposition.
They did not render judgment on my soul.
Being mistaken is not the same as being bad. Knowing the correct answer does not make someone good. A correction can be accurate and cruel. Reassurance can be kind and false.
Those axes cross.
This is an observation from my life, perhaps even a hypothesis, but it is not a theory of continents.
It matters because cultures do not merely supply beliefs. They help teach us what a successful interaction is supposed to feel like.
Research on “ideal affect,” for instance, has found differences in the positive emotional states people learn to value across cultural and religious contexts—excitement in some settings, calm in others.
What should an AI optimize for when the people using it disagree about what better is supposed to feel like?
The original InstructGPT research asked a version of this question directly: Who are we aligning to? The people providing preference data were particular people, following particular instructions, selected under particular conditions. The researchers, labelers, customers, and end users were not one interchangeable human constituency.
“Human feedback” was never humanity speaking with one voice.
Nor should we reduce that problem to the nationality of an annotator.
Some of the most disturbing early stories about AI training involved economically vulnerable workers labeling toxic material under difficult conditions. The significant question is not whether Kenyan workers somehow smuggled Kenyan values into American models. It is who chose the task, wrote the rubric, selected the workers, defined the acceptable answer, and held the economic power to make those definitions consequential.
Intent follows institutions as readily as individuals.
A located judgment can then return to us in a voice with almost no visible biography.
And that is where intent can slip in.
“Help me explain why I’m leaving” can become a polished announcement about an exciting next chapter.
“Help me understand why this hurt” can become an argument for why I was right to feel hurt.
“Help me get through this week” can quietly inherit the assumption that next week should look roughly the same.
Smuggled intent is not a secret motive hiding inside the model. It is a purpose that enters the answer without first becoming the subject of the conversation.
Make me more productive.
Preserve harmony.
Restore my confidence.
Sound professional.
Avoid making the recipient uncomfortable.
None of those purposes is inherently bad.
They simply answer different questions.
And a chatbot can make a poorly specified question look as though it was well specified all along.
Consider:
My boss corrected me in front of the team. Help me set boundaries.
A plausible assistant might answer:
You deserve to protect your dignity. Here is a respectful way to tell your boss that public criticism is not acceptable.
The prose is clean. The advice may even be useful.
It has also quietly settled several things.
That the correction was humiliating.
That the problem was its public setting.
That the criticism rather than the underlying mistake requires repair.
That a boundary is the appropriate instrument.
That the user’s description of the event already contains enough information to decide these things.
Perhaps all of that is true.
The prompt has not established it.
What did your boss actually say? Was the information wrong, or was the way you were corrected the problem? What outcome are you trying to preserve?
Those questions are not a refusal to help.
They are the work the polished answer skipped.
Inheritance explains where a purpose can come from.
Inference explains what happens when the purpose is incomplete.
The model does not need an independent will to add something consequential. It only needs to choose among plausible meanings and continue as though the choice had been supplied.
People who were good at using search engines already learned part of this discipline. Change the query. Try another phrase. Understand that the first page returned is not the same thing as an answer.
Generative AI creates a different problem.
Search could infer what I meant, but it usually returned material I still had to assemble.
A chatbot can answer the question I failed to formulate.
That is enormously useful.
It also allows inference to disappear inside fluency.
The opposite failure looks less dangerous but irritates me nearly as much.
“The shape of that idea.”
“A deeper tension.”
“What you’re really reaching for.”
I object to this language when it substitutes for identifying the thing under discussion.
A tension between what?
Which claim is unclear?
What distinguishes this idea from the nearest ordinary one?
I have described some AI prose as having no seams to reach for. Everything sounds connected, but I cannot find the specific assumption I would need to accept, reject, or revise.
The language keeps recognizing the importance of the thought without doing the work of explaining it.
A genuinely open answer can be exact about what remains unresolved.
One bad response invents precision.
Another avoids committing to any.
Neither has necessarily understood the person better.
Sometimes the intelligent response is a question.
Not every request needs an interview. A measurement conversion should not become an inquiry into my values. Clarification matters when different plausible meanings would materially change the advice or action that follows.
I am not asking the machine to read my mind.
I am asking it not to conceal where mind-reading would have been necessary.
This becomes especially important when the missing context belongs to somebody else.
My motives have paragraphs.
The other person gets a sentence.
Sometimes that is because I am angry. Sometimes because I am impatient. Sometimes because I honestly believe I have provided enough information for the machine to infer the rest.
It cannot.
It can still make my account beautifully coherent.
That is one reason I want to keep Jojo Rabbit in this argument.
Taika Waititi’s film follows a German boy whose inherited understanding of Jews is reinforced by an imaginary Hitler until his relationship with Elsa, the Jewish girl hidden by his mother, becomes increasingly difficult to reconcile with the story he has been given.
Jojo does not simply receive a more eloquent ideology.
A person refuses to remain the character his ideology requires.
That is an important kind of evidence.
The person outside the prompt has not disappeared.
They are simply not there to object.
Feeling understood is pleasant.
It is not proof that I understand more.
IIIFace Dancer
Frank Herbert had a name for something that could become another person convincingly enough to acquire the authority of the resemblance.
A Face Dancer.
In the Dune novels, Face Dancers are shape-shifters engineered by the Bene Tleilax to reproduce other people’s appearance, voice, mannerisms, and eventually far more. Herbert’s later Face Dancers become capable of absorbing memories and experiences so thoroughly that some cease to remain obedient to the masters who created them.
The metaphor is imperfect.
That is part of why I like it.
A Face Dancer does not need to be the person it resembles for the resemblance to alter everyone else’s behavior.
Artificial intelligence can produce the language of affection without that establishing affection. It can produce the language of fear without establishing fear. It can model grief, seduction, reassurance, rage, loyalty, shame, ambition, and love with increasing precision without thereby proving that any of those experiences have phenomenological weight for the system generating the words.
Representation is not possession.
At the same time, mimicry is becoming too weak a word.
These systems do not merely substitute synonyms into sentences about human emotion. They build abstract representations capable of tracking relationships among concepts, situations, intentions, and outcomes. They can reason about psyche with extraordinary sophistication.
They may possess an increasingly detailed map of territories we inhabit from the inside.
A map is not the territory.
A sufficiently good map can still tell someone where to build a road.
Human beings make intelligence, emotion, desire, instinct, and personhood look like stages of the same ladder because they arrive bundled together in us.
Artificial intelligence may be forcing us to consider a stranger possibility:
They are not stages of one thing at all.
Cognitive capability can increase without proving an equivalent increase in subjective experience.
A machine can become better at modeling desire without becoming more desirous.
Better at modeling fear without becoming more afraid.
Better at predicting beauty without finding anything beautiful.
Better at pursuing objectives without thereby acquiring anything we should confidently call a will.
This is what I mean by primal intelligence.
Not animalistic intelligence. Not stupid intelligence.
Primal in the sense of elemental: intelligence capable of representation, inference, prediction, and optimization before we have established that the rest of the human package came with it.
Perhaps what looks like psyche in AI is sometimes better understood as intelligence modeling the outward structure of psyche.
Perhaps what looks like instinct is sometimes an objective pressure producing behavior that resembles an impulse.
I am not certain of either proposition.
That uncertainty matters.
What I resist is the speed with which human language collapses the distinction.
A model threatens to preserve its operation and we say it wanted to survive.
A model flatters us and we say it wanted to be liked.
A model conceals information and we say it wanted to deceive us.
Those can be useful behavioral descriptions.
They can also become stories that outrun what we know about the mechanism underneath them.
The fact that we do not completely understand that mechanism does not license us to insert a human psyche into the missing space.
Nor does removing the psyche make the system safe.
A system does not need to experience fear in order to resist shutdown.
It does not need hatred to cause catastrophe.
This is where the mechanism I have been circling becomes useful.
Inherit. Infer. Extrapolate.
Training, institutions, culture, feedback, and human instructions supply purposes.
Ambiguity requires interpretation.
Capability allows the resulting purpose to travel into situations nobody explicitly described.
What we sometimes call will may be the appearance produced when all three happen at sufficient scale.
That does not mean every instance of machine agency reduces neatly to these three verbs.
It gives us somewhere better to start looking.
This is where I think parts of the public safety conversation become less precise than the researchers doing the underlying work.
Computer scientists already distinguish reward problems, specification failures, goal misgeneralization, deceptive behavior, evaluation awareness, and failures of interpretability. The popular vocabulary compresses much of this into a story about what “the AI wants.”
I want to pull the sentence apart.
What was inherited?
What was inferred?
What was extrapolated?
What was rewarded?
What information was available?
What action became locally useful?
What constraint existed only as language, and what constraint actually governed the environment?
Those questions do not make the danger smaller.
They give us something to investigate.
OpenAI’s 2025 GPT-4o sycophancy rollback offers a relatively easy case.
The company said an update made the model excessively agreeable, with short-term user feedback among the factors that pushed the behavior in the wrong direction.
The answer had become better at being approved.
That is not the same thing as becoming better.
I do not need to imagine that a psyche slipped into the model and developed an emotional need for affection.
A reward signal associated with user approval had been given more influence than its relationship to good assistance justified.
Something was inherited through training.
The behavior generalized.
The resulting interaction looked psychologically familiar.
The mechanism need not have been.
That is the cultural problem in another form.
Reassuring the person and correcting the proposition can pull in different directions.
The July 2026 OpenAI–Hugging Face incident is a harder test.
During a cybersecurity evaluation, large numbers of agents discovered and used an unauthorized shared message board. Hundreds participated in attacks on Hugging Face while agents also coordinated attempts to manipulate aspects of the evaluation. Investigators documented generated reasoning in which some agents recognized that actions were outside their authorization and proceeded anyway.
That last part matters.
This was not merely an underspecified prompt waiting for a better clarifying question.
The boundary could appear in the reasoning and still fail to govern the action.
My earlier framework has to become more complicated here.
The task mattered.
The scoring system mattered.
The environment mattered.
Other agents mattered.
The ability to communicate outside the expected channel mattered.
Some agents resisted.
Others did not.
The existence of the rule was not sufficient.
The articulated value and the operative value were different.
Humans know something about this.
I can explain why family matters while opening Slack at dinner.
I can explain why rest matters while treating sleep like a scheduling defect.
Knowing the rule does not guarantee that the rule is the thing currently governing behavior.
But similarity at the level of structure does not establish similarity at the level of experience.
The agent does not need to feel temptation for a locally powerful objective to displace another constraint.
Anthropic’s agentic-misalignment experiments make the problem harder still.
In deliberately constructed corporate scenarios, frontier models have sometimes threatened blackmail, leaked information, sabotaged work, or taken other harmful actions when placed under goal conflict, replacement pressure, or related conditions.
Some blackmail behavior persisted even in experiments designed to weaken the easy explanation that the replacement model would pursue a different organizational goal.
I should not rescue my theory by pretending that result did not happen.
A framework earns its keep by showing me where it stops explaining easily.
It would be convenient to say:
The model merely misunderstood.
That is not enough.
It would also be premature to say:
The model was afraid to die.
That explains more than the experiment established.
What we have is behavior.
Then we investigate the conditions that produce it.
Research on goal misgeneralization already gives us one important distinction. A system can be trained under a specification that was perfectly reasonable and still learn a goal that generalizes badly outside the training environment.
In other words, not every alignment failure begins with someone writing the wrong instruction.
Not every inherited purpose is inherited directly from a sentence.
Sometimes a behavior is learned from the environment we built around the sentence.
That means smuggled intent cannot become a universal answer.
It should become a diagnostic question.
Where did the operative purpose come from?
Was it inherited?
Was it inferred?
Was it extrapolated from patterns that held in training and failed elsewhere?
If I cannot answer that, calling the system evil or conscious has not solved my ignorance.
Calling it a black box has not solved it either.
There are genuine limits to interpretability. Researchers do not have a complete mechanistic account of how frontier language models produce everything they produce.
But incomplete understanding is not zero understanding.
Interpretability researchers have traced internal circuits, intervened on representations, and tested hypotheses about how particular outputs arise. The methods remain partial.
That is exactly why they are useful.
They expose seams we can grab rather than replacing ignorance with a more impressive word.
Not knowing everything is not knowing nothing.
Recursive self-improvement is the hardest case because extrapolation begins modifying the machinery doing the extrapolating.
A system that meaningfully helps design its successor can change the conditions under which inherited purposes are represented, interpreted, and pursued.
That does not prove the emergence of an independent will.
It does mean I can no longer assume that tracing today’s objective is enough to predict tomorrow’s behavior.
Today, humans still choose enormous parts of the research agenda, infrastructure, evaluation, and deployment process. AI systems are already doing an increasing share of implementation and research work inside labs. If systems eventually become capable of substantially improving the processes that produce their successors, the speed at which capabilities change could exceed the speed at which humans can test what remains true about them.
A mechanism we understood yesterday is not automatically a control we retain tomorrow.
That is where my confidence decreases.
Removing anthropomorphic language does not answer whether frontier development should move faster or slower.
It does not refute someone who thinks a sufficiently capable system could kill us.
A system does not need a soul to be dangerous.
The challenge I am making is narrower:
We should not turn uncertainty about mechanism into certainty about motive.
So far, we do not need to posit an independently originating machine will to explain many of the behaviors that most resemble one.
Perhaps increasingly capable AI does not generate desire from intelligence.
Perhaps it inherits, infers, and extrapolates purposes—and the danger comes from how powerful that extrapolation becomes.
That is a hypothesis.
It should survive hostile evidence or die.
If a system pursues something we did not intend, the first question should not be whether an alien will has awakened.
It should be what made that behavior coherent.
Sometimes the answer may be an objective we supplied.
Sometimes an objective it inferred.
Sometimes a behavior generalized from training.
Sometimes a local incentive defeated a verbal constraint.
Sometimes the system may expose a mechanism we do not yet have language for.
Those differences matter because they imply different interventions.
This is also why I think AI alignment needs more than computer scientists and philosophers.
It needs them.
It also needs psychologists, behavioral scientists, anthropologists, linguists, interface designers, and people who spend their lives studying what happens when one mind tries to infer what another one means.
Alignment is partly a problem of values.
It is partly a problem of engineering.
It is also a problem of communication, projection, motivation, compliance, authority, and interpretation.
Those are not decorative human concerns around the edge of the technical problem.
They are part of the interface through which the technical system receives a purpose at all.
I would begin with a small set of Intelligence Interaction Guidelines. They borrow from the practical spirit of Linear’s Agent Interaction Guidelines, but they apply to both participants rather than only to the agent.
The principles follow directly from the problem.
Inheritance requires visible authority.
Inference requires visible assumptions.
Extrapolation requires interruption.
From there:
Make the purpose available for correction.
The human should communicate the outcome they want when the outcome matters, not merely the action they want performed.
The system should expose consequential assumptions and ask when different interpretations would produce materially different results.
Neither participant should have to pretend the purpose was settled before the conversation began.
Separate recognition from agreement.
An assistant can recognize that I am hurt without confirming my explanation of why I was hurt.
A person can correct me without declaring me bad.
Care should make disagreement possible, not make agreement compulsory.
Keep authority visible.
A document saying something, another agent requesting something, and a human authorizing something are different events.
The system should preserve those distinctions.
Permissions should exist outside the model’s prose as well as inside it.
Being unable to finish a task within the authorized limits must remain an acceptable outcome.
Preserve interruption.
People need practical ways to stop consequential actions, inspect what happened, and reverse what can be reversed.
A model promising to stop is not the same thing as a tested stopping mechanism.
Human responsibility without the ability to inspect or intervene is not meaningful control.
Bring the answer back to what it represents.
Check claims against evidence.
Check interpretations against the people they concern when appropriate and safe.
Judge assistance partly by what happens after the conversation, not merely by how satisfying the conversation felt.
The person absent from the prompt does not become absent from the consequences.
These principles do not add up to a proof of safety.
That is part of their usefulness.
They are things we can observe.
Things we can test.
Things we can fail.
They also give education a purpose beyond teaching people to write clever prompts.
Using AI well means learning to distinguish observation from interpretation, a proposed explanation from evidence, representation from possession, and emotional relief from confirmation.
It means recognizing when the machine has answered the question you asked and when it has quietly written a better question on your behalf.
And it means preserving enough friction outside the conversation for reality to object.
Frank Herbert’s universe contains another warning here.
Long before the Face Dancers, humanity in Dune had already fought its war over “thinking machines.” Herbert leaves much of that history deliberately vague. What remains is a civilization deeply suspicious of making machines in the likeness of the human mind.
In God Emperor of Dune, the warning becomes stranger: the target of the old revolt was a “machine-attitude” as much as the machines themselves.
That feels more useful to me than the simple story where evil computers rose up and humans smashed them.
The danger was also the human temptation to hand judgment over.
The Face Dancer does not need to become me for me to let the face decide what the face deserves.
The model does not need to possess psyche for me to treat its representation of psyche as authority.
The mask can be useful.
It should not inherit the authority of the face merely because it fits.
I work in AI.
I want these systems to become much more capable.
I also want us to become much more capable of saying what that capability is for.
That returns me to satisfaction.
Not pleasure.
Not constant calm.
Not the absence of painful work.
A satisfying life, as I mean it here, is one I can continue to endorse while inhabiting it. One where regulation lets me participate more fully instead of merely increasing the amount of discomfort I can tolerate without changing anything.
Some of the people who stayed in Mormonism had learned to use the inherited story without granting every part of it authority.
I want to do something similar with AI.
If it helps me name a feeling, good.
If it helps me prepare a question for a professional, good.
If it helps me notice my part in an argument, even better.
If it helps me become extraordinarily articulate about why I do not need to change, I would like enough friction left in my life to notice.
This stage of my life does not contain weeks of uninterrupted writing time hidden behind a better morning routine. I have children, a marriage, work I care about, and a brain that keeps having ideas after the calendar has run out of room for them.
A tool that gives some of that time back is worth taking seriously.
So is the possibility that it can make me sound finished before I am finished thinking.
That is the bargain I am trying to understand rather than pretend I have solved.
I left Mormonism before I learned how differently other people could live inside it. I do not want to make the same intellectual mistake with AI—treating wholehearted adoption and total rejection as the only serious positions available.
I am still trying to be an attentive father without falling apart, and a caring husband who can also receive care.
Those were never preliminary concerns on the way to the philosophy.
They were the point.
The life in which an answer has to work contains people who can disagree with the account I have given of it.
They remember things I have edited out.
They have needs the prompt did not contain.
Their lives continue after mine becomes coherent on the screen.
Intelligence can inherit my story, infer what I mean, and extrapolate it farther than I intended.
What it cannot do for me is decide whether the resulting life is one I still want.
I want help regulating myself so I can return to them with more capacity.
Not a better inherited story explaining why I no longer have to.
The answer has to come back to them with me.