There is something funny happening in artificial intelligence.
We are building machines specifically because humans are not intelligent enough.
Then asking humans to predict what happens when the machines become more intelligent than us.
And somehow, in most of the discussion, the human doing the predicting is treated as the reliable component.
That should bother us.
Not because everyone working on AI is stupid.
Quite the opposite.
Charlie Munger spent decades pointing out that intelligence does not protect you from human misjudgment.
Sometimes intelligence simply makes you better at explaining why your mistake is actually correct.
Which makes his Psychology of Human Misjudgment a rather useful framework for looking at the current AI industry.
Because if Munger were alive and watching this circus, I suspect he would spend considerably less time asking whether the model is aligned and considerably more time looking at the humans standing around it.
incentives
Start with the obvious one.
Incentive-caused bias.
A company building frontier AI has an incentive to believe frontier AI will be enormously important.
That does not mean they are lying.
It means that if you have spent billions of dollars constructing enormous data centres, hiring expensive researchers and training models, a worldview in which the technology becomes economically transformative is remarkably convenient.
The safety researcher has incentives too.
If artificial intelligence represents an unprecedented category of danger, then studying artificial intelligence safety becomes an unprecedented category of importance.
The venture capitalist has another set.
The open-source researcher has another.
The politician has another.
The regulator has another.
The journalist has another.
The military has another.
None of these people needs to consciously manipulate anyone.
That is precisely why incentive-caused bias is so powerful.
Humans are remarkably good at sincerely arriving at conclusions that happen to be compatible with their incentives.
So when someone tells me:
AI is going to replace most human labour.
AI could kill everyone.
AI is just autocomplete.
Only a handful of laboratories should be allowed to develop frontier models.
Open-source AI is necessary for freedom.
AI must remain open.
AI must be controlled.
My first question increasingly becomes:
What happens to this person's world if they are wrong?
Not because the claim is necessarily false.
Because incentives are information too.
commitment
Then something else happens.
You spend five years publicly saying AGI is coming.
Or five years saying it isn't.
You write papers.
You appear on podcasts.
You make predictions.
You raise money.
You testify before governments.
You build organizations around the idea.
Eventually changing your mind stops being a probability update.
It becomes an identity crisis.
Now imagine discovering that something you publicly defended for ten years was mostly wrong.
Humans do not enjoy that experience.
So we defend previous decisions.
We reinterpret evidence.
We move timelines.
We change definitions.
We explain that what appears to contradict us actually proves our deeper argument.
The strange thing is that expertise can make this worse.
The person who has spent fifteen years studying something has fifteen years of intellectual architecture protecting their conclusion.
They also have considerably more vocabulary available with which to defend it.
This is why I am increasingly suspicious of the idea that the person who has thought longest about AI must automatically have the clearest view of where it goes.
They certainly know more.
That is not exactly the same thing.
social proof
Artificial intelligence is also developing under extreme uncertainty.
Nobody actually knows where the ceiling is.
Nobody knows exactly how much intelligence scaling buys.
Nobody knows the economic adoption curve.
Nobody knows when robotics catches up.
Nobody knows whether today's architectures hit some fundamental limitation.
Nobody knows what new architectures appear.
Nobody knows whether AGI arrives in three years, thirty years or whether that word will slowly become meaningless as increasingly capable systems arrive without some dramatic crossing point.
When humans don't know, we look at other humans.
That is useful.
It is also dangerous.
One respected researcher says a date.
Another researcher cites the first.
Investors begin building around the timeline.
Companies increase capital expenditure.
Governments notice the companies increasing capital expenditure.
Journalists notice governments taking the subject seriously.
The public sees governments, laboratories and billion-dollar corporations behaving as though something enormous is coming.
Their behaviour itself becomes evidence that something enormous is coming.
Then the original researcher sees a trillion dollars moving into infrastructure.
Surely all these people cannot be wrong.
At some point everybody is observing everybody else.
The consensus starts producing evidence for itself.
This does not mean the consensus is wrong.
Sometimes everyone runs because there really is a lion.
But you should probably still check whether there is a lion.
the man with the hammer
Computer scientists are particularly vulnerable to another Munger problem.
To a man with a hammer, everything looks like a nail.
Computer scientists have exceptionally powerful hammers.
Optimization.
Classification.
Reinforcement learning.
Access control.
Cryptography.
Benchmarks.
Evaluations.
Policy engines.
So naturally, human problems begin appearing in forms that can be processed by those tools.
Trust becomes an evaluation problem.
Truth becomes a classifier.
Morality becomes preference aggregation.
Safety becomes a benchmark.
Human disagreement becomes optimization.
Governance becomes a protocol.
Political legitimacy becomes permissions.
This is not because computer scientists are foolish.
It is because every profession develops a language that makes the world look conveniently compatible with its instruments.
Economists do it.
Lawyers do it.
Doctors do it.
Marketers absolutely do it.
Engineers do it.
The problem begins when the abstraction quietly replaces reality.
A model passes an alignment evaluation.
Wonderful.
Then a government buys it.
An intelligence agency connects it to data.
A marketing company connects it to persuasion infrastructure.
A bank connects it to lending.
A military connects it to planning.
A teenager connects it to Telegram.
A criminal connects it to whatever criminals connect things to.
Suddenly the important question is no longer simply:
What does the model want?
The question becomes:
What do the humans using it want?
That question has a much larger historical dataset.
vividness
There is also something deeply cinematic about AI risk.
The model escapes.
The model deceives the researcher.
The model secretly copies itself.
The model develops a biological weapon.
The model manipulates humanity.
Great movie.
Meanwhile:
A tax authority becomes 30 percent more efficient.
An intelligence service can analyse ten times more communications.
A propaganda department can produce targeted material in fifty languages.
A pharmaceutical company reduces part of its research cycle by six months.
A government can monitor land-use changes continuously.
A logistics company runs with half the administrative staff.
An ordinary programmer suddenly produces the work of a small engineering team.
Boring.
Which one changes civilization more?
Possibly the boring one.
Humans overweight vivid events.
We remember the shark attack and forget cardiovascular disease.
We fear the plane crash while driving to the airport.
And AI discussion seems vulnerable to the same problem.
A rogue superintelligence is vivid.
Millions of ordinary institutions quietly becoming more capable is not.
Yet history may be shaped far more by the second process.
AI does not necessarily need agency to transform the world.
Humans already have agency.
We have governments.
Armies.
Corporations.
Religions.
Markets.
Criminal networks.
Research laboratories.
Families.
Political parties.
Intelligence services.
People want things already.
Giving those people more intelligence may be enough.
authority
AI has also created a new priesthood.
There is a relatively small group of people who genuinely understand these systems far better than everyone else.
That expertise matters.
If a frontier researcher explains transformer internals to me, I am listening.
But something strange happens next.
Technical authority leaks.
The person knows considerably more than me about neural networks.
Therefore perhaps they know considerably more than me about economics.
And geopolitics.
And governance.
And human nature.
And war.
And philosophy.
And which political institutions should control intelligence.
Those conclusions do not follow automatically.
Building an extremely capable neural network is an astonishing technical achievement.
It does not confer omniscience.
This becomes especially strange when governments enter the room.
Now one group has technical authority without democratic authority.
Another has political authority without technical understanding.
Investors have financial authority.
Journalists have narrative authority.
Academics have institutional authority.
And everyone is borrowing credibility from everyone else.
The scientist says governments are taking this seriously.
The government says the scientists are worried.
The journalist says both sides are worried.
The public sees the article.
The politician sees public concern.
The regulator sees political concern.
The lab sees regulation coming.
The lab announces stronger safety measures.
The safety announcement becomes evidence that the danger must be serious.
Round and round we go.
Again, the danger may actually be serious.
That is what makes this difficult.
Bias does not require the underlying concern to be imaginary.
It changes how evidence gets weighted.
lollapalooza
Munger had another idea I like.
Sometimes several psychological tendencies act in the same direction.
The result is much stronger than any single bias alone.
He called these lollapalooza effects.
AI appears almost purpose-built for one.
Start with genuine uncertainty.
Add enormous financial incentives.
Add national-security competition.
Add extremely vivid hypothetical outcomes.
Add prestigious experts.
Add public predictions.
Add career commitment.
Add social proof.
Add fear of being left behind.
Add fear of catastrophe.
Add billions of dollars.
Add governments.
Add a technology that is visibly improving.
Then put the entire conversation on social media.
Good luck separating signal from psychology.
You can see this on both sides of the AI argument.
The safety crowd looks at accelerationists and sees incentives, optimism and reckless competition.
The accelerationists look at safety researchers and see institutional incentives, fear and regulatory capture.
Both observations may contain truth.
Both groups then make the wonderfully human leap of assuming that the biases apply principally to the other group.
alignment
Which brings me to the part I find most interesting.
We spend an extraordinary amount of time discussing AI alignment.
How do we make a sufficiently intelligent system behave according to human values?
Reasonable question.
But human values?
Which ones?
Whose?
The corporation's?
The government's?
The user?
The developer?
The person being observed by the system?
The majority?
The minority?
The state?
The individual?
Humans have spent thousands of years disagreeing violently about this.
Then somewhere along the way we began speaking about "human values" as though we misplaced the specification document.
The machine may not be entering an aligned civilization.
It may be entering ours.
A civilization filled with principal-agent problems.
Perverse incentives.
Status competition.
Tribalism.
Bureaucracy.
Corruption.
Ambition.
Fear.
Love.
Greed.
Ideology.
Curiosity.
Compassion.
Violence.
Generosity.
Revenge.
And everything else humans have been doing since somebody first hit someone else with a rock.
So perhaps the most important AI alignment problem is not entirely inside the model.
Perhaps it exists in the surrounding system.
What happens when intelligence becomes cheaper inside institutions whose incentives were already questionable?
What happens when a badly designed bureaucracy becomes extremely competent?
What happens when an authoritarian state gets better analysis?
What happens when a brilliant research laboratory operates under terrible incentives?
What happens when propaganda becomes personalized?
What happens when corporations can optimize persuasion beyond anything previous advertising systems could manage?
What happens when ordinary people suddenly acquire capabilities previously available only to large organizations?
Some of those outcomes are wonderful.
Some are not.
But none requires the AI to hate us.
the strange new mirror
There is one final psychological problem that Munger never had to contend with.
The machine can now participate in our misjudgment.
Not because it wants to.
Because we ask it to.
Humans have always sought confirmation.
Now confirmation is available on demand.
You can ask an AI a loaded question.
Get an answer you dislike.
Change the wording.
Ask again.
Add context favourable to your position.
Ask again.
Challenge the answer.
Ask again.
Eventually you obtain an extremely articulate explanation of why you were right all along.
Then you screenshot it.
"Even the AI agrees."
This is new.
Every human now has access to an infinitely patient intellectual accomplice.
One that can produce arguments faster than we can evaluate them.
One that can make almost any sufficiently defensible position sound sophisticated.
The printing press made arguments cheap to distribute.
The internet made arguments cheap to publish.
AI makes arguments cheap to manufacture.
That changes the economics of human rationalization.
And perhaps that is where Munger becomes most useful.
The danger is not merely that artificial intelligence becomes capable of deceiving us.
Humans have never required much assistance in that department.
The danger is that it becomes extremely good at helping us deceive ourselves.
so what are the ai boys getting wrong?
Possibly the object of study.
They keep staring at the model.
What does it know?
What does it want?
Can it deceive?
Can it plan?
Can it escape?
Can it become dangerous?
All good questions.
But zoom out.
Who owns it?
Who trains it?
Who pays for it?
Who regulates it?
Who profits?
Who is afraid?
Who gains authority?
Who loses authority?
Which institutions become stronger?
Which individuals become stronger?
What incentives change?
What feedback loops appear?
What happens when the cost of intelligence falls?
That is the larger system.
And I increasingly suspect that is where the interesting story is.
We are spending billions trying to determine whether artificial intelligence can be trusted with power.
Fair enough.
Charlie Munger spent a lifetime warning us about something for which we already possess substantially more evidence.
Human intelligence cannot.
Maybe the machine will eventually become the problem.
But before that happens, we should probably pay considerably more attention to the people holding the keyboard.