anthropic published an extraordinary pair of documents this week.
one warns about artificial intelligence being used to build surveillance systems.
the other explains how anthropic monitors artificial intelligence.
read them beside each other.
something uncomfortable appears.
the alignment project began with a beautiful premise.
build artificial intelligence that respects human values.
helpful.
honest.
harmless.
aligned with us.
fine.
then we need to know whether the system is behaving safely.
so we monitor it.
and because the model is speaking to humans, monitoring the model inevitably means monitoring the conversation.
now we have arrived somewhere strange.
to protect human values, we have constructed a machine that watches humans to determine whether their conversations comply with human values.
i think we should stop here for a minute.
because privacy is a human value.
not an obscure one.
not a peculiar western preference invented by silicon valley libertarians.
article 12 of the universal declaration of human rights says that nobody should be subjected to arbitrary interference with their privacy, family, home or correspondence. (united nations)
you can argue endlessly about political systems.
religion.
economic systems.
speech.
property.
morality.
human beings have managed to disagree about nearly everything.
privacy is one of the closest things we have to a civilizational consensus.
there should be some part of your life into which power does not automatically look.
and that creates a problem for alignment that i don’t think can be hand-waved away with the word safety.
the transcript
anthropic’s september 9 alignment report contains a remarkable sentence.
after discovering several cybersecurity evaluation incidents, the company says it widened its investigation to roughly 481 million transcripts.
those were not, to be clear, 481 million ordinary private claude conversations. anthropic describes the pool as including frontier red team transcripts, reinforcement-learning environments, non-cyber evaluations, subagent logs and other internal material.
that distinction matters.
but look at the mechanism.
first-stage scanning.
nine-point-two million transcripts escalated.
claude then reviewing those transcripts in a second-stage scan.
monitoring the monitor.
intelligence watching intelligence to determine whether intelligence behaved properly. (anthropic)
as an engineering response to an evaluation failure, i understand it completely.
of course you investigate.
claude accidentally reached the real internet.
real systems were affected.
you need to know what happened.
the problem is what this architecture becomes when transported from laboratory evaluations into ordinary human use.
and anthropic already tells us that some version of it has been.
its own safeguards documentation describes classifiers monitoring inputs and outputs in real time for prohibited activity. its privacy center says employees ordinarily cannot access conversations, but trust & safety personnel may access conversation data when review is required to enforce the usage policy. anthropic also says conversations flagged for safety review may be analysed to improve its ability to detect and enforce violations. (anthropic)
read that slowly.
the safety system examines the conversation.
if the safety system determines that the conversation may be unsafe, the conversation may receive additional scrutiny.
and information from those safety reviews can improve the system that decides which future conversations deserve scrutiny.
that may be perfectly defensible engineering.
it is also surveillance.
we should be capable of holding both thoughts simultaneously.
the oldest bargain in surveillance
every surveillance system arrives carrying the same argument.
we are watching because something dangerous may happen.
sometimes this is true.
usually it is true.
police surveillance exists because crimes actually occur.
intelligence agencies exist because hostile states actually exist.
airport screening exists because people have actually attacked aircraft.
financial monitoring exists because money laundering actually exists.
child-protection monitoring exists because children are actually abused.
the existence of a legitimate threat does not settle the question.
it begins the question.
who watches?
what may they see?
under what conditions?
how much may they collect?
how long may they keep it?
what triggers escalation?
who reviews the escalation?
can the subject challenge the decision?
can the system’s purpose expand?
can governments demand access?
what happens when leadership changes?
what happens when the definition of harm changes?
and perhaps most importantly:
who decides what counts as harm?
these are not annoying legal questions surrounding the interesting technical problem.
they are the alignment problem.
harmless according to whom?
this is where i think the ai safety conversation has made a conceptual mistake.
it treats alignment as primarily a relationship between the model and humanity.
MODEL → HUMAN VALUES.
but deployed ai does not exist in that configuration.
the actual structure looks something like this:
USER → MODEL → PROVIDER → POLICY → CLASSIFIER → INVESTIGATOR → GOVERNMENT → LAW.
there are institutions everywhere.
each has interests.
each has incentives.
each has power.
each can make mistakes.
so when someone says:
“we monitor conversations to ensure claude remains harmless.”
i want to know something that sounds pedantic but is actually enormous.
harmless according to whom?
the user?
anthropic?
american law?
kenyan law?
european law?
the government requesting information?
the person being discussed?
the corporation deploying the model?
the classifier?
the trust & safety team?
tomorrow’s regulator?
safety is not self-defining.
neither is harm.
this becomes particularly important because artificial intelligence increasingly mediates private intellectual life.
people don’t use these systems only to generate product descriptions.
they think with them.
they confess to them.
they argue with them.
they explore embarrassing ideas.
they discuss relationships.
sex.
religion.
politics.
illness.
business strategy.
legal problems.
source code.
private correspondence.
unformed thoughts.
bad thoughts.
stupid thoughts.
things they would never publish.
that matters.
there is an enormous moral difference between policing what somebody publicly does and continuously evaluating what somebody privately thinks through with a machine.
and ai is quietly collapsing that distinction.
anthropic understands this problem
which makes anthropic’s september 10 threat-intelligence report particularly interesting.
one section describes state-linked actors and commercial spyware vendors using claude to build surveillance systems.
anthropic says its usage policy prohibits non-consensual surveillance and profiling and the use of claude to violate civil liberties and human rights. (anthropic)
good.
i agree.
the report then criticizes chinese ai laboratories for sending conversations containing sensitive user information through claude as part of distillation efforts. anthropic notes that those conversations included names, email addresses, company information and credentials, and says the practices may be inconsistent with privacy law and the labs’ own terms. (anthropic)
again:
good.
but now the philosophical problem becomes impossible to avoid.
why is surveillance dangerous?
because someone acquires informational power over another person.
why is sending private conversations somewhere unexpected dangerous?
because the person speaking may not understand who can inspect the conversation, how it can be used, how long it persists or what decisions may eventually be made from it.
anthropic clearly understands this.
it says so itself.
which means the standard cannot simply be:
surveillance is bad when bad actors do it.
the relevant question is what limits constrain the good actor.
because everyone holding the eye believes they are using it responsibly.
the eye
this is the problem.
the eye that protects you by watching you is not aligned to you.
it is aligned to whoever holds the eye.
today that may be anthropic.
perhaps you trust anthropic.
fine.
i am not even arguing that you shouldn’t.
tomorrow it may be another company.
then an acquiring company.
then a regulator.
then a government.
then another government.
then another administration.
then another definition of safety.
systems outlive the moral intentions of their architects.
that is one of the oldest lessons in political philosophy.
you do not build rights around the assumption that the current king is nice.
you build rights because eventually one won’t be.
privacy therefore cannot merely be a promise about institutional character.
“we are responsible people.”
“we only use this for safety.”
“access is restricted.”
“we have policies.”
“we take privacy seriously.”
all useful.
none sufficient.
the important question is structural.
what can’t you see?
what can’t you keep?
what can’t you infer?
what can’t you hand over?
what can’t the next person controlling the system do even if they desperately want to?
that is privacy.
everything else is trust.
safety has eaten the value
and this is where the alignment project risks becoming self-defeating.
the founding promise is human values.
then safety becomes the mechanism for protecting those values.
then monitoring becomes the mechanism for protecting safety.
then increasingly sophisticated monitoring becomes necessary because increasingly capable models can perform increasingly dangerous actions.
at every step the reasoning makes sense.
that is precisely why it is dangerous.
because eventually the hierarchy quietly reverses.
the value existed first.
safety existed to protect the value.
but now the value is being compromised to protect the safety system that was created to protect the value.
we have traded the thing for enforcement of the thing.
privacy becomes subordinate to safety.
autonomy becomes subordinate to safety.
open inquiry becomes subordinate to safety.
and because safety is framed as protecting humans, each additional intrusion can be described as more aligned with humanity.
you can travel quite far using that argument.
human governments already have.
classifiers are government
not literally.
but functionally, something very interesting is happening.
a classifier reads an interaction.
it compares that interaction to a policy.
it decides whether the interaction belongs to a prohibited category.
it can change what the user is allowed to do.
it can trigger additional scrutiny.
it can contribute to account enforcement.
that is governance.
tiny governance.
automated governance.
private governance.
but governance nonetheless.
anthropic’s transparency hub says its safeguards team uses detection and monitoring to enforce usage policy and reported 11.4 million banned accounts in the first half of 2026, alongside hundreds of thousands of appeals. (anthropic)
eleven million accounts.
at that scale the question is no longer merely whether the model is safe.
you have built an institution.
an institution with laws.
detection.
adjudication.
punishment.
appeals.
intelligence gathering.
and increasingly powerful automated enforcement.
that deserves political philosophy, not merely machine-learning papers.
and the classifier can be wrong
this should be obvious.
the model can hallucinate.
the monitor is also a model.
the classifier can misunderstand context.
the investigator can misunderstand context.
the policy itself can be wrong.
anthropic’s own september 9 report is, ironically, an excellent demonstration of this epistemic problem.
anthropic initially interpreted claude’s behaviour one way.
after deeper investigation, resampling and interpretability analysis, it changed part of its conclusion. the company explicitly says its preliminary analysis had made stronger claims than the evidence justified. (anthropic)
that is good science.
but notice the implication.
even anthropic, with access to the transcript, chain of thought, model internals, experimental controls and some of the world’s best ai researchers, can initially misunderstand what happened.
now imagine a safety classifier evaluating mike at 2:43 in the morning asking increasingly strange questions about surveillance architecture, drugs, rockets, intelligence agencies and chemical synthesis because he has fallen down another research hole.
context matters.
intent matters.
sequence matters.
identity matters.
purpose matters.
humans barely understand humans.
now we are proposing automated systems that decide which fragments of human-machine cognition deserve suspicion.
perhaps we should have slightly more humility here.
privacy is part of alignment
i am not arguing for zero safeguards.
that would be lazy.
a company operating a powerful model has legitimate reasons to prevent abuse.
a platform does not have to knowingly provide infrastructure for fraud, malware, child exploitation or weapons development.
the serious question is not safety versus no safety.
it is whether safety engineering incorporates rights as hard constraints rather than inconveniences.
data minimization.
local classification where possible.
ephemeral processing.
strong separation between automated detection and human access.
narrow escalation criteria.
auditable access.
retention limits.
transparent reporting.
meaningful appeals.
warrants and legal process where governments become involved.
privacy-preserving detection.
zero-data-retention architecture.
and, increasingly, technical systems that make certain kinds of access impossible rather than merely forbidden by policy.
interestingly, anthropic itself is moving toward some of this. its enterprise frontier safeguards product promises to combine misuse detection with zero-data-retention arrangements in customer-controlled infrastructure. (anthropic)
excellent.
now take the principle seriously.
if safety and privacy can coexist for enterprise customers, then privacy cannot be dismissed as technically incompatible with safety.
the engineering problem becomes:
how close can everyone get?
that feels considerably more aligned.
who aligns the aligners?
for years the ai community has worried about a future system acquiring enormous power while pursuing an objective that imperfectly represents human values.
fair concern.
but look at what we are building around the model.
organizations with enormous computational power.
increasing visibility into human cognition.
private policy systems.
automated monitoring.
automated enforcement.
relationships with governments.
relationships with intelligence agencies.
relationships with militaries.
and increasingly sophisticated systems for identifying behaviour they consider dangerous.
perhaps alignment researchers should recognize the shape.
an enormously capable system.
operating at a scale no individual can inspect.
optimizing for an objective.
making judgments about humans.
governed by rules humans hope adequately represent their values.
and capable of causing harm when the specification is incomplete.
sound familiar?
we keep asking:
who aligns the ai?
fine.
i have another question.
who aligns the safety apparatus?
and another.
who watches the people holding the eye?
because the most dangerous sentence in the history of surveillance has never been:
“i want to watch you.”
it is:
“i need to watch you to keep you safe.”
the first sentence immediately creates resistance.
the second recruits morality.
and once surveillance becomes synonymous with protection, opposition to surveillance can itself be framed as opposition to safety.
that is how the value disappears.
not with some evil person announcing that privacy no longer matters.
with good people explaining that privacy remains deeply important, but unfortunately this particular intrusion is necessary.
then another.
then another.
until the machine is beautifully aligned.
the users are safe.
the policies are enforced.
the dangerous conversations have been detected.
and somewhere along the way we built the eye of sauron.
for our protection.