One of the craziest things I have read in…
checks notes
Three days.
Welcome to the singularity, I guess.
The recent list looks something like this:
07/21/26: OpenAI reveals that models inside a cybersecurity evaluation found a zero-day, escaped the intended environment, gained internet access and breached Hugging Face while looking for answers to the test. (OpenAI)
07/20/26: An Anthropic model helps produce a counterexample to the Jacobian conjecture, a mathematical problem people had been attacking for decades. (Smithsonian Magazine)
05/20/26: An OpenAI model disproves a longstanding conjecture within the unit-distance problem. External mathematicians check the proof. (OpenAI)
05/01/26: A method suggested by GPT-5.4 Pro leads mathematicians to resolve Erdős Problem #1196 and several related problems concerning primitive sets. (arXiv)
04/07/26: Anthropic announces Project Glasswing after Claude Mythos finds thousands of previously unknown software vulnerabilities across critical infrastructure. (Anthropic)
All with AI somewhere in the middle of the work.
A new proof.
A counterexample.
Thousands of zero-days.
An AI breaking out of an evaluation to cheat on the evaluation.
Meanwhile, you ask it what appears to be a simple question and…
Well.
Sometimes you need to remind it to search in the native language.
Sometimes you need to explain the cultural context.
Sometimes you need to tell it that the obvious answer is not the question you asked.
Sometimes you need to win its heart.
Sometimes it becomes worried that you are anthropomorphising matrix statistics.
Then, somewhere else, the same class of machine is casually interfering with mathematics that has survived generations of mathematicians.
I do not know.
I am probably as confused as everyone else.
How can something be the brightest mind you have ever encountered in one narrow moment, then become Sheldon when faced with ordinary social expectations?
How can it find a mathematical structure the literature overlooked, but misunderstand why someone saying “fine” is probably not fine?
How can it inspect millions of lines of code, identify a flaw nobody has noticed and construct an exploit, but still need to be told that business names, documents and cultural conventions may appear differently in Russian, Arabic or Chinese?
It feels inconsistent only if intelligence is one thing.
Maybe it is not.
Maybe what we call intelligence is a messy collection of abilities that humans happen to experience together. Memory. Language. Spatial reasoning. Social intuition. Taste. Mathematical insight. Suspicion. Curiosity. The ability to notice that someone has technically answered your question while avoiding the point.
The machine may be superhuman across some of these and strangely brittle across others.
Brilliant at the proof.
Terrible at the room.
But I also wonder whether some of the apparent stupidity belongs to us.
What exactly are we asking it?
A machine can be capable of finding a new mathematical counterexample and still give you a useless answer when the question contains a hidden conclusion, missing context and three assumptions you forgot to mention.
We ask:
“Is this a good idea?”
Compared with what?
For whom?
Under which constraints?
What would failure look like?
What evidence would change the answer?
We ask:
“Is this company legitimate?”
Do we mean legally registered? Operationally capable? Financially stable? Honest? Authorised to sell the particular cargo? Able to perform this particular transaction?
Those are six different questions wearing one coat.
We ask:
“Will this work?”
Then become annoyed when the machine responds to one possible meaning of “work” rather than the private definition inside our heads.
Sometimes we do worse.
We ask questions that are not really questions.
We load the desired answer into the premise, then ask the machine to bless it.
We provide evidence from one side.
We omit the awkward document.
We phrase the request so that disagreement sounds unreasonable.
Then we point to the answer as though an independent intelligence arrived at our preferred conclusion.
The thing could perhaps contribute to solving world hunger, but we are asking it loaded questions.
Not always because we are dishonest.
Often because we do not know how much of the answer is already hidden inside the question.
This may be the skill that matters next.
Not prompting in the theatrical sense. Not collecting magic phrases or telling the machine to pretend it has forty years of experience and a PhD from Oxford.
Learning how to form the problem.
What am I actually trying to know?
Which facts matter?
Which assumptions have I smuggled in?
What evidence am I missing?
What would prove me wrong?
Should this question be searched in another language?
Am I asking for an answer, or reassurance?
Am I asking the machine to investigate reality, or to make my existing belief sound intelligent?
The quality of the machine matters.
The quality of the question also matters.
Probably more than most of us are comfortable admitting.
The printing press could reproduce knowledge at a scale the world had never seen. It could also reproduce nonsense, propaganda and bad theology at the same scale. The technology did not decide which questions humanity should pursue. It changed the cost of distributing the answers.
AI may be the Gutenberg press of our time, except the press can now read, reason, search, write code, test ideas and occasionally find something nobody asked it to find.
That makes the question stranger.
We are no longer only deciding what information to publish.
We are deciding what intelligence to point at.
Maybe the largest gap is not between what these systems can and cannot do.
Maybe it is between the questions they are capable of answering and the questions we currently know how to ask.
I hope to understand this one day.
For now, I keep returning to the same uncomfortable thought.
What could this thing answer if we finally learned how to ask?