i am trying to build software.
this appears to come with a second job: keeping up with everyone explaining what i should or should not be able to build.
some developers say ai cannot manage a full production system. others say it is fully capable. i keep recognising parts of my own experience in both arguments, which is deeply inconvenient when the internet would prefer that i choose a side.
i am somewhere in the middle, with a rather boring question.
have you updated your developer docs?
because my agent needs to check whether what it remembers about your software agrees with what you have actually developed. which version are we using? what changed? can it find the current instructions?
if the documentation, the installed version and the model’s memory are describing three slightly different things, i would like to sort that out before interpreting the result.
maybe it still fails. but at least we will have a clearer idea of what failed.
this week, i have also started noticing senior developers saying something i wish came attached to more of these conversations: the prompt you copied is being used inside a whole setup.
without their harness, their tools, their instructions, their context and their feedback, the same words may produce very different work.
that makes sense to me.
i can copy the sentence you typed. reproducing the working arrangement around that sentence is another job. apparently i am collecting those.
it is also why i keep wanting to ask for the full transcript.
“claude says.”
“grok says.”
“astra says.”
okay. could we please see the conversation?
what did you ask? what had you already told it? what could it read? where did you challenge it? what did it change its mind about? what did you change your mind about?
because mine and i can disagree for hours. sometimes we end up agreeing to disagree.
there is a method to that madness. i want it to account for its claims, and i have to be willing to account for mine. sometimes that means going back to the documentation. sometimes it means inspecting the code. sometimes it means admitting that we have spent a very long time discussing something we still have not established.
agreement is pleasant. i still want something i can check.
a final answer can conceal a lot of that work. so can a screenshot of a spectacular failure. i want to know how we arrived there, because i am trying to work out what the result tells me.
and that is the question nagging me tonight.
when someone says ai cannot do this, or that, what exactly are they saying?
is it a skill issue? a harness issue? a money issue? an actual model limitation? am i getting carried away? is it ai psychosis? 😂
i am including myself in the questions. i have had enough impressive results and enough “what the fuck is this?” moments to be suspicious of my own certainty.
then there is the budget.
if you are telling me you exhausted the limits on a $500 plan in thirty minutes using astra, no fucking way am i casually entering that argument without a few more questions.
i hope you are solving cancer. 😂
how much work was happening? how many sessions? how many attempts? what were you asking it to do?
i would like to know what experience we are comparing before either of us delivers a verdict on the technology.
then i see nate berkopec describing matz merging 90 pull requests an hour.1
ninety.
elsewhere, i read developers saying they need to inspect and approve every pr themselves. someone else describes using a base model and finding the whole thing underwhelming.
i can imagine all of these people accurately describing their experience. i still have to work out what applies to mine.
how big are the changes? what happened before they reached the person merging them? which checks have already run? how much work has gone into making that volume manageable?
the number gets my attention. the process would help me understand it.
and i cannot pretend my own process has stayed still.
a few months ago, i swore by ralph loops.
my version was built around specifications and sprints. small, atomic tasks. something you could commit as a coherent change. tests or explicit validation. every sprint ending in software you could run and demonstrate.
an iteration picked up a defined piece of work. the next could start with fresh context, while the files and git carried the work forward.
i wanted to be able to point at what had been done, how it had been checked, and what i could actually use.
now, some of my sessions run for ten or twelve hours.
there is monitoring. there is evaluation. i am still looking at the work and steering it. occasionally the steer is simply:
“wtf.”
usually when internal instructions start appearing in the interface, as though the person using the app needs to see the agent’s homework. 😂
or when the session becomes very interested in satisfying the tests and i have to bring its attention back to what the person using the screen is supposed to accomplish.
i am building dashboards at the moment. that has made the question very immediate. i can look at a screen and ask whether the job makes sense there. what does the person see? what can they do? what happens next?
beautiful code still needs to become something a human can use.
ten hours tells you how long the session ran. i still have to look at what those hours produced.
this is part of why i am hesitant about announcing the correct way to work with these things. my own way keeps changing.
i have moved from insisting on tightly bounded iterations to allowing some sessions much longer runs, with supervision and correction continuing along the way. i am learning what i can leave running, what needs a closer look, and where i need to intervene.
if you had asked me a few months ago, i would have described the setup i had then. ask me now and you get a different answer.
both would be honest.
how much of this argument is people speaking from different stages of that process?
different models. different tools. different work. different budgets. different expectations of what “done” means.
i think a lot of what sounds like a universal verdict is someone describing an experience without enough of the conditions that produced it.
i do not know how to separate the signal from the noise when those conditions are missing.
show me what you were building. show me what the agent could access. show me where you had to step in. show me what failed, what you changed, and how you decided the result was ready.
then i have something i can learn from.
and while i am trying to make sense of all this, ci is still running.
i have had checks take about an hour. that is a substantial amount of time to sit with your allegedly revolutionary development speed.
then i read anthropic’s article on agentic coding and ci.
in september, they reported that their ci job volume had increased 25-fold over six months. they already had a service choosing which tests to run using package relevance and previous results. that service itself struggled to keep up. eventually they rebuilt it so workers could process results against shared state and scale more easily.
even the machinery deciding what to check needed more work.
that is the part i want to hear more about. once you have agents producing changes, a way of reviewing them and some confidence in your setup, how do you keep the checking machinery from consuming the day?
i am seeing people change when they run tests, reconsider which tests need to run, and look for more compute. i am trying to understand the decisions underneath those changes.
what runs on every change? what waits? what gets batched? how do you know the checks still tell you what you need to know?
so yes. i am thinking out loud. i can see why people arrive at very different conclusions, and i am still trying to understand my own.
but i do have one practical question.
how have you solved ci?
preferably with the setup. i have enough conclusions open in other tabs. 😂
references
- Nate Berkopec, X post dated 1 October 2026, quoting Yukihiro Matz. Source: the author’s supplied screenshot, IMG_7498.png. The figure describes merging pull requests and is attributed to Berkopec’s account. ↩