The cost of an agent is starting to show up somewhere other than the model bill: the machinery that checks, builds and merges its work.
On 2 October, developer Jan Wilmake reported that his team's GitHub CI bill had climbed from zero to $380 a month, with an $800 month on the horizon. His explanation was straightforward: coding agents were pushing roughly 300 times a day, while the test suite had grown fivefold since August.
Those numbers describe one team's experience. The mechanism is the part that travels: more pushes multiplied by more work per run creates a larger bill, even when the price of a runner stays the same.
Source: Jan Wilmake, 2 October. The $800 figure was a projection.
Wilmake said his response was to stop running CI on pull requests, keep development builds and staging deploys, and move the full suite to a nightly run. That changes when defects are discovered as well as what the pipeline costs. It is a case study in choosing the frequency of expensive checks, not evidence that every project should remove its merge gates.
Another developer, Kush Bhuwalka, described a rising CI/CD bill even after moving to Blacksmith. Changing the runner supplier and changing how often work is scheduled are two different interventions. Agent throughput makes that distinction harder to ignore.
Cheaper runners are only one lever
The two reports point to a useful distinction in the economics of agent work. A cheaper runner reduces the cost of a job. A smaller job reduces the work performed. Running fewer jobs reduces the frequency. Those choices can reinforce each other, but none tells a team which defects it can afford to discover later.
Consider a documentation edit and a change to a payment flow. Giving both the same expensive pipeline makes scheduling simple, but the risks are different. GitHub Actions supports branch and path filters, which let a team decide which changes should trigger a workflow. The hard part is deciding whether those boundaries match the application's dependencies.
Moving a full suite to the night creates another tradeoff. By morning, several changes may have landed since the last successful run. The saving needs to be weighed against the time spent finding the change that broke it. Fast feedback on small changes and broad checks on the integrated system each have a job to do.
This is our reading of the reports: the useful question is what each check buys at each point in the workflow. A runner invoice measures compute. It does not measure the cost of a delayed failure.
The queue after the tests
The pressure does not end with the build. On 1 October, GitHub's Sameen Karim announced an async merge API, aimed explicitly at developers and agents merging pull requests programmatically.
The significance is architectural: once software is doing the submitting, the merge path needs to behave like a dependable service. The handoff, waiting and completion matter as much as the request that starts it. Faster generation makes the seams between tools more visible.
A separate checkout is not a complete workspace
Parallel agents bring a second kind of plumbing problem. Git worktrees give each run its own checkout, keeping changes apart. But a tracked checkout does not bring along ignored dependencies or local environment files. John Rood's reminder is the unglamorous detail that determines whether the next agent can actually run the app.
Git's documentation describes worktrees as a way to check out more than one branch of a repository at once. That solves the problem of competing edits. It does not, by itself, make each checkout a running application.
A team can give three agents separate directories and still have all three waiting on dependency installation, missing configuration or the same local service. Preparing the workspace becomes part of the throughput calculation, just as testing and merging do.
The work is moving from “can the agent write this?” to “can the whole system absorb the work?” Test frequency, merge handling and workspace setup are different parts of that same question. The next productivity gain may come from making those handoffs predictable.