Most buyers of custom software assume the hard part of agentic coding is the writing. The writing is often the cheap part now. Much of the delivery pipeline downstream of it, including review, continuous integration, test suites, and staging, is suddenly carrying weight it wasn't sized for.

That shift is already visible to anyone shipping real software. An agent can open ten pull requests before lunch. The humans reviewing those PRs, and the CI pipeline running tests against them, cannot magically go ten times faster.

The bottleneck moved, and the budget, timeline, and risk profile of a custom build moved with it. Here is how the sequence plays out from the moment an agent picks up a ticket to the moment the code is in production.

Stage One: The Ticket Leaves the Human Backlog

The agent's day starts where the engineer's used to. It pulls an issue, drafts a plan, writes to a branch, runs local checks, and opens a pull request. In a look at agentic coding in custom software projects, Programming Insider frames this well: the prompt window has moved out of the editor and into the PR itself. The branch is the workspace, the diff is the deliverable.

That sounds like a small change in tooling, but it reshapes the whole delivery surface. The agent is now entering your codebase through the same door as a junior engineer, and every policy sitting on that door, including branch protection, required reviews, and status checks, is suddenly load-bearing in a way nobody stress-tested.

Stage Two: The Pull Request Volume Hits the Review Queue

Then the diffs start piling up. One engineer supervising a handful of agents can generate more pull requests in a day than a small team used to produce in a week. The review queue, which used to drain overnight, now fills faster than it empties.

Reviewers aren't the only ones under pressure. CI is too. Anthropic's engineering team wrote recently that CI jobs grew 25x over six months as Claude took over the bulk of their code authoring, forcing them to rebuild test impact analysis from scratch. That is where any team with real agent adoption lands within a quarter or two.

If you're commissioning custom software, the practical signal is simple: ask your vendor what their review queue looks like at the end of a heavy sprint. If the answer is a shrug, the agent isn't the risk. The queue is.

Stage Three: The Test Suite Earns or Loses Its Keep

This is where most pipelines break. A test suite that was tolerable at human throughput, flaky in a few spots and slow in a few others, becomes unworkable when it runs on every agent-authored branch all day.

Flakiness compounds, costs compound alongside it, and trust in the signal collapses soon after.

A few things usually have to change at once for the suite to keep pace with agent output:

  • Deterministic selection. Running the whole suite on every PR stops being affordable. Something has to decide which tests the diff actually affects, and that something has to be right.
  • Ephemeral environments. Shared staging becomes a traffic jam when agents and humans push in parallel. Each PR needs its own sandbox or the queue stalls.
  • A real contract for the agent. The agent needs to know which commands count as passing, which tests it is not allowed to delete or skip, and what 'done' means in exit-code terms, not vibes.
  • Zero tolerance for flakiness. When tests fail intermittently for reasons unrelated to the code, agents will happily re-run them until they pass. A build that goes green only after four retries is a laundered one, not a real signal.

What This Means When You Commission a Build

If you are the one writing the check, the questions worth asking a vendor have shifted. Headcount and velocity matter less than they used to. The pipeline matters more. A short list to run through before signing:

  1. Show me the review process. Who reads an agent-opened PR, how long does it take, and what are they allowed to merge without a second pair of eyes?
  2. Show me the test strategy. What percentage of the suite runs on every PR, how is flakiness tracked, and what happens when a test the agent wrote starts failing against code the agent wrote?
  3. Show me the environments. Can each PR spin up its own stack, or are a dozen branches fighting over one staging server?
  4. Show me the guardrails. Which parts of the codebase, including infrastructure config, auth, and billing, is the agent not allowed to touch unsupervised, and how is that enforced?

A vendor who can answer those four questions crisply has already done the hard work. One who talks mostly about how many agents they run has probably outsourced the problem to you without saying so. The agents are getting better fast. The pipeline around them is where the next year of custom software delivery will be won or lost.