AI/ML

·

·

8 min read

What if your CI pipeline ran itself?

Anthony Genova

Product Engineer, Juno Innovations

An engineer with a laptop crouched at an open server rack in a data center aisle

ON THIS PAGE

Most software teams have the same CI problem.

A pull request goes up. The pipeline runs. Something fails. A developer gets pinged, stops what they are doing, reads the logs, works out whether the failure is real or flaky, decides whether to retry or fix it, and loses twenty minutes of focus either way. Now multiply that by every PR, every day, across every developer on the team.

The pipeline is automated. The response to the pipeline still runs on people.

So we automated the response too. Agents watch the pipelines, triage failures, separate real failures from flaky tests, retry what should be retried, file issues for what needs a human, and ping the right person with enough context to act right away instead of starting from scratch.

Here is how it works and what it took to build.

One agent per pull request

The core setup is simple. When a pull request opens, the CI system spins up a dedicated agent for that PR. The agent has three jobs: watch the pipeline, respond to failures, and report status.

The agent hooks into the pipeline through a webhook. When a job finishes, pass or fail, the webhook fires and the agent gets the event. For a passing build, it logs the result and waits for the next event. For a failing build, it starts triage.

Triage is where the model earns its spot. The agent pulls the full build log, hands it to the local inference endpoint, and asks whether this is a deterministic failure or a flaky test. The model has seen enough CI logs to tell a real compilation error apart from a test that fails now and then because of timing or an external dependency.

For flaky failures, the agent triggers a retry and notes the flakiness in a PR comment. For deterministic failures, it pulls the root cause out of the log, writes it up as a structured report, files a GitHub issue with the context already filled in, and sends the PR author a summary that takes thirty seconds to read. Rebuilding the same picture from raw logs takes about five minutes.

What the agent does and what you still do

The human version goes like this. The notification arrives, you open the pipeline UI, you scroll the logs to find the failure, you decide if it is real, you decide what to do, and then you act. That is fifteen to forty minutes per failure depending on how gnarly it is, plus the cost of the context switch.

The agent version goes like this. The webhook fires, the agent reads the logs through the API, the model classifies the failure, and the agent either retries or opens a structured issue. The PR author gets a formatted summary with the root cause and a suggested fix in under two minutes. A human only gets involved when the agent escalates.

The agent does not replace the developer. It replaces the triage step between the raw failure and the developer’s judgment. You still decide how to fix the code. You just stop spending time figuring out what broke and whether it matters.

What runs underneath

Each PR agent runs as an Orion workload. When the PR opens, a webhook triggers a workflow that calls Orion’s API to launch a new agent workload from a template. The template spells out the container image, the resource allocation, the storage mount for logs and context, and the webhook endpoints the agent will use.

The agent keeps its state in persistent storage for the whole life of the PR. It can be interrupted, sit idle between pipeline events, and pick back up without losing context. When the PR closes, a second webhook fires and the agent workload shuts down. The compute goes back to the pool.

That lifecycle is what makes one agent per PR affordable. The agent is not a long-running process holding resources the whole time. It launches on demand, handles events as they come in, idles on very little in between, and shuts down when the job is done. The cost of each PR tracks the actual pipeline activity, not how many days the PR stays open.

The local inference endpoint handles classification and the write-ups. The model runs on shared hardware next to other workloads. The agent sends inference requests through a standard API and gets structured responses back. Nothing calls an external API, nothing gets billed per request, and no data leaves the infrastructure.

Where static rules run out

The usual alternative is rule-based automation. If the test name matches this pattern, retry. If the error message contains this string, file an issue. That works for failure modes somebody already wrote rules for.

It falls over on new failures, which is exactly when you need the help.

For example, an agent with a language model can read an error it has never seen, reason about the likely cause, and still produce a useful summary. When it is unsure, the summary gets less precise instead of disappearing. A rule-based system either matches or it does not.

Agents also get better over time in a way static rules do not. When a new failure mode keeps showing up, the team can add examples to the agent’s prompt so its answers get sharper. You can also tune prompts based on which summaries people actually found useful and which ones missed. The more you use it, the better it gets.

How this fits into the rest of our dev workflow

The CI agent is one piece of a bigger pattern we run in production. Agents take the operational overhead of software development so developers can focus on the work that needs a human.

The other agents in the same system: a code review agent that does a first pass before human reviewers, flagging obvious issues and summarizing what changed; a dependency monitoring agent that watches for security advisories and files PRs with updated versions; and a documentation agent that notices when code changes ship without matching doc updates and drafts the missing sections.

None of these agents make the final call on anything that matters. They handle the work that needs no judgment, surface the work that does, and hand over enough context that the judgment call takes thirty seconds instead of thirty minutes.

The infrastructure pattern is the same for all of them. They are Orion workloads launched on demand, sharing one inference endpoint, keeping state in persistent storage, driven by webhooks, and escalating to a human when they reach the edge of what they should decide alone.

This is the system we use every day to ship Orion itself.

See how Orion handles agentic development workloads. Get a Demo

Anthony Genova is a product engineer at Juno Innovations, building the automation stack and agent frameworks that run on top of Orion.

WRITTEN BY

Anthony Genova

Anthony Genova builds the things that make Orion work. As a product engineer at Juno Innovations, he has spent years writing the automation stack, agent frameworks, and infrastructure tooling that run on top of Kubernetes so researchers, developers, and creative teams do not have to think about what is underneath. Before Juno, he worked across software engineering and systems integration. He writes about agentic systems, workload automation, and the infrastructure patterns that make AI actually useful in production.