- Vishakha Sadhwani
- Posts
- System Design Interview: Design the CI/CD for an AI Agent | Part 2
System Design Interview: Design the CI/CD for an AI Agent | Part 2
The second of three questions worth preparing for ~
Hi Inner Circle!
This is part 2 of the system design questions I think you should be ready for this year.
Design the CI/CD for an AI agent.
You already know what an agent is ~ an LLM with tools, working through a task in a loop. That part gets explained constantly.
What doesn't get explained is how one gets deployed and maintained. That's what the interviewer is actually testing, and the one which can go in many directions.
There are five pieces to be ready with.
The pipeline, at a glance
The ARTIFACT ~ code, prompts, tool definitions, and model version as one bundle
A tool registry ~ owner, version, and schema for every tool
Evals as the release gate
Deployment ~ pinned versions, declared in the repo, rolled out in stages
Behavioral monitoring, with whole-bundle rollback

The whole pipeline on one page.
Why the thing you ship is four things
Start with what you're actually releasing.
Traditionally a release was code. For an agent, it's four parts:
Code - the agent's logic. How it loops, when it stops, what it does with a failed tool call.
Tool definitions - the descriptions of what each tool does and what it accepts. This is what the model reads when it decides which tool to call.
Prompts - system prompts and instruction templates. The text that decides how the agent behaves.
Model version - the exact model you're running on. The specific version, not "latest."
Change any one of those four and the agent behaves differently. Which means all four ship together, as one bundle, under one version number.
You build the bundle once, store it, and promote that same bundle from staging to production. You don't rebuild it along the way. You promote the exact thing you tested.
Why the tools need a registry
Every tool or API your agent calls at runtime gets an owner, a version, and a review before it goes live.
For example, let’s say you have an AI agent that books flights.
The tool originally accepts:
{"destination": "SFO", "date": "Oct 10"}
Then the backend team changes the field to:
{"arrival_airport": "SFO", "date": "Oct 10"}
But the agent still has the old tool schema.
So it keeps sending destination, the API rejects it, and the agent keeps retrying with the same outdated structure.
The model isn’t “wrong” ~ its view of the tool contract is stale.
So the registry holds the current shape of every tool, and the pipeline checks the bundle against it at build time. You catch the rename in CI, not from a support ticket.
This is also the layer where tool permissions belong. An agent should never hold broader access than the task requires, and the registry is where that's written down and reviewed.
Evals are the gate
This is the extension of your test suite, and it's the centerpiece of a good answer.
You keep a fixed set of cases where you already know the right answer. Every change to the bundle runs against all of them. Here’s the key difference between s/w testing and agent evals here:

You're not checking whether it crashed. You're checking whether it picked the right tool, produced the right output, and got there without wandering through twenty steps to do it.
If the score drops below your bar, the pipeline blocks the release.
Two things worth saying out loud in the interview:
The same input doesn't always produce the same output. So an eval isn't an assertion that passes or fails once ~ it's a pass rate across a set, sometimes across repeated runs, measured against a threshold.
And the eval set is a living thing. Every failure you find in production becomes a case in the set, so the same failure can never ship twice.
Deployment
Two cases again. Either you're calling a hosted model API, or you're running the model yourself on GPU servers.
If it's an API, the model version is pinned in the bundle ~ you name the exact version.
If you're hosting it, the weights and serving config ship with the deployment too.
Either way, the version running in production is written down in a file in your repo. Nobody deploys by hand. You change the file, and production updates to match.
Which means deploying is a commit. And so is rolling back.
Then you roll out in stages rather than all at once — a small slice of traffic first, then everyone.

Watch behavior, not servers
Last piece: how do you know something has gone wrong?
Not CPU. Not error rate. Because when an agent gives a bad answer, nothing errors. The request succeeds with a 200. It just did the wrong thing.
So you watch what the agent is doing instead:
Are tool calls failing more than usual?
How many steps is it taking to finish a task?
What does a task cost ~ if that doubles, it's going in circles
How often is it handing off to a human?
Those are the numbers you check during the staged rollout. If they get worse on the first slice of traffic, you stop and go back.
And when you go back, the whole bundle goes back with it. Reverting the code but leaving the new prompt and model live gives you a combination that was never tested — which is a worse position than the bug you were rolling back from.
Let's follow one change

Someone edits a system prompt. Here's the path:
The commit triggers a build. Code, prompts, tool definitions, and the pinned model version are packaged into one bundle with one version number.
The pipeline checks every tool definition in that bundle against the registry. A renamed field or a changed type fails the build here.
The eval set runs against the bundle. Correct tool selection, correct output, step count within bounds. The score has to clear the bar.
The bundle is promoted to staging ~ the same artifact, not a rebuild.
The version file in the repo is updated and merged. Production picks up the change and rolls it to a small slice of traffic.
Behavioral metrics are watched on that slice. If tool errors, step counts, or cost per task move the wrong way, the version file goes back to the previous bundle. All four parts, together.
One prompt edit. Four artifacts. Two gates before a user sees it.
Interview takeaways
Say why evals aren't tests. Non-determinism, thresholds instead of assertions, and a set that grows from production failures. This one sentence separates people who've shipped an agent from people who've read about one.
Prompts change far more often than code ~ and that creates a real tension. Teams want to hot-reload prompts without a full release. The moment you do, the running combination is one your evals never scored. If you allow it, say what you'd give up.
Someone has to own the eval set, and it usually isn't only engineers. Knowing the right answer for a support or finance agent is domain knowledge. Naming that ownership shows you've thought past the pipeline diagram.
Cost per task is a correctness signal, not just a finance metric. An agent that starts costing double is usually looping, retrying, or picking the wrong tool. It's often the first number that moves.
Rollback granularity is the whole point of the bundle. If your answer allows independent rollback of code, prompts, and model, you've designed a system where production can end up in a combination nobody tested.
Final thoughts
This isn’t an exact pattern that every team follows. The architecture will vary based on team requirements, existing infrastructure, scale, and the tools already in place.
Think of this as a framework to help you understand the moving pieces ~ and give you the right pointers to talk through these decisions in an interview.
Your CI system, eval harness, deployment targets or whatever else you use can all change. The important part is understanding the underlying pattern.
Once you get that, the rest of the answer starts to write itself.
Part 3 gets into what’s happening underneath all of this: the two phases of inference, the KV cache, and where your GPU time actually goes.
That’s coming next.
See you in the next one.