Frederik Rybansky

AI InfrastructureAI AgentsBratislava, SK

How to Introduce AI Agents Without Burning the IT Budget

Enterprise agent projects usually fail for organisational reasons, not technical ones. Here is the rollout sequence that has worked for me — narrow, measured, and reversible at every step.

Every enterprise I meet has the same ambition — an agent that handles work end to end — and the same problem: nobody can say what "handled" means, so nobody can say when it is finished.

Why agent projects stall

Agent work fails in three predictable places. Requirements are written as aspirations ("handle customer requests"), so nobody can measure success. The first demonstration is done by the people building it, against data they curated, which tells you nothing about production. Then autonomy is granted all at once, and the first expensive mistake costs the project its political capital.

None of these are about model quality. They are about how the work is sequenced.

Start with a queue, not a capability

The single most useful change I make is refusing to start with a capability and insisting on starting with a queue. Not "customer service" — "the 340 support tickets per week that arrive with a photo of a damaged unit and need a refund decision".

A queue has a volume, an arrival pattern, a measurable current handling cost, and — crucially — a definition of done. It also tells you what data is needed, which tells you what permissions the agent needs, which is where most of the actual engineering lives.

Pick the queue where the work is repetitive, the context is available in systems you already have, and a mistake is recoverable. That combination is rarer than people assume. Support triage is a good candidate. Anything involving a legal commitment or an irreversible payment is not — not yet.

Shadow mode before autonomy

Before the agent does anything, it does everything.

Run it on real traffic, in parallel with the humans who currently do the work, with the output going to a queue that nobody reads yet. This is the most valuable two weeks in the entire project, because it produces three things at once:

  • A precision number. What percentage of agent decisions would a human have accepted unchanged?
  • A taxonomy of failures. They are almost never random. They cluster into missing context, ambiguous policy, and cases where the correct answer was to ask a question.
  • A stakeholder artefact. Watching a dashboard of real cases is what converts scepticism into a budget conversation.

Shadow mode is cheap, embarrassingly honest, and routinely surprises people in both directions.

Draft, then do

Most agent projects deliver their value in the draft-only phase, and then push for autonomy too early.

A draft-only agent that proposes a refund, writes the reply, prepares the ticket update, or marks up the contract clause in a diff — that is already a large productivity gain, and it carries almost no risk, because a human is still reading every word. Escalate rate falls, review time falls, and the agent's quality improves with every review because the corrections are labelled data.

Autonomy should be granted per action, not per agent. Reading a CRM record can be unattended on day one. Creating a record can follow. Issuing a refund requires approval until the evaluation data says otherwise — and "otherwise" should be a number, not a feeling.

What actually needs a human

Keep a human in the loop when the action is irreversible, externally visible to a customer, legally consequential, ambiguous under your written policy, or unprecedented. That list is shorter than people fear, and keeping it short is what makes autonomy credible to a risk or audit function.

Everything else is a candidate — with limits. Cap spend per run, cap steps per run, cap concurrent runs. A run that hits a limit should stop and escalate rather than improvise, because an improvisation at step 40 is where the expensive incidents live.

Measuring it honestly

Measure four things, weekly, and resist adding a fifth:

  • Handled volume — cases closed end to end without a human touching them.
  • Quality — sampled by a human reviewer against a rubric agreed in week one.
  • Cost — per handled case, including tokens, tool calls and infrastructure.
  • Escalation rate — the share that goes to a person, and why.

Conversation counts and user satisfaction scores are not in that list for a reason. Both are easy to inflate and hard to interpret. If handled volume is flat while cost per case rises, the agent is getting worse, not better.

Multi-agent systems: mostly no

Most enterprise "multi-agent" architectures I am asked to build are a single agent with a durable loop and two or three tools. The pattern looks impressive in diagrams and is usually worse in production: more context to manage, more failure modes, harder debugging, and latency multiplied across hops.

Add agents when there is a real reason. Separate permissions that you do not want combined. Context budgets that genuinely cannot share a window. Work that is genuinely parallel and independently valuable. If the only reason is that "multi-agent" sounds more advanced, keep the loop.

A 90-day sequence

Days 1–30: choose the queue, write the rubric, build shadow mode, publish the precision number.

Days 31–60: ship draft-only into real review queues, instrument cost per case, fix the top three failure classes.

Days 61–90: grant unattended execution action by action, add a second queue only if the first is genuinely boring.

Most enterprises that do this end up with one production agent doing one thing well — and, crucially, with the organisation's belief system adjusted to what agents actually deliver. That adjustment is the real return on the first agent, and it is what makes the second one fundable.

Related writing