Agent architecture
Which steps need a model and which do not. Where to fan out in parallel, where to force a sequence, where to stop and ask. Usually this removes more cost and more risk than any prompt change.
AI InfrastructureAI AgentsBratislava, SK
Service 02
Multi-step systems that actually do work: reading the CRM, checking the contract, drafting the reply, filing the ticket — inside your existing tools, with the permissions and stop conditions a real company needs.
An agent is a model that can take actions instead of only returning text. That is the whole idea, and it is also where most "agentic AI" projects go wrong: an agent with access to everything and no boundaries will eventually do something expensive, embarrassing or irreversible.
I build agents the boring way. Narrow scope, explicit tools, hard budgets, deterministic steps wherever a language model is not needed, and a human in the loop for anything that touches money, legal commitments or customers.
Which steps need a model and which do not. Where to fan out in parallel, where to force a sequence, where to stop and ask. Usually this removes more cost and more risk than any prompt change.
Tools with typed contracts, narrow scopes and predictable errors. Exposed to the agent over MCP where useful, so the same capability serves your product and your internal agents.
Durable execution: long-running tasks survive restarts, every step is checkpointed, and nothing is retried blindly. Queues, concurrency limits and idempotency.
Review queues, inline diffs and approve/reject actions, risk-tiered autonomy — the agent drafts on low-risk work and acts alone only where you explicitly allow it.
Input and output filtering, prompt-injection defence, PII handling and an adversarial test set that runs on every prompt or model change.
Full trace per run: every prompt, tool call, argument and result. Retention you can defend to an auditor, and logs your security team can actually query.
A queue with a measurable volume and a clear definition of done. If nobody can say what "finished" looks like, it is not an agent project yet.
The agent runs on real data but a human does the work. You get a precision number before any customer or colleague is affected.
Ship the version that produces a proposal — a reply, a diff, a ticket update — and let a human approve. This is usually where the value already shows up.
Turn on unattended execution per tool, with limits, and keep approval on anything irreversible.
Track handled volume, quality, cost and escalation rate. Widen scope only where the numbers are boringly good.
TypeScript · Python · MCP · Anthropic / OpenAI tool use · Postgres · Redis queues · Docker · OpenTelemetry · AWS
A chatbot answers. An agent acts. The difference is the tool layer and the orchestration: an agent can read your CRM, run a query, create a record, and stop at a defined point — with permissions, budget and an audit trail for every action.
Enough to be useful, little enough to be recoverable. My default is: draft-only on anything a customer sees or a contract depends on, unattended execution only for reversible internal actions, and hard limits on spend and steps per run.
Only when there is a real reason — separate permissions, separate context budgets, or genuinely parallel work. Most enterprise agent problems are solved better by one well-tooled agent with a durable loop than by a swarm of specialised ones.
It is treated as a first-class threat, not an edge case. Untrusted content is separated from instructions, tools carry their own scopes, output is filtered, and there is an adversarial test set in CI that tries to break the agent before your users do.
Anything with an API, a database, a webhook or an MCP server — Salesforce, HubSpot, Dynamics, SAP, Zendesk, Jira, NetSuite, internal services, S3, email, calendars. Most of the work is in the contract and the permissions, not the connection.