Model gateway & routing
One interface in front of every provider. Per-task routing between models, automatic fallback when a provider degrades, retry and timeout policy, and streaming where it helps. Swap models without touching product code.
AI InfrastructureAI AgentsBratislava, SK
Service 01
The platform layer: everything between your product code and a language model. Built so that model changes are configuration, quality is measurable, and cost is a dashboard rather than an invoice shock.
AI projects rarely fail because the model is bad. They fail because the plumbing is missing — there is no way to retry, no way to see what the model saw, no way to switch providers, and no way to answer "what did this feature cost us last month".
AI infrastructure is the work that fixes all of that. It is unglamorous, it is the reason the rest of your AI roadmap is possible, and it is the part I spend most of my time on.
One interface in front of every provider. Per-task routing between models, automatic fallback when a provider degrades, retry and timeout policy, and streaming where it helps. Swap models without touching product code.
Ingestion pipeline, chunking strategy, embeddings and a vector store next to your relational data. Hybrid keyword plus dense search, reranking, incremental re-indexing and source citations on every answer.
Golden datasets from your own traffic, LLM-as-judge with human calibration, regression suites that run in CI, and a score you can put in front of a board instead of a vibe.
OpenTelemetry traces for every call: prompt, retrieved chunks, model, tokens, latency, cost. Log the exact input that produced a bad answer, reproduce it, and re-test the fix.
Prompt and context budgets, caching, small-model-first routing, batch and async paths. Most clients cut spend 40–70% in the first quarter, without touching answer quality.
Docker and Kubernetes, environments separated the way your security team expects, secrets management, PII redaction at the boundary, and optional fully on-premise inference for data that cannot leave.
Which features call a model, what data they may touch, which are customer-facing, what latency and cost you can tolerate. This becomes the written plan.
Single entry point for models, prompts and tools. Everything downstream talks to this, so nothing hard-codes a vendor SDK.
Ingestion, chunking, embedding, indexing, citations. Boring, testable, and the part users notice when it is missing.
Evals in CI, traces in production, cost per feature on a dashboard. Then tune: routing, caching, context length.
Load tests, failure drills, runbooks and a walkthrough with your team so you own it after I leave.
TypeScript · Python · Anthropic / OpenAI / Google APIs · pgvector or Qdrant · PostgreSQL · Redis · Kubernetes · Docker · OpenTelemetry · AWS
As the person who knows LLM-specific failure modes — non-determinism, retrieval drift, prompt and context changes, token economics, evaluation. I usually work inside your existing infrastructure and CI, not around it.
Usually not at the start. Hosted APIs plus a good gateway cover most enterprise use cases, and I will tell you when self-hosting genuinely pays off rather than defaulting to it. If your data cannot leave the network, self-hosting is the only option and I can size it.
Every platform engagement ships with an evaluation harness. We agree the metrics before we build, run them against a baseline, and re-run them on every change. It also becomes the regression suite for whatever we build next.
That is a routine operational cost, and it is exactly what the gateway is for. Model migrations are a mapping change plus a re-run of the eval suite; if a migration regresses quality, the evals catch it before customers do.