Frederik Rybansky

AI InfrastructureAI AgentsBratislava, SK

Service 05

LLM Evaluation & Observability

The difference between "we think the AI is good" and "we can prove it, and we know when it stops being true". Evals, guardrails and production tracing — the cheapest insurance in AI.

Non-deterministic systems need a different approach to quality. You cannot assert that a response equals a string. You need a test suite built around behaviour, run on every change, and a production trace for every real conversation.

This is the service I recommend to almost everyone who already has something running in production. It is a few weeks of work that pays back the first time a prompt change silently degrades a customer-facing feature.

Deliverables

Golden datasets

Curated sets of inputs with reference answers and acceptable-variation rules, drawn from real traffic and reviewed by the people who know the domain.

Automated scoring

Task-specific metrics, LLM-as-judge calibrated against human reviewers, and deterministic checks for the parts that must be exact.

Regression testing in CI

Every prompt, model or retrieval change runs the suite and reports a delta. Below threshold, the merge is blocked.

Production tracing

Full spans for every call — prompt, retrieved context, model, tokens, latency, cost — searchable after an incident instead of reconstructed from memory.

Guardrails

Input and output filters, refusal behaviour, PII masking and injection defence, each with its own test case.

Dashboards

Quality, latency, cost and failure rate per feature and per language, with alerting on drift. Weekly, not on request.

How I set it up

Agree the metric

Not accuracy — the thing your users would call correct. It has to be written down and agreed before any number is produced.

Sample real traffic

Build the dataset from actual conversations, filtered for the cases that matter and weighted to your real distribution.

Calibrate the judge

Run LLM-as-judge and human review side by side on a subset. Keep the judge only where it agrees.

Automate and gate

The suite runs on pull requests. Reports show the delta, the failing cases and the likely cause.

Watch production

Trace everything, sample and score a fraction of live traffic continuously, and alert on drift rather than on outages.

Outcome

  • Model or prompt changes stop being scary, because the evals tell you what moved.
  • You can quote a quality number to customers, auditors or your board.
  • Incidents start with a trace instead of an argument about what the model saw.
  • Cost and latency regressions are caught in CI, not in the monthly bill.
  • Guardrails have tests, so "we have a filter" is verifiable rather than assumed.

Typical stack

Python · TypeScript · pytest / Vitest · OpenTelemetry · Prometheus · Grafana · CI pipelines

Frequently asked questions

Is LLM-as-judge reliable enough to gate a release?

For a first filter, yes — it is dramatically better than eyeballing. But only if it is calibrated: I run human review against it on a subset, measure agreement, and keep it strictly as a gate where it agrees. Anything high-stakes stays human-reviewed.

We do not have test data.

Then we start by producing it: sample real production traffic, cluster it, and have your domain experts label the two or three hundred cases that matter most. That becomes the dataset, and it usually becomes the most valuable artefact of the project.

Does this slow us down?

It speeds you up, once the suite exists. Without it, every prompt change is a manual regression test performed by whoever is bravest that week.

More services