Golden datasets
Curated sets of inputs with reference answers and acceptable-variation rules, drawn from real traffic and reviewed by the people who know the domain.
AI InfrastructureAI AgentsBratislava, SK
Service 05
The difference between "we think the AI is good" and "we can prove it, and we know when it stops being true". Evals, guardrails and production tracing — the cheapest insurance in AI.
Non-deterministic systems need a different approach to quality. You cannot assert that a response equals a string. You need a test suite built around behaviour, run on every change, and a production trace for every real conversation.
This is the service I recommend to almost everyone who already has something running in production. It is a few weeks of work that pays back the first time a prompt change silently degrades a customer-facing feature.
Curated sets of inputs with reference answers and acceptable-variation rules, drawn from real traffic and reviewed by the people who know the domain.
Task-specific metrics, LLM-as-judge calibrated against human reviewers, and deterministic checks for the parts that must be exact.
Every prompt, model or retrieval change runs the suite and reports a delta. Below threshold, the merge is blocked.
Full spans for every call — prompt, retrieved context, model, tokens, latency, cost — searchable after an incident instead of reconstructed from memory.
Input and output filters, refusal behaviour, PII masking and injection defence, each with its own test case.
Quality, latency, cost and failure rate per feature and per language, with alerting on drift. Weekly, not on request.
Not accuracy — the thing your users would call correct. It has to be written down and agreed before any number is produced.
Build the dataset from actual conversations, filtered for the cases that matter and weighted to your real distribution.
Run LLM-as-judge and human review side by side on a subset. Keep the judge only where it agrees.
The suite runs on pull requests. Reports show the delta, the failing cases and the likely cause.
Trace everything, sample and score a fraction of live traffic continuously, and alert on drift rather than on outages.
Python · TypeScript · pytest / Vitest · OpenTelemetry · Prometheus · Grafana · CI pipelines
For a first filter, yes — it is dramatically better than eyeballing. But only if it is calibrated: I run human review against it on a subset, measure agreement, and keep it strictly as a gate where it agrees. Anything high-stakes stays human-reviewed.
Then we start by producing it: sample real production traffic, cluster it, and have your domain experts label the two or three hundred cases that matter most. That becomes the dataset, and it usually becomes the most valuable artefact of the project.
It speeds you up, once the suite exists. Without it, every prompt change is a manual regression test performed by whoever is bravest that week.