Frederik Rybansky

AI InfrastructureAI AgentsBratislava, SK

Service 04

RAG & Knowledge Systems

Retrieval-augmented generation done properly: the data pipeline, the search, the reranking and the citations. The part of your AI system that determines whether answers are right or confidently wrong.

RAG is called the solution to company knowledge because it is cheap and flexible. It is also the single most common reason enterprise AI disappoints: bad retrieval produces bad answers no matter how good the model is.

My work is usually not in the model at all. It is in making sure the right paragraph arrives, from the current version of the document, with a link the user can check.

Deliverables

Ingestion pipeline

Connectors for SharePoint, Confluence, Notion, Google Drive, Git, PDFs and HTML. Incremental updates, deletion handling, and a re-index path that does not require downtime.

Chunking strategy

Structure-aware splitting tuned to your documents — headings, tables, contract clauses — because the cheapest quality win in RAG is where you cut the text.

Hybrid retrieval

BM25 plus dense vectors, fused and reranked. Pure vector search misses product codes and legal references; pure keyword search misses paraphrase. You need both.

Reranking & query handling

A cross-encoder rerank stage, query rewriting and decomposition for multi-hop questions, and metadata filters that encode who is allowed to see what.

Freshness & citations

Every answer linked to the exact source and section, with the retrieval date visible. When a document is revised, the answer changes — not six weeks later.

Evaluation

Recall@k, faithfulness and answer-relevance scored against questions built from your real traffic, tracked per release so a regression is visible immediately.

The retrieval loop

Build the question set

A hundred real questions from support tickets, sales calls and internal Slack. These become the retrieval benchmark.

Baseline honestly

Pure vector search, measured. Most clients are at 60–75% recall@5 here, and they had no idea.

Improve in the right order

Chunking, then metadata filters, then hybrid search, then reranking. In that order — model changes help least and cost the most.

Add freshness

Incremental re-indexing and stale-answer detection, so the system is trustworthy a year in, not just in week one.

Hand over the numbers

You get a dashboard showing recall, faithfulness, latency and cost per query — and the test suite to catch regressions yourself.

Outcome

  • Answers cite the source document and section, so users can verify them in one click.
  • Revisions propagate: update a policy, and the assistant follows.
  • Permission filtering means people do not retrieve documents they should not see.
  • Measurable recall and faithfulness numbers per release.
  • A clear answer on which questions need content work rather than engineering.

Typical stack

Python · TypeScript · pgvector / Qdrant / OpenSearch · PostgreSQL · BM25 · cross-encoder rerankers · OpenTelemetry

Frequently asked questions

RAG or fine-tuning?

Almost always RAG for company knowledge. Fine-tuning changes how a model behaves; it does not teach it what your updated policy says. Use RAG for facts that change, fine-tuning for style, format and behaviour. I wrote a longer version of this in the article linked below.

Which vector database should we use?

Whichever is closest to your data. If you already run PostgreSQL, pgvector avoids a new operational dependency for most workloads. Standalone vector databases earn their place at higher scale or with specific filtering needs.

How do you handle permissioned documents?

Filters are applied at query time from the user's own authorisation, not from the prompt. The retriever never sees documents the requester cannot access, which also means the model cannot leak them.

What if our documents are a mess?

Then expect the first phase to be mostly content work: PDFs that are really scans, tables that break when split, duplicated pages across systems. It is common and it is fixable — usually it also improves life for humans.

More services