Ingestion pipeline
Connectors for SharePoint, Confluence, Notion, Google Drive, Git, PDFs and HTML. Incremental updates, deletion handling, and a re-index path that does not require downtime.
AI InfrastructureAI AgentsBratislava, SK
Service 04
Retrieval-augmented generation done properly: the data pipeline, the search, the reranking and the citations. The part of your AI system that determines whether answers are right or confidently wrong.
RAG is called the solution to company knowledge because it is cheap and flexible. It is also the single most common reason enterprise AI disappoints: bad retrieval produces bad answers no matter how good the model is.
My work is usually not in the model at all. It is in making sure the right paragraph arrives, from the current version of the document, with a link the user can check.
Connectors for SharePoint, Confluence, Notion, Google Drive, Git, PDFs and HTML. Incremental updates, deletion handling, and a re-index path that does not require downtime.
Structure-aware splitting tuned to your documents — headings, tables, contract clauses — because the cheapest quality win in RAG is where you cut the text.
BM25 plus dense vectors, fused and reranked. Pure vector search misses product codes and legal references; pure keyword search misses paraphrase. You need both.
A cross-encoder rerank stage, query rewriting and decomposition for multi-hop questions, and metadata filters that encode who is allowed to see what.
Every answer linked to the exact source and section, with the retrieval date visible. When a document is revised, the answer changes — not six weeks later.
Recall@k, faithfulness and answer-relevance scored against questions built from your real traffic, tracked per release so a regression is visible immediately.
A hundred real questions from support tickets, sales calls and internal Slack. These become the retrieval benchmark.
Pure vector search, measured. Most clients are at 60–75% recall@5 here, and they had no idea.
Chunking, then metadata filters, then hybrid search, then reranking. In that order — model changes help least and cost the most.
Incremental re-indexing and stale-answer detection, so the system is trustworthy a year in, not just in week one.
You get a dashboard showing recall, faithfulness, latency and cost per query — and the test suite to catch regressions yourself.
Python · TypeScript · pgvector / Qdrant / OpenSearch · PostgreSQL · BM25 · cross-encoder rerankers · OpenTelemetry
Almost always RAG for company knowledge. Fine-tuning changes how a model behaves; it does not teach it what your updated policy says. Use RAG for facts that change, fine-tuning for style, format and behaviour. I wrote a longer version of this in the article linked below.
Whichever is closest to your data. If you already run PostgreSQL, pgvector avoids a new operational dependency for most workloads. Standalone vector databases earn their place at higher scale or with specific filtering needs.
Filters are applied at query time from the user's own authorisation, not from the prompt. The retriever never sees documents the requester cannot access, which also means the model cannot leak them.
Then expect the first phase to be mostly content work: PDFs that are really scans, tables that break when split, duplicated pages across systems. It is common and it is fixable — usually it also improves life for humans.