skip to content

03 / retrieval · AI that cites its sources

RAG developer: AI search and extraction over your own documents

AI that answers from your documents instead of guessing. Every answer shows the document it came from, and "I could not find that" is a real answer with its own path through the code.

  • Retrieval from prototype to production at Milo Logic
  • Vectors in PostgreSQL, beside the relational data
  • An evaluation set built with the system, not after it
A builder lifting one highlighted paragraph out of a tall stack of documents and setting it into an answer card, with the document it came from pinned underneath.

what I build

Answers somebody can check in a second.

You have documents. People need to ask questions of them and trust what comes back. Most of that job happens before the model is involved.

  • RAG search

    search that understands the question, over documents only you have

  • document extraction

    named fields pulled out of contracts, forms and reports

  • question answering

    answers with their sources beside them, so anyone can check one

  • hybrid search

    meaning and filters in one query, because most real questions are both

  • ingestion pipelines

    PDFs, scans and exports read in, and kept current as they change

  • AI workflows

    a model doing one bounded step inside a process you already run

  • evaluation sets

    real questions with the sources that should answer them, rerun on every change

  • AI in an existing product

    retrieval added to software you already run, not a second product beside it

The same builder among floating retrieval panels — a search result list, a contract with fields lifted out of it, an answer with source chips, a filter row and a scored evaluation list — all fed from one document store.

how it goes

Corpus, retrieval, answer, evidence.

Generating the sentence is the last step and the smallest one. Everything before it decides whether the sentence is true.

  1. read

    your documents in — PDFs, scans, exports, and whatever else the last ten years left behind

  2. cut

    each document split into pieces that still make sense on their own, away from the rest

  3. find

    vector search over those pieces, running beside your ordinary filters in the same query

  4. answer

    the answer, the sources on screen next to it, and a real path for "I could not find that"

A document being cut into chunks, each chunk becoming a small numbered token, and the tokens dropping into the same database cylinder as ordinary table rows, with one query drawing from both.

why retrieval

A general model guesses. Retrieval makes it look first.

Fine-tuning teaches a model a style. Retrieval hands it the paragraph. Only one of those can show you where the answer came from.

your documents, not the internet
the model answers from your contracts and your tickets, not from what it read during training.
recall before precision
a wrong passage gets ignored. A missing one gets answered anyway, fluently, and it sounds the same as when it is right.
meaning and filters together
renewals agreed after March is half a search and half a date column. Both run in Postgres, in one query.
sources on the screen
an answer you cannot check is a rumour. The first wrong one nobody caught is what ends the project.
not sure is an answer
an endpoint that must always produce something will always produce something.
measured, not argued
every change to chunking, the embedding model or the prompt is scored against your own questions.

the problems I solve

You have probably said one of these out loud.

On the left, a confident-looking answer card floating alone with nothing attached to it. On the right, the same answer with a document pinned beneath it, and the builder attaching the thread.

“We tried a chatbot on our documents. It sounded confident, it was wrong twice, and now nobody uses it.”

It answered without the right passage in front of it. I fix the retrieval first, then put the source on screen so the next wrong answer is caught in a second rather than in a quarter.

“One person on the team finds the right paragraph and everybody else waits for them.”

That is a search problem with a person standing in for the index. The answers are already in the documents; nothing is there to ask them with.

“The answers are somewhere across ten years of contracts and tickets.”

Ingestion, chunking and vector search over the corpus you already have, with the filters your team already searches by.

“We need fields out of these documents, not a chat window.”

Same pipeline, different ending. You get named fields, so the rest of your software can treat a contract as data.

“How would we even know whether it got better?”

An evaluation set built alongside the system — real questions, the sources that should answer them, rerun on every change. The alternative is arguing about it in a meeting.

  • PostgreSQL
  • pgvector
  • Embeddings
  • Vector search
  • Document extraction
  • evals

The vectors live in PostgreSQL beside the rest of your data, so a question with a filter in it stays one query instead of a join written by hand in application code. I build the screen on top as well.

how I build it

The vectors live in PostgreSQL, beside the rest of your data, through an extension called pgvector.

Most real questions are half meaning and half filter. "What did we agree with this client about renewals, in contracts signed after March." The first half needs vector search. The client and the date are ordinary columns. Split those across two systems and you write that join by hand, in application code, every time somebody asks.

Retrieval aims for recall before precision. A wrong passage in front of the model usually gets ignored. A missing one is worse: the model does not tell you it came up empty. It answers anyway, fluently, and it sounds the same as when it is right. So the search step brings back more than it needs on purpose, and a ranking step narrows it down.

Every answer shows the documents it came from, on screen, next to the words. That one detail decides whether anyone is still using the system in month three. An answer you cannot check is a rumour. The first time somebody catches one being wrong, they stop trusting all of them.

"I could not find that" is a real answer with its own path through the code. An endpoint that must always produce something will always produce something.

when a project needs more than me

Retrieval, extraction and answering cover most of what teams actually need, and they share a useful property: every output can be checked against a source.

Some projects reach past that, into model fine-tuning, custom training or heavier machine learning. For those I bring in senior AI engineers I already work with. The pipeline, the data model and the interface stay with me, and the specialist work happens inside the same project.

what working together looks like

You bring it. I handle it. You get it back better.

you bring

  • the documents, and the questions your team cannot get answered
  • whoever knows which answer is the right one
  • the filters you already search by — client, date, status

I handle

  • ingestion, chunking and embedding
  • vector search in PostgreSQL, beside your relational data
  • ranking, and the path for when nothing is found
  • the interface that puts sources next to answers
  • the evaluation set, and the score on every change
  • deployment, inside your own infrastructure

you get

  • your code and your documents, in your infrastructure
  • answers with the source beside them, checkable in a second
  • a real number for retrieval quality on your corpus, instead of a promise about it
The builder pushing a sealed retrieval pipeline up a ramp into a cloud platform, through a gate holding a scored evaluation checklist.

Documents nobody can ask questions of?

Thirty minutes on the corpus and the questions your team keeps failing to answer. Bring the hard ones. I will tell you what retrieval would find in them and whether it is worth building.

Book a callEmail