03 / retrieval · AI that cites its sources
RAG developer: AI search and extraction over your own documents
AI that answers from your documents instead of guessing. Every answer shows the document it came from, and "I could not find that" is a real answer with its own path through the code.
- Retrieval from prototype to production at Milo Logic
- Vectors in PostgreSQL, beside the relational data
- An evaluation set built with the system, not after it

what I build
Answers somebody can check in a second.
You have documents. People need to ask questions of them and trust what comes back. Most of that job happens before the model is involved.
RAG search
search that understands the question, over documents only you have
document extraction
named fields pulled out of contracts, forms and reports
question answering
answers with their sources beside them, so anyone can check one
hybrid search
meaning and filters in one query, because most real questions are both
ingestion pipelines
PDFs, scans and exports read in, and kept current as they change
AI workflows
a model doing one bounded step inside a process you already run
evaluation sets
real questions with the sources that should answer them, rerun on every change
AI in an existing product
retrieval added to software you already run, not a second product beside it

how it goes
Corpus, retrieval, answer, evidence.
Generating the sentence is the last step and the smallest one. Everything before it decides whether the sentence is true.
read
your documents in — PDFs, scans, exports, and whatever else the last ten years left behind
cut
each document split into pieces that still make sense on their own, away from the rest
find
vector search over those pieces, running beside your ordinary filters in the same query
answer
the answer, the sources on screen next to it, and a real path for "I could not find that"

why retrieval
A general model guesses. Retrieval makes it look first.
Fine-tuning teaches a model a style. Retrieval hands it the paragraph. Only one of those can show you where the answer came from.
- your documents, not the internet
- the model answers from your contracts and your tickets, not from what it read during training.
- recall before precision
- a wrong passage gets ignored. A missing one gets answered anyway, fluently, and it sounds the same as when it is right.
- meaning and filters together
- renewals agreed after March is half a search and half a date column. Both run in Postgres, in one query.
- sources on the screen
- an answer you cannot check is a rumour. The first wrong one nobody caught is what ends the project.
- not sure is an answer
- an endpoint that must always produce something will always produce something.
- measured, not argued
- every change to chunking, the embedding model or the prompt is scored against your own questions.
the problems I solve
You have probably said one of these out loud.

“We tried a chatbot on our documents. It sounded confident, it was wrong twice, and now nobody uses it.”
It answered without the right passage in front of it. I fix the retrieval first, then put the source on screen so the next wrong answer is caught in a second rather than in a quarter.
“One person on the team finds the right paragraph and everybody else waits for them.”
That is a search problem with a person standing in for the index. The answers are already in the documents; nothing is there to ask them with.
“The answers are somewhere across ten years of contracts and tickets.”
Ingestion, chunking and vector search over the corpus you already have, with the filters your team already searches by.
“We need fields out of these documents, not a chat window.”
Same pipeline, different ending. You get named fields, so the rest of your software can treat a contract as data.
“How would we even know whether it got better?”
An evaluation set built alongside the system — real questions, the sources that should answer them, rerun on every change. The alternative is arguing about it in a meeting.
- PostgreSQL
- pgvector
- Embeddings
- Vector search
- Document extraction
- evals
The vectors live in PostgreSQL beside the rest of your data, so a question with a filter in it stays one query instead of a join written by hand in application code. I build the screen on top as well.
how I build it
The vectors live in PostgreSQL, beside the rest of your data, through an extension called pgvector.
Most real questions are half meaning and half filter. "What did we agree with this client about renewals, in contracts signed after March." The first half needs vector search. The client and the date are ordinary columns. Split those across two systems and you write that join by hand, in application code, every time somebody asks.
Retrieval aims for recall before precision. A wrong passage in front of the model usually gets ignored. A missing one is worse: the model does not tell you it came up empty. It answers anyway, fluently, and it sounds the same as when it is right. So the search step brings back more than it needs on purpose, and a ranking step narrows it down.
Every answer shows the documents it came from, on screen, next to the words. That one detail decides whether anyone is still using the system in month three. An answer you cannot check is a rumour. The first time somebody catches one being wrong, they stop trusting all of them.
"I could not find that" is a real answer with its own path through the code. An endpoint that must always produce something will always produce something.
when a project needs more than me
Retrieval, extraction and answering cover most of what teams actually need, and they share a useful property: every output can be checked against a source.
Some projects reach past that, into model fine-tuning, custom training or heavier machine learning. For those I bring in senior AI engineers I already work with. The pipeline, the data model and the interface stay with me, and the specialist work happens inside the same project.
what working together looks like
You bring it. I handle it. You get it back better.
you bring
- the documents, and the questions your team cannot get answered
- whoever knows which answer is the right one
- the filters you already search by — client, date, status
I handle
- ingestion, chunking and embedding
- vector search in PostgreSQL, beside your relational data
- ranking, and the path for when nothing is found
- the interface that puts sources next to answers
- the evaluation set, and the score on every change
- deployment, inside your own infrastructure
you get
- your code and your documents, in your infrastructure
- answers with the source beside them, checkable in a second
- a real number for retrieval quality on your corpus, instead of a promise about it

Documents nobody can ask questions of?
Thirty minutes on the corpus and the questions your team keeps failing to answer. Bring the hard ones. I will tell you what retrieval would find in them and whether it is worth building.