---
title: "RAG developer: AI search and extraction over your own documents"
description: "AI that answers from your documents instead of guessing. Every answer shows the document it came from, and \"I could not find that\" is a real answer with its own path through the code."
url: https://riteshkc.com.np/services/rag-and-document-ai
source: https://riteshkc.com.np/services/rag-and-document-ai.md
updated: 2026-08-27
site: "Ritesh KC"
---
# AI that cites its sources

RAG developer: AI search and extraction over your own documents

AI that answers from your documents instead of guessing. Every answer shows the document it came from, and "I could not find that" is a real answer with its own path through the code.

- Layer: 03 / retrieval
- Builds: RAG search (search that understands the question, over documents only you have); document extraction (named fields pulled out of contracts, forms and reports); question answering (answers with their sources beside them, so anyone can check one); hybrid search (meaning and filters in one query, because most real questions are both); ingestion pipelines (PDFs, scans and exports read in, and kept current as they change); AI workflows (a model doing one bounded step inside a process you already run); evaluation sets (real questions with the sources that should answer them, rerun on every change); AI in an existing product (retrieval added to software you already run, not a second product beside it)
- Tools: PostgreSQL, pgvector, Embeddings, Vector search, Document extraction, evals
- Retrieval from prototype to production at Milo Logic
- Vectors in PostgreSQL, beside the relational data
- An evaluation set built with the system, not after it

## You have probably said one of these out loud.

- We tried a chatbot on our documents. It sounded confident, it was wrong twice, and now nobody uses it. It answered without the right passage in front of it. I fix the retrieval first, then put the source on screen so the next wrong answer is caught in a second rather than in a quarter.
- One person on the team finds the right paragraph and everybody else waits for them. That is a search problem with a person standing in for the index. The answers are already in the documents; nothing is there to ask them with.
- The answers are somewhere across ten years of contracts and tickets. Ingestion, chunking and vector search over the corpus you already have, with the filters your team already searches by.
- We need fields out of these documents, not a chat window. Same pipeline, different ending. You get named fields, so the rest of your software can treat a contract as data.
- How would we even know whether it got better? An evaluation set built alongside the system — real questions, the sources that should answer them, rerun on every change. The alternative is arguing about it in a meeting.

## A general model guesses. Retrieval makes it look first.

Fine-tuning teaches a model a style. Retrieval hands it the paragraph. Only one of those can show you where the answer came from.

- your documents, not the internet: the model answers from your contracts and your tickets, not from what it read during training.
- recall before precision: a wrong passage gets ignored. A missing one gets answered anyway, fluently, and it sounds the same as when it is right.
- meaning and filters together: renewals agreed after March is half a search and half a date column. Both run in Postgres, in one query.
- sources on the screen: an answer you cannot check is a rumour. The first wrong one nobody caught is what ends the project.
- not sure is an answer: an endpoint that must always produce something will always produce something.
- measured, not argued: every change to chunking, the embedding model or the prompt is scored against your own questions.

## Corpus, retrieval, answer, evidence.

01. read — your documents in — PDFs, scans, exports, and whatever else the last ten years left behind
02. cut — each document split into pieces that still make sense on their own, away from the rest
03. find — vector search over those pieces, running beside your ordinary filters in the same query
04. answer — the answer, the sources on screen next to it, and a real path for "I could not find that"

## You bring it. I handle it. You get it back better.

- You bring: the documents, and the questions your team cannot get answered; whoever knows which answer is the right one; the filters you already search by — client, date, status
- I handle: ingestion, chunking and embedding; vector search in PostgreSQL, beside your relational data; ranking, and the path for when nothing is found; the interface that puts sources next to answers; the evaluation set, and the score on every change; deployment, inside your own infrastructure
- You get: your code and your documents, in your infrastructure; answers with the source beside them, checkable in a second; a real number for retrieval quality on your corpus, instead of a promise about it

The vectors live in PostgreSQL beside the rest of your data, so a question with a filter in it stays one query instead of a join written by hand in application code. I build the screen on top as well.

## how I build it

The vectors live in PostgreSQL, beside the rest of your data, through an extension called pgvector.

Most real questions are half meaning and half filter. "What did we agree with this client about renewals, in contracts signed after March." The first half needs vector search. The client and the date are ordinary columns. Split those across two systems and you write that join by hand, in application code, every time somebody asks.

Retrieval aims for recall before precision. A wrong passage in front of the model usually gets ignored. A missing one is worse: the model does not tell you it came up empty. It answers anyway, fluently, and it sounds the same as when it is right. So the search step brings back more than it needs on purpose, and a ranking step narrows it down.

Every answer shows the documents it came from, on screen, next to the words. That one detail decides whether anyone is still using the system in month three. An answer you cannot check is a rumour. The first time somebody catches one being wrong, they stop trusting all of them.

"I could not find that" is a real answer with its own path through the code. An endpoint that must always produce something will always produce something.

## when a project needs more than me

Retrieval, extraction and answering cover most of what teams actually need, and they share a useful property: every output can be checked against a source.

Some projects reach past that, into model fine-tuning, custom training or heavier machine learning. For those I bring in senior AI engineers I already work with. The pipeline, the data model and the interface stay with me, and the specialist work happens inside the same project.

## Proof

- none listed

## Related writing

- [In RAG, recall is the number that matters](https://riteshkc.com.np/blog/recall-beats-precision): Why the failure mode that sinks a retrieval system is the right document never showing up, and why that makes recall, not prompt tuning, the thing to optimize.
- [Endpoints that refuse to be oracles](https://riteshkc.com.np/blog/endpoints-that-refuse-to-be-oracles): A 404 on unsubscribe tells an attacker which tokens are real. A 409 on subscribe tells them who is on your list. Honest status codes leak, and the fix reads like a bug.
