What the student sees, what the office sees, from one set of records
part of Gradsy
A consultancy ran applications in spreadsheets and email. The platform gives a student one honest view of their own file, and gives the office one queue built from the same records.
- problem
- Students struggled to find suitable universities and manage applications, while counselors and admins had fragmented workflows for tracking students, documents, and applications.
- my role
- Full-stack engineer
- outcome
- Gradsy brought university discovery, student applications, and consultancy workflows into one platform - simplifying the experience for students, counselors and admins.
- stack
- Next.js
- NestJS
- TypeScript
- Prisma
- PostgreSQL
- Redis
- BullMQ
- AWS
- AWS Simple Email Service
- S3
- Lightsail
Context
Gradsy is an education consultancy in Nepal that places students in universities abroad. The business is a pipeline. A student arrives, builds a profile, uploads documents, picks programs, and applies. A counselor then moves that application through a series of states until there's an offer or there isn't.
Before the platform, that pipeline lived in spreadsheets, email threads and WhatsApp. Two questions were expensive, and both were expensive for the same reason.
A student's question is "where is my application?" Answering it meant somebody in the office finding the row, remembering the last thing that happened to it, and writing back. That is the highest-volume message a consultancy receives, and every answer costs staff time.
The office's question is "what needs doing, and who is doing it?" A spreadsheet doesn't say which applications are stuck waiting on a transcript, which counselor already has nine students, or whether the person who emailed twice today has been assigned to anyone at all.
Both questions are about the same records. They need opposite surfaces. A student is doing this once, under stress, and needs a guided path with nothing hidden. The office is doing it hundreds of times and needs a queue, a load balance, and a history of who changed what.
So the platform is one domain model with two products on top of it, plus one engineer to build both.
Architecture
Two repositories. A Next.js 15 App Router frontend carrying the public marketing site, a university explorer, and all three dashboards. A NestJS API owning the domain, the database, the queues, and every rule worth enforcing.
They talk over REST with session cookies. No shared package, no code generation, no monorepo.
Next.js 15 (App Router) NestJS 10
───────────────────────── ─────────────────────────
46 server / 46 client pages 24 modules, 42 controllers, 254 routes
TanStack Query (client data) Throttler → BetterAuth → Roles (default-deny)
fetch + RSC (public/SSG) Prisma, 19-file schema, 54 models
socket.io-client BullMQ + Redis · socket.io
│ │
└──────── REST + cookies ─────────┘
│
PostgreSQL 15 · S3 · SESTwo repos saved me setup time and cost me elsewhere. A solo project's bottleneck is rarely type safety. It's the number of things you have to set up before you can ship a feature. Two plain repos and hand-written types meant I could add an endpoint and consume it in about the time it takes to describe it. The cost arrived later, as constants that have to be mirrored by hand and a cache contract nothing validates.
Three properties of that diagram matter to everything below.
Every route is closed until it says otherwise. Every request passes three checks before it reaches my code: rate limit, then session, then role. A route carrying neither @Public() nor @Roles() is refused. Adding an endpoint and forgetting to protect it produces a 403, not a leak. I flipped that default partway through. It broke routes for a day. That was the point: I wanted the gaps to fail loudly instead of leaking quietly.
Only two routes are prerendered. The public university and article pages are built ahead of time and rebuilt on a timer, which Next calls ISR. Every dashboard is client-rendered behind auth. A page that's different for every user and stale by definition gains nothing from being prerendered.
Documents never touch the API. Uploads go straight from the browser to S3 with a presigned URL and are confirmed afterwards, so file bytes bypass Node entirely. The confirm step checks the returned key starts with applications/{applicationId}/documents/, because a client that could confirm any key could attach somebody else's file.
The student side
One document, uploaded once
A student applies to several universities. In the manual process that meant sending the same passport scan several times, and somebody in the office verifying the same file each time.
Here a document is stored once in the student's library, and an application either owns its file or points at the library copy. Attaching inherits the review status, so a passport verified last month is verified on the application it's attached to today. Attaching twice returns the existing attachment rather than a duplicate.
The exceptions are encoded as constants rather than left to judgment. Recommendation letters and transcripts can exist many times over; every other type is single-instance, and re-uploading replaces. The statement of purpose is written per application and is never saved to the library. Removing a document deletes the S3 object only when that row owns it, never when it's a reference to a shared one.

The submit button explains itself
An application can only leave draft when the student's documents are ready. That sentence sounds like a boolean. It isn't one.
The naive version — every required type is uploaded — is wrong three separate ways, and each way is either a student who can't submit or an incomplete application that reaches a partner university.
A document has to be attached to this application, not merely sitting in the student's library. Otherwise a verified passport unlocks an application it was never part of. English proficiency can't require attachment at all, because a test score satisfies it and a test score has nothing to attach to an application; requiring it would permanently lock out every student who proved English that way. The statement of purpose is written per application and reviewed after submission, so gating submission on its verification deadlocks the thing it's supposed to guard.
So the check is three rules with different shapes. The consequence worth noting is the error. It distinguishes not attached from not verified yet from not uploaded from rejected, because those are four different things for the student to do, and collapsing them into "documents incomplete" turns each one into a support ticket. The profile gate works the same way: it reports the exact percentage rather than the word incomplete.
A student who is still blocked gets a button that requests a consultation, which creates a real task for a real counselor. A gate with no way past it is a dead end, and a dead end becomes an email.
Written up in full: Ready to submit is not a boolean.

The history is the answer to "where is my application"
Every status change writes a history row and an audit row inside the same transaction as the change itself. The student reads the history the counselor writes, rejections included, with whatever reason somebody typed.
The transitions themselves live in a database table, not in if statements. Each status row names which statuses may follow it, and names a stricter set for the case where the actor is a student. A target status can be marked as requiring feedback, and then the transition is refused without a note. Changing the workflow is a catalog edit and a re-seed, so the ops team can reshape the pipeline without me.
That is the feature that removes the support question. The student doesn't ask, because the answer is already on the page, in the same words the office used.

More on the table-driven part: A status machine belongs in a table.
The community, and moderation that had to be split by certainty
The student community needed automated moderation. My first version checked everything on write and blocked anything suspicious. That produced a stream of outcomes I couldn't defend: enthusiastic posts blocked for being in caps, useful posts blocked for having four links.
The mistake was splitting the system by latency and then letting that split decide policy. Fast checks became blocking checks because they were fast.
The rewrite asks a different question: what has to be true before this content exists for even a second? That list is short — slurs, known-malicious domains. Everything else is evidence, and evidence deserves a judgment a second later. So the synchronous gate now blocks almost nothing and attaches its findings to the request. A queued worker then adds the two signals the request path can't afford, posting rate and the author's moderation history, before deciding approve, review, or remove.
Anything flagged for review is filed into the existing human reports queue under a system bot account, instead of getting its own admin screen. That is the first place the two sides meet: automated flags and student reports arrive in one place, so every future improvement to that queue applies to both.
Full write-up: The fast moderation layer does less on purpose.

The office side
The queue is the pipeline's states, in order
A counselor's board is the state machine seen from the other end. Where an application has got to is the column it's sitting in, and the count on each column is the answer to "what's blocked."

Work is routed, not remembered
When a student's required document set becomes complete, the system creates the document check itself. It first checks no actionable task already exists, then assigns to the eligible counselor with the least work in flight. When nobody is eligible, it messages every admin to say a student is ready and no counselor was available. The whole routine is wrapped so it can never throw: a failure to create the task must not fail the student's upload.
There are two least-loaded functions, kept separate on purpose. Delegated tasks measure load in open tasks and require delegation to be switched on. Appointment routing measures load in live appointments and doesn't require it, because a booked consultation is core counseling work rather than delegated admin work. Merging them would let booking volume distort the routing of document checks. Ties break on creation time so the result is deterministic.

Delegation is a capacity, not a job title
A counselor's ability to do a delegated task is checked three times: their account has to be approved, delegation has to be on with the task's slug in their permission set, and the assignment service re-checks the same conditions when it picks somebody. Verification calls additionally require a senior delegation level. A per-counselor cap limits how many tasks they can hold at once, so routing can't quietly overload the fastest person.

A review is a status change plus a note the student reads
There is no separate reviewer vocabulary. An admin reviewing an application moves it to a new status and types a reason, and that reason is the text on the student's timeline. The same document rules that block the student are shown to the reviewer as the explanation of why the application isn't submittable, so both people are looking at one answer.

Who changed this record
Most admin questions turn out to be questions about who last touched something. Every write records an actor, and the audit write runs inside the caller's transaction, so a change and its log entry commit together or not at all.
The log copies the actor's name, email and role onto the row rather than joining to the user. The actor reference is set to null when an account is deleted, and accounts do get hard-deleted, so a live join would silently lose attribution on exactly the records somebody is investigating. Machine-originated actions are marked as system, so auto-assignment and public bookings never look like a person did them. The dashboard's filter list is served from the backend's catalog of action names, so the two can't drift.

The catalog is edited once and read publicly
Universities and programs are edited by an admin and rendered on the public site from the same record. Bulk cataloguing goes through a CSV import that reports per-row errors with the real line number from the file. The preview shares the same parser as the import and writes nothing: an operator sees what would change before committing thousands of rows.

That shared record is where the two repositories collide. An admin edits a university, saves, opens the public page, and sees the old data. Nothing is broken. The page was built ahead of time and only rebuilds every 24 hours, and the process that changed the data isn't the process holding the page.
Lowering the window picks a smaller wrong number. The writer has to announce the write instead: after every mutation, the backend POSTs a set of cache tags to a secret-protected route on the frontend. The 24-hour timer stops being the freshness mechanism and becomes the backstop for a webhook that never arrives. It is the only call in the system running backwards, from API to frontend.
Three call shapes fell out of that, and I missed the third at first. A single edit busts one slug. A CSV import busts once for the whole batch instead of once per row. A rename has to bust both the new slug and the old one: the record moved, and its previous URL still holds a prerendered page that nothing will ever invalidate.
Full write-up: Cache invalidation when the cache is in another repo.
Where the two sides collide
Everything above assumes one person at a time. The places that took the most work are the ones where two people act on the same record at once, and the fix in each case was to make the rule a property of the database rather than a check in a service.
Consultations are bookable by anyone from the public site. Three seats per time slot.
The obvious implementation counts existing bookings and creates a row if there's room. It's wrong, and wrong in a way that never appears in development. Two people click book at the same moment. Postgres lets both requests read the booking count before either one saves, so both see one seat free and both insert. Four bookings in a three-seat slot, no error anywhere. That default behaviour has a name: Read Committed.
Instead of counting bookings, each booking now claims a numbered seat. A unique constraint on (slotDate, slotTime, seatIndex) makes "one booking per seat" a property of the schema rather than a check in a service. Two racers can both decide seat 1 is free; only one commits. The other gets a constraint violation, and the booking loop treats that as "try the next seat."
Cancellation sets seatIndex to NULL. Postgres treats NULLs as distinct in a unique index, so cancelled rows sit alongside live ones without blocking anything, and the seat is reusable immediately. No free-list table, no cleanup job, no lock.
I've written the full version of this up separately, including the transaction bug I hit on the way: Overbooking is a database problem.
The booking loop, with its original comment. Twelve lines that replaced a lock:
/**
* Claims the first free seat in the slot. A plain count-then-create
* overbooks under Read Committed (two racers both see capacity-1) so the
* unique index on (slotDate, slotTime, seatIndex) is the actual guard and
* P2002 just means "seat taken, try the next one".
*
* Each attempt is its own transaction: Postgres aborts a transaction on the
* first failed statement, so retrying a seat inside one would only produce
* "current transaction is aborted" on every subsequent create.
*/
private async createWithFreeSeat(
data: Omit<Prisma.AppointmentUncheckedCreateInput, 'seatIndex'>,
actorId: string | null,
) {
for (let seat = 0; seat < APPOINTMENT_SLOT_CAPACITY; seat++) {
try {
return await this.prisma.$transaction(async (tx) => {
const created = await tx.appointment.create({
data: { ...data, seatIndex: seat },
});
await this.audit.record(tx, { actorId, action: 'appointment.create', ... });
return created;
});
} catch (err) {
const isSeatTaken =
err instanceof Prisma.PrismaClientKnownRequestError && err.code === 'P2002';
if (!isSeatTaken) throw err;
}
}
throw new ConflictException(
'That time slot is fully booked. Please pick another time.',
);
}The audit write sits inside the transaction deliberately. If the booking commits, its audit row commits with it, or neither does. An audit log that can silently skip entries during a retry loop is worse than no audit log, because you would trust it.
The same shape appears elsewhere. A community vote toggles inside a transaction and then recomputes the score from the sum of votes, so the denormalised column can't drift from the votes it summarises. A report of the same content by the same person is caught as a constraint violation and treated as a no-op rather than an error.
Impact
No production numbers here. I don't have figures I could defend on a call, and a plausible-looking one is worse than none. What follows is what changed structurally, not what it measured.
A student's status question is answered on the page, in the same words the counselor typed, because the transition, its history row and its audit row are one write.
A document is verified once and reused, instead of being re-sent and re-checked per university.
An application that isn't ready names the four things that could be wrong with it, rather than saying incomplete.
Work routes itself to the counselor with the least in flight, and when nobody is eligible an admin is told rather than the task disappearing.
Booking correctness moved from application code into the schema. The rule survives someone rewriting the service, because it isn't in the service.
Admin edits appear on the public site immediately instead of on a timer, and the timer now covers only the case where the webhook fails.
Broadcast email stays under the provider's send quota by construction: one worker and a fixed pause, rather than a limiter that has to be tuned.
Moderation decisions are all replayable. Every one is logged, and a user's trust score is recomputed from that log rather than stored as a number someone has to trust.
Queue jobs are idempotent at the points where a retry would otherwise double-count. Campaign counters are deduplicated through a Redis set, and recipient status updates use a predicate, so a retry that finds the row already sent changes nothing.
Lessons
Constraints beat checks. Nearly every correctness win in this project came from moving a rule out of a service method and into the database or the queue. A check runs when someone remembers to call it. A constraint holds for every write, including the ones added later by someone who never read the service.
The error message is the feature. The document gate, the profile percentage and the full booked slot all took longer to word than to implement, and each one is a message the office would otherwise have to send by hand. On a two-sided product, a sentence written for the student is work removed from the admin.
Two repos without codegen cost more than it looks. Two constant files are mirrored by hand between the repos, with comments in both explaining that a drift shows the student an option the API will reject. The cache tag names are a contract with no schema, no types, and no failing test: rename one side and the webhook still returns 200 OK while the page silently stops updating. A small shared package would have cost an afternoon. I told myself it was ceremony.
Comments rot faster than code, especially alone. Two files in the frontend describe the revalidation webhook as "a webhook that never landed." It landed months ago and has fifteen call sites. I wrote that comment during the window where only one half existed, copied it into a second file, and never went back. No reviewer, no handoff, nobody to catch it.
The honest gaps. The backend's CI workflow is committed with every line commented out; deploys are still git pull and a process restart. The frontend has no test suite at all. The backend's thirty specs cover the moderation services and the application state machine and stop there, and the README says so out loud. The sitemap is a hardcoded list that omits both families of statically-generated dynamic routes. There's no observability beyond Nest's console logger, and because the queues run removeOnComplete, a job that failed yesterday left nothing to inspect. None of these are things I found while writing this page. Each one lost a prioritisation call against a feature.
Next
Codegen or a shared package for the mirrored constants and cache tags, so a drift is a compile error instead of a silent no-op.
Actual CI. Uncommenting the workflow is the smallest possible first step, and it isn't done.
A dynamic sitemap. The two prerendered route families are the ones worth crawling, and the sitemap is the one place they don't appear. Error boundaries were the other half of this item and have since landed: error.tsx, global-error.tsx and not-found.tsx are all in the tree now.
Observability. A console logger and no job history is fine until the first incident. Then it is the only thing that matters.
Web push. There's a complete cross-repo implementation spec sitting in the frontend repo, including the service worker source and the payload contract, marked "planned, not yet implemented." It would replace the current compromise, where a background tab gets a native browser notification only while the tab is still open.
Merge the two email queues. They independently rate-limit themselves to the same figure, so running both at once doubles the real send rate. The code says so in a comment. One queue with a job type is less machinery than two plus a coordinator.