# Ritesh KC — full site > Hire the best full-stack React developer and UI/UX designer in Nepal. I build lightning-fast, scalable web applications and enterprise platforms. Generated 2026-09-28. Every page of https://riteshkc.com.np as markdown, in reading order. The short index is at https://riteshkc.com.np/llms.txt. # Ritesh KC — Top React Developer & Software Engineer in Nepal | Ritesh KC Hire the best full-stack React developer and UI/UX designer in Nepal. I build lightning-fast, scalable web applications and enterprise platforms. Based in Kathmandu, Nepal. Available for remote work with US and Europe hours overlap. ## What I do - Product engineering end to end: React, Next.js and TypeScript on the front, Node and NestJS behind it, PostgreSQL underneath. - LLM features in production: retrieval-augmented generation, document extraction, and the ingestion and validation pipeline around them. - Multi-tenant SaaS: subdomain-scoped tenancy, role-based access, correctness under concurrency and retry. ## Proof - 5 shipped projects, listed at https://riteshkc.com.np/work. - 5 long-form case studies, at https://riteshkc.com.np/case-studies. - 12 engineering write-ups, at https://riteshkc.com.np/blog. - Three employers on the résumé (Milo Logic, Azminds Services, Lanceme Up) plus named freelance clients. ## Selected work - [Noveon](https://riteshkc.com.np/work/noveon) — Next.js and Strapi website for a Hong Kong investment firm (2026, Full-stack developer). A website for a Hong Kong investment firm, built so people who are not developers can run it. Next.js on the front, Strapi on the firm's own server behind it. - [Marti China](https://riteshkc.com.np/work/marti-china) — Bilingual corporate site and 49-page product catalogue on Next.js and Strapi (2026, Full-stack developer). A corporate site for the China arm of a Swiss tunnelling supplier. Next.js on the front, self-hosted Strapi behind it: 49 product pages, 17 project references, two languages. - [Gradsy](https://riteshkc.com.np/work/gradsy) — Turning a Manual Workflow into a Scalable Product (2026, Full-stack engineer). A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. - [Yacht Cloud](https://riteshkc.com.np/work/yacht-cloud) — YachtCloud: Engineering a Multitenant Platform for Luxury Charters (2024, Frontend Developer). Engineered a multitenant frontend and RAG AI planner that powers YachtCloud and multiple white-labeled broker platforms by instantly filtering massive luxury yacht inventories. - [Investment Banker/Analyst Portfolio](https://riteshkc.com.np/work/prajesh-sagar) — Next.js and Payload CMS portfolio website for a NEPSE portfolio manager (2026, Designer and developer). A portfolio site for a Kathmandu investment banker who manages private portfolios on the Nepal Stock Exchange. Next.js with Payload inside the same app, and a hero that drifts like a chart. ## What clients said - "Ritesh is the kind of engineer you want on your toughest problems. He brings strong technical depth, takes full ownership, and consistently turns complexity into solutions that work. Sharp, dependable, and always focused on getting the job done right." — CEO, Ancoda Labs - "Ritesh built my personal portfolio site and has been my go-to for anything technical since. He explains things clearly and actually follows through, which is rarer than it should be." — Senior Analyst, Salt Trading Corporation Limited - "He took ownership of what he was given and often went past what was asked. Solid engineer, easy to work with. Ritesh was a full-stack developer on our team and consistently delivered on time." — CEO, Milo Logic - "He took full ownership of our frontend work. Clean implementation, sensible decisions, and he pushed back when something didn't make sense rather than just building what he was told." — CEO, Azminds Services Pvt. Ltd. - "Alexa Lifesciences required internal tooling to handle customer inquiries, inventory management, and other operational workflows. Ritesh brought strong technical insight to the project and was a reliable, thoughtful partner throughout the process." — Managing Director, Alexa Life Sciences - "Ritesh's work ethic is unbelievable. He helped built a bulletproof backend and a perfectly sharp frontend. An absolute joy." — Account Executive, Momentum ## Frequently asked ### Are you senior enough to own a product on your own? I was the only engineer on Gradsy. Two repositories, three dashboards for students, counselors and admins, one API behind them that denies every request until a role allows it. Two counselors clicking the same appointment slot cannot both get it, because that rule lives in a database constraint rather than in application code. Background jobs can run twice without doing damage. Architecture, schema and deployment were mine, and so were the bugs. ### You're in Kathmandu. How does that work for a team in Europe/US? I am based in Nepal (UTC+5:45). My afternoon and evening naturally cover a full European workday, making live standups and pair programming seamless. For US teams, my evening perfectly aligns with East Coast mornings, and I am fully flexible to adjust my schedule to ensure we have all the overlapping hours you need. ### How does a company in Europe/US hire someone based in Nepal full time? The same way you hire anyone remote, plus one step: use a cross-border management platform like Niural or Deel. For a flat monthly fee, the platform handles all the legal compliance and tax forms, letting you generate a contract in minutes and pay me by approving a single monthly invoice. If your company prefers to officially hire me as a full-time employee rather than an independent contractor, that is completely possible too; these same platforms can act as an Employer of Record (EOR) to employ me locally on your behalf without you having to open a subsidiary here. ### What does full stack mean in your case? Frontend-heavy. React, Next.js and TypeScript are the strongest part: component libraries, design-system primitives, bundle size, data fetching. Behind that I write Node and NestJS APIs, design the PostgreSQL schema they sit on, and deploy to AWS, Cloudflare or Azure. Infrastructure is the thinnest of the three. I have not run Kubernetes in production. If a role needs it I will learn the cluster, but the platform is not what you would be hiring me for. ### You build with AI. Who is responsible for the code? Me. Agents write the first pass on a lot of what I ship. I read every line before it merges, and the parts that decide money, permissions or data shape I usually write myself. The review is most of the work and I schedule time for it. ### What are you like to work with? I ask what the constraint is before I ask what to build. If an estimate is going to slip, you hear it the day I know. I write documentation while I build, because I have been the only person who could read a system, and that is a bad place to put a team. ## Contact The ask is a 30-minute call: https://riteshkc.com.np/contact - Email: riteshkcdev@gmail.com - GitHub: https://github.com/kcritesh - LinkedIn: https://www.linkedin.com/in/riteshkc - Discord: @riteshkc --- # About Ritesh KC Full-stack developer with 4+ years shipping production web applications. Backend in Node.js, NestJS and PostgreSQL; frontend in React and Next.js. Recent work is LLM systems: RAG, embeddings, vector search and document extraction. Chapters on the page: 00 how i got here, 01 how i work, 02 where i've worked, 03 what i reach for, 04 what i want next. ## How I work ### 01 Start at the constraint Before any code, I map the domain and the failure modes. What breaks under concurrency, what has to stay correct when a job retries, what the deadline actually protects. The constraint decides the design; the feature list does not. ### 02 Ship the thin slice The smallest version that runs end to end, deployed, with the boring parts boring: typed, migrated, observable. No abstraction until the second time I need it. You see something working early enough to change your mind cheaply. ### 03 Own it in production Deployment, the bugs after it, and the documentation written while I build rather than promised for later. If an estimate is going to slip, you hear it the day I know instead of the week it lands. ## Where I have worked ### Freelance Full-Stack Developer, Independent (Dec 2025 - Present) - Gradsy: sole engineer on a study-abroad application platform in Next.js, NestJS and PostgreSQL on AWS. Built three role-scoped dashboards — student, counselor, admin — behind a default-deny API, enforced booking concurrency with database constraints, and made jobs retry-safe. - Noveon: site for a Hong Kong investment firm on Next.js and a self-hosted Strapi CMS. Every page is a CMS entry, so marketing and compliance edit regulator-facing copy without a developer, and six investor pages render from two templates. - Marti China: corporate site for an engineering company on Next.js with Strapi. ### Full-Stack Engineer, Milo Logic (Aug 2025 - Aug 2026) - Built enterprise multi-tenant architecture serving multiple organizations from one deployment, with tenant scoping at the data layer and custom domain support so each organization serves on its own domain. - Shipped a block-based website builder and a drag-and-drop form builder, letting non-technical users compose pages, define fields and validation, and publish without a developer. - Built revenue and analytics dashboards and master data models that keep reference data centrally managed while each organization maintains its own overrides. - Architected shared React and TypeScript component libraries and design-system primitives adopted across 7 products by a 9-person engineering team, and cut JavaScript bundle size to reduce page load times. ### Full-Stack Engineer, Azminds Services (Aug 2024 - Jul 2025) - Yachtcloud: live multi-tenant yacht booking platform where each operator runs as its own business under one admin panel. Built the remaining booking flows and third-party integrations, then used Hotjar recordings and funnel data to find checkout drop-offs and simplify the flow. Tripled sales. - Driving School Wiz: multi-tenant platform serving 5 driving schools, each on its own subdomain, with subdomain scoping on every request so one owner can run multiple schools without data crossover. Real-time instructor scheduling and assignment tracking. ### Associate Frontend Developer, Azminds Services (Apr 2023 - Jul 2024) - Built pension and revenue management modules for the Nepal Agricultural Research Council (NARC), a Government of Nepal agency, giving staff a single tracked record of pension disbursements and revenue entries. - Shipped production React and Next.js features with Redux Toolkit and Redux Saga for complex asynchronous flows, migrated JavaScript codebases to TypeScript, and optimized fetching and caching with React Query. ### Junior Frontend Developer (promoted from Intern), Lanceme Up (Jan 2022 - Mar 2023) - Delivered 4+ client React applications with Redux Toolkit and Material UI from Figma designs, including multi-locale builds with i18n over REST APIs. ## What I reach for ### languages TypeScript, JavaScript, SQL ### frontend React, Next.js, Redux Toolkit, Redux Saga, React Query, Tailwind CSS, Material UI, Astro ### backend & data Node.js, NestJS, Express, PostgreSQL, pgvector, REST APIs, Strapi ### ai & llm RAG, Embeddings, Vector search, Document extraction ### cloud & infra AWS, Cloudflare, Vercel, Azure, Docker ## Full résumé https://riteshkc.com.np/resume --- # Ritesh KC — résumé Top React Developer & Software Engineer in Nepal | Ritesh KC. Kathmandu, Nepal. Available for remote work, US and Europe hours overlap. - Email: riteshkcdev@gmail.com - GitHub: https://github.com/kcritesh - LinkedIn: https://www.linkedin.com/in/riteshkc - Site: https://riteshkc.com.np ## Summary Full-stack developer with 4+ years shipping production web applications. Backend in Node.js, NestJS and PostgreSQL; frontend in React and Next.js. Recent work is LLM systems: RAG, embeddings, vector search and document extraction. ## Experience ### Freelance Full-Stack Developer — Independent Dec 2025 - Present - Gradsy: sole engineer on a study-abroad application platform in Next.js, NestJS and PostgreSQL on AWS. Built three role-scoped dashboards — student, counselor, admin — behind a default-deny API, enforced booking concurrency with database constraints, and made jobs retry-safe. - Noveon: site for a Hong Kong investment firm on Next.js and a self-hosted Strapi CMS. Every page is a CMS entry, so marketing and compliance edit regulator-facing copy without a developer, and six investor pages render from two templates. - Marti China: corporate site for an engineering company on Next.js with Strapi. ### Full-Stack Engineer — Milo Logic Aug 2025 - Aug 2026 - Built enterprise multi-tenant architecture serving multiple organizations from one deployment, with tenant scoping at the data layer and custom domain support so each organization serves on its own domain. - Shipped a block-based website builder and a drag-and-drop form builder, letting non-technical users compose pages, define fields and validation, and publish without a developer. - Built revenue and analytics dashboards and master data models that keep reference data centrally managed while each organization maintains its own overrides. - Architected shared React and TypeScript component libraries and design-system primitives adopted across 7 products by a 9-person engineering team, and cut JavaScript bundle size to reduce page load times. ### Full-Stack Engineer — Azminds Services Aug 2024 - Jul 2025 - Yachtcloud: live multi-tenant yacht booking platform where each operator runs as its own business under one admin panel. Built the remaining booking flows and third-party integrations, then used Hotjar recordings and funnel data to find checkout drop-offs and simplify the flow. Tripled sales. - Driving School Wiz: multi-tenant platform serving 5 driving schools, each on its own subdomain, with subdomain scoping on every request so one owner can run multiple schools without data crossover. Real-time instructor scheduling and assignment tracking. ### Associate Frontend Developer — Azminds Services Apr 2023 - Jul 2024 - Built pension and revenue management modules for the Nepal Agricultural Research Council (NARC), a Government of Nepal agency, giving staff a single tracked record of pension disbursements and revenue entries. - Shipped production React and Next.js features with Redux Toolkit and Redux Saga for complex asynchronous flows, migrated JavaScript codebases to TypeScript, and optimized fetching and caching with React Query. ### Junior Frontend Developer (promoted from Intern) — Lanceme Up Jan 2022 - Mar 2023 - Delivered 4+ client React applications with Redux Toolkit and Material UI from Figma designs, including multi-locale builds with i18n over REST APIs. ## Projects - Noveon (2026) — Full-stack developer. A website for a Hong Kong investment firm, built so people who are not developers can run it. Next.js on the front, Strapi on the firm's own server behind it. Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi. https://riteshkc.com.np/work/noveon - Marti China (2026) — Full-stack developer. A corporate site for the China arm of a Swiss tunnelling supplier. Next.js on the front, self-hosted Strapi behind it: 49 product pages, 17 project references, two languages. Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi, Baidu Maps. https://riteshkc.com.np/work/marti-china - Gradsy (2026) — Full-stack engineer. A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. Stack: Next.js, NestJS, TypeScript, Prisma, PostgreSQL, Redis, BullMQ, AWS, AWS Simple Email Service, S3, Lightsail. https://riteshkc.com.np/work/gradsy - Yacht Cloud (2024) — Frontend Developer. Engineered a multitenant frontend and RAG AI planner that powers YachtCloud and multiple white-labeled broker platforms by instantly filtering massive luxury yacht inventories. Stack: Next.js, React, TypeScript, Tailwind CSS, Elasticsearch, RAG, LangChain. https://riteshkc.com.np/work/yacht-cloud - Investment Banker/Analyst Portfolio (2026) — Designer and developer. A portfolio site for a Kathmandu investment banker who manages private portfolios on the Nepal Stock Exchange. Next.js with Payload inside the same app, and a hero that drifts like a chart. Stack: Next.js, React, TypeScript, Tailwind CSS, Payload CMS, Vercel, Cloudflare R2, NeonDB. https://riteshkc.com.np/work/prajesh-sagar ## Skills ### languages TypeScript, JavaScript, SQL ### frontend React, Next.js, Redux Toolkit, Redux Saga, React Query, Tailwind CSS, Material UI, Astro ### backend & data Node.js, NestJS, Express, PostgreSQL, pgvector, REST APIs, Strapi ### ai & llm RAG, Embeddings, Vector search, Document extraction ### cloud & infra AWS, Cloudflare, Vercel, Azure, Docker ## Education - Bachelor of Computer Application (BCA), Information Technology — Kathmandu College of Technology, Tribhuvan University (Apr 2021 - Aug 2025) - Higher Secondary (10+2), Business and Commerce — Sainik Awasiya Mahavidyalaya, Bhaktapur (Jul 2018 - Aug 2020) ## Languages - English — Professional - Nepali — Native - Hindi — Native --- # Contact Ritesh KC The ask is a 30-minute call. Booking runs through Cal.com (riteshkc/30min), linked from https://riteshkc.com.np/contact. - Email: riteshkcdev@gmail.com - GitHub: https://github.com/kcritesh - LinkedIn: https://www.linkedin.com/in/riteshkc - Discord: @riteshkc (https://discord.com/users/425289494615556116) The same page carries a short inquiry form behind the booking card — name, email, phone, what it is about and a message. The options are read from the CMS rather than listed here, so this line cannot drift from the form: - new build - ongoing work - audit or second opinion - full-time role - something else Available for remote work, US and Europe hours overlap. Based in Kathmandu, Nepal. --- # Services What Ritesh KC is hired to build, in three layers. Most projects need parts of all three. Each entry links to its own page, which names the shipped projects behind it. ## Product interfaces Next.js and React developer for dashboards, admin panels and internal tools - Layer: 01 / interface - Builds: dashboards, admin panels, customer portals, data-heavy tables, AI product interfaces, onboarding flows, design systems, marketing sites - Tools: Next.js, React, TypeScript, Tailwind CSS, React Query, Node.js, Core Web Vitals - Page: https://riteshkc.com.np/services/product-interfaces (markdown: https://riteshkc.com.np/services/product-interfaces.md) The screens your customers and your team use all day. React, Next.js and TypeScript, built so the fortieth click is as fast as the first. ## The system under it Multi-tenant SaaS developer: scheduling, bookings, payments and roles - Layer: 02 / system - Builds: SaaS platforms, multi-tenant apps, booking systems, APIs, auth and roles, payments and billing, background jobs, third-party integrations - Tools: NestJS, Node.js, PostgreSQL, Prisma, Redis + BullMQ, AWS - Page: https://riteshkc.com.np/services/multi-tenant-systems (markdown: https://riteshkc.com.np/services/multi-tenant-systems.md) The half of your product nobody sees until it breaks. NestJS and PostgreSQL, where the database enforces the rules instead of trusting your code to remember them. ## AI that cites its sources RAG developer: AI search and extraction over your own documents - Layer: 03 / retrieval - Builds: RAG search, document extraction, question answering, hybrid search, ingestion pipelines, AI workflows, evaluation sets, AI in an existing product - Tools: PostgreSQL, pgvector, Embeddings, Vector search, Document extraction, evals - Page: https://riteshkc.com.np/services/rag-and-document-ai (markdown: https://riteshkc.com.np/services/rag-and-document-ai.md) AI that answers from your documents instead of guessing. Every answer shows the document it came from, and "I could not find that" is a real answer with its own path through the code. --- # Product interfaces Next.js and React developer for dashboards, admin panels and internal tools The screens your customers and your team use all day. React, Next.js and TypeScript, built so the fortieth click is as fast as the first. - Layer: 01 / interface - Builds: dashboards (the screen someone has open all day, where the second interaction decides everything); admin panels (internal tools for the people running the business rather than the customers); customer portals (the logged-in half of your product, where your customers do their own work); data-heavy tables (thousands of rows that filter, sort and page without the screen locking up); AI product interfaces (streaming answers with their sources beside them, and a real state for when the model is wrong); onboarding flows (the sequence between signing up and the first useful thing happening); design systems (the component library that stops a fourth engineer inventing a fifth button); marketing sites (static and fast, with no framework shipped to the browser to render text) - Tools: Next.js, React, TypeScript, Tailwind CSS, React Query, Node.js, Core Web Vitals - React in production since 2022 - Sole engineer on a three-role platform - Design-system primitives a frontend team builds on ## You have probably said one of these out loud. - The app works, but it feels terrible to use. Almost always a state and data-fetching problem wearing a visual disguise. I fix the interaction first, then the surface. - The UI is becoming impossible to maintain. Five buttons, three of them nearly identical, and none safe to change. I collapse them into one component system with every state drawn. - The dashboard is a pile of components nobody owns. I restructure it around the few patterns that actually repeat, so the next screen is assembled rather than invented again. - AI generated the prototype. Someone has to make it real. I read it, keep what works, and rebuild the rest with types, boundaries, and a deploy you can run twice. - There are designs in Figma nobody has built. I build them faithfully, including the states nobody drew — empty, loading, error, and far too much data. ## React is not the point. What it lets me promise is. A framework is a means of making guarantees. These are the ones worth having, and they are the reason the fourth screen costs less than the first. - reusable components: one button, one source. Your fourth screen becomes assembly rather than invention. - state you can reason about: an order that is both paid and unpaid stops being a bug when the type refuses to let it exist. - interaction that keeps up: filters that respond while your hand is still on the mouse, at row one and at row four thousand. - server-first data: data that lives on the server is fetched there, so the browser downloads a page instead of a program. - real-time surfaces: streams, subscriptions and optimistic updates, without the screen ever telling someone a comfortable lie. - architecture that survives: boundaries still findable in month nine, when the person editing them is not me. ## Idea, interface, product. 01. understand — who uses the screen, what they are trying to finish, and where the current one loses them 02. shape — information architecture and interaction patterns, decided before anything is drawn 03. build — React, Next.js and TypeScript in strict mode, on top of a component system rather than beside one 04. ship — deployed and measured — bundle size, Core Web Vitals, contrast, and every path a keyboard takes ## You bring it. I handle it. You get it back better. - You bring: an idea, a running product, or a prototype somebody generated; the designs, if they exist; the people who have the screen open all day - I handle: frontend architecture and component structure; UI implementation, down to the states nobody drew; state management and data fetching; API integration, and the API itself when it needs one; responsive behaviour from 360px up; bundle size, Core Web Vitals, accessibility - You get: your code, in your repository, from the first commit; a component library the next engineer can read; an interface that holds at every width, measured against the standard rather than judged by eye Interface first, but the line does not stop at the browser. When a screen needs an endpoint, a schema or a queue behind it, I build that too rather than filing a ticket and waiting for it. ## how I build it TypeScript in strict mode, with types that describe your business rather than restating the shape of a JSON response. Most frontend bugs I get called in to fix were a state nobody should have been able to reach. An order that is both paid and unpaid. A form that is submitting and still editable. Those stop being bugs when the type refuses to let them exist. Bundle size and data fetching get an owner on day one instead of a cleanup ticket the week before launch. Data that lives on the server is fetched on the server. Interactive components stay small and sit at the edges of the tree. Anything that wants to ship a library to the browser so it can render text has to make its case. I care what it looks like, and I have opinions I can defend. Spacing on a scale instead of whatever number felt right. A type ramp that still reads at 360 pixels and at 1600. Colour that means something, used sparingly, so that when one thing is highlighted you know why. That is not decoration. It is the difference between a screen someone tolerates and a screen someone trusts. Accessibility happens while I build, not in an audit afterwards. I read the markup and I tab the page. Focus rings stay visible. Semantics come from using the right element, not from an ARIA attribute patched over a `div`. ![The same builder surrounded by distinct interface panels — an analytics chart, a kanban board, a data table, a settings form and a stack of notifications — wired together with thin lines into one connected product.](https://pub-0d70e426aadb4f069a19cb045dfe20f4.r2.dev/media/product-interfaces-builds.avif) ## when a project needs more than me When a project wants brand work or deeper product design alongside the build, I bring in designers I already work with. One arrangement, one schedule, and one person answerable for how it turns out. ## Proof - none listed --- # The system under it Multi-tenant SaaS developer: scheduling, bookings, payments and roles The half of your product nobody sees until it breaks. NestJS and PostgreSQL, where the database enforces the rules instead of trusting your code to remember them. - Layer: 02 / system - Builds: SaaS platforms (one product, many paying customers, kept apart and billed separately); multi-tenant apps (one deployment, many organisations, and no data crossing between them); booking systems (two people cannot take the same slot, and the database is what stops them); APIs (endpoints your own frontend and someone else's client both call); auth and roles (who can do what, checked in one place at the edge of the system); payments and billing (subscriptions, invoices, and webhooks that arrive twice); background jobs (the work that runs at 3am, written so a second run does nothing); third-party integrations (someone else's API, retried through the hours it is down) - Tools: NestJS, Node.js, PostgreSQL, Prisma, Redis + BullMQ, AWS - Three multi-tenant products in production - Driving schools, yacht charter, study abroad - Sole engineer on Gradsy: schema, API, deploy ## You have probably said one of these out loud. - We sold to a second company, and now there is an if statement deciding whose data comes back. It works, and it works until someone forgets the if. I move the separation into the schema, where forgetting is not an option the code has. - Two customers booked the same slot. We refunded one and nobody knows why it happened. Your API read the slot, then wrote to it. The second booking got in between those two steps. A unique constraint closes that gap; no amount of checking in code does. - The payment provider sent the same webhook twice and we charged twice. At-least-once is how every queue and every webhook works. Every job I write does nothing on the second run. - Nobody on the team wants to own the boring half. The permissions, the billing edge cases, the job that runs at 3am and has to be right. That is the half I take. - One person understands this system and none of it is written down. I document while I build — the schema, the roles, the jobs and the reasons behind the awkward decisions. I have been that one person, and it is a bad place to leave a team. ## A rule belongs in the database if the database can hold it. Application code checks a rule. A database enforces one. The difference shows up the day two people click the same button in the same second. - one unique constraint: two people book the same slot. The second write fails, and that person is told it has gone. - states in a table: the moves an application may make are rows you can read, not an if chain spread across three services. - jobs that repeat safely: a queue delivers at least once, which sometimes means twice. Charge once, email once, count once. - tenancy in the schema: the subdomain scopes the query, so an if statement is never the thing keeping two customers apart. - one place for permissions: who is allowed to do this is a question you answer by opening one file. - migrations that reverse: a bad release rolls back the same way it went out, instead of becoming an incident. ## Constraint, schema, API, run. 01. constrain — what must never happen — a double booking, a crossed tenant, a card charged twice 02. model — the schema that makes those things impossible, written before a single endpoint exists 03. build — NestJS and Node on PostgreSQL, permissions checked at one edge, jobs safe to run twice 04. run — deployed to AWS with the pipeline that puts it there, so shipping is a Tuesday rather than an event ## You bring it. I handle it. You get it back better. - You bring: a product with a second customer, or a system that broke once; the rule that must never be violated; whoever answers the support ticket when it is - I handle: schema design, and migrations that run both ways; multi-tenant isolation and role-based access; APIs your frontend and someone else's client both call; queues, background jobs and retry safety; payments, invoices and duplicate webhooks; deployment and the pipeline that does it - You get: your code, in your repository, from the first commit; documentation written while I build, not promised for the end; a system where the rules are enforced by the database rather than remembered by people Schema first, then the API, then the interface. I build the screens on top as well, so there is no seam where my work stops and somebody else's is supposed to start. ## how I build it A rule belongs in the database if the database can hold it. Two counselors open the same appointment slot and both click book. Postgres lets both of them read that slot before either one saves. Both see it free. Both take it. No amount of checking in your API code closes that gap, because the gap sits between the read and the write, and it is as wide as your traffic is heavy. A unique constraint on the table does close it. The second write fails, and the second person is told the slot has gone. State works the same way. An application that moves through stages gets a table of states and the moves allowed between them, not a text column and a pile of if statements spread across three services. Then "can this go from submitted to accepted" is a row you can look at instead of a code review you have to run. Queues promise to deliver a message at least once. In practice that sometimes means twice. So every background job is written to do nothing on the second run. Charge once. Email once. Count once. Schema first, then the API, then the interface. Migrations run backwards as well as forwards. Permissions are checked in one place at the edge of the system, so "who is allowed to do this" is a question you answer by opening one file. ## when a project needs more than me I deploy and run the infrastructure most products need. Some projects need more than that: dedicated platform work, heavy data pipelines, an environment with real operational demands. For those I bring in senior infrastructure engineers I already work with. The database and the application stay with me. The platform goes to somebody who does it full time. You are still dealing with one arrangement rather than three, and one person answerable for how it turns out. ## Proof - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. Stack: Next.js, NestJS, TypeScript, Prisma, PostgreSQL, Redis, BullMQ, AWS, AWS Simple Email Service, S3, Lightsail. ## Related writing - [Overbooking is a database problem](https://riteshkc.com.np/blog/overbooking-is-a-database-problem): Counting rows before you insert one is not a capacity check. Here is how a unique index, a constraint violation, and a NULL ended up doing the work instead. - [A status machine belongs in a table](https://riteshkc.com.np/blog/a-status-machine-belongs-in-a-table): Nine statuses, two sets of allowed transitions, and two names for every state. Moving all of that into rows instead of an if chain got engineering out of the workflow business. - [Rate limiting a provider without a rate limiter](https://riteshkc.com.np/blog/rate-limiting-by-queue-shape): SES caps sends per second. The fix was worker concurrency of one and a sleep, plus an honest note in the code about the case where it stops working. --- # AI that cites its sources RAG developer: AI search and extraction over your own documents AI that answers from your documents instead of guessing. Every answer shows the document it came from, and "I could not find that" is a real answer with its own path through the code. - Layer: 03 / retrieval - Builds: RAG search (search that understands the question, over documents only you have); document extraction (named fields pulled out of contracts, forms and reports); question answering (answers with their sources beside them, so anyone can check one); hybrid search (meaning and filters in one query, because most real questions are both); ingestion pipelines (PDFs, scans and exports read in, and kept current as they change); AI workflows (a model doing one bounded step inside a process you already run); evaluation sets (real questions with the sources that should answer them, rerun on every change); AI in an existing product (retrieval added to software you already run, not a second product beside it) - Tools: PostgreSQL, pgvector, Embeddings, Vector search, Document extraction, evals - Retrieval from prototype to production at Milo Logic - Vectors in PostgreSQL, beside the relational data - An evaluation set built with the system, not after it ## You have probably said one of these out loud. - We tried a chatbot on our documents. It sounded confident, it was wrong twice, and now nobody uses it. It answered without the right passage in front of it. I fix the retrieval first, then put the source on screen so the next wrong answer is caught in a second rather than in a quarter. - One person on the team finds the right paragraph and everybody else waits for them. That is a search problem with a person standing in for the index. The answers are already in the documents; nothing is there to ask them with. - The answers are somewhere across ten years of contracts and tickets. Ingestion, chunking and vector search over the corpus you already have, with the filters your team already searches by. - We need fields out of these documents, not a chat window. Same pipeline, different ending. You get named fields, so the rest of your software can treat a contract as data. - How would we even know whether it got better? An evaluation set built alongside the system — real questions, the sources that should answer them, rerun on every change. The alternative is arguing about it in a meeting. ## A general model guesses. Retrieval makes it look first. Fine-tuning teaches a model a style. Retrieval hands it the paragraph. Only one of those can show you where the answer came from. - your documents, not the internet: the model answers from your contracts and your tickets, not from what it read during training. - recall before precision: a wrong passage gets ignored. A missing one gets answered anyway, fluently, and it sounds the same as when it is right. - meaning and filters together: renewals agreed after March is half a search and half a date column. Both run in Postgres, in one query. - sources on the screen: an answer you cannot check is a rumour. The first wrong one nobody caught is what ends the project. - not sure is an answer: an endpoint that must always produce something will always produce something. - measured, not argued: every change to chunking, the embedding model or the prompt is scored against your own questions. ## Corpus, retrieval, answer, evidence. 01. read — your documents in — PDFs, scans, exports, and whatever else the last ten years left behind 02. cut — each document split into pieces that still make sense on their own, away from the rest 03. find — vector search over those pieces, running beside your ordinary filters in the same query 04. answer — the answer, the sources on screen next to it, and a real path for "I could not find that" ## You bring it. I handle it. You get it back better. - You bring: the documents, and the questions your team cannot get answered; whoever knows which answer is the right one; the filters you already search by — client, date, status - I handle: ingestion, chunking and embedding; vector search in PostgreSQL, beside your relational data; ranking, and the path for when nothing is found; the interface that puts sources next to answers; the evaluation set, and the score on every change; deployment, inside your own infrastructure - You get: your code and your documents, in your infrastructure; answers with the source beside them, checkable in a second; a real number for retrieval quality on your corpus, instead of a promise about it The vectors live in PostgreSQL beside the rest of your data, so a question with a filter in it stays one query instead of a join written by hand in application code. I build the screen on top as well. ## how I build it The vectors live in PostgreSQL, beside the rest of your data, through an extension called pgvector. Most real questions are half meaning and half filter. "What did we agree with this client about renewals, in contracts signed after March." The first half needs vector search. The client and the date are ordinary columns. Split those across two systems and you write that join by hand, in application code, every time somebody asks. Retrieval aims for recall before precision. A wrong passage in front of the model usually gets ignored. A missing one is worse: the model does not tell you it came up empty. It answers anyway, fluently, and it sounds the same as when it is right. So the search step brings back more than it needs on purpose, and a ranking step narrows it down. Every answer shows the documents it came from, on screen, next to the words. That one detail decides whether anyone is still using the system in month three. An answer you cannot check is a rumour. The first time somebody catches one being wrong, they stop trusting all of them. "I could not find that" is a real answer with its own path through the code. An endpoint that must always produce something will always produce something. ## when a project needs more than me Retrieval, extraction and answering cover most of what teams actually need, and they share a useful property: every output can be checked against a source. Some projects reach past that, into model fine-tuning, custom training or heavier machine learning. For those I bring in senior AI engineers I already work with. The pipeline, the data model and the interface stay with me, and the specialist work happens inside the same project. ## Proof - none listed ## Related writing - [In RAG, recall is the number that matters](https://riteshkc.com.np/blog/recall-beats-precision): Why the failure mode that sinks a retrieval system is the right document never showing up, and why that makes recall, not prompt tuning, the thing to optimize. - [Endpoints that refuse to be oracles](https://riteshkc.com.np/blog/endpoints-that-refuse-to-be-oracles): A 404 on unsubscribe tells an attacker which tokens are real. A 409 on subscribe tells them who is on your list. Honest status codes leak, and the fix reads like a bug. --- # Work Projects Ritesh KC has shipped. Each entry links to its project page; a project with a long-form write-up also links to its case study. ## Noveon Next.js and Strapi website for a Hong Kong investment firm - Year: 2026 - Client: Noveon Capital - Role: Full-stack developer - Duration: 3 months - Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi - Page: https://riteshkc.com.np/work/noveon (markdown: https://riteshkc.com.np/work/noveon.md) - Case study: https://riteshkc.com.np/case-studies/noveon - Live: https://noveon.wecreatelabs.com.hk/ A website for a Hong Kong investment firm, built so people who are not developers can run it. Next.js on the front, Strapi on the firm's own server behind it. ## Marti China Bilingual corporate site and 49-page product catalogue on Next.js and Strapi - Year: 2026 - Role: Full-stack developer - Duration: 2 months - Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi, Baidu Maps - Page: https://riteshkc.com.np/work/marti-china (markdown: https://riteshkc.com.np/work/marti-china.md) - Case study: https://riteshkc.com.np/case-studies/marti-china - Live: https://marti.wecreatelabs.com.hk/ A corporate site for the China arm of a Swiss tunnelling supplier. Next.js on the front, self-hosted Strapi behind it: 49 product pages, 17 project references, two languages. ## Gradsy Turning a Manual Workflow into a Scalable Product - Year: 2026 - Role: Full-stack engineer - Duration: Handover Completed - Stack: Next.js, NestJS, TypeScript, Prisma, PostgreSQL, Redis, BullMQ, AWS, AWS Simple Email Service, S3, Lightsail - Page: https://riteshkc.com.np/work/gradsy (markdown: https://riteshkc.com.np/work/gradsy.md) - Case study: https://riteshkc.com.np/case-studies/gradsy - Live: https://dev.gradsy.io/ A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. ## Yacht Cloud YachtCloud: Engineering a Multitenant Platform for Luxury Charters - Year: 2024 - Client: AZMinds - Role: Frontend Developer - Duration: 7 months - Stack: Next.js, React, TypeScript, Tailwind CSS, Elasticsearch, RAG, LangChain - Page: https://riteshkc.com.np/work/yacht-cloud (markdown: https://riteshkc.com.np/work/yacht-cloud.md) - Case study: https://riteshkc.com.np/case-studies/yacht-cloud - Live: https://www.yachtcloud.net Engineered a multitenant frontend and RAG AI planner that powers YachtCloud and multiple white-labeled broker platforms by instantly filtering massive luxury yacht inventories. ## Investment Banker/Analyst Portfolio Next.js and Payload CMS portfolio website for a NEPSE portfolio manager - Year: 2026 - Client: Prajesh Sagar - Role: Designer and developer - Duration: 1 month - Stack: Next.js, React, TypeScript, Tailwind CSS, Payload CMS, Vercel, Cloudflare R2, NeonDB - Page: https://riteshkc.com.np/work/prajesh-sagar (markdown: https://riteshkc.com.np/work/prajesh-sagar.md) - Case study: https://riteshkc.com.np/case-studies/prajesh-sagar - Live: https://prajeshsagar.com.np A portfolio site for a Kathmandu investment banker who manages private portfolios on the Nepal Stock Exchange. Next.js with Payload inside the same app, and a hero that drifts like a chart. --- # Noveon Next.js and Strapi website for a Hong Kong investment firm - Year: 2026 - Client: Noveon Capital - Role: Full-stack developer - Duration: 3 months - Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi - Live: https://noveon.wecreatelabs.com.hk/ ## Summary A website for a Hong Kong investment firm, built so people who are not developers can run it. Next.js on the front, Strapi on the firm's own server behind it. ## Problem Noveon has no developer and did not want one on retainer. A marketing lead and a compliance officer had to be able to change any word on the site, including pages a regulator may read. ## Outcome Every page is an entry in a CMS the firm hosts itself. Six investor pages come from two templates, so a fourth investor type is a new entry rather than new code. Noveon is an investment firm in Hong Kong. It buys stakes in private equity funds from investors who want their money back before the fund pays out. > Screenshot (home hero): Noveon home page: the headline Asia's Global Private Equity Secondaries Platform over footage of the Hong Kong skyline and the IFC tower. The firm has no developer and did not want one on retainer, and that single fact shaped the build. Everything a visitor reads is an entry in a CMS running on the firm's own server, so the marketing lead can rewrite a page and the compliance officer can correct a disclosure without opening a ticket with me. > Screenshot (investor split): Home page section titled Empowering Tomorrow's Private Markets with two photographic panels, Investment Strategies on the left and Investors on the right. — Two doors, and behind them six pages built from two templates. The part worth showing is the investor section. Noveon runs three investment strategies and sells to three kinds of investor, which is six pages that a visitor compares closely. They come from two templates rather than six hand-built layouts, so they cannot drift apart, and a fourth investor type is a new entry rather than new code. > Screenshot (investor page): The Family Offices page: three numbered cards headed Bespoke Strategies designed for You, covering risk management, access to Asia and co-investment. — One of three investor pages, from the second of the two templates. The build is finished. It is not live yet: Noveon is still writing the words that go in it. The case study covers the CMS decision, how three kinds of insight share one card, and the two things I would change before launch. ## Screens ### Home - Noveon home page: the headline Asia's Global Private Equity Secondaries Platform over footage of the Hong Kong skyline and the IFC tower. - The site's opening frame: the Noveon wordmark and the line Mastery of Private Markets over slowed footage of a koi carp. — The first load plays this before the home page. It is the brand's own opening, not a loading screen. - Home page statement section: a large serif paragraph about pioneering private market solutions over green macro footage, with the first figures below it. - The core values rail: four photographic cards captioned client-aligned partnership, intelligence-driven investing, forward-looking innovation, stewardship. — Drag to advance, progress bar under it. Nothing in the component assumes four. - Home page section titled Empowering Tomorrow's Private Markets with two photographic panels, Investment Strategies on the left and Investors on the right. — Two doors, and behind them six pages built from two templates. - Home page insights band titled Insights Uncovered with three cards: two podcasts and a case study, each with its own image and label. — One card serves all three kinds of insight. The label is the content type. ### Details - The Investors menu open over the home page: an investment strategy column of three links, an investor column of three, and a photograph beside them. — The menu is the content model drawn: three strategies, three investor types. - The site's 404 page: the numerals 404 in the display face, the line We couldn't find the page you requested, and a back to homepage link. ### About - About page section titled Innovating Private Equity through Vision and Experience, with the firm's origin story in a column beside portraits. - About page mission block: three expandable rows, the first open on Unlock Asia's Private Markets, with two overlapping photographs at the right. ### Team - Team page: the heading The People Behind Noveon over two member portraits, each with a play control for a video introduction. - Our Office section: a tinted satellite map of east and southeast Asia with two pins, beside the Hong Kong address, phone number and email. — The office block is an entry too, pins included. ### Investors - The Asia General Partner Solutions page: a blue band headed Asia's Private Equity Moment with three figures for transactions, capital provided and GPs covered. — One of three strategy pages. All three are the same template with different entries. - The Family Offices page: three numbered cards headed Bespoke Strategies designed for You, covering risk management, access to Asia and co-investment. — One of three investor pages, from the second of the two templates. ### Partnerships - Partnerships page: the heading Noveon partners with you for innovative investment solutions beside four numbered reasons in a two by two grid. ### Labs - Noveon Labs page opening: the line One Platform. Complete Secondary Intelligence., a paragraph about the NAIP product and a request invite button. — The one page selling software rather than the firm. - What NAIP Delivers section: a video panel on the left, four numbered feature tiles on the right, the first one selected. — Four features, one selected, a video per feature. The count comes from the entry. - Noveon Labs FAQ: four collapsed questions about the product beside a card offering an email reply to anything not covered. ### Sustainability - Sustainability page section headed Embedding sustainability, with four coloured cards: ESG due diligence, active ownership, impact measurement, improvement. ### Insights - The all articles listing: a breadcrumb, the heading All Articles, a sort control at the right and the first article card below. — Each of the three insight types has this listing behind it. - An article page: a table of contents pinned at the left with the current section marked, the article's own sections and images at the right. — The contents list is built from the sections in the entry, so a new section adds a row. - Knowledge Center page: full-width cards for investment strategies and investment comparisons, each listing what the library holds. ### Contact - Contact page: the email address and Hong Kong office address on the left, a five-field message form on a dark green panel at the right. ### Mobile - The navigation at phone width: a full-screen sheet with a search field at the top and the nav items below, Investors and Insights carrying disclosure arrows. — Search moves to the top of the sheet on mobile, where the header has no room for it. - The Noveon Labs feature grid at phone width: the four numbered tiles reflowed into two columns under the heading. ## Case study Six investor pages, two templates, and no developer on retainer: https://riteshkc.com.np/case-studies/noveon (markdown: https://riteshkc.com.np/case-studies/noveon.md) --- # Marti China Bilingual corporate site and 49-page product catalogue on Next.js and Strapi - Year: 2026 - Role: Full-stack developer - Duration: 2 months - Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi, Baidu Maps - Live: https://marti.wecreatelabs.com.hk/ ## Summary A corporate site for the China arm of a Swiss tunnelling supplier. Next.js on the front, self-hosted Strapi behind it: 49 product pages, 17 project references, two languages. ## Problem Marti China needed a site its own staff in Zhongshan could keep current — a product catalogue, project references and service pages, in English and Chinese, with no developer in the loop. ## Outcome A Next.js site over self-hosted Strapi. Staff add a product, a project or a service and the page exists. The project map runs on Baidu, so it loads inside mainland China. Marti China makes the components that move excavated rock out of a tunnel. The reader is a contractor's engineer or a procurement lead, usually mid-project, usually after a specification and a reference before they email anyone. The site is built for that reader: what the parts are, what the company does around them, where the parts already run, and a form that asks for belt width rather than for a message. Nine page templates cover it — home, about, quality and sustainability, services and a page per service, projects, the catalogue and a page per product, contact. Every word on every one of them is a field in Strapi, in two languages, so the team in Zhongshan adds the fiftieth product without me. > Screenshot (marti cms): Marti China Strapi CMS — Strapi Home Page Eighty-six published entries behind the site, and the last four edited are a page and three products. That is the whole test of this build: the recent activity is the client's, not mine. > Screenshot (delivery model): Home page section titled The Marti China Delivery Model, four photographic cards for engineering, manufacturing, site support and quality control. — Four services, four cards, four pages. The cards are entries, not layout. Four services, four pages, one template. Each carries its own process, its own deliverables and the part of the business it sits in, which is what keeps four pages built from one file from reading as four copies of a page. > Screenshot (services grid): Services page showing four photographic cards: engineering and consultancy, advanced manufacturing, construction and site support, quality control. It holds down to phone width, where the category rail collapses above a single column and the project map keeps its list panel rather than losing it. Why Strapi rather than a hosted CMS, why the map is Baidu, and how two languages sit on one content model are in the case study. ## Screens ### Strapi - Marti China Strapi CMS — Strapi Home Page ### Home - Marti China home page over a photograph of an engineering meeting, with the headline Swiss Engineering Excellence and an Explore Services button. - Home page section titled The Marti China Delivery Model, four photographic cards for engineering, manufacturing, site support and quality control. — Four services, four cards, four pages. The cards are entries, not layout. - Home page product section with five category tabs above three product cards for carrying, return and impact idlers. — The five tabs are the five categories in Strapi. Adding a sixth adds a tab. ### Catalogue - Products page: a sticky rail of five categories on the left, the idler rollers grid on the right, and a download all datasheet button. — Five categories, forty-nine products, one template. The rail is an anchor per category. - Carrying idlers product page: a numbered image column on the left, the category label, name and description on the right. — The forty-ninth product page and the first one are the same file. ### Projects - Projects page: a Baidu map of Europe with Chinese labels, a floating panel listing tunnelling references by country and year. — Baidu, not Google. The primary reader is inside mainland China. - The same map after picking the Follo Line reference: the map has panned to Norway and the list row is highlighted. — Picking a reference pans the map. The list and the map are one selection. ### Services - Services page showing four photographic cards: engineering and consultancy, advanced manufacturing, construction and site support, quality control. - The engineering and consultancy page: a blue band titled The Engineering Process with four numbered steps and a timing note under each. — Each of the four services has its own page on this template. - Engineering deliverables section listing 3D system model, 2D layout and production drawings, structural calculations and bill of materials. ### Quality - Quality inspection section: four numbered cards over a photograph of two inspectors, covering receipt checks, in-process controls, verification and the report. - Safety section titled A Safe Facility, A Skilled Team, with cards for workplace safety, training, facility security, confidentiality and emergency preparedness. ### About - About page history timeline: years down the left from 1922 to 2023, the selected year 2007 enlarged beside a photograph of Marti, Swiss and Chinese flags. — The timeline advances with the scroll. Each year is an entry with a photograph. - About page fact block: the Marti Group founded 1922, over 80 companies, over 100 years of engineering, China established 2011 in Zhongshan. ### Contact - Contact page: email and phone on the left, a technical inquiry form on the right asking for name, company, email, project location, type and specification. — Six fields. The last one asks for capacity, belt width, incline and operating conditions. ### Details - The header's language control open, showing EN selected and SC beside it. — English and simplified Chinese, per field in the CMS rather than per site. - The site's 404 page: the numerals 404, the line Page Not Found, and a Back to Home Page button. ### Mobile - The projects page at phone width: the hero photograph of a conveyor installation with the map beginning below it. - The products page at phone width, the category rail collapsed above a single column of product cards. ## Case study Who edits it after I leave, and which map loads in China: https://riteshkc.com.np/case-studies/marti-china (markdown: https://riteshkc.com.np/case-studies/marti-china.md) --- # Gradsy Turning a Manual Workflow into a Scalable Product - Year: 2026 - Role: Full-stack engineer - Duration: Handover Completed - Stack: Next.js, NestJS, TypeScript, Prisma, PostgreSQL, Redis, BullMQ, AWS, AWS Simple Email Service, S3, Lightsail - Live: https://dev.gradsy.io/ ## Summary A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. ## Problem Students struggled to find suitable universities and manage applications, while counselors and admins had fragmented workflows for tracking students, documents, and applications. ## Outcome Gradsy brought university discovery, student applications, and consultancy workflows into one platform - simplifying the experience for students, counselors and admins. From a consultancy website to a product built around the real problems students face The project started with a simple request from [**UniHub**](https://unihubnetwork.com/), an education consultancy in Nepal - build a modern landing page for their business. We built the website, but as we worked closely with the team, we started seeing a bigger problem. Students, counselors, and admins were managing too much of the admission process manually. Information was scattered, communication was messy, and keeping track of students through different stages was difficult. That became **Gradsy**. ![unihub landing page hero](https://cdn.riteshkc.com.np/media/unihub-landing.avif) ## From landing page to product What started as a website turned into a platform designed around the actual workflow of an education consultancy. I worked as the full stack developer, taking care of the frontend, backend, authentication, file management, emails, and deployment. The goal wasn't to build another dashboard. It was to make the day-to-day work easier for students, counselors, and admins. ![An admin dashboard with four counters, quick actions, and an activity feed naming the actor and their role for each recent change.](https://cdn.riteshkc.com.np/media/admin-dashboard.avif) ## Why Next.js Gradsy had two very different sides - a public-facing experience and a highly interactive application. **Next.js** gave me the right balance. I could build fast, SEO-friendly pages for the public side while using the same application for the complex dashboard experience. It also kept the frontend architecture straightforward as the product grew. ![A student profile with personal and contact fields redacted, beside a sidebar tracking completion across personal, academic and preferences sections.](https://cdn.riteshkc.com.np/media/student-profile.avif) ## Why NestJS As the number of workflows increased, the backend needed structure. I chose **NestJS** because it gave me a clean way to organize the application around modules, services, controllers, guards, and business logic. This became especially useful when different users started having different roles and workflows. Students needed one experience. Counselors needed another. Admins needed a much broader view. The backend had to keep all of that organized. ![nestjs architecture](https://cdn.riteshkc.com.np/media/nestjs.avif) ## Authentication with Better Auth Authentication wasn't something I wanted to build from scratch. **Better Auth** gave us the foundation for secure authentication and session management while keeping the implementation flexible enough for Gradsy's different user roles. That meant less time maintaining authentication code and more time solving the actual product problems. ![A sign-in screen with an email field, a Cloudflare Turnstile notice, a button to email a sign-in link, and a password fallback below.](https://cdn.riteshkc.com.np/media/student-signin.avif) ![The delivered sign-in email, stating that the link works once and expires in 24 hours.](https://cdn.riteshkc.com.np/media/student-magic-link-email.avif) **Magic Link** signup using [Better Auth](https://better-auth.com/) and AWS Simple Email Service SES. ## AWS behind the scenes The application needed reliable infrastructure without making the project unnecessarily complicated. I used **Amazon S3** for storing documents and other uploaded files. **Amazon SES** handled transactional emails and communication from the platform. And **Lightsail** gave us a simple way to deploy and run the application while keeping infrastructure costs predictable. ![aws infra with lightsail SES](https://cdn.riteshkc.com.np/media/aws-infra.avif) ## What I built Gradsy became a full workflow platform for the consultancy. Students could manage their applications and information. Counselors could manage students and their progress. Admins could oversee the entire operation. And behind all of that was a backend, authentication system, file storage, email infrastructure, and deployment setup designed to keep everything connected. ## The part I liked most The interesting part of Gradsy wasn't choosing Next.js or NestJS. It was watching a simple landing page turn into a real product. We started by building what the client asked for. Then we spent enough time understanding how they actually worked. That changed what we built. ## Screens ### public - A public universities directory with filters for country, degree level, tuition, institution type and intake, beside one result card. — The public side of the same data admins edit. Students browse before they ever create an account. - A university profile page: photo gallery, a row of statistics, an about section, and a sidebar listing tuition, programmes and acceptance rate. — One university, assembled from the fields an admin fills in. Rank, cost and acceptance rate sit where a student compares them. ### student - A sign-in screen with an email field, a Cloudflare Turnstile notice, a button to email a sign-in link, and a password fallback below. — Sign-in is a magic link by default, with Turnstile in front of it and a password path kept for anyone who wants one. - The delivered sign-in email, stating that the link works once and expires in 24 hours. — The link is single-use and expires in 24 hours. Delivery goes out through SES on a queue rather than inline with the request. - A student profile with personal and contact fields redacted, beside a sidebar tracking completion across personal, academic and preferences sections. — Personal details, redacted here. The sidebar tracks completion across three sections and reports a single percentage. - A course preferences screen covering study level, fields of study, preferred programs, destination countries, budget range and intake. — Preferences are structured fields rather than free text: study level, subjects, countries, budget and intake. - A three-step new-application wizard on step one, choosing destination country, degree level and university. — Applying is three steps with a running summary and a save-draft escape hatch. A student doing this once, under stress, never sees the whole form at once. - A document library, six categories with four required, every file carrying a verified badge and a banner confirming the required set is complete. — Documents are uploaded once and reused across applications. The required set is a gate: until every one of them verifies, submission stays closed. - One application's status history, seven entries deep, each state change stamped with a date and the counselor's written reason. — The student sees the same state machine the counselor drives, rejections included. Every transition carries the reason somebody typed. - A student's articles list filtered by state, showing one article card marked under review. — Students write articles too. The same review states apply: draft, under review, published, changes requested, archived. - A community feed with vote counts, tag filters, a pinned announcement from the Gradsy team, and a sidebar of tags and guidelines. — One feed, tagged by topic and voted on. The pinned post carries a team badge so official answers are distinguishable from replies. ### counselor - A counselor's applications board with four columns - draft, pending review, sent to university, decided - each showing a count. — The counselor's queue is the pipeline's states, in order. Where an application has got to is the column it is sitting in. - A delegated task queue with one open document check, filters by state, and claim, start, complete and escalate actions on the detail panel. — Tasks are claimed off a shared queue, so the claim is the operation that has to hold when two counselors click at once. Escalation hands one back. ### admin - An admin dashboard with four counters, quick actions, and an activity feed naming the actor and their role for each recent change. — The activity feed names who did what, role badge attached. Most admin questions turn out to be questions about who last touched a record. - A reports screen in light theme: applications by status, a country breakdown, top universities by application count, and an activity feed. — Reporting, in the light theme. The counts are from a staging database, not production traffic. - An admin reviewing one application: verified documents, a warning that a missing statement of purpose blocks submission, and a status-change form. — Review is a status change plus a note, and the note is what the student reads. The missing-document rule blocks the student, not the reviewer. - An admin delegation panel setting a counselor's delegation level, maximum concurrent tasks, and which categories of work they may take. — Delegation is per-permission and capacity-capped: a level, a task ceiling, and a checklist of what this counselor may do. Seniority gates the rest. - The tasks tab of the same counselor record, listing an assigned task with its schedule and the comment thread between admin and counselor. — The same counselor record, tasks tab. Assignment, schedule and the thread live together so a handover has one place to look. - A university editor with tabs for overview, academics, costs, media and programs, showing basic information, location and record metadata. — Universities are edited here and rendered publicly from the same record. The slug is visible because it is a URL somebody may already have shared. - An article review queue filtered to under review, with author-type filters and approve, reject and preview actions on the card. — Articles from students and counselors land in one queue. Approve, request changes, or preview it as published. - A crop dialog over the article editor, with a zoom slider, rotate controls, a rule-of-thirds grid and a live preview of the cover. — Cover images are cropped in the browser before upload, so the stored file is already the shape the card needs. ## Case study What the student sees, what the office sees, from one set of records: https://riteshkc.com.np/case-studies/gradsy (markdown: https://riteshkc.com.np/case-studies/gradsy.md) ## Related writing - [A status machine belongs in a table](https://riteshkc.com.np/blog/a-status-machine-belongs-in-a-table): Nine statuses, two sets of allowed transitions, and two names for every state. Moving all of that into rows instead of an if chain got engineering out of the workflow business. - [Endpoints that refuse to be oracles](https://riteshkc.com.np/blog/endpoints-that-refuse-to-be-oracles): A 404 on unsubscribe tells an attacker which tokens are real. A 409 on subscribe tells them who is on your list. Honest status codes leak, and the fix reads like a bug. - [Every reset link returned 400](https://riteshkc.com.np/blog/every-reset-link-returned-400): The reset endpoint was fine. The token was fine. The email was fine. The bug was one path in a captcha config, and the word doing the damage was "includes". - [Ready to submit is not a boolean](https://riteshkc.com.np/blog/ready-to-submit-is-not-a-boolean): A submit gate that looked like one rule and needed three, plus an error message that names which of four things the student actually has to fix. - [Six socket events, and not one refetch](https://riteshkc.com.np/blog/six-socket-events-and-not-one-refetch): Realtime usually arrives as a second copy of your state. Treating every socket event as a write into the query cache instead of a signal to refetch keeps it to one. - [A trust score you can rebuild from the log](https://riteshkc.com.np/blog/a-score-you-can-rebuild-from-the-log): Ten lines that replay a moderation history into a score, and why clamping inside the fold instead of after it is the difference between working and quietly drifting. - [Cache invalidation when the cache is in another repo](https://riteshkc.com.np/blog/cache-invalidation-in-another-repo): An admin edits a university and the public page keeps the old number for a day. The backend now calls the frontend, over tag names nothing checks. - [Overbooking is a database problem](https://riteshkc.com.np/blog/overbooking-is-a-database-problem): Counting rows before you insert one is not a capacity check. Here is how a unique index, a constraint violation, and a NULL ended up doing the work instead. - [Rate limiting a provider without a rate limiter](https://riteshkc.com.np/blog/rate-limiting-by-queue-shape): SES caps sends per second. The fix was worker concurrency of one and a sleep, plus an honest note in the code about the case where it stops working. - [The fast moderation layer does less on purpose](https://riteshkc.com.np/blog/the-fast-moderation-layer-does-less): A synchronous gate on five routes that blocks almost nothing, and an async worker that decides everything else. Splitting them by latency was the wrong axis. --- # Yacht Cloud YachtCloud: Engineering a Multitenant Platform for Luxury Charters - Year: 2024 - Client: AZMinds - Role: Frontend Developer - Duration: 7 months - Stack: Next.js, React, TypeScript, Tailwind CSS, Elasticsearch, RAG, LangChain - Live: https://www.yachtcloud.net ## Summary Engineered a multitenant frontend and RAG AI planner that powers YachtCloud and multiple white-labeled broker platforms by instantly filtering massive luxury yacht inventories. ## Problem High net worth users demand lightning fast browsing. We needed a scalable system for franchise brokers to sell our yacht inventory under their own brands without duplicating code bases. ## Outcome Built a centralized frontend engine. It delivers instant inventory filtering and easily deploys standalone branded broker sites like Exclusive Gulets from a single codebase. **How I Engineered a Digital Marina for Ultra-Luxury Yachts** Selling a $50,000-a-week yacht charter isn't like selling a t-shirt. Buyers expect a flawless digital experience. If the site lags? They bounce. If the UI feels cheap? They book elsewhere. While working as a frontend developer at AZminds, I took on the challenge of building **YachtCloud**—a digital harbor for ultra-premium Mediterranean charters. > Screenshot (yacht page hero explaination): yacht page hero explained — Every pixel on this screen has a single job: eliminate friction, establish elite trust, and funnel high-net-worth traffic into a seamless, high-ticket conversion engine. I didn't just write code. I engineered the frontend to convert. I built high-converting landing pages, structured sleek blog architectures, and developed a dynamic, lightning-fast Yacht Listing interface so clients could browse massive inventories without a single hiccup. > Screenshot (fleet search filters): Fleet listing showing 263 yachts with price, length and cabin filters down the left and a sort control. — 263 yachts. Every filter here is a query against Elasticsearch, not a pass over a fetched list. But building one luxury site wasn't enough. The real challenge? Other premium brokers wanted this exact system, but with their own branding. > So, I implemented a robust **multi-tenant frontend architecture**. > Screenshot (Side-by-side comparison of YachtCloud and Exclusive Gulets ): Side-by-side comparison of YachtCloud and Exclusive Gulets — Building a single booking site is standard. Engineering a multitenant frontend that allows franchise brokers to launch their own bespoke platforms is how you scale a business. I designed the system to spin up fully branded, white-labeled platforms running flawlessly off a single core engine. One unified, highly scalable codebase. Multiple elite broker sites. I deployed platforms like [***Exclusive Gulets***](https://) using this exact multi-tenant setup, allowing them to leverage the powerful backend while maintaining their unique brand identity on the frontend. ![exclusive gulets](https://pub-0d70e426aadb4f069a19cb045dfe20f4.r2.dev/media/exclusive-gulets.avif) > My core focus was the frontend, but building a seamless multitenant system required me to cross over and architect the backend data pipelines. **How I Killed the Traditional Search Filter (Using AI)** If someone is dropping $50,000 on a yacht charter, they aren't going to click through 12 different dropdown menus. They don't have time for clunky UI. They want answers. That is why I replaced the standard search experience with a custom RAG (Retrieval-Augmented Generation) pipeline Agent Chat. > Screenshot (AI powered Search): AI powered search of Yachts on Exculsive Gulets — AI powered search implemented using RAG Pipelines of our existing Yachts Fleet I built an AI-powered planner Agent that understands natural language. Instead of forcing users to check boxes for dates, locations, and guest counts, they just type exactly what they want. - "Plan an anniversary sailing trip in Greece for September." The system takes that simple prompt and goes to work. It acts like a high-end digital broker. Here is what the backend delivers in seconds: - **Total Market Access:** It instantly scans our private fleet and hundreds of ***global MYBA network yachts*** to find the perfect boat. - **Local Insights:** It pulls geographical data to recommend quiet coves and premium anchorages. - **Custom Itineraries:** It builds a personalized route based on the specific group size, season, and interests. This is not a basic, frustrating chatbot. It is a dynamic sales tool. It removes all friction from the browsing process. It keeps wealthy buyers highly engaged with the platform. Most importantly, it turns a boring search bar into a ***24/7 digital concierge*** that actually drives high-ticket inquiries. Try It Here : [Exclusive Gulets Agent](https://www.exclusivegulets.com/gulets) **The Impact** I built more than a static website. I created a scalable frontend ecosystem that sells the luxury dream before the client ever steps on board. ## Screens ### Yacht - Yacht detail page for S Nur Taylan: gallery mosaic, guests, cabins, crew, length and refit year, and an availability card. — The yacht page after the gallery, the specification strip and the enquiry card. - Pricing section listing what the weekly rate includes, beside a card totalling one week plus 20% VAT. — The week priced as a figure, VAT included, rather than a rate with the tax underneath it. - yacht page hero explained - Side-by-side comparison of YachtCloud and Exclusive Gulets — Whitelabeled site Exclusive Gulets using our system ### Search - Fleet listing showing 263 yachts with price, length and cabin filters down the left and a sort control. — 263 yachts. Every filter here is a query against Elasticsearch, not a pass over a fetched list. - AI powered search of Yachts on Exculsive Gulets — AI powered search implemented using RAG Pipelines of our existing Yachts Fleet ### Landing - Yacht Cloud home page over a video of a motor yacht, with a four-field charter search bar. - The same home page at phone width, the four-field search collapsed to a single search pill. — At phone width the four-field search collapses to one control. ### Search pages - Greece charter landing page: best months, main ports, popular routes and yacht types in a fact strip. — One of around a hundred pages that exist for search. This is where most of the QA tickets landed. - Interactive Mediterranean chart with five coastlines picked out in colour beside a country list. ## Case study The yacht page, the fleet search, and the pages Google sees: https://riteshkc.com.np/case-studies/yacht-cloud (markdown: https://riteshkc.com.np/case-studies/yacht-cloud.md) --- # Investment Banker/Analyst Portfolio Next.js and Payload CMS portfolio website for a NEPSE portfolio manager - Year: 2026 - Client: Prajesh Sagar - Role: Designer and developer - Duration: 1 month - Stack: Next.js, React, TypeScript, Tailwind CSS, Payload CMS, Vercel, Cloudflare R2, NeonDB - Live: https://prajeshsagar.com.np ## Summary A portfolio site for a Kathmandu investment banker who manages private portfolios on the Nepal Stock Exchange. Next.js with Payload inside the same app, and a hero that drifts like a chart. ## Problem Prajesh sells portfolio management to people who have to hand him money first. He needed a page that earns that in one scroll, and that he can rewrite himself without a developer. ## Outcome One page ordered around the questions a buyer asks, ending in a form that asks what a first call would. Every section of it is an entry in a Payload admin running inside the same app. Prajesh Sagar runs private portfolios on the Nepal Stock Exchange. His buyer is somebody holding anywhere from a few lakh to a few crore of NEPSE stock, deciding whether to hand it to a stranger. That decision does not get made on a call. It gets made on this page, in about a minute, by a reader who has never met him. So the page had to look like money and not like a pitch. This category defaults to stock photography: handshakes, glass towers, a man in a suit pointing at a rising line. The artwork here comes from the currency instead. Line drawings of Nepali rupee notes and a candlestick chart sit behind the headline in gold on navy, a bull and a bear frame the video, and every figure is set in a serif so it reads like a number in a report. > Screenshot (about): The about section: two paragraphs about the practice, a video of Prajesh presenting to a room, and line drawings of a bull and a bear either side of it. The design on this one is mine as well as the build. One month, one page, plus a blog and a projects section waiting for him to fill. > Screenshot (partners orbit): The trusted by section: six client logos sitting on two faint rings that rotate slowly around the lion monogram at the centre. — One rotation takes 26 seconds. The logos counter-rotate so they stay upright. He edits all of it. A headline, a statistic, a service, a client logo. Enquiries land in the same admin he writes from, each with a box to tick once he has answered. The case study covers the order of the page, how the hero animation is built, and what the form asks before it asks for a call. ## Screens ### Landing page - The home page: the headline Your Portfolio, Professionally Managed over line drawings of Nepali banknotes and a candlestick chart, above a gold button. - The same hero on a phone: headline, one line of subheading, the gold button, and the first two stat cards below it. — The two heaviest drawings are never requested at this width. - The proof band: four dark stat cards two by two on the left, the line Where Capital Becomes Real Growth on the right, and client logos under it. — The cards start above the fold. The first scroll lands on a number. - The about section: two paragraphs about the practice, a video of Prajesh presenting to a room, and line drawings of a bull and a bear either side of it. - The services grid: two wide cards, portfolio management and investment strategy, over four narrow ones for planning, audit and pooled capital. — Two sizes, one item type. A checkbox decides which a service gets. - The core values rail: three columns headed trust, innovation and strategic thinking, each with a large faint icon behind the text. - The testimonials band: two rows of quote cards running in opposite directions, each card carrying a quote, an avatar and a name. - The trusted by section: six client logos sitting on two faint rings that rotate slowly around the lion monogram at the centre. — One rotation takes 26 seconds. The logos counter-rotate so they stay upright. - Three portfolio cards side by side, each with a summary, a performance sparkline, a risk meter, and the strategy and result written under it. - The enquiry form: name, email and phone, then rows of chips asking what brings you here, what you are managing today and how soon you want to move. ### Other pages - The projects page with nothing published: the heading Every mandate, on the record over a bordered panel reading no projects published yet. — The empty state is written copy, not a blank page. ### CMS - The Payload dashboard, grouped into content, system, settings and sections, with one card per collection and one per section of the landing page. - The hero section in Payload: the gold half of the headline and the white half as separate fields, the subheading, the button labels, and one stat. — The help text under each field says where the words come out. - A service entry in Payload: an icon picker set to Briefcase, a title, a description, a featured checkbox reading render as a large card, an accent. - The contact section in Payload: the heading fields, then contact detail rows carrying an icon, a title, a value and the link the value opens. - The enquiries list in Payload, two submissions, with the name, email and phone of both blacked out, and the submitted date and a handled flag beside them. — Every personal cell is painted out in the capture itself, not in CSS. ## Case study One page, one ask, and a hero made out of banknotes: https://riteshkc.com.np/case-studies/prajesh-sagar (markdown: https://riteshkc.com.np/case-studies/prajesh-sagar.md) --- # Case studies Long-form write-ups on selected projects. Each one is attached to a project on https://riteshkc.com.np/work. - [One page, one ask, and a hero made out of banknotes](https://riteshkc.com.np/case-studies/prajesh-sagar) — A month on a one-page site for a NEPSE portfolio manager. The order of the page, how the hero animation is built in CSS, and what the form asks before it asks for a call. Project: https://riteshkc.com.np/work/prajesh-sagar. Markdown: https://riteshkc.com.np/case-studies/prajesh-sagar.md - [Who edits it after I leave, and which map loads in China](https://riteshkc.com.np/case-studies/marti-china) — Two months on a corporate site and product catalogue for the China arm of a Swiss tunnelling supplier. The decisions that mattered were the CMS, the catalogue template and the map provider. Project: https://riteshkc.com.np/work/marti-china. Markdown: https://riteshkc.com.np/case-studies/marti-china.md - [Six investor pages, two templates, and no developer on retainer](https://riteshkc.com.np/case-studies/noveon) — Three months on a website for a Hong Kong investment firm. Why the CMS runs on the client's own server, how six investor pages come from two templates, and what is left before it launches. Project: https://riteshkc.com.np/work/noveon. Markdown: https://riteshkc.com.np/case-studies/noveon.md - [What the student sees, what the office sees, from one set of records](https://riteshkc.com.np/case-studies/gradsy) — A consultancy ran applications in spreadsheets and email. The platform gives a student one honest view of their own file, and gives the office one queue built from the same records. Project: https://riteshkc.com.np/work/gradsy. Markdown: https://riteshkc.com.np/case-studies/gradsy.md - [The yacht page, the fleet search, and the pages Google sees](https://riteshkc.com.np/case-studies/yacht-cloud) — Seven months on Yacht Cloud as one of the frontend developers. Three surfaces needed work: the yacht page, the fleet search over 263 yachts, and the landing pages. Here is what each was missing. Project: https://riteshkc.com.np/work/yacht-cloud. Markdown: https://riteshkc.com.np/case-studies/yacht-cloud.md --- # One page, one ask, and a hero made out of banknotes A month on a one-page site for a NEPSE portfolio manager. The order of the page, how the hero animation is built in CSS, and what the form asks before it asks for a call. - Project: [Investment Banker/Analyst Portfolio](https://riteshkc.com.np/work/prajesh-sagar) - Year: 2026 - Role: Designer and developer - Stack: Next.js, React, TypeScript, Tailwind CSS, Payload CMS, Vercel, Cloudflare R2, NeonDB - Published: 2026-08-25 Prajesh Sagar is an investment banker in Kathmandu who runs private portfolios on the Nepal Stock Exchange. What he sells is a person: a stranger reads a page, then decides whether to hand over money and a phone number. So the site is one page with one ask, and everything on it is arranged around that decision. I designed it and built it, in a month. > Screenshot (hero): The home page: the headline Your Portfolio, Professionally Managed over line drawings of Nepali banknotes and a candlestick chart, above a gold button. ## The page is ordered the way the questions arrive A visitor asks the same things in roughly the same order. What is this. Does he really do it. Who else trusted him. What would he do for me. What has he done before. What happens if I get in touch. The sections run in that order and nothing sits between them. Headline and one button. Four numbers. The logos of firms he has worked with. The services. The values and the quotes. Three past portfolios with the strategy and the result written under each. Then the form. The hero is not a full screen tall. It is sized to the viewport minus a fixed strip, so the stat cards start inside the fold and are cut off by it. That is the whole trick: the first scroll lands on a number instead of on more headline. On a phone, where most of these visitors stop within one screen, that strip does more work than any rewrite of the sentence above it. > Screenshot (stats): The proof band: four dark stat cards two by two on the left, the line Where Capital Becomes Real Growth on the right, and client logos under it. — The cards start above the fold. The first scroll lands on a number. There is one ask and it never changes. The button in the nav, the button in the hero, the button under the video and the form all point at the same place. A second offer would give a reader a way to feel they have acted without acting. ## The hero moves like the thing it is selling Behind the headline are six line drawings: Nepali banknotes, a rising wire, a candlestick chart. Each one is its own layer, and each layer carries its own distance, rotation, duration and start time as four custom properties. A single keyframe reads those four and moves the layer, so the drawing at the back travels six pixels across and eight up over seventeen seconds, and the one nearest the reader travels sixteen and twenty-two over nine. That ladder is the parallax. Nothing measures the pointer and nothing measures the scroll; depth here is the fact that near things move further and faster than far ones. Two more details keep it from reading as a loop. The animation runs `alternate`, so a layer eases back along its own path instead of snapping to the start. And every layer has a negative delay, between one and eight seconds, so each one is already part way through its pass when the page loads and the six are never in step. It is CSS. There is no canvas, no animation library and no JavaScript in it, and every property being animated is one the compositor can handle on its own. The cost is weight, not frames. The drawings are vector, and two of them are large. Those two are behind a `` element with a desktop media query, so a phone requests four files instead of six and gets a quieter version of the same idea. > Screenshot (hero mobile): The same hero on a phone: headline, one line of subheading, the gold button, and the first two stat cards below it. — The two heaviest drawings are never requested at this width. Under `prefers-reduced-motion` every animation on the page stops, including this one and the two marquees, and the drawings stay exactly where they are. The hero still looks like the hero. It just holds still. The button does the other half of the work. It is the only saturated object on a dark page, and a single arc of light travels its edge once every three seconds. One arc, one colour: a button that pulses or changes colour reads as an advertisement, and this one has to read as the thing you press. ## Six services, two sizes, and a checkbox that decides which The services grid is two wide cards over four narrow ones. That is not two components. It is one item in the CMS with a checkbox on it, described in the admin as "render as a large card at the top of the grid". > Screenshot (services): The services grid: two wide cards, portfolio management and investment strategy, over four narrow ones for planning, audit and pooled capital. — Two sizes, one item type. A checkbox decides which a service gets. The icon is a picker, the accent is a picker, and both write into a fixed set that the page knows how to render. Prajesh can add a seventh service, tick the box, choose an icon and get a grid that still looks like a grid. He cannot pick a colour that does not exist in the palette or an icon that does not ship, which is the point of a picker over a text field. > Screenshot (admin services): A service entry in Payload: an icon picker set to Briefcase, a title, a description, a featured checkbox reading render as a large card, an accent. ## The form asks four questions before it asks for a call Under the phone number are three rows of chips: what brings you here, roughly what are you managing today, and how soon do you want to move. Then one open box. That is the first fifteen minutes of a discovery call, asked in advance. A portfolio review for someone just starting out and a review for someone holding five crore are different conversations, and knowing which one it is before replying is the difference between a useful first message and a request for a meeting to find out. The line under the form says the details stay with him, which is the objection anyone typing a portfolio size into a stranger's website is having. They are chips rather than dropdowns because the options are the argument. A reader who sees "Rs 5 crore+" in the list learns that clients at that size exist here. > Screenshot (contact form): The enquiry form: name, email and phone, then rows of chips asking what brings you here, what you are managing today and how soon you want to move. Every submission lands in the CMS with a handled flag on it, so the admin is also the inbox. > Screenshot (admin enquiries): The enquiries list in Payload, two submissions, with the name, email and phone of both blacked out, and the submitted date and a handled flag beside them. — Every personal cell is painted out in the capture itself, not in CSS. ## Payload instead of Strapi The last two sites I built ran on self-hosted Strapi: [Marti China](https://riteshkc.com.np/work/marti-china), because hosted APIs are unreliable inside mainland China, and [Noveon](https://riteshkc.com.np/work/noveon), because a firm with a compliance officer wanted the content on a server it owned. Both reasons are about where the CMS sits. Neither applies here. This is one person with one page, no ops budget and nobody to keep a second server patched. Payload runs inside the Next.js app itself: one repository, one deploy, one admin at `/admin`, and media on a CDN domain of the site's own. > Screenshot (admin home): The Payload dashboard, grouped into content, system, settings and sections, with one card per collection and one per section of the landing page. Every section of the landing page is a separate entry in there, and so are the navigation, the footer and the SEO fields. The hero is not a rich text field; it is the gold half of the headline, the white half, the subheading, both button labels and four counters, each with a value, a suffix and a label. The help text under each field says where the words come out, in the words a person who is not a developer would use. > Screenshot (admin hero): The hero section in Payload: the gold half of the headline and the white half as separate fields, the subheading, the button labels, and one stat. — The help text under each field says where the words come out. The other consequence of Payload living in the app is that the page is rendered on the server. The HTML that arrives at a browser already has the headline, the services and the quotes in it, which is not true of the two Strapi sites, where the browser fetches the content after the shell loads. ## What is not on the page yet The projects section is published and empty. It has a written empty state and not a blank panel, because a page that says nothing is added yet is a page that says the practice is real and the list is coming. > Screenshot (projects empty): The projects page with nothing published: the heading Every mandate, on the record over a bordered panel reading no projects published yet. — The empty state is written copy, not a blank page. The blog has one post in it and it is a test. Both of those are Prajesh's to fill, and the fields are waiting. What he has is a page he can rewrite on a Sunday, and an inbox that tells him what somebody is holding before he picks up the phone. --- # Who edits it after I leave, and which map loads in China Two months on a corporate site and product catalogue for the China arm of a Swiss tunnelling supplier. The decisions that mattered were the CMS, the catalogue template and the map provider. - Project: [Marti China](https://riteshkc.com.np/work/marti-china) - Year: 2026 - Role: Full-stack developer - Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi, Baidu Maps - Published: 2026-07-01 Marti China is the Chinese arm of Marti Technik AG, part of a Swiss construction group that has been family-owned since 1922. It makes conveyor systems and steel components for tunnelling contractors — idlers, frames, conveyor modules, fabricated steel — out of a works in Zhongshan, Guangdong. The site had to do four jobs: say what the company sells, say what it does, show where its parts are already running, and take a technical enquiry. The design arrived finished. Two months, one developer, frontend and CMS. > Screenshot (home hero): Marti China home page over a photograph of an engineering meeting, with the headline Swiss Engineering Excellence and an Explore Services button. ## Why Strapi, and what it had to beat The first requirement was not a feature. It was that nobody should have to call me to add a product, correct a specification, or fix a typo. That decides the stack before anything else does. Four things had to be true at the same time. Staff in Zhongshan edit everything themselves. Every field exists twice, once in English and once in Chinese. The catalogue is relational — a product belongs to a category, a project carries a country and a coordinate — so it needs content types rather than a page builder. And all of it has to be reachable from inside mainland China. That last condition removes the hosted options. Contentful and Sanity would model this content well. Their APIs and their image CDNs sit outside the country, and I cannot promise a buyer in Shenzhen that they answer. If they do not, the page renders with holes in it. Markdown files in the repository fail the first condition instead: a corrected typo becomes a pull request and a deploy, which is the thing the client was paying to avoid. WordPress could carry the content, but the frontend is Next.js, so WordPress would be a second rendering stack kept alive for its admin screen, with the second language bolted on by plugin. Strapi is self-hosted, so it runs on the client's own machine — the API sits behind nginx on its own subdomain, next to the site it feeds. Its localisation is a core, per-field feature rather than a plugin. And its admin is a plain form, which is what you hand to someone who is never going to read the documentation. ## Forty-nine product pages that nobody wrote Five categories: idler rollers, idler frames and stations, conveyor modules, general steel fabrication, custom engineered equipment. Forty-nine products under them. Nobody hand-builds forty-nine pages. There is one content type and one template: an image set, a category label, the description, and related products pulled from the same category. The catalogue page stacks all five categories with a sticky rail down the left and an anchor per category, so a link can point at conveyor modules and land on conveyor modules. Each category header carries a datasheet download, because the person reading wants a PDF to forward to procurement. > Screenshot (product catalogue): Products page: a sticky rail of five categories on the left, the idler rollers grid on the right, and a download all datasheet button. — Five categories, forty-nine products, one template. The rail is an anchor per category. When the client adds the fiftieth product, the page is already written. > Screenshot (product page): Carrying idlers product page: a numbered image column on the left, the category label, name and description on the right. — The forty-ninth product page and the first one are the same file. ## The map has to be Baidu The projects page is the sales argument. Marti components have run in seventeen tunnelling projects across twelve countries — Gotthard, Brenner, the Follo Line, Grand Paris Express, Changi Terminal 5, Coca Codo Sinclair. Read as a list that is a wall of names. On a map it is a claim about reach. Google Maps is not a dependable choice for a page whose primary reader is inside mainland China. So the map is Baidu's BMapGL. A panel of references floats over it; picking one pans the map and highlights the row, so the list and the map are a single selection rather than two widgets that happen to share a page. > Screenshot (project map): Projects page: a Baidu map of Europe with Chinese labels, a floating panel listing tunnelling references by country and year. — Baidu, not Google. The primary reader is inside mainland China. Baidu brings two costs worth naming. Its tiles are labelled in Chinese, which is right for the primary reader and a curiosity for the European one. And it uses its own coordinate system, so a latitude and longitude copied out of Google sits a few hundred metres from the pin you meant. Coordinates are content here — they go into Strapi with the project — so they have to be correct at the point of entry, not corrected in the browser. > Screenshot (project selected): The same map after picking the Follo Line reference: the map has panned to Norway and the list row is highlighted. — Picking a reference pans the map. The list and the map are one selection. The contact page runs the same map with one pin on the Zhongshan works. ## Two languages, and the toggle says EN and SC Chinese is not a translation layer over an English site. In Strapi it is a second value on every localised field, so the whole site exists twice and the routes are live under both languages today. The Chinese entries are the client's to write. That is the shape of the work: I built the surface for both languages, they own the words that go in it. > Screenshot (language toggle): The header's language control open, showing EN selected and SC beside it. — English and simplified Chinese, per field in the CMS rather than per site. ## One stylesheet, no component library The design was finished before I started, so the CSS job was to name it rather than invent it: a type scale, a fixed palette — charcoal, cloud white, a hairline grey, Marti blue — and Tailwind utilities generated from those names. No component kit. A library would have meant overriding somebody else's defaults on every screen to arrive at a design that had already been decided. > Screenshot (quality inspection): Quality inspection section: four numbered cards over a photograph of two inspectors, covering receipt checks, in-process controls, verification and the report. ## The pages that are not in the brief A brief lists the home page and the product pages. The rest is what makes it a site somebody can work in. A 404 that offers the way back. Privacy and terms. A footer that repeats the whole catalogue, because that is where a reader who has scrolled to the bottom is looking. One work with us band at the foot of every route, so the enquiry is never more than a scroll away. > Screenshot (not found): The site's 404 page: the numerals 404, the line Page Not Found, and a Back to Home Page button. The floating contact widget is a WeChat QR code rather than a chat bubble from a Western SaaS, for the same reason the map is Baidu. ## Where it ended up Nine page templates, a 404, privacy and terms. Forty-nine product pages, four service pages, seventeen project references, two languages, and a contact form that asks for belt width and incline angle instead of a message box. > Screenshot (inquiry form): Contact page: email and phone on the left, a technical inquiry form on the right asking for name, company, email, project location, type and specification. — Six fields. The last one asks for capacity, belt width, incline and operating conditions. Adding a product, a project reference or a service is a form submission now, not a deploy. Everything above is what that one requirement cost. --- # Six investor pages, two templates, and no developer on retainer Three months on a website for a Hong Kong investment firm. Why the CMS runs on the client's own server, how six investor pages come from two templates, and what is left before it launches. - Project: [Noveon](https://riteshkc.com.np/work/noveon) - Year: 2026 - Role: Full-stack developer - Stack: Next.js, React, TypeScript, Tailwind CSS, Strapi - Published: 2026-04-01 Noveon is an investment firm in Hong Kong. It buys stakes in private equity funds from investors who need their money back before the fund matures, and it is building software to price those deals. I built the website: Next.js on the front, Strapi behind it, three months, one developer. The design arrived finished, from another studio. > Screenshot (home hero): Noveon home page: the headline Asia's Global Private Equity Secondaries Platform over footage of the Hong Kong skyline and the IFC tower. One requirement sat under all the others. Noveon has no developer. After I left, a marketing lead and a compliance officer had to be able to change any word on the site, including the words on pages a regulator may one day read. ## Why the CMS runs on the client's own server Contentful and Sanity would both model this content well. I ruled them out for a reason that has nothing to do with their features. The content would live in a vendor's account, and everyone who needs to fix a sentence needs a paid seat in it. The person correcting a disclosure here is a compliance officer who logs in twice a month. Charge that person a seat fee and the edits go by email to somebody else instead. Markdown files in the repository fail sooner. A corrected job title becomes a pull request, and the firm is back to needing me on the day it wants the fix. WordPress would hold the content well enough. The front end is Next.js, so I would be running a second system for the sake of its admin screen. Strapi runs on a server the firm controls, on a subdomain beside the site it feeds. Editors are unlimited, the database is theirs, and the admin screen is a plain form. This is the second site I have built that way. The first was a product catalogue for [Marti China](https://riteshkc.com.np/work/marti-china), where the deciding factor was that hosted APIs are unreliable inside mainland China. Different reason, same answer. ## Six investor pages built from two templates Noveon runs three investment strategies and sells to three kinds of investor: family offices, institutions and private wealth. That is six pages, and they are the pages a serious visitor opens first. Building them as six separate page entries would have been quicker for about a week. Then somebody edits one. A heading moves on one page and not the other five, a figure gets a different label, and nobody notices until a prospect has two of them open side by side. So there are two content types instead, one for a strategy and one for an investor. A strategy page leads with numbers. An investor page leads with what the firm will do for that reader. Each one is a single template with a single set of fields, so the sixth page renders from the same code as the first, and adding a fourth investor type is a new entry rather than a new page. > Screenshot (strategy page): The Asia General Partner Solutions page: a blue band headed Asia's Private Equity Moment with three figures for transactions, capital provided and GPs covered. — One of three strategy pages. All three are the same template with different entries. The dropdown in the header shows the same split: strategies in one column, investor types in the other. > Screenshot (nav menu): The Investors menu open over the home page: an investment strategy column of three links, an investor column of three, and a photograph beside them. — The menu is the content model drawn: three strategies, three investor types. ## Articles, podcasts and case studies are three types, not one Noveon publishes three kinds of thing on three different schedules. They could have been one post type with a category field on it. They are not, because the fields are genuinely different: an article carries an author, a date and its own numbered sections, and the other two carry none of that. > Screenshot (home insights): Home page insights band titled Insights Uncovered with three cards: two podcasts and a case study, each with its own image and label. — One card serves all three kinds of insight. The label is the content type. What they do share is the card, the listing layout and the sort control, so the three feeds read as one publication. An article's contents list is generated from the sections in the entry, so a writer who adds a section gets a new row in the sidebar without asking anyone for it. > Screenshot (article): An article page: a table of contents pinned at the left with the current section marked, the article's own sections and images at the right. — The contents list is built from the sections in the entry, so a new section adds a row. ## One page sells software instead of the firm Noveon Labs is the firm's own product. Its page needed a vocabulary the rest of the site does not use: numbered features with a video for each, four short user stories, an FAQ, and a request for an invitation where every other page has a contact form. > Screenshot (labs hero): Noveon Labs page opening: the line One Platform. Complete Secondary Intelligence., a paragraph about the NAIP product and a request invite button. — The one page selling software rather than the firm. It also has to look like the company that publishes the disclosures page two clicks away, so it uses the same header, the same type scale and the same colours. A separate product microsite would have been easier to build and would have read as a different business. > Screenshot (labs faq): Noveon Labs FAQ: four collapsed questions about the product beside a card offering an email reply to anything not covered. ## The design is motion, and the copy comes out of a form The studio delivered the design as movement. The site opens with a full-screen brand sequence, sections assemble as you reach them, and one row of cards is dragged sideways instead of clicked through. > Screenshot (splash): The site's opening frame: the Noveon wordmark and the line Mastery of Private Markets over slowed footage of a koi carp. — The first load plays this before the home page. It is the brand's own opening, not a loading screen. Animation like that is easy to build against copy that never changes and awkward once an editor owns the copy. A card holding two lines today holds six after somebody rewrites it. A grid of four becomes a grid of three the moment an entry is deleted, and an entrance timed to four items then plays a beat of empty space. So each animated section takes whatever the CMS hands it and stages that, instead of animating a layout I measured once. The values row is the clearest case. Four cards today, dragged sideways, progress bar underneath. Delete one in Strapi and it is three cards, and the row still behaves. > Screenshot (core values): The core values rail: four photographic cards captioned client-aligned partnership, intelligence-driven investing, forward-looking innovation, stewardship. — Drag to advance, progress bar under it. Nothing in the component assumes four. ## What is left before it launches The site is not live. noveoncapital.com is still closed, and the address I worked against is the review build. The client's remaining work is words. Two of the strategy pages still show placeholder figures, the insight entries are test rows, and one policy page has not been written. Every field those need is built and waiting. Mine is narrower. The pages fetch their content from Strapi in the browser, which is fine behind a review URL and wrong for a public marketing site: a crawler that does not run JavaScript sees an empty shell, and every route currently shares the title "Noveon". Moving those fetches to the server is route-by-route work in the App Router rather than a rebuild, and it is the first thing I do when a launch date is set. > Screenshot (contact): Contact page: the email address and Hong Kong office address on the left, a five-field message form on a dark green panel at the right. What the firm has now is a site two of its own staff can run: two logins, and no deploy between either of them and a correction. --- # What the student sees, what the office sees, from one set of records A consultancy ran applications in spreadsheets and email. The platform gives a student one honest view of their own file, and gives the office one queue built from the same records. - Project: [Gradsy](https://riteshkc.com.np/work/gradsy) - Year: 2026 - Role: Full-stack engineer - Stack: Next.js, NestJS, TypeScript, Prisma, PostgreSQL, Redis, BullMQ, AWS, AWS Simple Email Service, S3, Lightsail - Published: 2026-02-01 ## Context Gradsy is an education consultancy in Nepal that places students in universities abroad. The business is a pipeline. A student arrives, builds a profile, uploads documents, picks programs, and applies. A counselor then moves that application through a series of states until there's an offer or there isn't. Before the platform, that pipeline lived in spreadsheets, email threads and WhatsApp. Two questions were expensive, and both were expensive for the same reason. A student's question is "where is my application?" Answering it meant somebody in the office finding the row, remembering the last thing that happened to it, and writing back. That is the highest-volume message a consultancy receives, and every answer costs staff time. The office's question is "what needs doing, and who is doing it?" A spreadsheet doesn't say which applications are stuck waiting on a transcript, which counselor already has nine students, or whether the person who emailed twice today has been assigned to anyone at all. Both questions are about the same records. They need opposite surfaces. A student is doing this once, under stress, and needs a guided path with nothing hidden. The office is doing it hundreds of times and needs a queue, a load balance, and a history of who changed what. So the platform is one domain model with two products on top of it, plus one engineer to build both. ## Architecture Two repositories. A Next.js 15 App Router frontend carrying the public marketing site, a university explorer, and all three dashboards. A NestJS API owning the domain, the database, the queues, and every rule worth enforcing. They talk over REST with session cookies. No shared package, no code generation, no monorepo. ```text Next.js 15 (App Router) NestJS 10 ───────────────────────── ───────────────────────── 46 server / 46 client pages 24 modules, 42 controllers, 254 routes TanStack Query (client data) Throttler → BetterAuth → Roles (default-deny) fetch + RSC (public/SSG) Prisma, 19-file schema, 54 models socket.io-client BullMQ + Redis · socket.io │ │ └──────── REST + cookies ─────────┘ │ PostgreSQL 15 · S3 · SES ``` Two repos saved me setup time and cost me elsewhere. A solo project's bottleneck is rarely type safety. It's the number of things you have to set up before you can ship a feature. Two plain repos and hand-written types meant I could add an endpoint and consume it in about the time it takes to describe it. The cost arrived later, as constants that have to be mirrored by hand and a cache contract nothing validates. Three properties of that diagram matter to everything below. **Every route is closed until it says otherwise.** Every request passes three checks before it reaches my code: rate limit, then session, then role. A route carrying neither `@Public()` nor `@Roles()` is refused. Adding an endpoint and forgetting to protect it produces a `403`, not a leak. I flipped that default partway through. It broke routes for a day. That was the point: I wanted the gaps to fail loudly instead of leaking quietly. **Only two routes are prerendered.** The public university and article pages are built ahead of time and rebuilt on a timer, which Next calls ISR. Every dashboard is client-rendered behind auth. A page that's different for every user and stale by definition gains nothing from being prerendered. **Documents never touch the API.** Uploads go straight from the browser to S3 with a presigned URL and are confirmed afterwards, so file bytes bypass Node entirely. The confirm step checks the returned key starts with `applications/{applicationId}/documents/`, because a client that could confirm any key could attach somebody else's file. ## The student side ### One document, uploaded once A student applies to several universities. In the manual process that meant sending the same passport scan several times, and somebody in the office verifying the same file each time. Here a document is stored once in the student's library, and an application either owns its file or points at the library copy. Attaching inherits the review status, so a passport verified last month is verified on the application it's attached to today. Attaching twice returns the existing attachment rather than a duplicate. The exceptions are encoded as constants rather than left to judgment. Recommendation letters and transcripts can exist many times over; every other type is single-instance, and re-uploading replaces. The statement of purpose is written per application and is never saved to the library. Removing a document deletes the S3 object only when that row owns it, never when it's a reference to a shared one. > Screenshot (student-documents-library): A document library, six categories with four required, every file carrying a verified badge and a banner confirming the required set is complete. — The library. Six categories, four of them required, each file carrying its verification state. ### The submit button explains itself An application can only leave draft when the student's documents are ready. That sentence sounds like a boolean. It isn't one. The naive version — every required type is uploaded — is wrong three separate ways, and each way is either a student who can't submit or an incomplete application that reaches a partner university. A document has to be attached to *this* application, not merely sitting in the student's library. Otherwise a verified passport unlocks an application it was never part of. English proficiency can't require attachment at all, because a test score satisfies it and a test score has nothing to attach to an application; requiring it would permanently lock out every student who proved English that way. The statement of purpose is written per application and reviewed *after* submission, so gating submission on its verification deadlocks the thing it's supposed to guard. So the check is three rules with different shapes. The consequence worth noting is the error. It distinguishes not attached from not verified yet from not uploaded from rejected, because those are four different things for the student to do, and collapsing them into "documents incomplete" turns each one into a support ticket. The profile gate works the same way: it reports the exact percentage rather than the word incomplete. A student who is still blocked gets a button that requests a consultation, which creates a real task for a real counselor. A gate with no way past it is a dead end, and a dead end becomes an email. Written up in full: [Ready to submit is not a boolean](https://riteshkc.com.np/blog/ready-to-submit-is-not-a-boolean). > Screenshot (student-application-new): A three-step new-application wizard on step one, choosing destination country, degree level and university. — Applying is three steps, saved at each boundary so a drop-off survives, and resumed at the first gap. ### The history is the answer to "where is my application" Every status change writes a history row and an audit row inside the same transaction as the change itself. The student reads the history the counselor writes, rejections included, with whatever reason somebody typed. The transitions themselves live in a database table, not in `if` statements. Each status row names which statuses may follow it, and names a stricter set for the case where the actor is a student. A target status can be marked as requiring feedback, and then the transition is refused without a note. Changing the workflow is a catalog edit and a re-seed, so the ops team can reshape the pipeline without me. That is the feature that removes the support question. The student doesn't ask, because the answer is already on the page, in the same words the office used. > Screenshot (student-application-progress): One application's status history, seven entries deep, each state change stamped with a date and the counselor's written reason. — Seven transitions on one application, each stamped with a date and the counselor's reason. More on the table-driven part: [A status machine belongs in a table](https://riteshkc.com.np/blog/a-status-machine-belongs-in-a-table). ### The community, and moderation that had to be split by certainty The student community needed automated moderation. My first version checked everything on write and blocked anything suspicious. That produced a stream of outcomes I couldn't defend: enthusiastic posts blocked for being in caps, useful posts blocked for having four links. The mistake was splitting the system by latency and then letting that split decide policy. Fast checks became blocking checks because they were fast. The rewrite asks a different question: what has to be true before this content exists for even a second? That list is short — slurs, known-malicious domains. Everything else is evidence, and evidence deserves a judgment a second later. So the synchronous gate now blocks almost nothing and attaches its findings to the request. A queued worker then adds the two signals the request path can't afford, posting rate and the author's moderation history, before deciding approve, review, or remove. Anything flagged for review is filed into the existing human reports queue under a system bot account, instead of getting its own admin screen. That is the first place the two sides meet: automated flags and student reports arrive in one place, so every future improvement to that queue applies to both. Full write-up: [The fast moderation layer does less on purpose](https://riteshkc.com.np/blog/the-fast-moderation-layer-does-less). > Screenshot (admin-article-review): An article review queue filtered to under review, with author-type filters and approve, reject and preview actions on the card. — The queue those flags land in. It already existed for human reports, which is why automated flags were filed into it rather than beside it. ## The office side ### The queue is the pipeline's states, in order A counselor's board is the state machine seen from the other end. Where an application has got to is the column it's sitting in, and the count on each column is the answer to "what's blocked." > Screenshot (counselor-applications): A counselor's applications board with four columns - draft, pending review, sent to university, decided - each showing a count. — Draft, pending review, sent to university, decided. The counselor's queue is the same states the student reads. ### Work is routed, not remembered When a student's required document set becomes complete, the system creates the document check itself. It first checks no actionable task already exists, then assigns to the eligible counselor with the least work in flight. When nobody is eligible, it messages every admin to say a student is ready and no counselor was available. The whole routine is wrapped so it can never throw: a failure to create the task must not fail the student's upload. There are two least-loaded functions, kept separate on purpose. Delegated tasks measure load in open tasks and require delegation to be switched on. Appointment routing measures load in live appointments and doesn't require it, because a booked consultation is core counseling work rather than delegated admin work. Merging them would let booking volume distort the routing of document checks. Ties break on creation time so the result is deterministic. > Screenshot (counselor-dashboard): A delegated task queue with one open document check, filters by state, and claim, start, complete and escalate actions on the detail panel. — Tasks are claimed off a shared queue rather than handed out. Escalation hands one back. ### Delegation is a capacity, not a job title A counselor's ability to do a delegated task is checked three times: their account has to be approved, delegation has to be on with the task's slug in their permission set, and the assignment service re-checks the same conditions when it picks somebody. Verification calls additionally require a senior delegation level. A per-counselor cap limits how many tasks they can hold at once, so routing can't quietly overload the fastest person. > Screenshot (admin-counselor-delegation): An admin delegation panel setting a counselor's delegation level, maximum concurrent tasks, and which categories of work they may take. — A level, a task ceiling, and a checklist of permitted work. The cap is what keeps auto-routing honest. ### A review is a status change plus a note the student reads There is no separate reviewer vocabulary. An admin reviewing an application moves it to a new status and types a reason, and that reason is the text on the student's timeline. The same document rules that block the student are shown to the reviewer as the explanation of why the application isn't submittable, so both people are looking at one answer. > Screenshot (admin-application-review): An admin reviewing one application: verified documents, a warning that a missing statement of purpose blocks submission, and a status-change form. — Verified documents, the rule that is currently blocking submission, and the status form that writes to the student's timeline. ### Who changed this record Most admin questions turn out to be questions about who last touched something. Every write records an actor, and the audit write runs inside the caller's transaction, so a change and its log entry commit together or not at all. The log copies the actor's name, email and role onto the row rather than joining to the user. The actor reference is set to null when an account is deleted, and accounts do get hard-deleted, so a live join would silently lose attribution on exactly the records somebody is investigating. Machine-originated actions are marked as system, so auto-assignment and public bookings never look like a person did them. The dashboard's filter list is served from the backend's catalog of action names, so the two can't drift. > Screenshot (admin-dashboard): An admin dashboard with four counters, quick actions, and an activity feed naming the actor and their role for each recent change. — The activity feed names the actor and their role. Attribution is copied onto the row, not joined at read time. ### The catalog is edited once and read publicly Universities and programs are edited by an admin and rendered on the public site from the same record. Bulk cataloguing goes through a CSV import that reports per-row errors with the real line number from the file. The preview shares the same parser as the import and writes nothing: an operator sees what would change before committing thousands of rows. > Screenshot (admin-university-edit): A university editor with tabs for overview, academics, costs, media and programs, showing basic information, location and record metadata. — One university record. The slug is visible because it is a URL somebody may already have shared. That shared record is where the two repositories collide. An admin edits a university, saves, opens the public page, and sees the old data. Nothing is broken. The page was built ahead of time and only rebuilds every 24 hours, and the process that changed the data isn't the process holding the page. Lowering the window picks a smaller wrong number. The writer has to announce the write instead: after every mutation, the backend POSTs a set of cache tags to a secret-protected route on the frontend. The 24-hour timer stops being the freshness mechanism and becomes the backstop for a webhook that never arrives. It is the only call in the system running backwards, from API to frontend. Three call shapes fell out of that, and I missed the third at first. A single edit busts one slug. A CSV import busts once for the whole batch instead of once per row. A **rename** has to bust both the new slug and the old one: the record moved, and its previous URL still holds a prerendered page that nothing will ever invalidate. Full write-up: [Cache invalidation when the cache is in another repo](https://riteshkc.com.np/blog/cache-invalidation-in-another-repo). ## Where the two sides collide Everything above assumes one person at a time. The places that took the most work are the ones where two people act on the same record at once, and the fix in each case was to make the rule a property of the database rather than a check in a service. Consultations are bookable by anyone from the public site. Three seats per time slot. The obvious implementation counts existing bookings and creates a row if there's room. It's wrong, and wrong in a way that never appears in development. Two people click book at the same moment. Postgres lets both requests read the booking count before either one saves, so both see one seat free and both insert. Four bookings in a three-seat slot, no error anywhere. That default behaviour has a name: Read Committed. Instead of counting bookings, each booking now claims a numbered seat. A unique constraint on `(slotDate, slotTime, seatIndex)` makes "one booking per seat" a property of the schema rather than a check in a service. Two racers can both decide seat 1 is free; only one commits. The other gets a constraint violation, and the booking loop treats that as "try the next seat." Cancellation sets `seatIndex` to `NULL`. Postgres treats NULLs as distinct in a unique index, so cancelled rows sit alongside live ones without blocking anything, and the seat is reusable immediately. No free-list table, no cleanup job, no lock. I've written the full version of this up separately, including the transaction bug I hit on the way: [Overbooking is a database problem](https://riteshkc.com.np/blog/overbooking-is-a-database-problem). The booking loop, with its original comment. Twelve lines that replaced a lock: ```ts /** * Claims the first free seat in the slot. A plain count-then-create * overbooks under Read Committed (two racers both see capacity-1) so the * unique index on (slotDate, slotTime, seatIndex) is the actual guard and * P2002 just means "seat taken, try the next one". * * Each attempt is its own transaction: Postgres aborts a transaction on the * first failed statement, so retrying a seat inside one would only produce * "current transaction is aborted" on every subsequent create. */ private async createWithFreeSeat( data: Omit, actorId: string | null, ) { for (let seat = 0; seat < APPOINTMENT_SLOT_CAPACITY; seat++) { try { return await this.prisma.$transaction(async (tx) => { const created = await tx.appointment.create({ data: { ...data, seatIndex: seat }, }); await this.audit.record(tx, { actorId, action: 'appointment.create', ... }); return created; }); } catch (err) { const isSeatTaken = err instanceof Prisma.PrismaClientKnownRequestError && err.code === 'P2002'; if (!isSeatTaken) throw err; } } throw new ConflictException( 'That time slot is fully booked. Please pick another time.', ); } ``` The audit write sits inside the transaction deliberately. If the booking commits, its audit row commits with it, or neither does. An audit log that can silently skip entries during a retry loop is worse than no audit log, because you would trust it. The same shape appears elsewhere. A community vote toggles inside a transaction and then recomputes the score from the sum of votes, so the denormalised column can't drift from the votes it summarises. A report of the same content by the same person is caught as a constraint violation and treated as a no-op rather than an error. ## Impact No production numbers here. I don't have figures I could defend on a call, and a plausible-looking one is worse than none. What follows is what changed structurally, not what it measured. A student's status question is answered on the page, in the same words the counselor typed, because the transition, its history row and its audit row are one write. A document is verified once and reused, instead of being re-sent and re-checked per university. An application that isn't ready names the four things that could be wrong with it, rather than saying incomplete. Work routes itself to the counselor with the least in flight, and when nobody is eligible an admin is told rather than the task disappearing. Booking correctness moved from application code into the schema. The rule survives someone rewriting the service, because it isn't in the service. Admin edits appear on the public site immediately instead of on a timer, and the timer now covers only the case where the webhook fails. Broadcast email stays under the provider's send quota by construction: one worker and a fixed pause, rather than a limiter that has to be tuned. Moderation decisions are all replayable. Every one is logged, and a user's trust score is recomputed from that log rather than stored as a number someone has to trust. Queue jobs are idempotent at the points where a retry would otherwise double-count. Campaign counters are deduplicated through a Redis set, and recipient status updates use a predicate, so a retry that finds the row already sent changes nothing. ## Lessons **Constraints beat checks.** Nearly every correctness win in this project came from moving a rule out of a service method and into the database or the queue. A check runs when someone remembers to call it. A constraint holds for every write, including the ones added later by someone who never read the service. **The error message is the feature.** The document gate, the profile percentage and the full booked slot all took longer to word than to implement, and each one is a message the office would otherwise have to send by hand. On a two-sided product, a sentence written for the student is work removed from the admin. **Two repos without codegen cost more than it looks.** Two constant files are mirrored by hand between the repos, with comments in both explaining that a drift shows the student an option the API will reject. The cache tag names are a contract with no schema, no types, and no failing test: rename one side and the webhook still returns `200 OK` while the page silently stops updating. A small shared package would have cost an afternoon. I told myself it was ceremony. **Comments rot faster than code, especially alone.** Two files in the frontend describe the revalidation webhook as "a webhook that never landed." It landed months ago and has fifteen call sites. I wrote that comment during the window where only one half existed, copied it into a second file, and never went back. No reviewer, no handoff, nobody to catch it. **The honest gaps.** The backend's CI workflow is committed with every line commented out; deploys are still `git pull` and a process restart. The frontend has no test suite at all. The backend's thirty specs cover the moderation services and the application state machine and stop there, and the README says so out loud. The sitemap is a hardcoded list that omits both families of statically-generated dynamic routes. There's no observability beyond Nest's console logger, and because the queues run `removeOnComplete`, a job that failed yesterday left nothing to inspect. None of these are things I found while writing this page. Each one lost a prioritisation call against a feature. ## Next **Codegen or a shared package** for the mirrored constants and cache tags, so a drift is a compile error instead of a silent no-op. **Actual CI.** Uncommenting the workflow is the smallest possible first step, and it isn't done. **A dynamic sitemap.** The two prerendered route families are the ones worth crawling, and the sitemap is the one place they don't appear. Error boundaries were the other half of this item and have since landed: `error.tsx`, `global-error.tsx` and `not-found.tsx` are all in the tree now. **Observability.** A console logger and no job history is fine until the first incident. Then it is the only thing that matters. **Web push.** There's a complete cross-repo implementation spec sitting in the frontend repo, including the service worker source and the payload contract, marked "planned, not yet implemented." It would replace the current compromise, where a background tab gets a native browser notification only while the tab is still open. **Merge the two email queues.** They independently rate-limit themselves to the same figure, so running both at once doubles the real send rate. The code says so in a comment. One queue with a job type is less machinery than two plus a coordinator. --- # The yacht page, the fleet search, and the pages Google sees Seven months on Yacht Cloud as one of the frontend developers. Three surfaces needed work: the yacht page, the fleet search over 263 yachts, and the landing pages. Here is what each was missing. - Project: [Yacht Cloud](https://riteshkc.com.np/work/yacht-cloud) - Year: 2024 - Role: Frontend Developer - Stack: Next.js, React, TypeScript, Tailwind CSS, Elasticsearch, RAG, LangChain - Published: 2024-06-01 I joined an existing Next.js site as one of the frontend developers. The site was live and selling. That shapes everything below: nothing here is a rewrite, and every change had to land on a page that was already taking enquiries. ## The yacht page The yacht page is where the decision happens. Someone has already picked a country and a week. They are on one yacht now, and they are deciding whether to send the enquiry. It was not set up to carry that. There was no gallery, so an eighteen-guest gulet was represented by one photograph. There was no itinerary, so nobody could see where the week actually went. There was no clear place to enquire. And the page moved under you while it loaded. I added the three things the decision needs, in the order it needs them. A gallery mosaic first, because a charter is sold on the interior and one exterior shot does not show a cabin. Then the specification strip — guests, cabins, crew, length, built and refit — as one row, because those six numbers are the filter someone has already been using and they should not have to hunt for them again. Then a sample itinerary, day by day, because "seven days from Marmaris" is an abstraction until it is a list of bays. > Screenshot (yacht page hero): Yacht detail page for S Nur Taylan: gallery mosaic, guests, cabins, crew, length and refit year, and an availability card. — The yacht page after the gallery, the specification strip and the enquiry card. The enquiry card stays beside all of it. It carries the weekly rate, the check-in and check-out dates, and a total — one week, plus VAT at 20%, as a figure rather than a footnote. A price with the tax hidden underneath it is a price the reader has to redo, and someone comparing yachts is going to redo it wrong. > Screenshot (yacht pricing calculator): Pricing section listing what the weekly rate includes, beside a card totalling one week plus 20% VAT. — The week priced as a figure, VAT included, rather than a rate with the tax underneath it. Then the page had to stop moving. Content that arrives after first paint pushes everything below it down, so a thumb aimed at the enquiry button lands on whatever slid into its place. That is a conversion bug wearing the costume of a polish item, and it is worst on the device where the target is a thumb rather than a cursor. The same pass got the page working at phone width, which it had not been. ## The fleet search 263 yachts, filtered by price, length, cabins, guests, type, destination and dates, then sorted. The search and the filters were returning the wrong yachts. The index is not in the browser. It is Elasticsearch, on the backend, and the frontend's job is narrow: turn what someone clicked into a query, send it, render what comes back. So a filter returning the wrong yachts is not something the frontend can settle on its own. I worked it through with the backend team until the filters returned what the panel claims. The alternative was available and worse — filter the fetched page of results in the browser, and the grid would agree with itself while disagreeing with the fleet. > Screenshot (fleet search filters): Fleet listing showing 263 yachts with price, length and cabin filters down the left and a sort control. — 263 yachts. Every filter here is a query against Elasticsearch, not a pass over a fetched list. ## The pages Google sees Around a hundred pages exist for search: destination guides, yacht-type pages, cost and booking guides. They are the top of the funnel. Someone reads one of those before they ever reach a yacht page. Most of my work here arrived as QA tickets against page speed and on-page SEO. None of it is clever work. All of it is the kind that fails quietly: nothing throws, no page breaks, the score just drops, and you find out because somebody ran a report. > Screenshot (destination landing page): Greece charter landing page: best months, main ports, popular routes and yacht types in a fact strip. — One of around a hundred pages that exist for search. This is where most of the QA tickets landed. Speed carries the same weight here as content does. A landing page that takes its time is a page someone leaves before they have seen a single yacht, and the enquiry never happens on any of the pages downstream. ## Working on a site that is already selling None of this shipped on a quiet branch. The site was live and taking enquiries for the whole seven months, which is what decided the order: the yacht page first, because that is where the enquiry is sent, then the search that feeds it, then the pages that feed the search. --- # Writing Engineering write-ups by Ritesh KC. - 2026-09-04 — [How to Get a Claude Pro Free Trial in October 2026](https://riteshkc.com.np/blog/claude-pro-guest-pass): Anthropic removed the public trial button, but you can still get 7 days of Claude Pro for free. Click here to use an active referral pass. Tags: Claude Pro free trial, Anthropic Claude, Claude AI premium, Claude referral link, AI tools 2026, Claude guest pass, ChatGPT alternative. Markdown: https://riteshkc.com.np/blog/claude-pro-guest-pass.md - 2026-08-05 — [A status machine belongs in a table](https://riteshkc.com.np/blog/a-status-machine-belongs-in-a-table): Nine statuses, two sets of allowed transitions, and two names for every state. Moving all of that into rows instead of an if chain got engineering out of the workflow business. Tags: Architecture, PostgreSQL, Prisma. Markdown: https://riteshkc.com.np/blog/a-status-machine-belongs-in-a-table.md - 2026-08-05 — [Endpoints that refuse to be oracles](https://riteshkc.com.np/blog/endpoints-that-refuse-to-be-oracles): A 404 on unsubscribe tells an attacker which tokens are real. A 409 on subscribe tells them who is on your list. Honest status codes leak, and the fix reads like a bug. Tags: Security, NestJS, Architecture. Markdown: https://riteshkc.com.np/blog/endpoints-that-refuse-to-be-oracles.md - 2026-08-05 — [Every reset link returned 400](https://riteshkc.com.np/blog/every-reset-link-returned-400): The reset endpoint was fine. The token was fine. The email was fine. The bug was one path in a captcha config, and the word doing the damage was "includes". Tags: Security, Auth, Debugging. Markdown: https://riteshkc.com.np/blog/every-reset-link-returned-400.md - 2026-08-05 — [Ready to submit is not a boolean](https://riteshkc.com.np/blog/ready-to-submit-is-not-a-boolean): A submit gate that looked like one rule and needed three, plus an error message that names which of four things the student actually has to fix. Tags: Architecture, NestJS, Prisma. Markdown: https://riteshkc.com.np/blog/ready-to-submit-is-not-a-boolean.md - 2026-08-05 — [Six socket events, and not one refetch](https://riteshkc.com.np/blog/six-socket-events-and-not-one-refetch): Realtime usually arrives as a second copy of your state. Treating every socket event as a write into the query cache instead of a signal to refetch keeps it to one. Tags: Next.js, Architecture, Realtime. Markdown: https://riteshkc.com.np/blog/six-socket-events-and-not-one-refetch.md - 2026-07-29 — [A trust score you can rebuild from the log](https://riteshkc.com.np/blog/a-score-you-can-rebuild-from-the-log): Ten lines that replay a moderation history into a score, and why clamping inside the fold instead of after it is the difference between working and quietly drifting. Tags: Architecture, Moderation, PostgreSQL. Markdown: https://riteshkc.com.np/blog/a-score-you-can-rebuild-from-the-log.md - 2026-07-29 — [Cache invalidation when the cache is in another repo](https://riteshkc.com.np/blog/cache-invalidation-in-another-repo): An admin edits a university and the public page keeps the old number for a day. The backend now calls the frontend, over tag names nothing checks. Tags: Next.js, Caching, Architecture. Markdown: https://riteshkc.com.np/blog/cache-invalidation-in-another-repo.md - 2026-07-29 — [Overbooking is a database problem](https://riteshkc.com.np/blog/overbooking-is-a-database-problem): Counting rows before you insert one is not a capacity check. Here is how a unique index, a constraint violation, and a NULL ended up doing the work instead. Tags: PostgreSQL, Concurrency, Prisma. Markdown: https://riteshkc.com.np/blog/overbooking-is-a-database-problem.md - 2026-07-29 — [Rate limiting a provider without a rate limiter](https://riteshkc.com.np/blog/rate-limiting-by-queue-shape): SES caps sends per second. The fix was worker concurrency of one and a sleep, plus an honest note in the code about the case where it stops working. Tags: Queues, BullMQ, Infrastructure. Markdown: https://riteshkc.com.np/blog/rate-limiting-by-queue-shape.md - 2026-07-29 — [The fast moderation layer does less on purpose](https://riteshkc.com.np/blog/the-fast-moderation-layer-does-less): A synchronous gate on five routes that blocks almost nothing, and an async worker that decides everything else. Splitting them by latency was the wrong axis. Tags: Architecture, Moderation, NestJS. Markdown: https://riteshkc.com.np/blog/the-fast-moderation-layer-does-less.md - 2025-06-01 — [In RAG, recall is the number that matters](https://riteshkc.com.np/blog/recall-beats-precision): Why the failure mode that sinks a retrieval system is the right document never showing up, and why that makes recall, not prompt tuning, the thing to optimize. Tags: AI, RAG, Retrieval. Markdown: https://riteshkc.com.np/blog/recall-beats-precision.md --- # How to Get a Claude Pro Free Trial in October 2026 Anthropic removed the public trial button, but you can still get 7 days of Claude Pro for free. Click here to use an active referral pass. - Published: 2026-09-04 - Updated: 2026-09-26 - Tags: Claude Pro free trial, Anthropic Claude, Claude AI premium, Claude referral link, AI tools 2026, Claude guest pass, ChatGPT alternative - Author: Ritesh KC (https://riteshkc.com.np) **How to Unlock a 7 Day Claude Pro Free Trial in 2026** Are you trying to test Claude Pro without paying upfront? I completely get it. You want to see if this AI is actually better for your workflow before you hand over your credit card. ![claude code: terminal claude pro opus 5](https://cdn.riteshkc.com.np/media/claude%20terminal.avif) Here is the reality. Anthropic removed the public free trial button from their main website. You cannot just sign up and get a free week by default anymore. But you can still bypass the paywall legally. **The Guest Pass Loophole** Anthropic relies on a private referral system. They give existing subscribers exclusive Guest Passes to share. ![claude pro: guest invitation page](https://cdn.riteshkc.com.np/media/claude%20pro%20guest.avif) When you click a verified referral link, you instantly unlock 7 full days of premium access. This gives you everything. You get the massive context window. You get the expanded message limits. You get priority access when the servers are busy. There is only one requirement. You must be a new customer and enter your payment details to activate the promotion. You can easily cancel before the week ends to avoid any unexpected charges. **Stop Searching for Expired Codes** You could spend hours scouring forums for a working link. Most public codes hit their claim limits within minutes and expire. You do not need to waste your time doing that. ![claude pro free trial activated](https://cdn.riteshkc.com.np/media/claude%20free%20trial%20activated.avif) I have an active referral pass ready for you to use right now. Grab your free 7 days of Claude Pro right here: [Claude 1 Week Pass](https://claude.ai/referral/etIZD2Clzg?s=claude_ai) --- # A status machine belongs in a table Nine statuses, two sets of allowed transitions, and two names for every state. Moving all of that into rows instead of an if chain got engineering out of the workflow business. - Published: 2026-08-05 - Tags: Architecture, PostgreSQL, Prisma - Author: Ritesh KC (https://riteshkc.com.np) Every workflow app has a status column. And somewhere near it, almost always, sits a function that decides which status is allowed to follow which. Mine started exactly where yours probably did: ```ts if (current === 'DRAFT' && next === 'PENDING') return true; if (current === 'PENDING' && next === 'SENT_TO_UNIVERSITY') return true; // ... ``` Perfectly fine code. It works right up until the business changes its mind, which is the one thing you can count on it doing. Applications in this system start as a draft, go through internal review, get sent out to a university, and then land in one of several endings. Nine states. Within a month of shipping the first version, three requests came in, and each one meant opening that same function: slot a new state between two existing ones, let students withdraw from a place they previously couldn't, and stop calling one of the statuses what it was called. Three deploys. Three sentences of business logic. Not one of them an engineering problem. ## That if chain was holding five things, not one The useful exercise wasn't rewriting it. It was sitting down and writing out everything the transition rules actually had to know. I expected one thing. I got five. **Which transitions are legal.** The obvious one, and the only one the `if` chain was honest about. **Which of those a *student* may trigger.** Staff can move an application from pending to sent-to-university. A student can't. But a student *can* withdraw from nearly anywhere. So this was never one graph with a permission check bolted onto the side. It's two graphs, and one is a subset of the other. **What the state is called, and to whom.** Internally, a status is `REJECTED_BY_GRADSY`. Now imagine showing a student "Rejected by Gradsy" when the actual meaning is "we're not forwarding this yet, here's what to fix." That's a support ticket, and quite possibly a lost customer. Staff need the blunt name to do their jobs. Students need the accurate one. Same row, two audiences. **Whether the transition needs an explanation attached.** Dropping an application into a rejection state without telling the student why is the fastest way I know to generate an angry email. Notice where that requirement lives, though: it's a property of the destination state, not of the person clicking the button. **Where the state sits in the workflow.** Sort an application list "by status" alphabetically and `APPROVED_BY_UNIVERSITY` lands above `DRAFT`, which tells a user precisely nothing. What people actually want is workflow order, and workflow order isn't something you can derive from the name. None of those five fit inside a boolean function. And four of them aren't about transitions at all. They're plain facts about a state that the rest of the app kept needing to know. ## So the status became a row Each status turned into a record, seeded from a catalog file: ```ts { code: S.PENDING, studentLabel: 'Under review', staffLabel: 'Pending approval', description: 'Submitted; awaiting Gradsy counselor/admin approval.', studentVisible: true, isTerminal: false, sortOrder: 1, allowedNextCodes: [ S.SENT_TO_UNIVERSITY, S.REJECTED_BY_GRADSY, S.REVISION_REQUESTED, S.WITHDRAWN, ], studentAllowedNextCodes: [S.WITHDRAWN], requiresFeedback: false, } ``` `allowedNextCodes` is the staff graph. `studentAllowedNextCodes` is the narrower student one: from under-review, a student's only available move is to withdraw. The two label fields are the same state, described to two different readers. `requiresFeedback` is set to true on the rejection states, and the service flatly refuses the transition if notes are missing. Here's the load-bearing detail, though: the application table's status column is a foreign key into this catalog, not an enum. That single choice is what pays for everything else. A status list query can `order by statusDef.sortOrder` and come back in workflow order. The label a user sees becomes a join instead of a `switch` buried in the frontend. And adding a new state is an insert. Validation, meanwhile, shrank to almost nothing: read two rows, check membership, narrow the set if the actor is a student, and enforce the feedback requirement on the destination. That's the whole thing. ## What it bought What I actually wanted was to stop being the bottleneck for workflow changes. That part worked. Changing which transitions are legal, renaming a state for students, making feedback mandatory somewhere new: all of it is a catalog edit and a re-seed. The backend README spells this out, because the instinct of anyone new on the codebase is to go hunting for the `if` chain: > To change transitions, edit the catalog and re-seed: don't look for an `if` chain in the service. The part I didn't see coming was how much every *other* status-aware feature benefited. The deadline reminder scheduler only nags about applications that aren't finished, so it joins on `isTerminal` instead of carrying a hardcoded list of ending states: the kind of list that goes quietly stale the second someone adds a tenth status. The admin filter pills are generated. The student-facing timeline renders `studentLabel`. None of them needed a second definition of anything, which means none of them can drift from the first. Then the article CMS came along with its own workflow, and the same pattern dropped straight in: `authorAllowedNextCodes` playing the role that `studentAllowedNextCodes` plays here. That's the real evidence the shape was right. The second use didn't require bending the abstraction to fit. ## What it costs Two things. I'd pay both again, but they're real, and pretending otherwise would be dishonest. **The catalog is a deploy artifact wearing a data costume.** It lives in a TypeScript file and gets seeded. So "operations can change this without engineering" is true in the sense that no service code changes, and false in the sense that somebody still has to run a seed. A proper admin UI over these rows is the version where that claim holds up completely, and it isn't built yet. **Referential integrity has teeth.** Statuses are foreign-keyed from applications *and* from the status history table, which is exactly what makes history survive. The flip side: you cannot delete a status that any application has ever passed through. That's correct behaviour, and it still surprises people every time. Retiring a state means marking it retired, not removing it. ## The rule I'd carry forward When a `switch` on an enum starts showing up in more than two files, the enum is asking to become a table. The `switch` was never the problem. The problem is five separate files each holding a partial copy of the same domain knowledge, all of them free to disagree with one another the moment someone edits four of them. Moving it into rows doesn't make the logic any cleverer. It just makes there be one of it. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Endpoints that refuse to be oracles A 404 on unsubscribe tells an attacker which tokens are real. A 409 on subscribe tells them who is on your list. Honest status codes leak, and the fix reads like a bug. - Published: 2026-08-05 - Tags: Security, NestJS, Architecture - Author: Ritesh KC (https://riteshkc.com.np) An oracle is any endpoint that answers a question you didn't intend to expose. Usually it does it through a status code, and usually the status code is the correct one. Three of these turned up in a newsletter feature, which is not where I expected to spend a day thinking about information leakage. A subscribe form, an unsubscribe link, and an admin list. All small. All leaked something. ## Subscribe: the 409 that enumerates your list The natural implementation of a mailing list signup: look for the address, create it if it's new, return a conflict if it's already there. That conflict is a lookup service. Point a script at it with a list of email addresses and it will tell you, one 409 at a time, exactly which of them are on your list. For a general newsletter that's mildly bad. For a list that implies something about the person: a study-abroad consultancy's list implies you're planning to emigrate. It's worse, and it's the kind of thing that is nobody's business by default. So subscribe is idempotent and says the same thing either way: ```ts /** * Idempotent by design: unlike the waitlist, a repeat subscribe is a no-op * rather than a 409, and the response never reveals whether the address was * already on the list (that would be an email-enumeration oracle). */ async subscribe(dto: SubscribeDto, meta: RequestMeta) { ``` The comment names the sibling case on purpose. The waitlist *does* return a conflict, because "you're already on the waitlist, position 340" is information the person is entitled to and actively wants. Same shape, opposite answer, because the two lists mean different things. That's the part worth copying, not "never 409", but *check what a repeated request reveals about the first one*. ## Unsubscribe: the 404 that validates tokens Unsubscribe links carry a token. A token that doesn't resolve is, on the face of it, a 404. But the endpoint is unauthenticated by necessity: the whole point is that it works from an email client, in one click, with no session. So a 404 is a free token validity check for anyone who wants to brute-force the space, and every successful guess unsubscribes a real person. ```ts /** * Always reports success, including for an unknown token: a 404 here would * turn the endpoint into a token oracle. */ async unsubscribe(token: string) { await this.prisma.newsletterSubscriber.updateMany({ where: { unsubscribeToken: token, status: NewsletterStatus.SUBSCRIBED }, data: { status: NewsletterStatus.UNSUBSCRIBED, unsubscribedAt: new Date(), }, }); return { unsubscribed: true }; } ``` `updateMany` is doing the work. `update` throws when nothing matches, and now the exception handler is the oracle instead of the status code. `updateMany` updates zero rows and returns quietly, so a valid token and a garbage token are indistinguishable from outside: same status, same body, same shape. This looks like a bug in review. It reads as swallowing a failure. That's exactly why the comment states the threat rather than the behaviour, without it, someone helpfully "fixes" this back into a 404 within a year. The user-facing side stays honest, incidentally: there's a separate preview endpoint that shows a masked address before you confirm, so a real unsubscribe still feels like it did something specific. ## The admin list: a capability in a column The unsubscribe token isn't an identifier. It's a capability: anyone holding it can unsubscribe that person, no login required. Which means the admin subscriber list, an ordinary paginated table behind an admin role, must not return it. Not because admins are untrusted, but because the value ends up in a JSON response, a browser cache, a log, a support screenshot. A capability's blast radius is wherever it has ever been copied. ```ts // The token is a capability. It never leaves the server. ``` An explicit `select` rather than a default include, so adding a field to the model doesn't silently publish it. This is the one of the three that isn't really about attackers. It's about the value escaping through ordinary, well-intentioned plumbing. ## Honeypots: lying on purpose The last one goes further than staying quiet. Public forms (newsletter, appointment booking) carry a hidden `company` field. Humans never see it. Bots fill everything. The obvious response to a filled honeypot is to reject the submission. But a rejection is feedback: a script that gets a 400 knows the submission failed, and whoever wrote it will try again with different field handling until it stops failing. You've built them a test harness. ```ts // Honeypot filled → a bot. Mimic a successful booking without persisting // anything, so the script gets no failure signal to adapt to. if (dto.company) { this.logger.warn(`Honeypot tripped from ${meta.ipAddress ?? 'unknown'}`); return { id: null, slotDate: dto.date, slotTime: dto.time, status: AppointmentStatus.PENDING, assigned: false, }; } ``` It returns a booking-shaped object that was never written. The bot records a success and moves on. The `id: null` is the tell for anyone reading the code, and it's server-side only: the bot has no reason to check. This is the one I'd flag as a tradeoff rather than a straight win. Silent fake success means a false positive is invisible: a browser extension autofilling that field, or an accessibility tool exposing it, produces a person who thinks they booked an appointment and didn't show up in anyone's calendar. The log line exists precisely so that's discoverable. Name the field carefully, keep it genuinely hidden, and watch the warnings. ## The common thread Every one of these started as the textbook-correct response. 409 for a duplicate. 404 for a missing resource. Reject invalid input. Return the model. The question that changes the answer isn't "is this the right status code." It's **what does an attacker learn from the difference between two responses**, because an endpoint that responds differently to a hit and a miss is a lookup service for whatever distinguishes them, no matter what it was built for. And once you've decided to answer identically in both cases, you have to write down why. Every fix here looks like a mistake: a repeat that doesn't conflict, a miss that doesn't 404, a rejection that reports success. The comments aren't documentation. They're the only thing standing between the behaviour and a well-meaning cleanup. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Every reset link returned 400 The reset endpoint was fine. The token was fine. The email was fine. The bug was one path in a captcha config, and the word doing the damage was "includes". - Published: 2026-08-05 - Tags: Security, Auth, Debugging - Author: Ritesh KC (https://riteshkc.com.np) Forgot-password worked. You typed your email, the form said check your inbox, and the email arrived: correct branding, correct link, correct token. Click the link and you got a 400. Not a "token expired" page. Not a redirect to login. A raw 400 from the API, in a browser tab, with a JSON body, which is roughly the worst thing a user can be shown in the middle of trying to get back into their account. ## What I checked first, and what was fine The token. Freshly issued, unexpired, present in the URL, present in the database. The handler. It ran correctly against curl. Against the exact same token from the exact same email, a manual request succeeded and the password changed. That last part is the interesting one, and I sat on it for longer than I should have. The same token, the same route, the same server: succeeding from my terminal and failing from Gmail. Which meant the request was being rejected before it ever reached the handler, based on something about *how* it arrived rather than *what* it carried. ## One header, one list The site puts Cloudflare Turnstile behind a single header (`x-captcha-response`) on both the auth endpoints and the public marketing forms, so there is one thing to send and one thing to verify no matter which surface you're on. On the auth side that's better-auth's captcha plugin, and it's configured with a list of paths: ```ts captcha({ provider: "cloudflare-turnstile", secretKey: process.env.TURNSTILE_SECRET_KEY, endpoints: [ "/sign-in/email", "/sign-up/email", "/forget-password", "/reset-password", ], }); ``` Read that list. It names four endpoints. It is four endpoints in the sense that a human means four endpoints. The plugin does not match by equality. It matches by containment: is the request path *in* one of these. So `/reset-password` on that list does not protect one route. It protects everything with `/reset-password` in front of it, and the route that inherited the protection was `/reset-password/:token`: the GET a user performs by clicking a link in their email client. An email client cannot attach a custom header to a link. There is no request the browser could possibly send that carries `x-captcha-response` here, because nothing on my page issued it: the user clicked a plain anchor tag from Gmail. The captcha layer looked for the header, found nothing, and returned 400 before any of my code ran. The POST that requests a reset was genuinely protected and genuinely working. That's why the flow looked half-healthy instead of broken: the half a bot would abuse worked perfectly, and the half a locked-out human needs did not. ## Why it never showed up locally Unsetting the Turnstile keys disables enforcement on both the client and the server, deliberately, so a new developer can run the whole stack without a Cloudflare account. Locally I had no keys. Locally there was no captcha layer at all, and the reset link worked every single time. The bug existed only where the keys did. ## The fix is a deletion Delete `/reset-password` from the list. Nothing replaces it, because a captcha was always the wrong control here: it exists to make a submission expensive to repeat at volume, and the thing worth stopping on a token-consume route is guessing, not volume. Guessing gets a rate limit: ```ts "/forget-password": { window: 300, max: 3 }, "/reset-password": { window: 300, max: 5 }, ``` Five attempts per five minutes, keyed on IP. Against the token space that's not a defence so much as a statement that brute force isn't going to be pleasant, and unlike the captcha it costs a real user holding a real link exactly nothing. Then the line that matters more than either of them: ```ts // /reset-password is deliberately absent from the captcha endpoints: the // plugin matches by substring, so listing it also gates the GET email-link // callback /reset-password/:token, which cannot carry x-captcha-response. // Every reset link 400'd. Rate limiting covers this path instead. ``` Because the shipped state of that config is now a list of protected auth routes with one obvious hole in it, and holes get filled. Someone tightening security a year from now adds the missing line in good faith, in a one-line diff nobody blocks, and breaks password reset for everyone who isn't already logged in, which is, by definition, everyone who needs it. The comment is the whole fix. The deletion just stops the bleeding. ## Two things I took from this **Find out how your path lists match.** Equality, prefix, or containment: they read identically in the config and behave nothing alike. That list looked like an enumeration of four routes. It was four match rules, and a `:param` route sitting underneath one of them inherited a protection nobody chose for it. **A protection that requires a header only covers requests your own JavaScript makes.** Email links, OAuth callbacks, unsubscribe clicks, anything a user reaches by clicking outside your app: none of it goes through your fetch wrapper, so none of it can be defended by something the wrapper attaches. That's not a gap to plug harder. It's a category of request that needs a control it can actually carry. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Ready to submit is not a boolean A submit gate that looked like one rule and needed three, plus an error message that names which of four things the student actually has to fix. - Published: 2026-08-05 - Tags: Architecture, NestJS, Prisma - Author: Ritesh KC (https://riteshkc.com.np) A student fills in a profile, uploads a passport and a transcript, picks a university, and presses submit. Somewhere behind that button is a function that decides whether the application is allowed to leave draft. I wrote the obvious version first. Every required document type is uploaded, so return true. It was wrong in three different ways, and none of them showed up in testing, because testing an application flow means uploading every document in one sitting as one person. The failures all live in the gaps between those uploads. ## The first gap: uploaded where? Documents in this system live in two places on purpose. A student has a library: one passport, uploaded once, reused across every application they ever make. That's the whole point of it. Re-uploading the same identity document for six universities is the kind of small indignity that makes people abandon a product. An application has its own documents, and each one either owns its file or points at a library document. Attaching is a reference, not a copy, so the same S3 object serves six applications and nobody is billed for six copies of a passport. Which means "the student has a verified passport" and "this application has a verified passport" are different sentences. The naive check asks the first one and lets a student submit an application they never attached anything to. The library is a shelf; putting a book on the shelf isn't the same as putting it in the envelope. So the rule is attachment plus verification, on this application: ```ts const documents = await this.prisma.applicationDocument.findMany({ where: { applicationId, documentType: { in: attachedTypes } }, select: { documentType: true, status: true }, }); const notVerified = attachedTypes.filter( (type) => !documents.some( (d) => d.documentType === type && d.status === DocumentStatus.VERIFIED, ), ); ``` The comment above it in the repo says the part that matters: *a verified passport the student never attached should not unlock an application it isn't part of.* ## The second gap: a requirement with no document English proficiency is required. English proficiency is also not necessarily a document. A student can satisfy it with an IELTS or TOEFL score, which is a row in a test-score table with a number in it. There's no file, and more to the point there's nothing to attach: a test score has no attach-to-this-application concept, because it isn't a document, it's a fact about the student. Run the attachment rule over it and you get a student who has proved their English, sees the requirement listed as unmet, and can never submit anything. Not a slow path. A permanent one. So English proficiency is checked at the student level, against a rule that accepts either a test score or a library-verified document: ```ts const englishReq = getEnglishProficiencyRequirement(); const englishOk = await this.prisma.student.findFirst({ where: { id: studentId, ...requirementWhere(englishReq, 'has') }, select: { id: true }, }); if (!englishOk) missing.push(DocumentType.ENGLISH_PROFICIENCY); ``` `requirementWhere` is the piece I'd reach for again. Requirements are declarative objects that emit Prisma `where` fragments, in either a `has` or a `missing` mode: the same definition drives this gate, the admin filter that lists students missing English proficiency, and the notification group that emails them. One definition, three consumers, and adding a fourth requirement means adding an entry rather than editing three services. ## The third gap: reviewed after the fact The statement of purpose is required too, and it inverts the rule again. An SOP is written for one specific university and program. It never goes in the library: there's nothing reusable about it. And it's read by a counselor *after* submission, as part of reviewing the application, not before it as a precondition. Gate it on verification and you've built a deadlock: the student can't submit until it's verified, and nobody verifies it until it's submitted. So SOP is gated on presence. Upload one and you may submit. With one exception that took a second pass to notice, if every copy the student uploaded was rejected, presence is technically satisfied and the application is definitely not ready. Rejected-only still blocks. Three required things, three different rules, because they are three different kinds of claim: a document you supply, a fact about you, and a thing you write for this application specifically. ## The error message is the feature Here's what I actually changed my mind about while building this. The gate is maybe fifty lines. The message it throws is about forty more, and I initially resented writing them. Then I thought about who reads it. A student, on a deadline, at eleven at night, who cannot submit and does not know why. If the API says `documents incomplete`, that student emails the consultancy. A human reads the email, opens the admin panel, looks at four document rows, and writes back. That's twenty minutes of staff time to communicate something the server already knew precisely. So the failure is sorted into four buckets, each meaning a different action: ```ts if (missing.length) { parts.push(`${labels(missing)} ${missing.length === 1 ? 'is' : 'are'} not attached`); } if (awaiting.length) { parts.push(`${labels(awaiting)} ${awaiting.length === 1 ? 'is' : 'are'} not verified yet`); } if (notUploaded.length) { parts.push(`${labels(notUploaded)} ${notUploaded.length === 1 ? 'is' : 'are'} not uploaded`); } if (rejected.length) { parts.push( `${labels(rejected)} ${rejected.length === 1 ? 'was' : 'were'} rejected: upload a revised copy`, ); } throw new BadRequestException(`Your documents aren't ready yet. ${parts.join('; ')}.`); ``` Which produces: *Your documents aren't ready yet. Passport is not attached; Transcripts are not verified yet.* Not attached means go to your library and attach it: thirty seconds. Not verified yet means wait, nobody is blocked on you. Not uploaded means go find the file. Rejected means look at the reason and send a new one. Four sentences, four different next actions, and only one of them is "wait." The pluralisation is in there because "Transcripts is not verified" reads like the system is broken, and a student who thinks the system is broken emails anyway, which defeats the entire point. ## What I'd take to the next project The rule that generalises isn't about documents. When a gate guards something a user genuinely wants to do, every branch of it is a message, and the messages are the part with the business value. The boolean is for the code. The four-way split is for the person stuck behind it. And when one check needs three different rules, that's usually not the check being messy. It's the domain telling you that the three things you grouped together aren't the same kind of thing. I grouped them because they were all rows in one table. That's a storage detail, not a rule. The version that works asks a different question per requirement: *how would someone prove this?* A document you attach. A fact we already hold. A thing you write for this application, which we'll read later. Three answers, three rules, and the shape of each rule follows from the answer rather than from where the data happens to be stored. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Six socket events, and not one refetch Realtime usually arrives as a second copy of your state. Treating every socket event as a write into the query cache instead of a signal to refetch keeps it to one. - Published: 2026-08-05 - Tags: Next.js, Architecture, Realtime - Author: Ritesh KC (https://riteshkc.com.np) The default way to add realtime to an app that already fetches data is to bolt a socket onto the side, listen for events, and refetch whatever changed. It works. It also means every piece of state now has two sources: the fetch that owns it and the socket that pokes it, and the two disagree for as long as the refetch takes. Add a few features and you have a second state tree, unversioned, with no cache policy, racing the first one. The version I ended up with has six event handlers and almost none of them fetch anything. The socket writes into the same cache the rest of the app reads from. Realtime stops being a parallel system and becomes a second writer to the one you already have. ## The unread badge Start with the smallest case, because it makes the principle obvious. A notification bell shows an unread count. The count changes when a notification arrives, when you read one, and when you read one *in another tab*. The refetch version listens for an event and invalidates the count query. That's a round trip to learn a number the server just sent you. ```ts // Authoritative unread count: covers multi-tab and post-mutation sync. socket.on(UNREAD_COUNT_EVENT, (count: number) => { queryClient.setQueryData(notificationKeys.unreadCount(), count); }); ``` The server emits the count because the server knows the count. Writing it straight into the cache is one line, zero requests, and it fixes multi-tab for free: every tab holding a socket gets the same authoritative number at the same time, without any of them asking. The word doing the work is *authoritative*. This only holds if the server sends the real value rather than a nudge. An event that says "something changed" forces a refetch by construction. An event that says "it's four" doesn't. That's a payload design decision, and it's the one that determines whether the rest of this is possible. ## The row that only updates if you're looking at it An admin sends a broadcast to a few thousand people. There's a progress card, and a table you can expand to see per-recipient delivery status. The backend emits one event per recipient as it works through the queue. If each of those invalidated the campaign detail query, an admin who happened to have the table open would fire thousands of requests at their own API while it was already busy sending email. So the handler patches, and only if there's something to patch: ```ts // Patch one recipient's row in an open drill-down. Only touches the detail // cache if it's already loaded (table expanded): no fetch on its own. socket.on(CAMPAIGN_RECIPIENT_EVENT, (event: CampaignRecipientEvent) => { queryClient.setQueryData( campaignKeys.detail(event.campaignId), (prev) => prev ? { ...prev, recipients: prev.recipients.map((r) => r.userId === event.userId ? { ...r, inApp: event.inApp, emailStatus: event.emailStatus } : r, ), } : prev, ); }); ``` The `prev ? ... : prev` is the whole trick. `setQueryData` with an updater that returns `undefined` would seed an entry; returning `prev` unchanged when there's nothing cached means the handler is a no-op for every admin who *isn't* staring at that table. No fetch, no cache entry, no work. An event whose only effect is on data nobody has loaded should cost nothing. That's hard to arrange with invalidation, which is why invalidation is the wrong default here. ## Invalidating something the event didn't mention One handler does invalidate, and it invalidates more than you'd expect. When a counselor approves or rejects a student's document, the obvious move is to refresh the document library. But documents also gate application submission: an application sitting in draft is blocked until its attached documents are verified. A student watching that draft while their document gets approved should see the submit button unlock, without reloading. So `document:changed` busts the applications cache too. The reason is in the code, because six months later it looks like a copy-paste mistake: > application documents mirror the library's review outcome and gate submission: refresh them too so an open draft unlocks live This is the case where invalidation is right. The event carries a document id; the consequence is spread across an unknown number of application records the client may or may not hold. There's no value to write. So you refetch, and you write down why you're refetching something the event didn't name. The handler is also role-scoped: an admin receives these events for every student, so the key it invalidates includes the `studentId` from the payload rather than the current user's. One handler, two audiences, different cache keys. ## The one that isn't cache at all The last piece isn't a cache write, and it's the one users actually notice. A notification arrives while the tab is in the background. The app fires a toast into a document nobody is looking at, the toast times out, and the notification is gone. The user finds out later, or doesn't. ```ts // Tab in the background → fire a native OS notification instead of an // in-app toast the user can't see. Foreground keeps the sonner toast so // the two never fire at once. if ( typeof window !== "undefined" && "Notification" in window && Notification.permission === "granted" && document.hidden ) { const native = new Notification(ui.title, { body: ui.message, tag: ui.id, icon: "/logo/gradsy-mascot.png", }); native.onclick = () => { window.focus(); if (ui.actionUrl) router.push(ui.actionUrl); native.close(); }; return; } ``` `document.hidden` picks the channel. The early `return` is what stops both from firing. `tag: ui.id` dedupes, so the same notification arriving twice replaces rather than stacks. And the click handler focuses the window before routing, because a deep link into a background tab is a page change nobody sees. The honest limit: this only works while the tab exists. Close it and there's no page to run the handler. Real push needs a service worker, which is written up in that repo and not built. ## What the rule turned out to be I didn't set out with a principle. It emerged from getting the campaign handler wrong first. **If the event carries the new value, write it. If it carries a fact whose consequences you can't enumerate, invalidate. If it changes something nobody has loaded, do nothing.** Most events are the first kind, and most codebases treat all of them as the second. That's the difference between a socket layer that stays small and one that turns into a second state tree. The prerequisite is the payload. All of this rests on the server sending values instead of nudges, which is a backend decision made before any of this frontend code exists. If your events say `{ type: "changed" }`, you don't have this option, and getting it back means changing both sides. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # A trust score you can rebuild from the log Ten lines that replay a moderation history into a score, and why clamping inside the fold instead of after it is the difference between working and quietly drifting. - Published: 2026-07-29 - Tags: Architecture, Moderation, PostgreSQL - Author: Ritesh KC (https://riteshkc.com.np) Gradsy scores community content for harm. Profanity, spam signals, link volume, shouting, posting bursts: each contributes points, and the total decides whether a post is approved, queued for a human, or removed. Then there's one more term, and it's the one that makes this interesting: ```ts // Higher trust reduces the harm score. total -= (input.userTrustScore / 100) * w.trustReduction; ``` A user's trust score discounts their content's harm score. And trust itself moves based on past moderation decisions: up a little when content is approved, down when it's flagged or blocked. So the score depends on trust, and trust depends on previous scores. That's a feedback loop, and every feedback loop eventually needs the same thing: a way to recompute the current state from nothing, because sooner or later the running value will be wrong and you'll need to prove what it should have been. ## The scoring function is boring on purpose ```ts const DEFAULT_WEIGHTS: ScoringWeights = { profanity: 40, spam: 20, links: 15, caps: 10, repetition: 10, rate: 25, trustReduction: 15, }; // Pure, synchronous. Higher score = more harmful, clamped to [0, 100]. score(input: ScoringInput): number { const w = this.weights; let total = 0; if (input.profanityFlagged) total += w.profanity; if (input.spamFlagged) total += w.spam; if (input.linksFlagged) total += w.links; if (input.repetitionFlagged) total += w.repetition; if (input.rateFlagged) total += w.rate; // Caps contributes proportionally to how far over it leaned. if (input.capsRatio > 0) total += w.caps * Math.min(1, input.capsRatio); total -= (input.userTrustScore / 100) * w.trustReduction; return Math.max(0, Math.min(100, Math.round(total))); } ``` No database calls, no async, no injected clock. Everything it needs arrives as an argument. That's not architectural purity for its own sake. It's the property that makes the whole thing testable without a running Postgres, and it's why I could reason about the clamping problem later on paper instead of in a debugger. Two details that took a second pass to get right. **Caps is the only proportional signal.** Everything else is a boolean that either fires or doesn't. Caps uses the ratio, because "35% uppercase" and "100% uppercase" are genuinely different behaviours and collapsing them to one bit threw away the distinction that mattered. The `Math.min(1, ...)` is defensive: the ratio should never exceed 1, and if a bug upstream makes it 3, I'd rather cap the contribution than let one signal blow through the whole scale. **Trust is worth 15 points out of 100.** That number is a policy decision disguised as a constant. A perfectly trusted user gets a 15-point discount; someone with zero trust gets none. It's a thumb on the scale, not a pardon. Set it to 40 and a long-standing member could post almost anything; set it to 5 and history stops meaning anything. Fifteen means trust can move a borderline post across a threshold but can't rescue content that trips several signals at once. Then the thresholds: ```ts // score < approve => APPROVE; >= block => BLOCK; otherwise REVIEW // (REVIEW is the inclusive lower bound at the approve threshold). decide(score: number): ModerationDecision { if (score < this.approveThreshold) return ModerationDecision.APPROVE; if (score >= this.blockThreshold) return ModerationDecision.BLOCK; return ModerationDecision.REVIEW; } ``` Below 30, approve. At or above 70, block. In between, a human looks. The forty-point middle band is the point. An automated system that only decides yes or no has to be right, and this one isn't good enough to be right. It's a weighted sum of heuristics. Giving it a third answer, "I don't know, ask someone," is what lets the thresholds be conservative at both ends instead of split down the middle. ## Two ways to know a score Trust moves after every decision: ```ts private deltaForDecision(decision: ModerationDecision): number { switch (decision) { case ModerationDecision.APPROVE: return this.config.get('moderation.trust.cleanPost', 2); case ModerationDecision.REVIEW: return this.config.get('moderation.trust.review', -5); case ModerationDecision.BLOCK: return this.config.get('moderation.trust.block', -10); default: return 0; } } ``` Everyone starts at 50. Clean content earns +2. A review costs −5, a block −10. Losses hurt more than wins help, and clean posting takes a while to dig you out: five good posts to undo one block. That asymmetry is deliberate: the cost of a false negative in this system is content nobody catches, and the cost of a false positive is one annoyed student and a review queue. That's the incremental path: read, add delta, clamp, write. The other path replays everything: ```ts // Recomputes a score from moderation history: starts at default and folds in // each logged decision's delta. Clamped to [0, 100]. async recalculate(userId: string): Promise { const logs = await this.prisma.moderationLog.findMany({ where: { userId }, select: { decision: true }, }); const score = logs.reduce( (acc, { decision }) => clamp(acc + this.deltaForDecision(decision)), this.defaultScore, ); const row = await this.prisma.userTrustScore.upsert({ where: { userId }, create: { userId, score }, update: { score }, }); return row.score; } ``` Ten lines, one `reduce`. The moderation log already existed. It's the audit trail for every decision, written whether or not anyone ever reads it. Trust turned out to be a projection of data I was already keeping, which meant the rebuild function was almost free. This is event sourcing in the only dose I've ever wanted it. No framework, no event store, no aggregate roots. One append-only table that exists for audit reasons, and one function that folds it. Why you need it: someone will change a weight. A block will be reversed by an admin. A bug will double-apply a delta during a retry. Without a rebuild, every one of those leaves a number in a column that nobody can justify and nobody dares touch. With it, the stored score is a cache and you can always regenerate the truth. ## The clamp goes inside the fold Look at where `clamp` sits. It wraps each step, not the final total. That looked like a stylistic choice when I wrote it. It isn't. It changes the answer. Take a user with six blocks followed by six clean posts. Clamping per step: ```text 50 → 40 → 30 → 20 → 10 → 0 → 0 (sixth block clamps, -10 becomes -0) → 2 → 4 → 6 → 8 → 10 → 12 ``` Final: **12**. Clamping only at the end: ```text 50 + (6 × -10) + (6 × +2) = 50 - 60 + 12 = 2 ``` Final: **2**. Same events, same order, six-fold difference. The clamp isn't a display concern. It's part of the arithmetic, because it discards magnitude. Once you're at zero, further blocks cost nothing, and that "wasted" negative is exactly what the end-clamped version keeps and re-applies. Which is right? The per-step version, and not because it's nicer, because **it has to match the incremental path.** `updateScore` clamps every single write. If `recalculate` clamped only at the end, a rebuilt score would silently differ from a live one for any user who ever hit a bound. You'd have two functions claiming to compute the same value, agreeing on most users, disagreeing on precisely the users you most want to be sure about. If you keep a running value and a rebuild function, they aren't two features. They're one invariant, and it's worth a test that asserts they agree, including a case that saturates a bound, which is the only place they can drift. There's a consequence I didn't think through at the time. Clamping per step makes the fold **order-dependent**. Six blocks then six approves gives 12; six approves then six blocks gives 2. Reordering the same events changes the result. Which is fine: the log is a history, histories have an order, and rebuilding it in order is the correct thing to do. Except. ## The bug I found writing this post ```ts const logs = await this.prisma.moderationLog.findMany({ where: { userId }, select: { decision: true }, }); ``` There is no `orderBy`. A `SELECT` without `ORDER BY` has no guaranteed row order in Postgres. It usually comes back in physical order, which usually resembles insertion order, which is why this has never visibly misbehaved. But "usually" is doing all the work in that sentence: an index-only scan, a plan change after a vacuum, or a parallel sequential scan can each hand back a different order, and any of them would produce a different score for the same user with nothing in the logs to indicate why. An order-dependent fold over an unordered query. The fix is one line: ```ts orderBy: { createdAt: 'asc' }, ``` I did not find this by testing. I found it writing the paragraph above, where I typed "rebuilding it in order is the correct thing to do" and then went to check that it actually was. That's the second time explaining code to an imagined reader has caught something reading the code didn't. I don't have a clean theory for why, beyond this: reviewing code asks "is this right?", which your brain answers by pattern-matching. Explaining it asks "why is this right?", which it can only answer by actually deriving it. ## One weight that barely fires While I'm being honest about things I noticed too late: profanity is the heaviest signal in the table at 40 points, and as far as I can tell it almost never fires. The synchronous gate in front of these routes already hard-blocks profanity with a `422` before the content is saved. Every job that reaches the scorer came through that gate: I traced all five enqueue sites, which means `profanityFlagged` is false by construction on the path that matters. It's not dead code. It's the correct weight for a signal that a *different* layer currently intercepts, and if the gate ever loosens, it's already right. But someone tuning these numbers should know that raising or lowering 40 will change nothing about how the system behaves today, and I'd rather say that than let them spend an afternoon on it. ## What I'd add next **A reason column on trust changes.** Right now the score moves and the log records the decision, but reconstructing *why* a specific user is at 18 means reading their whole history and doing the arithmetic by hand. Storing the delta and the resulting score per event would make the rebuild verifiable instead of merely repeatable. **A bound on the replay.** `findMany` loads every log row for the user. Fine at current volumes, unbounded in principle. The usual fix is a periodic checkpoint: store a score plus a watermark, replay only what came after. That reintroduces exactly the drift problem the rebuild exists to solve, so I'd want the agreement test in place first. **Admin overrides don't participate.** When an admin reverses a bot decision, the trust delta is applied but the original log row stays as it was. A rebuild replays the bot's original call, not the human's correction. That's arguably a bug and definitely a surprise, and it's the next thing I'd fix. --- Derived state is a promise you make about a number. The running value is the convenient version of that promise and the rebuild is the honest one, and they only stay the same promise if you're careful about the places where arithmetic stops being arithmetic: every clamp, every floor, every saturating bound. Those are the points where "recompute it from the log" quietly becomes "recompute it *the same way* from the log," which is a much stronger requirement than it looks like when you're writing the convenient version first. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Cache invalidation when the cache is in another repo An admin edits a university and the public page keeps the old number for a day. The backend now calls the frontend, over tag names nothing checks. - Published: 2026-07-29 - Tags: Next.js, Caching, Architecture - Author: Ritesh KC (https://riteshkc.com.np) An admin edits a university's tuition range, saves, and sees a success toast. They open the public page in another tab. The old number is still sitting there. They refresh. Still old. They refresh again, harder, the way people do. Then they message me. Nothing was broken. The public page at `/university/[slug]` is built once and then served from that build for a day: `revalidate = 86400`. The copy they were staring at had been rendered that morning, and Next.js would not build a new one until 24 hours were up. It was doing what I told it to. What I told it had a 24-hour blast radius, and I had built an admin panel next to it that implies edits show up immediately. ## No timer was going to be right First instinct: lower the window. A day is too long. An hour? An hour is a smaller wrong number. It is still fifty-nine minutes of an admin not trusting the tool. And every reduction costs regenerations on pages nobody edited. University pages change a few times a month, so a one-hour window means roughly seven hundred pointless renders a month for every page, to catch a handful of real edits. That ratio is the argument against using a timer at all. The window I want is zero right after a write and infinity the rest of the time, and no single number does that. Time is standing in for the thing I actually care about, which is whether someone changed the data. Next.js has the right primitive for this: `revalidateTag`. Label a fetch with a string, tell Next that string is stale, and the page rebuilds on the next request. Straightforward. Except the write doesn't happen in the app that owns the cache. ## The write happens in another process Gradsy is two repos. The admin UI is Next.js. The mutation is a `PATCH` to a NestJS API in a different process, on a different host, with its own database. So the sequence is: the browser calls the backend, the backend writes to Postgres, and the frontend is never told. The frontend is the thing holding the stale prerendered page, and from where it sits, nothing happened at all. There are three ways out. Only one of them survives contact with the rest of the system. **Have the frontend call `revalidateTag` after its own mutation succeeds.** Tempting, because it keeps everything in one repo. It breaks the moment anything writes without going through that UI: the CSV importer, a script, a background job, a second client. The cache would be correct only for writes that started in one place. **Poll.** Some process asks the backend "changed since?" every N seconds. That adds a second staleness window and a job to operate. **Let the writer announce the write.** The backend knows which university changed and when, because it is the thing that changed it. It just has to say so. So: a webhook, pointing backwards. The frontend usually calls the backend. Here the backend calls the frontend, and it is the only route in the app that works that way. ## The receiving end ```ts // Shared with the backend (REVALIDATE_SECRET there too). Unlike the set-title // webhook this is mandatory: an open endpoint lets anyone dump the cache. const SECRET = process.env.REVALIDATE_SECRET; export const dynamic = "force-dynamic"; export async function POST(req: NextRequest) { if (!SECRET) { return NextResponse.json( { error: "REVALIDATE_SECRET is not configured" }, { status: 500 }, ); } if (req.headers.get("x-webhook-secret") !== SECRET) { return NextResponse.json({ error: "Unauthorized" }, { status: 401 }); } const body = await req.json().catch(() => ({})); const tags = Array.isArray(body.tags) ? body.tags.filter(isNonEmpty) : []; const paths = Array.isArray(body.paths) ? body.paths.filter(isNonEmpty) : []; if (tags.length === 0 && paths.length === 0) { return NextResponse.json( { error: "Provide at least one of `tags` or `paths`" }, { status: 400 }, ); } for (const tag of tags) revalidateTag(tag); for (const path of paths) revalidatePath(path); return NextResponse.json({ ok: true, tags, paths }); } ``` A missing secret returns 500 instead of skipping the check. That is deliberate. There is another webhook in this app, one that sets a dashboard title, where the secret is optional: if it isn't configured, the check is skipped. That's fine there. The worst case is that someone changes a string. This one fails closed. An unauthenticated revalidation endpoint isn't a data leak. It is a free denial-of-service: anyone who finds the URL can bust every tag in a loop and force a regeneration storm against the origin. "Secret not set" has to mean "refuse," not "allow everyone." ## Tags are a vocabulary, and my first version was wrong My first pass tagged each fetch with its own URL. `university:harvard`, `university:mit`, one tag per page. Then someone edited a *program* — a degree attached to forty universities — and I had forty tags to bust and no list of which forty without querying for it. Tags are not page identifiers. They are reasons a page might be wrong. Once I named them that way, the grouping fell out: ```ts // The backend busts these tags on every university/program mutation via // POST /api/revalidate, so the timer is only a backstop for a webhook that // never landed. Keep the tag names in sync with the backend payload. const UNIVERSITY_TTL = 86400; function universityCache(tags: string[]) { return { next: { revalidate: UNIVERSITY_TTL, tags } }; } ``` Every university fetch carries `universities`. The detail fetch carries `university:${slug}` on top of that. The facet endpoints carry `university-countries` or `university-intakes`. The program list carries `programs`. A single university edit busts `universities` and that one slug. A program edit busts `universities` and nothing else. That is coarse, and correct, and one line: ```ts // Programs render inside every linked university page, so there is no // single slug to target: bust the whole collection. this.revalidation.revalidateUniversityCollection(); ``` I'd rather regenerate more pages than maintain a dependency graph I have to keep accurate. The graph would be more precise, and it would go wrong eventually, silently, in the direction of serving stale data. Over-invalidating goes wrong in the direction of a few extra renders. ## Three call shapes, and the one I missed Single-entity edits were obvious. Two others weren't. **The CSV import.** Universities get bulk-loaded from a scraper, hundreds of rows at a time. The first version fired a webhook per row: a few hundred POSTs and a few hundred regeneration triggers for one logical operation. Now the import collects slugs and fires once at the end: ```ts // One webhook for the whole import instead of one per row. this.revalidation.revalidateUniversities(touchedSlugs); ``` **Renames.** This is the one I missed, and it's the one worth remembering: ```ts // A rename mints a new slug, so the old URL has to stop serving the // prerendered page too. this.revalidation.revalidateUniversities([slug, existing.slug]); ``` When a university's slug changes, the new URL has no cache entry, so it renders fresh. Fine. The *old* URL still has a good prerendered page sitting in the cache, and it keeps serving it. The record moved and nothing told its old address. Busting a tag for the new slug does nothing about that. You have to bust both. Any time an identifier can change, invalidation takes two arguments: where the record went and where it was. I've been bitten by this in a router, in a CDN, and in Next's route cache. Same shape every time. ## Failures are swallowed on purpose ```ts /** * Fire-and-forget: a cold cache must never fail an admin write, so failures * are logged and swallowed. Callers do not await. */ ``` Callers don't `await`. The service returns `void`. If the frontend is down, mid-deploy, or slow, the admin's save still succeeds. They get their toast, the row is in Postgres, and the page catches up on the 24-hour revalidate timer instead. This is where the 24-hour window earns its keep. It stopped being a freshness strategy and became a failure backstop: if the webhook never fires, the page is stale for up to a day instead of forever. Same line of config, different job. The tradeoff isn't free. A dropped webhook is invisible. It logs a warning on a box nobody watches, and the admin never learns their edit didn't propagate. If staleness cost money here, I'd want a retry and an alert. For a university page showing tuition ranges, the worst case is a day-old number and a manual re-save. ## The stale comment Read that cache comment again: > the timer is only a backstop for **a webhook that never landed** The webhook landed. It has been running for months. Two services in the backend call it, from fifteen different sites. I wrote that comment during the window where the frontend side existed and the backend side didn't, copied it verbatim into a second file, and never went back. Two repos, one person, and the code drifted from its own documentation inside a week: no reviewer, no handoff, and nobody to blame for not reading the other repo. The tag names are a contract. The backend sends `["universities", "university:harvard"]`, and the frontend has to have tagged its fetches with those exact strings. Nothing checks that. Not TypeScript, not a test, not a build step. Rename a tag on one side and the request still returns `200 OK` with `{ ok: true }`, because `revalidateTag` on a tag nobody uses is a legal no-op. You get a successful webhook, a green log line, and a page that never updates. The comment above the TTL says *"keep the tag names in sync with the backend payload."* That is a comment asked to do a type system's job. It works until the person reading it is in a hurry, and the person in a hurry was me. ## What I'd do differently **Share the tag builders.** A small package, or even a generated file, exporting `universityTags(slug)` and used by both repos. A rename becomes a compile error instead of a silent no-op. I skipped it because a shared package for one function felt like ceremony on a solo project. I was wrong about which of those was cheaper. **Return what actually got busted.** The route already echoes `{ ok: true, tags, paths }`. If the backend logged a warning when it sent a tag the frontend didn't recognise, drift would surface in minutes instead of months. That needs the frontend to know its own tag vocabulary, which loops back to the shared package. **Keep fire-and-forget.** Coupling an admin's save to the availability of a different service, to save a few hours of staleness, is a bad trade. --- Cache invalidation is one of the two hard problems. Doing it across a process boundary adds a third. The tag names become a contract between two services, and unlike the API between them it has no schema, no types, and no failing test: two files in two repos that have to keep agreeing with each other, and nothing that tells you when they stop. Mine stopped agreeing with its own comment in under a week. The code kept working. The comment didn't. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Overbooking is a database problem Counting rows before you insert one is not a capacity check. Here is how a unique index, a constraint violation, and a NULL ended up doing the work instead. - Published: 2026-07-29 - Tags: PostgreSQL, Concurrency, Prisma - Author: Ritesh KC (https://riteshkc.com.np) I was building a consultation booking flow. Three seats per time slot, seven slots a day, a public wizard anyone can walk through without an account. Not a hard feature. I wrote the obvious version in about twenty minutes: ```ts const booked = await prisma.appointment.count({ where: { slotDate, slotTime, status: { not: 'CANCELLED' } }, }); if (booked >= CAPACITY) throw new ConflictException('Slot is full'); return prisma.appointment.create({ data }); ``` Read it. It looks right. It reads like the sentence you'd say out loud: *check if the slot is full, and if it isn't, book it.* It's wrong, and it's wrong in a way that no amount of testing on my laptop was ever going to show me. ## The gap between the count and the insert Postgres runs at Read Committed by default. Prisma doesn't change that. So each of those two statements sees a snapshot of the database taken at the moment that statement began, not at the moment the request began. Two people click "book" on the last remaining seat, eighty milliseconds apart. Both counts run before either insert commits. Both see two rows. Both conclude there's one seat left. Both insert. Now you have four bookings in a three-seat slot, two of them belonging to people who each think they got the last one. The window is small. That's precisely what makes it nasty. It will not show up in development, it will not show up in a smoke test, and when it does show up you'll get a support email that reads "we had four people on a call meant for three" and no error anywhere in your logs. Nothing failed. The code did exactly what it said. I want to be honest about how I found this, because the version where I discovered it in production would make a better story and it isn't what happened. I found it while writing the *availability* endpoint: the one that tells the wizard which slots to grey out. I was writing a `groupBy` to count bookings per slot, and it occurred to me that I was about to use the same count for two completely different jobs: deciding what to show a user, and deciding whether a write is legal. Those are not the same job. One of them is allowed to be a little stale. The other one absolutely is not. ## The fixes I didn't take **Serializable isolation.** This does work. Postgres would detect the conflict and abort one of the transactions. But it raises the isolation level for a whole transaction to solve one constraint, and it hands you serialization failures that you have to catch and retry anyway. You end up writing retry logic regardless, so you may as well write retry logic against something narrower. **`SELECT ... FOR UPDATE` on the slot.** There's no slot row to lock. Slots aren't entities in this schema; they're a date plus a label from a constant array. I could have created a `Slot` table purely to have something to lock, and then every booking serializes on that row. That's a table whose only purpose is to be a mutex. I've built that before. It's fine until someone adds a second reason to touch the table. **Advisory locks.** `pg_advisory_xact_lock(hashtext(...))` on the slot key. Genuinely a good option, and I'd use it if the constraint were more complicated than "at most N rows." It's also invisible: nothing in the schema tells the next person the lock exists. Delete one line and the guarantee is gone with no error, no failing test, nothing. That last point is what pushed me. I wanted the rule to live somewhere it couldn't be casually removed. ## Give each seat a name The move is to stop thinking about *how many* bookings exist and start thinking about *which seat* each booking occupies. A slot with capacity 3 has seats 0, 1, and 2. A booking doesn't increment a counter: it claims a specific seat. And "one booking per seat" is a thing a database can enforce on its own: ```prisma model Appointment { slotDate DateTime @map("slot_date") @db.Date slotTime String @map("slot_time") // Capacity N occupies seats 0..N-1. Set to NULL on cancel to release the // seat: Postgres treats NULLs as distinct in a unique index, so cancelled // rows never block a rebooking. seatIndex Int? @map("seat_index") @@unique([slotDate, slotTime, seatIndex]) @@index([slotDate, slotTime]) } ``` That's the whole guard. Not a service method: an index. The capacity check is now structural. Two concurrent requests can both decide seat 1 is free; only one of them gets to commit a row that says so. The other one gets a unique-constraint violation, which Prisma surfaces as error code `P2002`. ## P2002 is not an error here This is the part that felt wrong to write and turned out to be the whole idea. `P2002` normally means something went wrong: you tried to register an email that already exists, you double-submitted a form. Here it means something completely mundane: *that seat is taken, try the next one.* It's not an exception in the "exceptional" sense. It's the answer to a question. So the booking path is a loop over the seats, and a constraint violation is how you advance it: ```ts /** * Claims the first free seat in the slot. A plain count-then-create * overbooks under Read Committed (two racers both see capacity-1) so the * unique index on (slotDate, slotTime, seatIndex) is the actual guard and * P2002 just means "seat taken, try the next one". * * Each attempt is its own transaction: Postgres aborts a transaction on the * first failed statement, so retrying a seat inside one would only produce * "current transaction is aborted" on every subsequent create. */ private async createWithFreeSeat( data: Omit, actorId: string | null, ) { for (let seat = 0; seat < APPOINTMENT_SLOT_CAPACITY; seat++) { try { return await this.prisma.$transaction(async (tx) => { const created = await tx.appointment.create({ data: { ...data, seatIndex: seat }, }); await this.audit.record(tx, { actorId, action: 'appointment.create', entityType: 'Appointment', entityId: created.id, }); return created; }); } catch (err) { const isSeatTaken = err instanceof Prisma.PrismaClientKnownRequestError && err.code === 'P2002'; if (!isSeatTaken) throw err; } } throw new ConflictException( 'That time slot is fully booked. Please pick another time.', ); } ``` No lock. No isolation-level change. No counting. The loop runs at most `CAPACITY` times, and it only iterates when a seat is genuinely occupied, which, with three seats, means the pathological case is three round trips. If capacity were 500 this would be a bad design and I'd go back to advisory locks. It's 3. I'll take the three round trips. ## The bug inside the fix The first version of that loop had one transaction wrapping the whole thing, with the retry inside it. It looked tidier. It also failed on every seat after the first, with an error I hadn't seen before: ```text current transaction is aborted, commands ignored until end of transaction block ``` Postgres aborts a transaction on the *first* failed statement. Once seat 0 raises a unique violation, that transaction is dead: every subsequent statement inside it fails with the same message regardless of what it is. Catching the error in application code doesn't resurrect anything. The connection is sitting in a poisoned state waiting for a `ROLLBACK`. (Savepoints would let you recover: `SAVEPOINT` before each attempt, `ROLLBACK TO` on failure. That works and is what I'd reach for if the surrounding transaction had to stay open for other reasons. Here it didn't, so a fresh transaction per attempt is simpler and I don't have to explain savepoint semantics to whoever reads this next.) The audit write is inside the transaction on purpose. If the appointment row commits, the audit row commits with it, or neither does. An audit log that can silently miss entries during a retry loop is worse than no audit log, because you'll trust it. The thing I'd tell my past self: when a retry loop lives inside a transaction, the transaction is usually the thing that needs to move, not the loop. ## Cancellation, and a NULL doing real work Then the obvious follow-up question: what happens when someone cancels? A cancelled booking is still a row. If that row keeps `seatIndex = 1`, seat 1 is occupied forever and the slot quietly shrinks to two seats. Deleting the row is not an option: cancellations are history, and staff need to see them. The answer is a single field assignment: ```ts // Cancelling frees the seat; the unique index treats NULL seats as // distinct, so the slot immediately reopens. if (dto.status === AppointmentStatus.CANCELLED) { data.seatIndex = null; data.cancelledAt = new Date(); } ``` In a Postgres unique index, `NULL` is not equal to `NULL`. Two rows with a NULL in the indexed column do not collide. So every cancelled booking for a slot can carry `seatIndex = null`, all of them coexisting happily, none of them blocking a new booking from claiming that seat number. The free list is the absence of a value. No status column in the index, no partial index, no cleanup job. Cancel a booking and the seat is back in circulation on the next insert. I like this more than I probably should. It's also the part I'd flag hardest in review, because it depends on a piece of SQL semantics that reads as a footnote until it's load-bearing. Postgres 15 added `UNIQUE NULLS NOT DISTINCT`: opt-in, so the default behaviour is unchanged, but it means the guarantee this design rests on is now a thing someone can switch off in a migration. The comment sitting directly above the constraint is not decoration. ## Two counts, two jobs The availability endpoint still counts rows. That's fine, because it answers a different question: ```ts const grouped = await this.prisma.appointment.groupBy({ by: ['slotTime'], where: { slotDate, status: { notIn: RELEASED_APPOINTMENT_STATUSES } }, _count: { _all: true }, }); ``` This drives the wizard's slot picker, which times render as available, which render as full. It's allowed to be a few hundred milliseconds stale, because being stale here costs a user one failed booking attempt and a clear error message. Being stale in the write path costs you a fourth person on a three-person call. Worth noting the two views agree only because cancelled rows are excluded from the count *and* carry a NULL seat. Those two facts have to stay in sync. `COMPLETED` and `NO_SHOW` keep their seat and keep being counted, which is correct: a no-show still consumed the slot. ## What I'd carry to the next one **Constraints beat checks.** A check is a statement about a moment. A constraint is a statement about the data. If the rule matters, put it where it survives someone rewriting your service method. **Name the resource you're allocating.** Almost every "N at a time" problem gets easier when you stop counting the things and start identifying the slots they go into. Seats, ports, shard ids, worker lanes. Once each one has a name, uniqueness does the coordinating. **Some errors are answers.** Treating `P2002` as control flow felt like a hack for about a day. It isn't. The database is telling you the current state of the world faster and more reliably than any query you could have written to ask. The version of this code I'm happiest about is the one that's mostly not code. Twelve lines of loop, one line of schema, and a constraint the database enforces whether or not the next person to touch this file understands why it's there. That last part is the only kind of correctness that survives contact with a team. --- *Backend on this is NestJS, Prisma, and PostgreSQL 15: the [Gradsy](https://riteshkc.com.np/work/gradsy) build. Slot capacity is 3, seven slots a day, and yes, all of it is wall-clock time in a timezone that's 45 minutes off the hour. That's a different post.* ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # Rate limiting a provider without a rate limiter SES caps sends per second. The fix was worker concurrency of one and a sleep, plus an honest note in the code about the case where it stops working. - Published: 2026-07-29 - Tags: Queues, BullMQ, Infrastructure - Author: Ritesh KC (https://riteshkc.com.np) Amazon SES gives you a sending quota measured per second. Go over it and you don't get a queue. You get `Throttling` errors, one per rejected message, and the messages are gone. Gradsy sends broadcasts. An admin picks a group of students, writes a message, hits send, and somewhere between fifty and a few thousand emails need to go out. Every one of them is also an in-app notification and a websocket event, so the send is already fanned out into a job per recipient on a BullMQ queue. Which meant, on the first version, a few thousand jobs hitting SES as fast as the workers could drain them. ## Where I started, and why I stopped I reached for a rate limiter. That's the named solution to the named problem. `bottleneck`, or a token bucket in Redis, or BullMQ's own `limiter` option: pick one, wrap the send, done. I got about as far as reading the docs before the shape of it started bothering me. A token bucket is a concurrency control. It exists because you have N things happening at once and need to hold them to a rate. But look at what's actually true here: the jobs are already serialized by a queue. There is exactly one thing happening at a time if I want there to be. I'd be adding a mechanism to control parallelism I hadn't introduced yet. The queue already has the knob: ```ts this.emailSendDelayMs = this.config.get('ses.sendDelayMs', 200); ``` ```ts // Pause after each email the worker actually sends, to stay under the SES // per-second send rate on large broadcasts. Worker concurrency is 1, so // this deterministically spaces sends. ~200ms ≈ 5/sec. ``` Concurrency of one, plus a sleep after each send. That's the whole rate limiter. The word doing the work in that comment is **deterministically**. With one worker and a fixed pause, the interval between sends has a floor you can compute on paper. There's no bucket to refill, no burst allowance to reason about, no window boundary where a hundred tokens become available at once. Sends are spaced by construction. A token bucket sized for five per second will, correctly and by design, let you fire five in the same millisecond. That's usually the feature. Against a provider that measures per-second, it's the exact thing you were trying to avoid. ## Concurrency is a rate limiter you already have The general form, which I now reach for before reaching for a library: > A queue with concurrency `C` and a post-task delay `D` sends at most `C / D` per unit time. Set `C = 1` and you're just choosing `D`. Set `D = 0` and you're just choosing `C`, bounded by how fast the task runs. Between them you can hit most rate targets without introducing a third mechanism. This isn't better than a token bucket at everything. It's worse at bursts: a bucket lets you use your full quota when you have headroom, and this doesn't, so a fifty-recipient broadcast that could finish in a second takes ten. I don't care. Nobody is watching a progress bar for a broadcast, and "finishes slower than strictly necessary" is a category of problem I'll take over "loses mail." What it *is* better at is being obvious. There is no state. Two lines of code, and the second one explains itself: ```ts // Space out real SES sends (worker concurrency is 1) so a large broadcast // stays under the per-second send rate. Skipped emails don't wait. if (emailOutcome !== 'SKIPPED' && this.emailSendDelayMs > 0) { await new Promise((resolve) => setTimeout(resolve, this.emailSendDelayMs)); } ``` ## Only sleep for sends that happened That condition is the part I got wrong first, and it's the part that makes the mechanism actually work. Not every job sends an email. A notification carries an email policy, and the most common one is `IF_OFFLINE`: don't email someone who's currently staring at the app, they already got the toast: ```ts private async maybeEmail(...): Promise { if (emailPolicy === 'NEVER') return 'SKIPPED'; if (emailPolicy === 'IF_OFFLINE' && (await this.redis.isOnline(userId))) { return 'SKIPPED'; } // ... look up address, send } ``` So the worker returns `SENT`, `FAILED`, or `SKIPPED`, and only the first two wait. The version that slept unconditionally was technically safe. It can only under-use the quota, never over-use it, but it made the delay's meaning wrong. Sleeping after a skip isn't rate limiting, it's just being slow. A broadcast to a mostly-online group would crawl for no reason at all, and a future reader trying to work out why 800 recipients took three minutes would find a sleep with no corresponding network call. The rule I'd extract: **pace the thing you're pacing, not the loop it lives in.** Attach the delay to the side effect, not the iteration. (`FAILED` waits too, which is deliberate. A failed send still consumed an API call, and a failure under throttling is exactly when you least want to immediately try again.) ## The hole, which is in the code in writing There are two queues that send email. Notifications is one. The newsletter is the other, and it made the same choice: ```ts // Same knob as the notifications worker. Both run at concurrency 1, so a // simultaneous broadcast and newsletter send doubles the real SES rate. // lower this if that ever trips the send quota. this.sendDelayMs = this.config.get('ses.sendDelayMs', 200); ``` Two workers, each correctly limiting itself to five per second, running at the same time. Ten per second at the provider. The mechanism is local. A token bucket in Redis, shared by both queues, would be global, and this is the case where the thing I argued against is straightforwardly the better tool. Concurrency-as-rate- limit works precisely because there's one consumer, and the moment there are two it silently stops being what it claims to be. It's still there. It hasn't tripped the quota, because a newsletter and an admin broadcast going out in the same minute hasn't happened yet. When it does, the fix is either a shared limiter or merging both onto one queue, and I'd probably merge: one email queue with a job type is less machinery than one queue each plus a coordinator. I'm writing this down partly because I think the comment is the right call. Not "TODO: fix", not silence. A statement of exactly when the design breaks and what to reach for. When it does break at 2am, the person reading that line has the whole answer, and there's a decent chance the person is me having forgotten. An assumption a design depends on should be written where the design lives. "Concurrency is 1" is load-bearing here, and nothing enforces it. Someone can change it in config and quietly triple the send rate with no error anywhere. ## Retries would have doubled sends, except for one `where` Rate limiting isn't the only thing a queue changes about email. BullMQ retries failed jobs three times with exponential backoff, which is great for transient SES failures and terrible if "failure" means "the send worked and the database write didn't." The newsletter worker handles it with a predicate rather than a flag: ```ts // The `status: PENDING` predicate makes a retried job idempotent: a // recipient already marked SENT is never rewritten. await this.prisma.newsletterRecipient.updateMany({ where: { id: recipientId, status: NewsletterDelivery.PENDING }, data: { status: outcome.status, error: outcome.error ?? null, processedAt: new Date() }, }); ``` `updateMany` with a status predicate, so a retry that finds the row already `SENT` matches zero rows and changes nothing. No read-then-write, no race between checking and updating. The condition is evaluated by the database as part of the write. The same file has one more thing worth stealing, which is that the worker re-checks subscription status at send time rather than trusting the snapshot taken when the campaign was queued: ```ts // Someone can unsubscribe between the snapshot and their turn in the queue. ``` A five-thousand-recipient send takes a while at five per second. Someone unsubscribing eight minutes in, at position 3,000, must not receive it, and the list you queued eight minutes ago says they should. Anything queued in bulk needs to ask "is this still true?" at execution time, not just at enqueue time. ## Would I do it again For this, yes. One consumer, one provider, a rate that has to be respected rather than optimized: concurrency and a sleep is proportionate and it has no moving parts. The moment there's a second consumer, no. And I'd argue the failure I actually made wasn't picking the simple mechanism, it was building the second queue without going back and reconsidering the first one's assumption. The design was right when I wrote it and became wrong when I wrote something else, which is the ordinary way designs go wrong. --- Most rate limiting problems are really parallelism problems wearing a costume. Before you add a limiter, look at how many things can call the API at once, if the answer is "one, because a queue says so," you already have a limiter and it's a `setTimeout`. Just write down that the answer is one. That's the part that expires. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # The fast moderation layer does less on purpose A synchronous gate on five routes that blocks almost nothing, and an async worker that decides everything else. Splitting them by latency was the wrong axis. - Published: 2026-07-29 - Tags: Architecture, Moderation, NestJS - Author: Ritesh KC (https://riteshkc.com.np) Community moderation has an obvious shape. Check the content when it arrives, block the bad stuff, let the rest through. One function, called on write. That's what I built. It worked, and then it started producing outcomes I couldn't defend. A student wrote a genuinely useful post about visa interviews, in caps, because they were excited. Blocked. Someone pasted four links to university funding pages. Blocked. Meanwhile the checks that actually mattered: is this person posting the same thing eight times a minute, do they have a history of removals: weren't in the gate at all, because you can't ask those questions in the two hundred milliseconds a `POST` is willing to wait for. The gate was strict about the things it could measure instantly and blind to the things that mattered. ## The wrong axis My first attempt at fixing this was to make the gate smarter. Tune the caps threshold. Raise the link limit. Add exceptions. Every adjustment traded one class of false positive for another, and none of it addressed the real issue: I'd split the system by *latency* and then let that split decide *policy*. Fast checks became blocking checks, purely because they were fast. That's an implementation detail choosing the product behaviour. The question I should have asked first is the one that made everything else fall out: **What actually has to be true before this content is allowed to exist for even a second?** Not "what's suspicious." Not "what might be spam." What is so unambiguously unacceptable that showing it to one person for one second is a real harm. That list is short. Slurs. Links to domains known to be malicious. That's about it. Everything else (caps, link volume, repetition, posting bursts, a user's history) is *evidence*. Evidence deserves a judgment, and judgment can happen a second later. So the split isn't fast versus slow. It's **certain versus probable**. The fast layer got the certain things, which turned out to be a much smaller set than the things it could physically check. ## Layer 1: five routes, two rules The synchronous gate is Express middleware, mounted on exactly the routes that create or edit community content: ```ts consumer .apply(ModerationGateMiddleware) .forRoutes( { path: 'community/posts', method: RequestMethod.POST }, { path: 'community/posts/:id', method: RequestMethod.PATCH }, { path: 'community/posts/:id/comments', method: RequestMethod.POST }, { path: 'community/comments/:id/replies', method: RequestMethod.POST }, { path: 'community/comments/:id', method: RequestMethod.PATCH }, ); ``` Five routes, listed explicitly. Not a global guard with an opt-out, because the failure modes point in opposite directions: a global guard that someone forgets to exempt breaks an unrelated endpoint loudly, which is annoying but visible. This list, if someone adds a sixth write route and forgets, lets content through ungated, which is quiet, and worse. I took the loud failure. The comment above the block says which routes and why, and a new content route is a rare enough event that reading it is realistic. Inside, the middleware runs all the checks (profanity, repetition, caps, links) and then blocks on almost none of them: ```ts // Layer 1 hard-blocks only unambiguous violations: profanity/slurs and // known-malicious link domains. Caps, link *count* and repetition are soft // signals: they pass the gate (attached below) and are scored by Layer 2, // so the two layers no longer overlap. const reasons: string[] = []; if (profanity.flagged) reasons.push('profanity'); if (linkResult.blockedDomains.length > 0) reasons.push('blocked-domain'); if (reasons.length > 0) { res.status(422).json({ blocked: true, reasons }); return; } req.moderationFlags = { profanity, spam: { repetition, caps }, links: linkResult, }; next(); ``` The checks it doesn't act on aren't wasted. They ride along on the request object and get handed to the async layer, which is the part I like: the expensive text analysis happens once, in the place that already has the text, and the slow layer inherits the results instead of re-deriving them. `422` rather than `400`. The request is well-formed: the server understood it perfectly and is refusing on content grounds. The frontend needs to tell those apart to show the right message, and `{ blocked: true, reasons }` gives it something to render besides "something went wrong." ## The bug that made me strip HTML Posts are rich text from a TipTap editor, so the body arrives as HTML. The profanity filter splits on whitespace and normalizes each word. Which means this sails straight through: ```html

asshole

``` The token is `passholep` after stripping non-letters, not a word in any dictionary, profane or otherwise. Wrap a slur in any tag and it stops being a word. I found this by accident, testing formatting. It had been live. ```ts // Concatenates the text-bearing fields of a community write payload. // `body` is rich HTML: strip tags so words don't glue to markup (e.g. // "

asshole

") and slip past the word-level filters. ``` The general version of this is worth internalising: **any filter that tokenizes has an encoding attack against it**, and the fix is always to normalize into the filter's domain before filtering, never to make the filter cleverer. Strip the markup, then match words. Not: teach the word matcher about markup. ## The dictionary problem, and where I stopped `leo-profanity` ships a 253-word list and matches whole words exactly. No stemming. So "fuck" is caught and "fucker" isn't. The obvious fix is substring matching. The obvious problem with substring matching is Scunthorpe: "class" contains a slur if you squint, "assessment" starts with one, "cockpit" and "hello" are casualties of the naive version. What I landed on is boring and I think correct: ```ts // leo-profanity only matches whole words exactly (no stemming), so inflected // forms like "fucker" leak. These roots are matched as substrings to catch // derivations. Curated to avoid false positives: only stems that never occur // inside clean English words (NOT "ass"/"dick"/"cock"/"hell"). const STRONG_ROOTS = [ 'fuck', 'shit', 'cunt', 'bitch', 'nigg', 'slut', 'whore', 'pussy', 'bastard', 'motherfuck', 'dumbass', 'jackass', ]; ``` Two mechanisms, chosen per word. Roots that can't appear inside innocent English are matched as substrings. Slurs that *can* (and a second list of terms missing from the default dictionary) are added as exact matches instead. The parenthetical is the important part of that comment. It's not documenting what the code does, it's documenting the rule for editing it. The next person adding a word needs to know which list it belongs in, and the answer is a question they can actually answer: *does this string ever appear inside a clean word?* I want to be straight about the limits. This catches lazy profanity. It does not catch `f u c k`, or `fµck`, or someone determined. Leetspeak normalization, homoglyph folding, and an ML classifier all exist and all cost either latency, money, or a false-positive budget I didn't want to spend on a student forum where the realistic threat is frustration, not coordinated abuse. If the community grows into a target, this is the first thing I'd replace. Right now it's proportionate, and I'd rather ship a filter I can explain than one I can only tune. ## Layer 2: the part that judges Everything that got through goes onto a BullMQ queue with the Layer 1 findings attached. The worker adds the two things the request path couldn't afford: the user's trust score, and their recent posting rate: scores the whole picture, and maps the score to `APPROVE`, `REVIEW`, or `BLOCK`. The scoring itself is a separate post. What matters here is the plumbing decision on the other side: what do you *do* with a `REVIEW`? The tempting answer is a bot-review queue in the admin panel. A new page, a new list, new filters, new empty state. I didn't build any of that, because a queue of flagged content already existed: the one humans fill by pressing "report": ```ts // Funnels an auto-flagged item into the Report table so it shows up in the // admin reports view, attributed to the system Moderation Bot. Never fails // the job: content may have been hard-deleted by the author meanwhile. ``` The bot files a report the same way a user does, under a fixed system user id. Admins see one queue, sorted the same way, with the same actions. The only difference is the reporter's name. This turned out to be the decision that paid off most in the whole feature, and it took ten minutes. Every improvement to the reports view (filters, bulk actions, deep links) now applies to automated flags for free, permanently, without anyone remembering to make it apply to both. When an automated process needs a human in the loop, look for the human queue that already exists before building it one. A second inbox is a second thing to check, and the one nobody checks is whichever one is emptier. ## Two swallowed errors, both deliberate The worker has two `try/catch` blocks that log and continue rather than fail the job. Both look sloppy in review and both are load-bearing. **Filing the report.** The bot may have already filed on a previous pass: an edit re-triggers scoring, and the reporter table has a uniqueness constraint of one filing per user per item: ```ts .catch((err) => { // P2002: bot already filed on a prior moderation pass; ignore. if (!(err instanceof Prisma.PrismaClientKnownRequestError && err.code === 'P2002')) { throw err; } }); ``` Note it re-throws anything else. Swallowing a specific known-benign error is fine; swallowing every error is how you lose a database outage. **Hiding the content.** Between the write and the worker running, the author can delete their own post. Then the update targets a row that no longer exists and throws: ```ts // Soft-hides blocked content. Content may have been hard-deleted by the // author meanwhile, so swallow errors rather than fail the job. ``` Without the catch, BullMQ retries the job three times, each one failing on a row that will never come back, and the queue accumulates permanent failures for content that resolved itself in the best possible way. Async moderation means the world changes under you. The content you were asked to judge might be gone, edited, or already handled by a human. Every step has to tolerate that, and "tolerate" mostly means deciding in advance which errors are just the world moving on. ## What I'd change **Blocked content still isn't reviewable in one click.** A `BLOCK` soft-hides the post, files the report as `removed`, and notifies the author. If the bot was wrong, an admin can undo it, but the author's notification has already gone out. I'd hold that notification behind a short delay, or send it only after the report has aged without an admin touching it. **One threshold set for all content types.** A comment and a long post get scored identically, which means a two-word reply hits the caps ratio far more easily than an essay does. Per-type thresholds would be a config change, not a rewrite. I just haven't seen it hurt anyone yet. --- The version of this I'd defend in review isn't the one that catches the most bad content. It's the one where I can point at any block and name the specific rule that fired. Making the fast layer do less was what bought that. It only ever runs two rules, so when it says no, there's no ambiguity about why, and everything ambiguous got moved to a place where being wrong costs a second look instead of a rejection. ## Related work - [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail. --- # In RAG, recall is the number that matters Why the failure mode that sinks a retrieval system is the right document never showing up, and why that makes recall, not prompt tuning, the thing to optimize. - Published: 2025-06-01 - Tags: AI, RAG, Retrieval - Author: Ritesh KC (https://riteshkc.com.np) Most RAG demos are convincing for the wrong reason. You ask a question, the answer comes back fluent and correct, and it feels solved. Then you put it in front of real documents and real questions, and it starts confidently answering from nothing, because the chunk it needed was never retrieved. That's the failure mode that matters. Not a slightly-off phrasing, not a suboptimal prompt: the correct source simply not being in the context window. ## Retrieval is the ceiling The generation step can only work with what retrieval hands it. If the right passage isn't in the retrieved set, the model has two options: refuse, or invent. Neither is what you want. So the ceiling on answer quality is set at retrieval, not generation, which means recall is the metric to chase. Precision (how much of what you retrieved was relevant) matters too, but it's recoverable. You can rerank a noisy set down. You cannot rerank a document that was never retrieved back into existence. ## What that changes Optimizing for recall changes concrete decisions: - **Chunk for standalone meaning.** A chunk that loses its context when isolated hurts recall, because it no longer matches the queries it should. - **Retrieve more, then narrow.** Pulling a wider set and reranking beats pulling a tight set and hoping. - **Measure the thing that fails.** Track whether the known-correct source was retrieved at all. That's the number that predicts trust. I wrote about the production version of this in the [rag-pipeline](https://riteshkc.com.np/work/rag-pipeline) case study. The short version: make the right document show up, and the rest of the system gets a lot easier to trust.