---
title: "The fast moderation layer does less on purpose"
description: "A synchronous gate on five routes that blocks almost nothing, and an async worker that decides everything else. Splitting them by latency was the wrong axis."
url: https://riteshkc.com.np/blog/the-fast-moderation-layer-does-less
source: https://riteshkc.com.np/blog/the-fast-moderation-layer-does-less.md
updated: 2026-07-29
site: "Ritesh KC"
---
# The fast moderation layer does less on purpose

A synchronous gate on five routes that blocks almost nothing, and an async worker that decides everything else. Splitting them by latency was the wrong axis.

- Published: 2026-07-29
- Tags: Architecture, Moderation, NestJS
- Author: Ritesh KC (https://riteshkc.com.np)

Community moderation has an obvious shape. Check the content when it arrives, block the bad stuff, let the rest through. One function, called on write.

That's what I built. It worked, and then it started producing outcomes I couldn't defend.

A student wrote a genuinely useful post about visa interviews, in caps, because they were excited. Blocked. Someone pasted four links to university funding pages. Blocked. Meanwhile the checks that actually mattered: is this person posting the same thing eight times a minute, do they have a history of removals: weren't in the gate at all, because you can't ask those questions in the two hundred milliseconds a `POST` is willing to wait for.

The gate was strict about the things it could measure instantly and blind to the things that mattered.

## The wrong axis

My first attempt at fixing this was to make the gate smarter. Tune the caps threshold. Raise the link limit. Add exceptions.

Every adjustment traded one class of false positive for another, and none of it addressed the real issue: I'd split the system by *latency* and then let that split decide *policy*. Fast checks became blocking checks, purely because they were fast. That's an implementation detail choosing the product behaviour.

The question I should have asked first is the one that made everything else fall out:

**What actually has to be true before this content is allowed to exist for even a second?**

Not "what's suspicious." Not "what might be spam." What is so unambiguously unacceptable that showing it to one person for one second is a real harm.

That list is short. Slurs. Links to domains known to be malicious. That's about it.

Everything else (caps, link volume, repetition, posting bursts, a user's history) is *evidence*. Evidence deserves a judgment, and judgment can happen a second later.

So the split isn't fast versus slow. It's **certain versus probable**. The fast layer got the certain things, which turned out to be a much smaller set than the things it could physically check.

## Layer 1: five routes, two rules

The synchronous gate is Express middleware, mounted on exactly the routes that create or edit community content:

```ts
consumer
  .apply(ModerationGateMiddleware)
  .forRoutes(
    { path: 'community/posts', method: RequestMethod.POST },
    { path: 'community/posts/:id', method: RequestMethod.PATCH },
    { path: 'community/posts/:id/comments', method: RequestMethod.POST },
    { path: 'community/comments/:id/replies', method: RequestMethod.POST },
    { path: 'community/comments/:id', method: RequestMethod.PATCH },
  );
```

Five routes, listed explicitly. Not a global guard with an opt-out, because the failure modes point in opposite directions: a global guard that someone forgets to exempt breaks an unrelated endpoint loudly, which is annoying but visible. This list, if someone adds a sixth write route and forgets, lets content through ungated, which is quiet, and worse. I took the loud failure. The comment above the block says which routes and why, and a new content route is a rare enough event that reading it is realistic.

Inside, the middleware runs all the checks (profanity, repetition, caps, links) and then blocks on almost none of them:

```ts
// Layer 1 hard-blocks only unambiguous violations: profanity/slurs and
// known-malicious link domains. Caps, link *count* and repetition are soft
// signals: they pass the gate (attached below) and are scored by Layer 2,
// so the two layers no longer overlap.
const reasons: string[] = [];
if (profanity.flagged) reasons.push('profanity');
if (linkResult.blockedDomains.length > 0) reasons.push('blocked-domain');

if (reasons.length > 0) {
  res.status(422).json({ blocked: true, reasons });
  return;
}

req.moderationFlags = {
  profanity,
  spam: { repetition, caps },
  links: linkResult,
};
next();
```

The checks it doesn't act on aren't wasted. They ride along on the request object and get handed to the async layer, which is the part I like: the expensive text analysis happens once, in the place that already has the text, and the slow layer inherits the results instead of re-deriving them.

`422` rather than `400`. The request is well-formed: the server understood it perfectly and is refusing on content grounds. The frontend needs to tell those apart to show the right message, and `{ blocked: true, reasons }` gives it something to render besides "something went wrong."

## The bug that made me strip HTML

Posts are rich text from a TipTap editor, so the body arrives as HTML.

The profanity filter splits on whitespace and normalizes each word. Which means this sails straight through:

```html
<p>asshole</p>
```

The token is `passholep` after stripping non-letters, not a word in any dictionary, profane or otherwise. Wrap a slur in any tag and it stops being a word.

I found this by accident, testing formatting. It had been live.

```ts
// Concatenates the text-bearing fields of a community write payload.
// `body` is rich HTML: strip tags so words don't glue to markup (e.g.
// "<p>asshole</p>") and slip past the word-level filters.
```

The general version of this is worth internalising: **any filter that tokenizes has an encoding attack against it**, and the fix is always to normalize into the filter's domain before filtering, never to make the filter cleverer. Strip the markup, then match words. Not: teach the word matcher about markup.

## The dictionary problem, and where I stopped

`leo-profanity` ships a 253-word list and matches whole words exactly. No stemming. So "fuck" is caught and "fucker" isn't.

The obvious fix is substring matching. The obvious problem with substring matching is Scunthorpe: "class" contains a slur if you squint, "assessment" starts with one, "cockpit" and "hello" are casualties of the naive version.

What I landed on is boring and I think correct:

```ts
// leo-profanity only matches whole words exactly (no stemming), so inflected
// forms like "fucker" leak. These roots are matched as substrings to catch
// derivations. Curated to avoid false positives: only stems that never occur
// inside clean English words (NOT "ass"/"dick"/"cock"/"hell").
const STRONG_ROOTS = [
  'fuck', 'shit', 'cunt', 'bitch', 'nigg', 'slut',
  'whore', 'pussy', 'bastard', 'motherfuck', 'dumbass', 'jackass',
];
```

Two mechanisms, chosen per word. Roots that can't appear inside innocent English are matched as substrings. Slurs that *can* (and a second list of terms missing from the default dictionary) are added as exact matches instead.

The parenthetical is the important part of that comment. It's not documenting what the code does, it's documenting the rule for editing it. The next person adding a word needs to know which list it belongs in, and the answer is a question they can actually answer: *does this string ever appear inside a clean word?*

I want to be straight about the limits. This catches lazy profanity. It does not catch `f u c k`, or `fµck`, or someone determined. Leetspeak normalization, homoglyph folding, and an ML classifier all exist and all cost either latency, money, or a false-positive budget I didn't want to spend on a student forum where the realistic threat is frustration, not coordinated abuse.

If the community grows into a target, this is the first thing I'd replace. Right now it's proportionate, and I'd rather ship a filter I can explain than one I can only tune.

## Layer 2: the part that judges

Everything that got through goes onto a BullMQ queue with the Layer 1 findings attached. The worker adds the two things the request path couldn't afford: the user's trust score, and their recent posting rate: scores the whole picture, and maps the score to `APPROVE`, `REVIEW`, or `BLOCK`.

<FigureImage src="/blog/moderation/two-layer-flow-light.webp" srcDark="/blog/moderation/two-layer-flow-dark.webp" width={2752} height={1536} alt="Flow diagram. A community write passes through the synchronous gate, which blocks only profanity and malicious domains with a 422, and otherwise attaches its findings to the request and continues. The handler saves the content and enqueues a scoring job, which loads trust and rate, scores, and decides approve, review, or block." caption="Layer 1 answers 'is this certainly unacceptable'. Layer 2 answers 'is this probably a problem', a second later, with more context."/>

The scoring itself is a separate post. What matters here is the plumbing decision on the other side: what do you *do* with a `REVIEW`?

The tempting answer is a bot-review queue in the admin panel. A new page, a new list, new filters, new empty state.

I didn't build any of that, because a queue of flagged content already existed: the one humans fill by pressing "report":

```ts
// Funnels an auto-flagged item into the Report table so it shows up in the
// admin reports view, attributed to the system Moderation Bot. Never fails
// the job: content may have been hard-deleted by the author meanwhile.
```

The bot files a report the same way a user does, under a fixed system user id. Admins see one queue, sorted the same way, with the same actions. The only difference is the reporter's name.

This turned out to be the decision that paid off most in the whole feature, and it took ten minutes. Every improvement to the reports view (filters, bulk actions, deep links) now applies to automated flags for free, permanently, without anyone remembering to make it apply to both.

When an automated process needs a human in the loop, look for the human queue that already exists before building it one. A second inbox is a second thing to check, and the one nobody checks is whichever one is emptier.

## Two swallowed errors, both deliberate

The worker has two `try/catch` blocks that log and continue rather than fail the job. Both look sloppy in review and both are load-bearing.

**Filing the report.** The bot may have already filed on a previous pass: an edit re-triggers scoring, and the reporter table has a uniqueness constraint of one filing per user per item:

```ts
.catch((err) => {
  // P2002: bot already filed on a prior moderation pass; ignore.
  if (!(err instanceof Prisma.PrismaClientKnownRequestError && err.code === 'P2002')) {
    throw err;
  }
});
```

Note it re-throws anything else. Swallowing a specific known-benign error is fine; swallowing every error is how you lose a database outage.

**Hiding the content.** Between the write and the worker running, the author can delete their own post. Then the update targets a row that no longer exists and throws:

```ts
// Soft-hides blocked content. Content may have been hard-deleted by the
// author meanwhile, so swallow errors rather than fail the job.
```

Without the catch, BullMQ retries the job three times, each one failing on a row that will never come back, and the queue accumulates permanent failures for content that resolved itself in the best possible way.

Async moderation means the world changes under you. The content you were asked to judge might be gone, edited, or already handled by a human. Every step has to tolerate that, and "tolerate" mostly means deciding in advance which errors are just the world moving on.

## What I'd change

**Blocked content still isn't reviewable in one click.** A `BLOCK` soft-hides the post, files the report as `removed`, and notifies the author. If the bot was wrong, an admin can undo it, but the author's notification has already gone out. I'd hold that notification behind a short delay, or send it only after the report has aged without an admin touching it.

**One threshold set for all content types.** A comment and a long post get scored identically, which means a two-word reply hits the caps ratio far more easily than an essay does. Per-type thresholds would be a config change, not a rewrite. I just haven't seen it hurt anyone yet.

---

The version of this I'd defend in review isn't the one that catches the most bad content. It's the one where I can point at any block and name the specific rule that fired.

Making the fast layer do less was what bought that. It only ever runs two rules, so when it says no, there's no ambiguity about why, and everything ambiguous got moved to a place where being wrong costs a second look instead of a rejection.

## Related work

- [Gradsy](https://riteshkc.com.np/work/gradsy): A centralized platform for students, counselors, and admins with role-based access - built with Next.js and NestJS, and deployed on AWS using S3, SES, and Lightsail.
