Ein Trust Score, den man aus dem Log neu bauen kann
Gradsy scores community content for harm. Profanity, spam signals, link volume, shouting, posting bursts: each contributes points, and the total decides whether a post is approved, queued for a human, or removed.
Then there's one more term, and it's the one that makes this interesting:
// Higher trust reduces the harm score.
total -= (input.userTrustScore / 100) * w.trustReduction;A user's trust score discounts their content's harm score. And trust itself moves based on past moderation decisions: up a little when content is approved, down when it's flagged or blocked.
So the score depends on trust, and trust depends on previous scores. That's a feedback loop, and every feedback loop eventually needs the same thing: a way to recompute the current state from nothing, because sooner or later the running value will be wrong and you'll need to prove what it should have been.
The scoring function is boring on purpose
const DEFAULT_WEIGHTS: ScoringWeights = {
profanity: 40,
spam: 20,
links: 15,
caps: 10,
repetition: 10,
rate: 25,
trustReduction: 15,
};
// Pure, synchronous. Higher score = more harmful, clamped to [0, 100].
score(input: ScoringInput): number {
const w = this.weights;
let total = 0;
if (input.profanityFlagged) total += w.profanity;
if (input.spamFlagged) total += w.spam;
if (input.linksFlagged) total += w.links;
if (input.repetitionFlagged) total += w.repetition;
if (input.rateFlagged) total += w.rate;
// Caps contributes proportionally to how far over it leaned.
if (input.capsRatio > 0) total += w.caps * Math.min(1, input.capsRatio);
total -= (input.userTrustScore / 100) * w.trustReduction;
return Math.max(0, Math.min(100, Math.round(total)));
}No database calls, no async, no injected clock. Everything it needs arrives as an argument. That's not architectural purity for its own sake. It's the property that makes the whole thing testable without a running Postgres, and it's why I could reason about the clamping problem later on paper instead of in a debugger.
Two details that took a second pass to get right.
Caps is the only proportional signal. Everything else is a boolean that either fires or doesn't. Caps uses the ratio, because "35% uppercase" and "100% uppercase" are genuinely different behaviours and collapsing them to one bit threw away the distinction that mattered. The Math.min(1, ...) is defensive: the ratio should never exceed 1, and if a bug upstream makes it 3, I'd rather cap the contribution than let one signal blow through the whole scale.
Trust is worth 15 points out of 100. That number is a policy decision disguised as a constant. A perfectly trusted user gets a 15-point discount; someone with zero trust gets none. It's a thumb on the scale, not a pardon. Set it to 40 and a long-standing member could post almost anything; set it to 5 and history stops meaning anything. Fifteen means trust can move a borderline post across a threshold but can't rescue content that trips several signals at once.
Then the thresholds:
// score < approve => APPROVE; >= block => BLOCK; otherwise REVIEW
// (REVIEW is the inclusive lower bound at the approve threshold).
decide(score: number): ModerationDecision {
if (score < this.approveThreshold) return ModerationDecision.APPROVE;
if (score >= this.blockThreshold) return ModerationDecision.BLOCK;
return ModerationDecision.REVIEW;
}Below 30, approve. At or above 70, block. In between, a human looks.
The forty-point middle band is the point. An automated system that only decides yes or no has to be right, and this one isn't good enough to be right. It's a weighted sum of heuristics. Giving it a third answer, "I don't know, ask someone," is what lets the thresholds be conservative at both ends instead of split down the middle.
Two ways to know a score
Trust moves after every decision:
private deltaForDecision(decision: ModerationDecision): number {
switch (decision) {
case ModerationDecision.APPROVE: return this.config.get('moderation.trust.cleanPost', 2);
case ModerationDecision.REVIEW: return this.config.get('moderation.trust.review', -5);
case ModerationDecision.BLOCK: return this.config.get('moderation.trust.block', -10);
default: return 0;
}
}Everyone starts at 50. Clean content earns +2. A review costs −5, a block −10. Losses hurt more than wins help, and clean posting takes a while to dig you out: five good posts to undo one block. That asymmetry is deliberate: the cost of a false negative in this system is content nobody catches, and the cost of a false positive is one annoyed student and a review queue.
That's the incremental path: read, add delta, clamp, write.
The other path replays everything:
// Recomputes a score from moderation history: starts at default and folds in
// each logged decision's delta. Clamped to [0, 100].
async recalculate(userId: string): Promise<number> {
const logs = await this.prisma.moderationLog.findMany({
where: { userId },
select: { decision: true },
});
const score = logs.reduce(
(acc, { decision }) => clamp(acc + this.deltaForDecision(decision)),
this.defaultScore,
);
const row = await this.prisma.userTrustScore.upsert({
where: { userId },
create: { userId, score },
update: { score },
});
return row.score;
}Ten lines, one reduce. The moderation log already existed. It's the audit trail for every decision, written whether or not anyone ever reads it. Trust turned out to be a projection of data I was already keeping, which meant the rebuild function was almost free.
This is event sourcing in the only dose I've ever wanted it. No framework, no event store, no aggregate roots. One append-only table that exists for audit reasons, and one function that folds it.
Why you need it: someone will change a weight. A block will be reversed by an admin. A bug will double-apply a delta during a retry. Without a rebuild, every one of those leaves a number in a column that nobody can justify and nobody dares touch. With it, the stored score is a cache and you can always regenerate the truth.
The clamp goes inside the fold
Look at where clamp sits. It wraps each step, not the final total.
That looked like a stylistic choice when I wrote it. It isn't. It changes the answer.
Take a user with six blocks followed by six clean posts.
Clamping per step:
50 → 40 → 30 → 20 → 10 → 0 → 0 (sixth block clamps, -10 becomes -0)
→ 2 → 4 → 6 → 8 → 10 → 12Final: 12.
Clamping only at the end:
50 + (6 × -10) + (6 × +2) = 50 - 60 + 12 = 2Final: 2.
Same events, same order, six-fold difference. The clamp isn't a display concern. It's part of the arithmetic, because it discards magnitude. Once you're at zero, further blocks cost nothing, and that "wasted" negative is exactly what the end-clamped version keeps and re-applies.
Which is right? The per-step version, and not because it's nicer, because it has to match the incremental path. updateScore clamps every single write. If recalculate clamped only at the end, a rebuilt score would silently differ from a live one for any user who ever hit a bound. You'd have two functions claiming to compute the same value, agreeing on most users, disagreeing on precisely the users you most want to be sure about.
If you keep a running value and a rebuild function, they aren't two features. They're one invariant, and it's worth a test that asserts they agree, including a case that saturates a bound, which is the only place they can drift.
There's a consequence I didn't think through at the time. Clamping per step makes the fold order-dependent. Six blocks then six approves gives 12; six approves then six blocks gives 2. Reordering the same events changes the result.
Which is fine: the log is a history, histories have an order, and rebuilding it in order is the correct thing to do.
Except.
The bug I found writing this post
const logs = await this.prisma.moderationLog.findMany({
where: { userId },
select: { decision: true },
});There is no orderBy.
A SELECT without ORDER BY has no guaranteed row order in Postgres. It usually comes back in physical order, which usually resembles insertion order, which is why this has never visibly misbehaved. But "usually" is doing all the work in that sentence: an index-only scan, a plan change after a vacuum, or a parallel sequential scan can each hand back a different order, and any of them would produce a different score for the same user with nothing in the logs to indicate why.
An order-dependent fold over an unordered query. The fix is one line:
orderBy: { createdAt: 'asc' },I did not find this by testing. I found it writing the paragraph above, where I typed "rebuilding it in order is the correct thing to do" and then went to check that it actually was.
That's the second time explaining code to an imagined reader has caught something reading the code didn't. I don't have a clean theory for why, beyond this: reviewing code asks "is this right?", which your brain answers by pattern-matching. Explaining it asks "why is this right?", which it can only answer by actually deriving it.
One weight that barely fires
While I'm being honest about things I noticed too late: profanity is the heaviest signal in the table at 40 points, and as far as I can tell it almost never fires.
The synchronous gate in front of these routes already hard-blocks profanity with a 422 before the content is saved. Every job that reaches the scorer came through that gate: I traced all five enqueue sites, which means profanityFlagged is false by construction on the path that matters.
It's not dead code. It's the correct weight for a signal that a different layer currently intercepts, and if the gate ever loosens, it's already right. But someone tuning these numbers should know that raising or lowering 40 will change nothing about how the system behaves today, and I'd rather say that than let them spend an afternoon on it.
What I'd add next
A reason column on trust changes. Right now the score moves and the log records the decision, but reconstructing why a specific user is at 18 means reading their whole history and doing the arithmetic by hand. Storing the delta and the resulting score per event would make the rebuild verifiable instead of merely repeatable.
A bound on the replay. findMany loads every log row for the user. Fine at current volumes, unbounded in principle. The usual fix is a periodic checkpoint: store a score plus a watermark, replay only what came after. That reintroduces exactly the drift problem the rebuild exists to solve, so I'd want the agreement test in place first.
Admin overrides don't participate. When an admin reverses a bot decision, the trust delta is applied but the original log row stays as it was. A rebuild replays the bot's original call, not the human's correction. That's arguably a bug and definitely a surprise, and it's the next thing I'd fix.
Derived state is a promise you make about a number. The running value is the convenient version of that promise and the rebuild is the honest one, and they only stay the same promise if you're careful about the places where arithmetic stops being arithmetic: every clamp, every floor, every saturating bound.
Those are the points where "recompute it from the log" quietly becomes "recompute it the same way from the log," which is a much stronger requirement than it looks like when you're writing the convenient version first.