Rate limiting a provider without a rate limiter
Amazon SES gives you a sending quota measured per second. Go over it and you don't get a queue. You get Throttling errors, one per rejected message, and the messages are gone.
Gradsy sends broadcasts. An admin picks a group of students, writes a message, hits send, and somewhere between fifty and a few thousand emails need to go out. Every one of them is also an in-app notification and a websocket event, so the send is already fanned out into a job per recipient on a BullMQ queue.
Which meant, on the first version, a few thousand jobs hitting SES as fast as the workers could drain them.
Where I started, and why I stopped
I reached for a rate limiter. That's the named solution to the named problem. bottleneck, or a token bucket in Redis, or BullMQ's own limiter option: pick one, wrap the send, done.
I got about as far as reading the docs before the shape of it started bothering me.
A token bucket is a concurrency control. It exists because you have N things happening at once and need to hold them to a rate. But look at what's actually true here: the jobs are already serialized by a queue. There is exactly one thing happening at a time if I want there to be. I'd be adding a mechanism to control parallelism I hadn't introduced yet.
The queue already has the knob:
this.emailSendDelayMs = this.config.get<number>('ses.sendDelayMs', 200);// Pause after each email the worker actually sends, to stay under the SES
// per-second send rate on large broadcasts. Worker concurrency is 1, so
// this deterministically spaces sends. ~200ms ≈ 5/sec.Concurrency of one, plus a sleep after each send. That's the whole rate limiter.
The word doing the work in that comment is deterministically. With one worker and a fixed pause, the interval between sends has a floor you can compute on paper. There's no bucket to refill, no burst allowance to reason about, no window boundary where a hundred tokens become available at once. Sends are spaced by construction.
A token bucket sized for five per second will, correctly and by design, let you fire five in the same millisecond. That's usually the feature. Against a provider that measures per-second, it's the exact thing you were trying to avoid.
Concurrency is a rate limiter you already have
The general form, which I now reach for before reaching for a library:
A queue with concurrency
Cand a post-task delayDsends at mostC / Dper unit time.
Set C = 1 and you're just choosing D. Set D = 0 and you're just choosing C, bounded by how fast the task runs. Between them you can hit most rate targets without introducing a third mechanism.
This isn't better than a token bucket at everything. It's worse at bursts: a bucket lets you use your full quota when you have headroom, and this doesn't, so a fifty-recipient broadcast that could finish in a second takes ten. I don't care. Nobody is watching a progress bar for a broadcast, and "finishes slower than strictly necessary" is a category of problem I'll take over "loses mail."
What it is better at is being obvious. There is no state. Two lines of code, and the second one explains itself:
// Space out real SES sends (worker concurrency is 1) so a large broadcast
// stays under the per-second send rate. Skipped emails don't wait.
if (emailOutcome !== 'SKIPPED' && this.emailSendDelayMs > 0) {
await new Promise((resolve) => setTimeout(resolve, this.emailSendDelayMs));
}Only sleep for sends that happened
That condition is the part I got wrong first, and it's the part that makes the mechanism actually work.
Not every job sends an email. A notification carries an email policy, and the most common one is IF_OFFLINE: don't email someone who's currently staring at the app, they already got the toast:
private async maybeEmail(...): Promise<EmailOutcome> {
if (emailPolicy === 'NEVER') return 'SKIPPED';
if (emailPolicy === 'IF_OFFLINE' && (await this.redis.isOnline(userId))) {
return 'SKIPPED';
}
// ... look up address, send
}So the worker returns SENT, FAILED, or SKIPPED, and only the first two wait.
The version that slept unconditionally was technically safe. It can only under-use the quota, never over-use it, but it made the delay's meaning wrong. Sleeping after a skip isn't rate limiting, it's just being slow. A broadcast to a mostly-online group would crawl for no reason at all, and a future reader trying to work out why 800 recipients took three minutes would find a sleep with no corresponding network call.
The rule I'd extract: pace the thing you're pacing, not the loop it lives in. Attach the delay to the side effect, not the iteration.
(FAILED waits too, which is deliberate. A failed send still consumed an API call, and a failure under throttling is exactly when you least want to immediately try again.)
The hole, which is in the code in writing
There are two queues that send email. Notifications is one. The newsletter is the other, and it made the same choice:
// Same knob as the notifications worker. Both run at concurrency 1, so a
// simultaneous broadcast and newsletter send doubles the real SES rate.
// lower this if that ever trips the send quota.
this.sendDelayMs = this.config.get<number>('ses.sendDelayMs', 200);Two workers, each correctly limiting itself to five per second, running at the same time. Ten per second at the provider.
The mechanism is local. A token bucket in Redis, shared by both queues, would be global, and this is the case where the thing I argued against is straightforwardly the better tool. Concurrency-as-rate- limit works precisely because there's one consumer, and the moment there are two it silently stops being what it claims to be.
It's still there. It hasn't tripped the quota, because a newsletter and an admin broadcast going out in the same minute hasn't happened yet. When it does, the fix is either a shared limiter or merging both onto one queue, and I'd probably merge: one email queue with a job type is less machinery than one queue each plus a coordinator.
I'm writing this down partly because I think the comment is the right call. Not "TODO: fix", not silence. A statement of exactly when the design breaks and what to reach for. When it does break at 2am, the person reading that line has the whole answer, and there's a decent chance the person is me having forgotten.
An assumption a design depends on should be written where the design lives. "Concurrency is 1" is load-bearing here, and nothing enforces it. Someone can change it in config and quietly triple the send rate with no error anywhere.
Retries would have doubled sends, except for one where
Rate limiting isn't the only thing a queue changes about email. BullMQ retries failed jobs three times with exponential backoff, which is great for transient SES failures and terrible if "failure" means "the send worked and the database write didn't."
The newsletter worker handles it with a predicate rather than a flag:
// The `status: PENDING` predicate makes a retried job idempotent: a
// recipient already marked SENT is never rewritten.
await this.prisma.newsletterRecipient.updateMany({
where: { id: recipientId, status: NewsletterDelivery.PENDING },
data: { status: outcome.status, error: outcome.error ?? null, processedAt: new Date() },
});updateMany with a status predicate, so a retry that finds the row already SENT matches zero rows and changes nothing. No read-then-write, no race between checking and updating. The condition is evaluated by the database as part of the write.
The same file has one more thing worth stealing, which is that the worker re-checks subscription status at send time rather than trusting the snapshot taken when the campaign was queued:
// Someone can unsubscribe between the snapshot and their turn in the queue.A five-thousand-recipient send takes a while at five per second. Someone unsubscribing eight minutes in, at position 3,000, must not receive it, and the list you queued eight minutes ago says they should. Anything queued in bulk needs to ask "is this still true?" at execution time, not just at enqueue time.
Would I do it again
For this, yes. One consumer, one provider, a rate that has to be respected rather than optimized: concurrency and a sleep is proportionate and it has no moving parts.
The moment there's a second consumer, no. And I'd argue the failure I actually made wasn't picking the simple mechanism, it was building the second queue without going back and reconsidering the first one's assumption. The design was right when I wrote it and became wrong when I wrote something else, which is the ordinary way designs go wrong.
Most rate limiting problems are really parallelism problems wearing a costume. Before you add a limiter, look at how many things can call the API at once, if the answer is "one, because a queue says so," you already have a limiter and it's a setTimeout.
Just write down that the answer is one. That's the part that expires.