Cache-Invalidierung, wenn der Cache in einem anderen Repo liegt
An admin edits a university's tuition range, saves, and sees a success toast. They open the public page in another tab. The old number is still sitting there.
They refresh. Still old. They refresh again, harder, the way people do. Then they message me.
Nothing was broken. The public page at /university/[slug] is built once and then served from that build for a day: revalidate = 86400. The copy they were staring at had been rendered that morning, and Next.js would not build a new one until 24 hours were up. It was doing what I told it to. What I told it had a 24-hour blast radius, and I had built an admin panel next to it that implies edits show up immediately.
No timer was going to be right
First instinct: lower the window. A day is too long. An hour?
An hour is a smaller wrong number. It is still fifty-nine minutes of an admin not trusting the tool. And every reduction costs regenerations on pages nobody edited. University pages change a few times a month, so a one-hour window means roughly seven hundred pointless renders a month for every page, to catch a handful of real edits.
That ratio is the argument against using a timer at all. The window I want is zero right after a write and infinity the rest of the time, and no single number does that. Time is standing in for the thing I actually care about, which is whether someone changed the data.
Next.js has the right primitive for this: revalidateTag. Label a fetch with a string, tell Next that string is stale, and the page rebuilds on the next request. Straightforward.
Except the write doesn't happen in the app that owns the cache.
The write happens in another process
Gradsy is two repos. The admin UI is Next.js. The mutation is a PATCH to a NestJS API in a different process, on a different host, with its own database.
So the sequence is: the browser calls the backend, the backend writes to Postgres, and the frontend is never told. The frontend is the thing holding the stale prerendered page, and from where it sits, nothing happened at all.
There are three ways out. Only one of them survives contact with the rest of the system.
Have the frontend call revalidateTag after its own mutation succeeds. Tempting, because it keeps everything in one repo. It breaks the moment anything writes without going through that UI: the CSV importer, a script, a background job, a second client. The cache would be correct only for writes that started in one place.
Poll. Some process asks the backend "changed since?" every N seconds. That adds a second staleness window and a job to operate.
Let the writer announce the write. The backend knows which university changed and when, because it is the thing that changed it. It just has to say so.
So: a webhook, pointing backwards. The frontend usually calls the backend. Here the backend calls the frontend, and it is the only route in the app that works that way.

The receiving end
// Shared with the backend (REVALIDATE_SECRET there too). Unlike the set-title
// webhook this is mandatory: an open endpoint lets anyone dump the cache.
const SECRET = process.env.REVALIDATE_SECRET;
export const dynamic = "force-dynamic";
export async function POST(req: NextRequest) {
if (!SECRET) {
return NextResponse.json(
{ error: "REVALIDATE_SECRET is not configured" },
{ status: 500 },
);
}
if (req.headers.get("x-webhook-secret") !== SECRET) {
return NextResponse.json({ error: "Unauthorized" }, { status: 401 });
}
const body = await req.json().catch(() => ({}));
const tags = Array.isArray(body.tags) ? body.tags.filter(isNonEmpty) : [];
const paths = Array.isArray(body.paths) ? body.paths.filter(isNonEmpty) : [];
if (tags.length === 0 && paths.length === 0) {
return NextResponse.json(
{ error: "Provide at least one of `tags` or `paths`" },
{ status: 400 },
);
}
for (const tag of tags) revalidateTag(tag);
for (const path of paths) revalidatePath(path);
return NextResponse.json({ ok: true, tags, paths });
}A missing secret returns 500 instead of skipping the check. That is deliberate.
There is another webhook in this app, one that sets a dashboard title, where the secret is optional: if it isn't configured, the check is skipped. That's fine there. The worst case is that someone changes a string.
This one fails closed. An unauthenticated revalidation endpoint isn't a data leak. It is a free denial-of-service: anyone who finds the URL can bust every tag in a loop and force a regeneration storm against the origin. "Secret not set" has to mean "refuse," not "allow everyone."
Tags are a vocabulary, and my first version was wrong
My first pass tagged each fetch with its own URL. university:harvard, university:mit, one tag per page.
Then someone edited a program — a degree attached to forty universities — and I had forty tags to bust and no list of which forty without querying for it.
Tags are not page identifiers. They are reasons a page might be wrong. Once I named them that way, the grouping fell out:
// The backend busts these tags on every university/program mutation via
// POST /api/revalidate, so the timer is only a backstop for a webhook that
// never landed. Keep the tag names in sync with the backend payload.
const UNIVERSITY_TTL = 86400;
function universityCache(tags: string[]) {
return { next: { revalidate: UNIVERSITY_TTL, tags } };
}Every university fetch carries universities. The detail fetch carries university:${slug} on top of that. The facet endpoints carry university-countries or university-intakes. The program list carries programs.
A single university edit busts universities and that one slug. A program edit busts universities and nothing else. That is coarse, and correct, and one line:
// Programs render inside every linked university page, so there is no
// single slug to target: bust the whole collection.
this.revalidation.revalidateUniversityCollection();I'd rather regenerate more pages than maintain a dependency graph I have to keep accurate. The graph would be more precise, and it would go wrong eventually, silently, in the direction of serving stale data. Over-invalidating goes wrong in the direction of a few extra renders.
Three call shapes, and the one I missed
Single-entity edits were obvious. Two others weren't.
The CSV import. Universities get bulk-loaded from a scraper, hundreds of rows at a time. The first version fired a webhook per row: a few hundred POSTs and a few hundred regeneration triggers for one logical operation. Now the import collects slugs and fires once at the end:
// One webhook for the whole import instead of one per row.
this.revalidation.revalidateUniversities(touchedSlugs);Renames. This is the one I missed, and it's the one worth remembering:
// A rename mints a new slug, so the old URL has to stop serving the
// prerendered page too.
this.revalidation.revalidateUniversities([slug, existing.slug]);When a university's slug changes, the new URL has no cache entry, so it renders fresh. Fine. The old URL still has a good prerendered page sitting in the cache, and it keeps serving it. The record moved and nothing told its old address.
Busting a tag for the new slug does nothing about that. You have to bust both.
Any time an identifier can change, invalidation takes two arguments: where the record went and where it was. I've been bitten by this in a router, in a CDN, and in Next's route cache. Same shape every time.
Failures are swallowed on purpose
/**
* Fire-and-forget: a cold cache must never fail an admin write, so failures
* are logged and swallowed. Callers do not await.
*/Callers don't await. The service returns void. If the frontend is down, mid-deploy, or slow, the admin's save still succeeds. They get their toast, the row is in Postgres, and the page catches up on the 24-hour revalidate timer instead.
This is where the 24-hour window earns its keep. It stopped being a freshness strategy and became a failure backstop: if the webhook never fires, the page is stale for up to a day instead of forever. Same line of config, different job.
The tradeoff isn't free. A dropped webhook is invisible. It logs a warning on a box nobody watches, and the admin never learns their edit didn't propagate. If staleness cost money here, I'd want a retry and an alert. For a university page showing tuition ranges, the worst case is a day-old number and a manual re-save.
The stale comment
Read that cache comment again:
the timer is only a backstop for a webhook that never landed
The webhook landed. It has been running for months. Two services in the backend call it, from fifteen different sites.
I wrote that comment during the window where the frontend side existed and the backend side didn't, copied it verbatim into a second file, and never went back. Two repos, one person, and the code drifted from its own documentation inside a week: no reviewer, no handoff, and nobody to blame for not reading the other repo.
The tag names are a contract. The backend sends ["universities", "university:harvard"], and the frontend has to have tagged its fetches with those exact strings. Nothing checks that. Not TypeScript, not a test, not a build step. Rename a tag on one side and the request still returns 200 OK with { ok: true }, because revalidateTag on a tag nobody uses is a legal no-op.
You get a successful webhook, a green log line, and a page that never updates.
The comment above the TTL says "keep the tag names in sync with the backend payload." That is a comment asked to do a type system's job. It works until the person reading it is in a hurry, and the person in a hurry was me.
What I'd do differently
Share the tag builders. A small package, or even a generated file, exporting universityTags(slug) and used by both repos. A rename becomes a compile error instead of a silent no-op. I skipped it because a shared package for one function felt like ceremony on a solo project. I was wrong about which of those was cheaper.
Return what actually got busted. The route already echoes { ok: true, tags, paths }. If the backend logged a warning when it sent a tag the frontend didn't recognise, drift would surface in minutes instead of months. That needs the frontend to know its own tag vocabulary, which loops back to the shared package.
Keep fire-and-forget. Coupling an admin's save to the availability of a different service, to save a few hours of staleness, is a bad trade.
Cache invalidation is one of the two hard problems. Doing it across a process boundary adds a third. The tag names become a contract between two services, and unlike the API between them it has no schema, no types, and no failing test: two files in two repos that have to keep agreeing with each other, and nothing that tells you when they stop.
Mine stopped agreeing with its own comment in under a week. The code kept working. The comment didn't.
