Search

↑↓ selectEnter openAdvanced search

Guides

Fast and still fresh: edge caching EmDash on Cloudflare

Our public pages sit in Cloudflare's edge cache for five minutes, yet a published fix shows up right away, because EmDash purges by cache tag when you publish, and editors keep their Edit button. The five minutes are just the safety net for a purge that doesn't happen. How much faster the pages get, I'll measure after launch.

By Published 12 min readchecked with 1.1 · current 1.2

A blog is read far more often than it's written. Every anonymous visitor to theweeklydash.com gets the same page, so there's no reason for the Worker to render it again and ask D1 for the same rows each time. Cloudflare can keep it at the edge, in the data center closest to the reader.

Anyone who has put a cache in front of a CMS knows the catch. You fix a mistake in an article, and readers keep seeing the old version. Fast and fresh pull in opposite directions. So we hold ourselves to three rules: a published fix is visible immediately, editors keep their Edit button on cached pages, and a request for /wp-login.php doesn't cost a database query.

Here's how theweeklydash.com pulls that off, with the actual code from our repository. First we'll sort out what's allowed into the cache at all, then take the three rules one at a time. I checked everything against EmDash 1.1.0, Astro 7.3 and @astrojs/cloudflare 14.3.

What goes to the edge, and what doesn't

The cache is Cloudflare's Workers Cache. It sits in front of the Worker, and when a matching response is already there, the Worker doesn't run at all. You decide in Astro which pages get cached and for how long, using route rules. The cacheCloudflare() provider turns them into headers Cloudflare understands:

// astro.config.mjs (excerpt)
cache: { provider: cacheCloudflare() },
routeRules: {
	"/": { maxAge: 300, swr: 86400 },
	"/en": { maxAge: 300, swr: 86400 },
	"/archiv": { maxAge: 300, swr: 86400 },
	// Posts and pages.
	"/[slug]": { maxAge: 300, swr: 86400 },
	"/en/[slug]": { maxAge: 300, swr: 86400 },
	// Rules match paths, not routes, so the catch-all rules above would also cover these.
	// A rule with tags only and no maxAge keeps them out of the edge cache.
	"/search": { tags: [] },
	"/en/search": { tags: [] },
	"/robots.txt": { tags: [] },
	"/404": { tags: [] },
	"/rubrik/[slug]": { maxAge: 300, swr: 86400 },
	// … topics, authors and the English listings alike
	"/_image": { maxAge: 300, swr: 2592000 },
	"/og/[locale]/[slug].png": { maxAge: 300, swr: 2592000 },
	"/rss.xml": { maxAge: 300, swr: 86400 },
	"/sitemap.xml": { maxAge: 3600, swr: 86400 },
},

Home pages, archive, articles, sections, topics and author pages are fresh for five minutes (maxAge: 300). After that, the edge keeps serving the old copy for up to a day while it re-renders in the background (swr: 86400, stale-while-revalidate). Of the two numbers, I'd say swr matters more. The cache does have two tiers: if the data center near the reader doesn't have the page, an upper-tier one fills in. But on a small site, a single article can go hours without a single request. With a short swr, nearly every first reader would sit through a full render.

Images and link-preview images get 30 days of swr, because rendering them again is expensive. Feeds follow the pages, and the sitemaps get by with an hour and no purge at all.

The comment lines in the middle mark the trap we nearly walked into. Route rules match paths, not routes, so /[slug] also catches /search and /robots.txt. A rule that sets only tags and no maxAge takes them back out.

A few things stay out on purpose:

  • Search. Every query is its own URL. Storing a thousand variants that almost never come up twice isn't worth it to me.
  • Admin and API. EmDash sends private, no-store there itself, and we have no route rule for /_emdash/.
  • Preview and edit views. Behind ?_preview= are unpublished drafts, behind ?_edit= the editor mode. Neither has any business in a shared cache.
  • Anything that isn't GET or HEAD, plus redirects and error pages. The exception is 404s, more on those below.

Cache tags: getting a fix out right away

Five minutes sounds short until you spot a typo in a headline. So every response carries cache tags, and when you publish, EmDash purges exactly the tags involved.

The tags come from your queries. Besides the data, getEmDashEntry() and getEmDashCollection() return a cacheHint, which you pass to Astro.cache.set() in the page's frontmatter:

---
// src/pages/[slug].astro (excerpt)
const found = await findEntry("de", Astro.params.slug);
if (!found) return Astro.rewrite("/404");
if (Astro.cache?.enabled) Astro.cache.set(found.result.cacheHint);
---

A single entry is tagged with its ULID, a list with the collection name, posts. Where you make the call matters: in the page's frontmatter, not in a layout or component, because those render inside the stream, after Astro has already written the headers. A cache.set() there is lost after the first await, and the page doesn't get cached at all.

Navigation, settings and sections show up on every page, so a menu change has to clear every page. EmDash has cache-hint variants for these, and our middleware registers them before every cached render:

// src/middleware.ts (excerpt)
const dependencies = await Promise.all([
	getSiteSettingsWithCacheHint(),
	getMenuWithCacheHint("primary", { locale }),
	getTaxonomyTermsWithCacheHint("category", { locale, includeCounts: false }),
	getTaxonomyTermsWithCacheHint("tag", { locale, includeCounts: false }),
]);
for (const dependency of dependencies) context.cache.set(dependency.cacheHint);

The plain getMenu() and getSiteSettings() return no hint. Use them on cached pages and a menu change only shows up once the TTL runs out.

Purging happens in the API routes the admin calls: publish, unpublish, schedule, delete, restore and discard draft. Menus, settings, taxonomies and widget areas purge their own tags. The provider calls cache.purge() from cloudflare:workers and needs neither a zone ID nor an API token.

Autosave on a published post doesn't purge, and that's deliberate. The change goes into the draft, and nothing public has changed yet. Publish changes is what triggers the purge. SQL straight against D1 purges nothing at all, which is why we only ever change content in the admin.

With the purge doing all that, the pages could just as well stay cached for a whole day. I still stick with five minutes, because a purge can fail. Workers Cache belongs to the Worker, not the zone, so it always gets the Free plan purge limits: five requests a minute, with a burst of 25. A rejected purge is reported in the return value, and Astro's provider doesn't check it. For a two-person team that limit is plenty. If something does go wrong anyway, maxAge caps the damage at five minutes, plus one request that still gets the old copy and kicks off the re-render.

Two headers: one for the edge, one for the browser

The purge only reaches the edge, though, not the browser. That's why our pages send two different instructions:

Cloudflare-CDN-Cache-Control: public, max-age=300, stale-while-revalidate=86400
Cache-Control: public, max-age=0, must-revalidate

Only Cloudflare reads the first one, and it never makes it to the browser. The browser reads the second. It may keep the page, but has to check back on every visit. If the browser cached the HTML for five minutes on its own, no purge could reach it. And after a deploy, the old page would point at CSS files that no longer exist.

We started out sending browsers no-store. That was safe, but it had two downsides. Pages weren't eligible for the back/forward cache, and Astro's link prefetching went to waste, because a browser isn't allowed to hold on to a no-store response until the click. max-age=0, must-revalidate allows both. And EmDash folds the build time into Last-Modified, so a deploy never gets a stale 304.

The headers are set in a middleware of their own. A response that isn't the same for everyone, or whose status isn't 200, gets no-store and stays out of the cache. Regular HTML pages get the browser version from above:

// src/cache-response.ts (excerpt)
const personalized = !!context.locals.user;
const shared = !privateView && !personalized && ["GET", "HEAD"].includes(context.request.method);

if (shared && response.status === 404 && CATCH_ALL_ROUTES.has(context.routePattern)) {
	// A slug with nothing published under it, see below.
} else if (!shared || response.status !== 200) {
	context.cache.set(false);
	response.headers.set("Cache-Control", personalized ? "private, no-store" : "no-store");
} else if (response.headers.get("Content-Type")?.startsWith("text/html") && !response.headers.has("Cache-Control")) {
	response.headers.set("Cache-Control", "public, max-age=0, must-revalidate");
}

This middleware has to run after EmDash, since EmDash still touches the HTML stream, for the toolbar among other things. That's why we register it as middleware.outer in the EmDash integration. Sounds backwards, since the outer middleware starts before EmDash, but it gets the finished response back from await next().

Editors keep their Edit button

That takes care of readers. The second rule is about the people who write the articles. Workers Cache runs before the Worker and doesn't look at cookies, so a logged-in editor gets the same cached page as everyone else. The EmDash docs are upfront about it: with the default toolbar, the editor's toolbar is missing whenever an anonymous visitor filled the cache.

The fix is a single line in astro.config.mjs, toolbar: "client":

emdash({
	toolbar: "client",
	middleware: { outer: "./src/outer-middleware.ts" },
	// …
}),

In client mode, every visitor gets the same HTML. A small script checks localStorage for a flag the admin sets once someone has signed in on that browser. If it finds one, the page shows an Edit button at the bottom. Clicking it asks the server whether the session is still valid and then reloads the page with ?_edit=1. EmDash always renders that URL fresh, with the full toolbar and private, no-store. All it costs anonymous readers is one localStorage read.

It works the other way round too: nothing personal goes into the cache. When the Worker renders a page for a signed-in user, the comment form carries their name and email. Those responses get private, no-store and cache.set(false). I checked this on the live site with a fresh, never-requested URL and an admin session: three requests in a row, all cf-cache-status: BYPASS. A page that's already cached, on the other hand, comes back as a HIT even when you're signed in.

404s without the database

That leaves the third rule, which has nothing to do with readers at all. Every public site gets bots probing for /wp-login.php, /.env or /xmlrpc.php. Our articles live right at the root, at /<slug> and /en/<slug>. Without precautions, every one of those probes would boot EmDash and ask D1 for a post that can't exist.

So the outer middleware checks the slug before EmDash even starts. A slug no entry could ever have gets a bare 404 straight away. That's any slug that doesn't look like what EmDash generates (a dot, an uppercase letter, a double hyphen, more than 120 characters), or one a route already owns, like archive or rss.xml.

// src/outer-middleware.ts (excerpt)
if (CATCH_ALL_ROUTES.has(context.routePattern)) {
	const slug = decodeSlugParam(context.params.slug);
	if (!slug || slugIssue(slug)) return invalidSlug(context);
}

function invalidSlug(context: APIContext): Response {
	if (context.cache?.enabled) {
		context.cache.set(false);
		context.cache.set({ maxAge: INVALID_SLUG_MAX_AGE }); // 3600
	}
	// …
	return new Response(notFoundPage(locale), { status: 404, headers });
}

This 404 does without a layout, because the menu and settings would need the database. It stays cached for an hour. Only new code can change the answer, and every deploy starts with an empty cache anyway. On the live site, /wp-login.php returns a 404 with cf-cache-status: HIT. The bot gets turned away right at the edge, and the Worker never even notices.

Bare English 404 page in dark mode: This page doesn't exist, with links To the current issue and To the archive

A slug that looks valid but has nothing published under it, say /en/edge-caching-part-2, is a different story. This one is back to the first rule, because an article could go live at exactly that address tomorrow. EmDash queries the database and renders the regular 404 page. That page is cached for just 60 seconds and tagged posts and pages:

// src/cache-response.ts (excerpt)
if (shared && response.status === 404 && CATCH_ALL_ROUTES.has(context.routePattern)) {
	context.cache.set(false);
	context.cache.set({ maxAge: MISSING_ENTRY_MAX_AGE, tags: ["posts", "pages"] }); // 60
	response.headers.set("Cache-Control", "public, max-age=0, must-revalidate");
}

When something does get published under that slug, EmDash purges the posts tag and the 404 is gone. The 60 seconds only cap the case where that purge doesn't happen. To make sure nobody publishes an article as archive and hides the route, a small site plugin refuses to publish reserved or malformed slugs.

Scheduled posts: publishing without a click

So far, every purge has hung off a click in the admin. Scheduled posts go live without one. On Cloudflare, the Cron Trigger publishes them, and ours runs every five minutes (*/5 * * * *), so a post scheduled for 9:00 goes live by 9:05. That delay comes from the cron, not the cache.

According to the source, the cron purges afterwards too. The scheduled handler in @emdash-cms/cloudflare 1.1.0 gets the cache provider through Astro's manifest. After each collection it published something in, it purges the collection's tag and the IDs of the new entries. Those are the same tags you purge when you click Publish. Our src/worker.ts takes that handler over unchanged with createScheduledHandler(). Write your own, and that purge is gone until you rebuild it.

I haven't watched it happen on the live site yet. If that purge fails, the safety nets from above kick in. Home page and listings pick the post up within five minutes plus one request, and a cached 404 for its URL lasts a minute at most. The headers will tell you whether it all works.

How to check it

On the live site, cf-cache-status tells you what happened. Run this twice in a row:

curl -sI https://theweeklydash.com/en | grep -i -E 'cf-cache-status|^age|cache-control'

The first request after a deploy or purge shows MISS, the second HIT along with an age header. Once the five minutes are up, you'll see UPDATING or STALE, which is stale-while-revalidate doing its job. BYPASS means the Worker ran and the response wasn't allowed into the cache.

You won't see Cloudflare-CDN-Cache-Control on the live site, because Cloudflare strips it before delivery. So the real check runs locally against the built Worker:

pnpm build
pnpm exec wrangler dev --config dist/server/wrangler.json --port 4322
node scripts/check-cache.mjs http://localhost:4322

scripts/check-cache.mjs checks headers, tags and the client toolbar on 14 public pages in both languages. It also goes through everything that has to stay uncached, meaning search, admin, preview, the edit URL and POST, plus both kinds of 404, the old redirects, feeds, sitemaps and preview images. Add a public page and forget its route rule, and this is where you find out, not when the page starts dragging.

I'll add measured hit and miss figures from the live site after launch.

Whatever those numbers turn out to be, one thing already holds: Workers Cache includes the Worker version in its cache key. Every deploy starts with an empty cache, and the first readers pay for a full render, D1 queries included. How fast a miss is still depends on where your database lives.

Ours is in Western Europe, with read replication turned on. With d1({ session: "auto" }), EmDash reads from the nearest copy for anonymous visitors, so a reader overseas doesn't wait on the trip to Europe for every query. Why the location matters so much, and what to do when it's wrong, I cover in the article on D1 regions.

Conclusion

Edge caching EmDash on Cloudflare takes a few lines of config. The work is in the exceptions. For a fix to show up right away, cache hints go in the page's frontmatter, menus and settings need the WithCacheHint variants or nothing gets purged, and the browser gets a different header than the edge. Since route rules match paths, search and friends have to be taken out explicitly. And on a cached site, toolbar: "client" isn't optional, or the Edit button comes and goes depending on who got there first. For the bots, a look at the slug is enough.

Which brings us back to the title. For us, the five minutes at the edge aren't how long a fix takes to show up; the purge on publish takes care of that. They're just the ceiling for when a purge doesn't happen.

About the author

Daniel Müller

Daniel MüllerEmDash maintainer

Daniel Müller builds websites that get by without him. Sounds like bad business, turns out it's his best pitch. And if something does come up, he answers the phone himself, no hold music. Eisbachcode in Munich, maintainer at EmDash.

Comments

No comments yet

First-time comments appear once we've approved them. How we handle your details