Our services

If you can think it, we can make it brainsoft.

Caching the expensive and stable parts of a request

Written By: BrainSoft In Backend

Most slow endpoints aren't slow everywhere. They're slow in one or two spots: a report query that scans a million rows, a call to a pricing service, a rendered fragment that takes 300ms to assemble. The rest of the request is cheap. If you cache the whole response you inherit a mess of auth headers, per-user data and stale personalisation. If you cache nothing you pay that 300ms on every hit.

Cache the expensive, stable parts individually. Identify the costly computation, give it a key derived only from the inputs it actually depends on, store the result with a TTL, and invalidate explicitly when the underlying data changes. The cheap parts stay live and correct.

Find the part worth caching

Before touching a cache, measure. Wrap the suspicious calls in timers and log them for a day. You're looking for two properties at once: high cost and low change rate. A query that takes 800ms but returns different results every second is a bad candidate. A query that takes 40ms and returns the same thing for a week is also not worth the complexity.

The sweet spot is something like a category tree, a currency rate table, a permissions matrix, or an aggregate over historical data. These are expensive to build and change on a human timescale, not a request timescale.

  • Cost: how long does it take, and how often is it called per minute?
  • Stability: what events actually change the result?
  • Cardinality: how many distinct keys will exist? A thousand is fine, a million per day is not.
  • Blast radius: if the cached value is wrong, what breaks?

Key it on inputs, not on the request

This is where most caching bugs come from. If you key on the URL, you've keyed on query strings, tracking parameters, and headers that have nothing to do with the computation. The cache hit rate collapses and you cache ten thousand near-identical entries.

Key on the actual inputs to the function. If the expensive call is buildCategoryTree(tenantId, locale), the key is exactly those two values, nothing else. Not the user, not the session, not the request id.

function cacheKey(tenantId, locale) {
  return `cat-tree:v3:${tenantId}:${locale}`;
}

async function getCategoryTree(tenantId, locale) {
  const key = cacheKey(tenantId, locale);
  const hit = await redis.get(key);
  if (hit) return JSON.parse(hit);

  const tree = await buildCategoryTree(tenantId, locale);
  await redis.set(key, JSON.stringify(tree), 'EX', 3600);
  return tree;
}

Include a version segment in the key. When the shape of the cached value changes after a deploy, bump v3 to v4 and the old entries become unreachable garbage that expires on its own. No migration script, no flush.

TTL as a safety net, invalidation as the real mechanism

A TTL is not a correctness strategy. It's a bound on how long a mistake can survive. Set it based on how stale the value is allowed to be if nothing else works, not on how often it changes. If a category tree changes twice a day, a one-hour TTL is fine as a backstop.

Then delete the key explicitly when the data changes. The write path already knows: whoever updates a category calls the invalidation. Keep that in one place so it's hard to forget.

async function updateCategory(tenantId, categoryId, patch) {
  await db.categories.update({ tenantId, categoryId }, patch);
  const locales = await db.locales.forTenant(tenantId);
  const keys = locales.map(l => cacheKey(tenantId, l.code));
  if (keys.length) await redis.del(...keys);
}

Delete rather than update in place. Writing a new value into the cache means you have to construct it identically to the read path, and any drift between the two shows up as a bug six weeks later. Deleting is one line and always correct.

One more thing: guard against the stampede. When a hot key expires, ten concurrent requests will all miss and all rebuild it. A short lock, or a slightly stale serve, avoids hammering the database. If the value is a category tree, serving the old one for another two seconds while one worker refreshes is invisible to users. This is the kind of thing we build into most read-heavy services, and it's a common topic when clients get in touch about performance work.

What not to cache

Anything user-specific that changes with the session, anything with a permission check baked into the value, and anything where being wrong for a minute causes real harm. If you cache a permissions matrix, cache the raw matrix keyed by role, not the resolved answer keyed by user. The resolution step is cheap; the matrix lookup is the expensive part.

Also skip caching things you haven't measured. A cache adds a dependency, an invalidation path, and a class of bugs where the app is correct but the data is old. That's worth it for an 800ms query. It isn't worth it for a 15ms one.

Frequently asked questions

Should I cache the whole HTTP response instead?

Only if the response is genuinely identical for every caller. The moment auth, personalisation or per-tenant data enters the response, whole-response caching forces you to key on all of it, which usually means no cache hits. Caching the expensive inner call keeps the cheap, variable parts live.

How do I choose a TTL?

Pick it as a maximum acceptable staleness if invalidation fails, not as a guess at the change rate. If you invalidate explicitly on writes, the TTL only matters when something goes wrong, so an hour or a day is usually fine.

What happens when the cache is down?

The read should fall through to the source and log the failure. A cache outage should be a performance incident, not an availability one. Wrap the get and set in try/catch and treat any error as a miss.


#Backend