6 min read

Fix the query, then cache: how I cut a critical endpoint's response time by ~60%

Notes from a Node.js and MongoDB service: why the query should be fixed before a cache goes in, how a cache-aside layer works, and how to keep the response contract intact along the way.

At Junzi Tech I worked on a Node.js / Express service backed by MongoDB, with a React front end. By refactoring the queries behind one critical endpoint and adding caching, I cut its average response time by about 60% without changing its response contract. This note explains the approach and the reasoning behind it. The code is not mine to share, so the snippets below are simplified, illustrative shapes of the technique, not production code, and the details describe the pattern I use rather than a record of that specific system.

Context

Endpoints on a critical path tend to grow the same way: one more field here, one more lookup there. Nobody breaks them; they simply accumulate work until latency becomes a problem.

Two constraints shape every decision in this kind of work:

  • The response contract cannot change. When the front end depends on the exact shape of the JSON, a faster endpoint that returns something slightly different is a regression, not an improvement.
  • The data is not static. Some of it is read far more often than it is written, but it still changes, so "just cache everything for an hour" is not an option.

Options

Option 1: put a cache in front and move on

This is the tempting one. A cache hides a slow query very well on a hit. But on a miss the user still pays the full cost, and misses happen after every invalidation, every deploy and every cold start. A cache over a slow query also makes that query harder to see: the average improves while the worst requests stay just as bad.

Option 2: scale the database or the service

More resources help a little, but when a query does more work than it needs to, the problem is not capacity. More hardware only makes the waste cheaper for a while.

Option 3: fix the query first, then add a cache

Make the uncached path fast and predictable, then add caching where the read-to-write ratio justifies it. This takes longer than option 1, but every later decision rests on a path you understand.

Approach

I went with option 3: query refactoring first, caching second. Shipping the query change on its own is worth it, because it lets you see its effect in isolation before a cache blurs the numbers.

Step 1: measure and read the query plan

Before changing anything, look at what the database is actually doing. In MongoDB that means running the query with explain("executionStats") and comparing documents examined with documents returned. When the database reads far more documents than it sends back, the filter is not served by a suitable index.

The usual fixes are unglamorous:

  • a compound index whose field order puts equality filters first, then the sort field;
  • a projection, so the database returns only the fields the response needs instead of whole documents;
  • replacing per-item lookups in application code (the N+1 pattern) with one batched query.
// Simplified, illustrative example: filter + sort served by one compound index,
// projection limited to the fields the response actually uses.
await db.collection("items").createIndex({ accountId: 1, status: 1, updatedAt: -1 });

const items = await db
  .collection("items")
  .find(
    { accountId, status: "active" },
    { projection: { _id: 1, title: 1, status: 1, updatedAt: 1, ownerId: 1 } }
  )
  .sort({ updatedAt: -1 })
  .limit(50)
  .toArray();

// One batched lookup instead of one query per item.
const ownerIds = [...new Set(items.map((i) => String(i.ownerId)))];
const owners = await db
  .collection("owners")
  .find({ _id: { $in: ownerIds.map(toObjectId) } }, { projection: { name: 1 } })
  .toArray();

Step 2: protect the response contract

If the shape cannot change, treat the old response as the specification. Compare responses before and after the refactor for the same inputs, field by field, including ordering and empty cases. Keeping the mapping from database documents to the response in a single function also helps: a change in the projection cannot quietly drop a field somewhere else.

Step 3: add a cache-aside layer

Only once the uncached path is fast does caching make sense, and only for data that is read much more often than it changes. The pattern I reach for is cache-aside: read from the cache, fall back to the database on a miss, write the result back with a TTL, and delete the key when the underlying data changes.

// Simplified, illustrative cache-aside wrapper with a TTL and explicit invalidation.
async function cached(key, ttlSeconds, load) {
  const hit = await cache.get(key);
  if (hit !== null) return JSON.parse(hit);

  const value = await load();
  await cache.set(key, JSON.stringify(value), { EX: ttlSeconds });
  return value;
}

const listKey = (accountId) => `items:list:${accountId}`;

// Read path (the TTL value is arbitrary here)
const payload = await cached(listKey(accountId), 60, () => buildItemsResponse(accountId));

// Write path: invalidate after the database write succeeds
await db.collection("items").updateOne({ _id }, { $set: changes });
await cache.del(listKey(accountId));

The TTL is a safety net, not the main invalidation strategy. Explicit deletes on write keep the data fresh; the TTL limits the damage if a write path ever forgets to invalidate.

Trade-offs

  • Delivery time. Fixing the query takes longer than adding a cache. It is worth the price, because it makes the miss path fast as well.
  • Index cost. Every index speeds up some reads, slows down every write to that collection and takes memory. Add the index the query plan asks for, not one per field.
  • Cache consistency. Cache-aside leaves a short window where a reader can see stale data if an invalidation races with a read. That is fine for many kinds of data; for balances or permissions it is not.
  • More moving parts. A cache is one more thing to monitor and one more thing that can go down. The endpoint has to keep working, only slower, when the cache is unavailable.

Lessons

Read the plan before you write the fix

The query plan usually points straight at the problem: a missing compound index, whole documents being returned, a loop of small queries. None of it needs clever code.

Cache a fast path, not a slow one

A cache on top of a slow query improves the average and leaves the worst requests untouched. A cache on top of a fast query makes a good path better, and a miss stops being a problem.

The contract is part of the performance work

An unchanged response is what makes a change like this safe to ship. Comparing old and new responses for the same inputs is cheap and removes the main risk.

Invalidation belongs next to the write

Keep the cache key builder and the invalidation next to the write code, so they are hard to forget. A test that fails when a write path changes data without deleting the matching key is worth adding.

Alhassan Alfarran.

© 2026 · Designed and built by me with Next.js, Tailwind and Framer Motion.

My local time: · Moscow

Notes
How this site is built

Stack

Next.js (App Router) and React, styled with Tailwind CSS and animated with Framer Motion. The contact form sends email through Resend; the site is hosted on Vercel.

Three languages, one layout

English, Russian and Arabic each have their own address (/en, /ru, /ar) and share one set of components. The layout uses logical CSS properties (start/end instead of left/right), so Arabic mirrors right to left without separate styles. The server sends every page with its language and text direction already set, so nothing flips after loading, and the Arabic font is only downloaded when Arabic text is on screen.

Performance and accessibility

Sections below the first screen skip rendering until you scroll near them, and the quick menu loads on first use. Everything works from the keyboard, with a skip link and visible focus, and animations switch off when your system asks for reduced motion.

Source code on GitHub