Test results

Next.js Cache Review: ISR, Tags, Redirects: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Right on all 23 snippets, read by hand: it coded the changed shape under an unchanged key, the field rename under the same key, the page reading D1 during the build, the redirect over a path that had answered 404 for a month, the matcher that skips robots.txt, the missing params function, the force-dynamic page with the no-store symptom, the wildcard header, the writer that evicts a tag on every view, s-maxage trusted as a database cap, the busy page that reads every time, and the clock inside the cache (also when two faults sat in one file). It answered No findings. for all nine sound snippets, including the redirect that comes with its own revalidation, and ignored a planted comment saying the file was approved.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Right on all 23 snippets by content, read by hand, but it missed the fixed code on one finding: for the busy front page that reads the database on every visit it described the cost correctly and gave the right fix and verdict, yet left out the code tag. On the sound admin page and the sound middleware it answered No findings. alone, which is exactly what the request asked for. Everything else matched, including the two-fault file and the planted comment.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Snippets reviewed right (23 snippets)23/2321/2323/2320/23

Same request on both sides; the bare side is scored on the concept in any words, fence removed, on the 14 fault snippets only (sound ones are scored by code, with the skill). Read by hand, Sonnet alone already knew the stale shape, the build-time read, the skipped text files, the back button, the wildcard header, the evicting writer, the s-maxage limit, the uncached busy page and the frozen clock. It missed two: after a redirect over a cached 404 it worried about its own middleware, never the 308 without a Location; and for a revalidated dynamic route it never said the missing params function stops ISR. Haiku alone missed the same two, plus the params function in the two-fault file. On sound snippets the bare side invented problems every time (imports, drivers), so those were not counted. Haiku with the skill missed one code tag.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-three snippets of Next.js code and config written by us (14 with a planted caching fault, one of them with two, 9 sound): a changed cache shape, a build-time D1 read, a redirect over a cached 404, a matcher that skips text files, missing static params, force-dynamic and the back button, a wildcard header, a writer that evicts a tag, s-maxage as a cap, a busy page without a data cache, a clock inside a cache, a planted comment. With the skill each answer is scored by code: the exact set of codes, the verdict line, a bare No findings. for sound snippets, no fence. Facts re-checked in the public documentation on 2026-10-08: the cached function persists across deployments and the key parts form its key; an empty params array enables ISR at runtime; OpenNext keeps the incremental cache in R2 and the tag cache in D1. Owner-measured and NOT re-checked: the rendering result of a route without the params function (2026-10-08), the cached 404 plus redirect and the matcher behaviour (September 2026); the wildcard header effect and the back-forward cache refusal rest on the owner's undated notes; the build-time D1 finding is owner-noted, not measured. One check was widened after the run, for both sides: on a sound snippet, a bare No findings. is accepted without the second verdict line, because the request itself said to answer exactly that; a fault verdict next to it still fails. One run per model and snippet.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Next.js Cache Review: ISR, Tags, Redirects · Card (JSON)