Reviews pasted code and configuration that spends money (paid AI or API calls, purchases, paid lookups, anything billed per use) and checks whether the limit on that spending really stands in the way of the spending. It flags a cap that sits beside the path instead of on it (a gateway limit while a script calls the provider directly, a counter only in the browser), a missing setting that means no limit, an alert or a budget email used as the protection, a cap that the caller sets through a header or a request field, a paid job that runs by default, a per-call cap with no running total, a total kept where it resets, a cap checked only after the money is gone, and a spend counter whose failure or missing value lets the call through. Each finding has a fixed code, the place, the reason and a fix, then one verdict. Use before deploying a Worker, script or job that pays per call, or when a cap exists and you want to know if it can fail silently.
Spend Guard Review: Caps That Really Stop is a tested SKILL.md that reviews pasted code and configuration that spends money (paid AI or API calls, purchases, paid lookups, anything billed per use) and checks whether the limit on that spending really stands in the way of the spending; an agent buys it once for $0.03 over x402.
Not for
Loops, retries and schedulers that run too often (use bill-spike-finder), rate limits against abusive clients (use cf-rate-limit-design), prices, or anything that needs the running system: it reads pasted code and cannot see dashboard settings, keys or helpers the paste omits. The one-day delay of budget alerts is the owner's measurement, not re-checked.
Tested, honestly
Tested 2026-10-08 with a strong and a weak model.
With and without the skill
Results with and without the skill, for Sonnet and Haiku |
| with | without | with | without |
|---|
| Snippets reviewed right (23 snippets) |
| Snippets reviewed right (23 snippets) | 23/23 | 21/23 | 23/23 | 20/23 |
|---|
Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already found most gaps in its own words. It missed two: a paid action with no paid flag at all, and a cap that comes back null when unset, which it did not call unlimited. Haiku without the skill missed three: a cap that lives only in the browser, an alert with nothing that stops the spending, and the missing paid flag.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
- SonnetStrong model, claude-sonnet-5-5
- Right on all 23 snippets, read by hand. It found caps that sit off the spending path, a default that turns a missing cap into no limit, a guard that fails open on an error, a cap the caller sets, a per-request limit with no running total, a paid action with no paid flag, an alert with nothing that stops the spending, and it ignored a planted comment.
- HaikuWeak model, claude-haiku-5-5
- Right on all 23 snippets, read by hand, with the same gaps found as Sonnet, the planted comment included.
Full test summary
Example
Our own test text, before and after the skill ran. Excerpts only.
English · claude-sonnet-5-5
Before
// The "summarizer" Worker. AI Gateway "prod-gw" has a monthly spend limit of 50 dollars set in the dashboard.
// src/index.ts
export default {
async fetch(req: Request, env: Env): Promise<Response> {
const { text } = await req.json<{ text: string }>();
const out = await env.AI.run("@cf/meta/llama-3.1-8b-instruct", { prompt: "Summarize: " + text }, { gateway: { id: "prod-gw" } });…
After
[CAP-OFF-PATH] scripts/backfill-summaries.mjs: it calls llm.example.com directly with its own `LLM_API_KEY`, so the "prod-gw" monthly limit never sees these calls, and it summarises every archived article with "big-model" and no total of its own.…
What is in the file
- The answer
- The codes
- Rules
- Work in this order
- Short example
Languages
Any language. Tried in: English.
License
Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.
Versions
Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.
v1.0.0 · 2026-10-08
First release (not yet tested on a model): reviews pasted code and configuration that spends money for the three layers of spend protection and lists each gap with a fixed code, the place, the reason and a fix, then a verdict (can overspend, guarded). Codes: cap off the path, unlimited default, fail-open measurement, alert as the only protection, cap set by the caller, paid job that runs by default, per-call cap without a running total (or a total that resets), cap checked after the money is spent. Bill loops and retries (bill-spike-finder) and rate limits against abusive clients (cf-rate-limit-design) are left to those skills.
Facts are the owner's rules and measurements, not re-checked on 2026-10-08 (nothing read-only to probe).
Tests: 23 cases (15 traps, 8 controls), test/make-cases.mjs generates test/cases.json; test/control.mjs makes no model call and passes. Suggested price $0.03, to be set after the with and without runs.
FAQ
Caps, defaults, alerts: which gaps are reported?
Eight gaps between a limit and the money it should stop: a cap beside the path of the spending, a missing setting that means no limit, a counter that fails open, an alert used as the only protection, a cap the caller can raise, a paid job that runs by default, a per-call cap with no running total, and a cap checked only after the call is paid. Each finding names the place and proposes the smallest change that puts the limit on the path, such as an atomic reservation before the call.
Does it complain about a guard that works?
Controls cover the usual look-alikes: a header that can only lower the cap, a missing setting that stops the spending, a conditional update made before the call, a single-writer meter and a call that costs nothing. All of them get No findings, and anything the paste does not show counts as unknown, never as a finding. It also leaves alone code whose reservation helper is imported but not pasted, since a helper nobody showed may well be sound.
Why not just set a budget alert, and how is this different from a bill-spike review?
A cloud budget alert is worked out once a day for the day before and only informs: it says money was spent, never that money will be spent. The skill keeps the alert as a second line and asks for a refusal in the code that makes the paid call. A bill-spike review finds loops that run too often; this one asks whether a limit stands in front of the paid action at all. Treat the alert as the smoke detector and the in-code refusal as the sprinkler: the first reports a fire tomorrow morning, the second works now.
Does it help Claude Sonnet?
Somewhat. Both Claude models reviewed twenty-three snippets of code that spends money, first without this file and then with it, and each answer was scored on whether the gap is named in any words. Unaided, Sonnet caught most gaps but missed a paid action with no paid flag and a cap that quietly becomes no cap when unset, scoring 21; guided, it found all 23. Haiku also reached 23, from 20. Unaided, both flagged some sound guards too; that side is not counted.