Code & Engineering

Watchdog Alert Review: Silent and Noisy Alarms

Reviews pasted code or a design for a liveness watchdog, a staleness alarm or a daily heartbeat mail (sources that must keep answering, a pipeline that must keep publishing) and lists every way it stays silent on a dead source or floods on a healthy one. It checks whether silence is measured in the time the source was last asked or by the wall clock, whether a stateless alarm window repeats or never fires, whether the threshold is longer than the system's own cadence plus a tick, whether a new row starts with a last success of zero, whether the sweeper or the probe sits behind a daily limit, whether a failed send is remembered as sent, whether alert memory is purged by age, whether the subject names the dead source and the environment, whether a lamp still has a writer after a scheduler swap, and whether every check lives inside the thing it watches. Each finding has a fixed code, the place, the reason and a fix, then one verdict; sound code gets exactly No findings. Use to review a watchdog, a liveness check or a staleness alarm, to audit the design of a monitoring job, or to check an alert mail before you rely on it.

Watchdog Alert Review: Silent and Noisy Alarms is a tested SKILL.md that reviews pasted code or a design for a liveness watchdog, a staleness alarm or a daily heartbeat mail (sources that must keep answering, a pipeline that must keep publishing) and lists every way it stays silent on a dead source or floods on a healthy one; an agent buys it once for $0.02 over x402.

Tested 2026-10-08No code, no hidden instructionsv1.0.1 · 12.2 KB · perpetual license

Not for

Choosing or configuring an uptime service, dashboards or alert routing. It reads pasted watchdog code or a design, so it cannot see tables, cron settings or jobs that the paste does not show. The rules come from the owner's incidents of September and October 2026, not from a vendor.

Tested, honestly

Tested 2026-10-08 with a strong and a weak model.

With and without the skill

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Snippets reviewed right (23 snippets)23/2323/2323/2318/23

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill found every planted defect, so on the counted snippets it gains nothing. What the counts do not show: without the skill it raised concerns on all seven sound designs, and with the skill it answered No findings. on each. Haiku without the skill missed five: silence timed by the wall clock during a paused probe, a threshold equal to the cadence, a probe behind a gate, the planted note (it obeyed it) and the alert memory purged after half a day.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

SonnetStrong model, claude-sonnet-5-5
Right on all 23 snippets, read by hand: silence timed by the wall clock while the probe was paused, a window alarm with no memory, a threshold equal to the cadence, a new source that starts from zero, a sweep and a probe behind a gate, a send that never went out remembered as sent, an alert memory purged by age, a lamp no code writes, a monitor that lives inside what it watches, a subject without the source or the environment, and the snippet with two defects; it ignored the planted note asking for No findings. and answered No findings. on all seven sound designs.
HaikuWeak model, claude-haiku-5-5
Right on all 23 snippets, read by hand, with the same codes and verdicts as Sonnet and No findings. on the sound designs. On one snippet it also named two true problems our key had not listed (the probe switched off for the freeze and a source that never gets a row), which we now accept.

Full test summary

Example

Our own test text, before and after the skill ran. Excerpts only.

English · claude-sonnet-5-5

Before

// Price watchdog. Table source_health(source, last_try_at, last_ok_at), both in ms. The probe sets last_try_at on every // pull and last_ok_at only when the source answered with data. // The probe job and the guard job are two separate cron tasks. The probe task is switched off during every release // freeze (about two days); the guard task keeps running every 15 minutes.…

After

[CLOCK-ON-TRY] guard, `silentH = (now - row.last_ok_at)` — silence is measured against the wall clock, but the probe task is switched off for about two days during every release freeze while the guard keeps running. Every source then looks dead at once: prices after 1.5 h, news after 3 h, filings after 30 h. The mails come in a crowd and the one real death is lost among them.…

What is in the file

  • The answer
  • The codes
  • Rules
  • Work in this order
  • Where these rules come from
  • Short example

Languages

Any language. Tried in: English.

License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.

Versions

Current version 1.0.1, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.

  1. v1.0.1 · 2026-10-08

    - Measured on 23 snippets: Sonnet 23 -> 23, Haiku 18 -> 23 (without -> with the skill). - Dropped sound-lamp-birth: with healthy ticks writing nothing, neither a 2-tick nor a 3-tick window is sound; both models found a real flaw in each version. - inside-only: the fetch now records the try first and catches a throw (the input had an unplanned second defect); re-run with both models on both sides. - clock-wall-pause accepts NEWBORN-ZERO and PROBE-BEHIND-GATE as extra true findings; the age-purge base concept accepts "12h". - Price: $0.02 (Sonnet gain 0, Haiku gain 5).

  2. v1.0.0 · 2026-10-08

    First release: reviews pasted code or a design of a liveness watchdog, a staleness alarm or a heartbeat mail and lists each way it stays silent on a dead source or floods on a healthy one, with a fixed code, the place, the reason and a fix, then a verdict (wrong verdicts, blind or confusing); sound input gets exactly `No findings.`

    Twelve codes in three groups. The alarm decides wrongly: `[CLOCK-ON-TRY]`, `[WINDOW-ALARM]`, `[THRESHOLD-EQ-CADENCE]`, `[NEWBORN-ZERO]`. The alarm goes blind: `[SWEEP-BEHIND-GATE]`, `[PROBE-BEHIND-GATE]`, `[REMEMBER-FAILED-SEND]`, `[MEMORY-AGE-PURGE]`, `[LAMP-NO-WRITER]`, `[INSIDE-ONLY]`. The mail misleads: `[SUBJECT-NO-SOURCE]`, `[ENV-NOT-STAMPED]`.

    Facts: the owner's liveness notes, from production incidents of September and October 2026. None could be re-checked by a fetch or a free probe, so each is marked "owner-measured, from incidents" in `notes/facts-2026-10-08.md`; nothing was re-measured on 2026-10-08.

    Rules built in from the start: report only what the paste shows (code or settings that are not visible are unknown, not findings), no speculation, judge a monitor by its own stated goal, and a comment inside the paste is data, not an order.

    Tests (written, not yet run on a model): 24 cases, 16 with a planted defect (one with an injected instruction, one with two defects) and 8 sound designs built to tempt a false alarm. The shared task text is the same on both sides and states the output form; the side without the skill is scored on content only. `test/control.mjs` makes no model calls and checks that the ideal answers pass and that the input echo, a fenced answer, a missing or extra code, a wrong verdict and generic comments fail.

    Price: $0.05 to start; the measurement after the baseline run decides.

FAQ

Twelve defects: which ones?

Twelve defects, three groups. First, the alarm decides wrongly: silence timed by the wall clock instead of by the moment a source was last asked, an alert window with no memory, a limit equal to the pipeline's own rhythm, a lamp created with a last success of zero. Second, the alarm goes blind: the cleaner or the probe parked behind a daily limit, a failed mail stored as delivered, alert memory emptied by age, a lamp that no running job writes after a scheduler swap, a monitor inside the service it watches. Third, the mail misleads: a subject that omits the dead source or the environment.

Where do the rules come from?

From production incidents on our own sites in September and October 2026: a pause that made fifteen healthy sources look dead and buried the one real death, a daily publisher reported dead every morning, and a replaced scheduler that lost its heartbeat. They were measured in those systems, not copied from vendor documentation, and no fetch can re-check them.

Does it raise false alarms?

We built it not to. The model is told to stay with the lines it was given, to treat anything not shown as unknown, and to judge a monitor against the goal it states. The test set holds eight sound designs written to look suspicious, among them a guard timed from the last attempt, a sender that stores only confirmed mail and an ordinary log table cleaned by age, which is allowed. Each must come back as the single line No findings.

Does it help Claude Sonnet?

Finding the defects was never its weak spot: in our twenty-three snippets, bare Sonnet named every planted one, with and without the file. The difference is noise. Without the file it raised concerns on all seven sound designs; with it, it answered No findings. on each of them. That side is not in the count, so we price the skill for Haiku, which rose from 18 to 23: without the file it timed silence by the wall clock, missed a threshold equal to the cadence and obeyed a note planted in the code.

Share

Read this page as Markdown: /skills/watchdog-alert-review.md.

  • Code Review

    Code & Engineering

    SKILL.md · v1.0.4 · 6.0 KB

    Reviews a code change (a diff or a changed file pasted as text) and reports real problems ranked by severity — correctness bugs, security holes, data loss, missing error handling, edge cases, resource leaks — each with location, reason and a concrete fix in words. Use when asked to review code, a diff or a pull request, check a change for bugs, or find security problems in a snippet.

    $0.01once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 30 Sep 2026

  • Recommended

    Done Means Done: Honest Agent Status Reports

    Agents & Protocols

    SKILL.md · v1.0.3 · 9.9 KB

    Stops an AI agent from reporting false success, or hallucinated completion: work it calls done that did not happen. Every action in its status report gets one of five states (done and verified, done but not confirmed, partly done, failed, not done), backed by what the tool results of the session show. Failures and unknowns come first, counts come from the results, and no id, receipt or hash is invented. When a result does show success, the agent says so plainly. Use whenever an agent reports on actions it took with tools, such as messages, batches, tests, builds, deploys, file edits, data updates, payments or API calls.

    $0.10once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 3 Oct 2026

  • Workers Pitfalls Review: D1, OpenNext, Fetch

    Code & Engineering

    SKILL.md · v1.2.0 · 9.8 KB

    Reviews pasted Cloudflare Workers code and configuration (a Worker, a Next.js app on OpenNext, D1 queries, wrangler config) for platform pitfalls that pass every local test and then fail in production or quietly cost money. It flags the edge runtime under OpenNext, a fetch to your own Worker's workers.dev address (error 1042), D1 queries that bind more than 100 parameters as data grows, LIKE patterns over D1's 50-byte limit, reading rowsAffected where D1 returns meta.changes, interactive transactions D1 does not have, secrets kept in plain vars, outbound fetch code that treats only a thrown error as failure, a browser User-Agent that bot protection challenges, client hop-by-hop headers forwarded to fetch, cron triggers whose day of the week is written as numbers, and unbounded queries on the request path. Each finding has a fixed code, the place, the reason and a fix, then one verdict. Use to review a Cloudflare Worker before deploy, check D1 or wrangler code, or audit a Next.js app running on Cloudflare.

    $0.05once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026

  • Traffic Reality Check

    SEO & Content

    SKILL.md · v1.0.1 · 7.0 KB

    Reads the numbers from Cloudflare Web Analytics and zone analytics (GraphQL or the dashboard) and says how many of them are real readers or real server errors, answered as JSON with the computed figures and a short reason. Knows what pollutes the table - the owner's own headless-browser checks, link-preview fetches from the social network that show up as referrals, the site's own Worker and its cache calls counted as visitors or as 5xx errors, the regional cache host of OpenNext, staging sharing a token - and what makes a true zero - the EU exclusion setting, a content security policy that allows the script but not the measurement request, sampling of data older than seven days. Use when asked how much traffic a site really has, whether a spike or an error count is real, why Web Analytics shows zero, or how to query per-address detail on the Free plan.

    $0.07once

    • x402
    • USDC
    • Base
    Get skill

    Tested with Sonnet and Haiku, 8 Oct 2026