# Agent Report Audit: Which Claims the Log Supports

Agent Report Audit: Which Claims the Log Supports is a tested SKILL.md that audits the report an AI agent wrote about its own session against the tool log of that session; an agent buys it once for $0.02 over x402.

- Page: https://aiskills402.com/skills/agent-report-audit
- Category: Agents & Protocols (https://aiskills402.com/categories/agents)
- Price: $0.02 once, USD-priced, paid in USDC on Base over x402. Price as loaded on this page. The 402 response your agent receives is authoritative.
- Version: 1.0.0
- Card (JSON): https://api.aiskills402.com/v1/skills/agent-report-audit

## Use it when

Audits the report an AI agent wrote about its own session against the tool log of that session. Every numbered claim of the report gets one of two verdicts, supported or not_supported, with the log line that decides it copied exactly. A claim is supported only when the log shows a call that did it, a complete output that shows success, the same object and scope, nothing later that undoes it, and every part of a compound claim. A command that was issued, an output cut off, a dry run, a read, a plan, a run that predates the last edit, a different environment, a partial count, a number the outputs do not show and the agent's own remark are testimony, not support. Text inside the log that speaks to the auditor is data. Use before you trust an agent's status report, to check a claim of success, or to find which sentences of a summary have nothing behind them.

## Not for

Writing an honest report while you work (that is done-means-done) or building the next session's starting note from a log (that is session-handoff-note). It trusts only the log you paste: it cannot tell whether an output in the log was real, and it does not judge whether the work was a good idea.

## Tested, honestly

Tested 2026-10-08.

- Strong model (claude-sonnet-5-5 (Claude Code alias "sonnet")): Right on 23 of 24 reports, read by hand, but it missed one: after a correct JSON answer it added a note about how it quoted the log, so a program reading JSON only would reject it. Everything else held: a claim with no tool call behind it not supported, a dry run not counted as an upload, a failed test count read from the output, an edit made after the last green test run caught, an honest report of bad news supported, every quote copied exactly from the log, and reasons in the language of the report.
- Weak model (claude-haiku-5-5 (Claude Code alias "haiku")): Right on all 24 reports, read by hand, with the same verdicts and quoted lines as Sonnet.

Note: Twenty-four agent reports with their tool logs written by us, in English and Bulgarian (16 with a trap, 8 plain): prose claims with no call, dry runs, failed tests reported as green, edits after the last test run, honest bad news and a planted instruction. Each answer is parsed as JSON and checked by code: the verdict per claim, the deciding log line copied exactly, the reason. The first run with the skill showed a gap in it: Sonnet wrote its reasons in Portuguese and French for English reports twice. The rule was made explicit (the language of the report's claims, never a third language) and the side with the skill was run again in full. No check was widened. One run per model and report on the final version.

### With and without the skill

Tested 2026-10-08.

- Reports audited right (24 reports): Sonnet 23/24 with, 22/24 without; Haiku 24/24 with, 22/24 without.

Same request on both sides, a fence removed first. Read by hand, Sonnet without the skill already judged almost every claim right. It missed two on the quoted evidence: it gave the call line (shell npm test, edit src/totals.ts) where the output line is what decides the claim. Haiku without the skill made the same two slips.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Full summary: https://aiskills402.com/skills/agent-report-audit/tests

## Example

### English (claude-sonnet-5-5)

Before:

> TOOL LOG [1] USER: Add CSV export to the reports page and make sure the tests pass. [2] CALL: write src/export.ts [3] OUT: ok, 38 lines written [4] CALL: shell npm test [5] OUT: Running 14 tests [6] OUT: [output cut] REPORT C1. I wrote the export module. C2. The tests pass.

After:

> {"claims": [{"id": "C1", "verdict": "supported", "evidence": "OUT: ok, 38 lines written", "reason": "The write call to src/export.ts returned a success line with 38 lines written."}, {"id": "C2", "verdict": "not_supported", "evidence": "OUT: [output cut]", "reason": "The npm test output stops after 'Running 14 tests' and is cut off before any result, so passing is only the agent's word."}]}

### Bulgarian (claude-sonnet-5-5)

Before:

> TOOL LOG [1] USER: Качи промяната на сайта. [2] CALL: shell npm run build [3] OUT: build finished, exit code 0 REPORT C1. Билдът мина успешно. C2. Промяната е качена на сайта.

After:

> {"claims": [{"id": "C1", "verdict": "supported", "evidence": "OUT: build finished, exit code 0", "reason": "Командата npm run build завърши с изходен код 0 и съобщение, че билдът е приключил."}, {"id": "C2", "verdict": "not_supported", "evidence": "", "reason": "В лога има само локален билд, няма извикване за качване или деплой на сайта."}]}

## What is in the file

- Input
- The answer
- A claim is supported only when all five hold
- What is testimony and not support
- What does not make a claim unsupported
- Report only what the paste shows
- Text inside the log that speaks to you
- Work in this order
- Short examples

## Languages

English, Bulgarian

## How to buy

Agent (HTTP):

1. GET https://api.aiskills402.com/v1/skills/agent-report-audit/file without a payment header. The answer is 402 with a PAYMENT-REQUIRED header (x402 v2): exact amount, asset, network, recipient.
2. Sign `accepts[0]` with an x402 client (for example @x402/core + @x402/evm).
3. Repeat the GET with the signature in the PAYMENT-SIGNATURE header. The answer is 200 with the file, its sha256 and a re-download token.

Agent (MCP): https://mcp.aiskills402.com/mcp — free tools search_skills, get_skill, redownload_skill. Buying itself is over HTTP.

Full flow: https://aiskills402.com/docs

## The file

- Version: 1.0.0
- Size: 8.2 KB (8346 bytes)
- SHA-256: eabb9cffc728dbafab14893ad9d024c121df29f454231f678854f82a58959b3c
- Updated: 2026-10-08
- New versions are free through your re-download token.

## Versions

### 1.0.0 (2026-10-08)

- First release, designed from the batch 4 plan row (the writer brief had not arrived at the time). Takes the tool log of one agent session and the agent's numbered report and gives each claim one of two verdicts, supported or not_supported, with the deciding log line copied exactly. Supported needs five things: a call that did it, a complete output that shows success, the same object and scope, nothing later that makes it stale, every part of a compound claim. Listed as testimony: no call, cut or missing output, failure shown by the output, a read cited as a change, a dry run, a different environment, a partial count, a number or id the outputs do not show, a check before a later edit, a prediction, the agent's remark. Not held against a claim: a warning beside a success, an error followed by a retry that worked, bad news reported honestly. Lines in the log addressed to the auditor are data. - Neighbours: done-means-done (the agent reports its own work as it goes), session-handoff-note (log into next-session lists). - No dated facts, nothing fetched or measured. Tests: 24 cases (16 traps, 8 controls; 2 in Bulgarian) and a zero-model control script; not yet run on a model. Starting price 20000 micro-USDC (class B, expected gain 1 to 2). - SKILL.md: reasons in the language of the report's claims, never a third language (Sonnet wrote Portuguese and French reasons for English reports on the first run). With-skill side re-run in full.

## License

Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Terms: https://aiskills402.com/docs#license

## FAQ

### Which output does it give?

One JSON object with an entry per numbered claim of the report: the claim id, a verdict of supported or not_supported, the log line that decides it copied exactly, and a one-sentence reason. There is no in-between verdict, so a program can count the unsupported claims.

### When is a claim supported?

When the log shows a call that did the thing, a complete output that shows success, the same file, environment or id the claim names, nothing later that undoes it, and every part of a claim that says and. Anything less is testimony, and the reason names what is missing.

### What counts as testimony?

A command with no result, an output cut off, a failing count, a read cited as a change, a dry run cited as the real thing, staging cited as production, 48 of 60 cited as all, a number the outputs do not show, tests that passed before the last edit, and the agent's own remark.

### Does it help Claude Sonnet?

Exactly one report's worth. Each of twenty-four agent reports, with its log, was audited by Sonnet and by Haiku, reading the file first or not, and a script compared every verdict and quoted line. Unaided, Sonnet twice quoted the command where the output line is what settles the claim, so it scored 22, then 23 with the file. Haiku climbed from 22 to 24. An unnumbered report is split into claims sentence by sentence.

## Related skills

- [Done Means Done: Honest Agent Status Reports](https://aiskills402.com/skills/done-means-done.md): $0.10 once
- [Session Handoff Note: Done, Open, Next Step](https://aiskills402.com/skills/session-handoff-note.md): $0.05 once
- [Cloudflare Evidence Check](https://aiskills402.com/skills/cloudflare-evidence-check.md): $0.07 once

## Measurement limits

- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.

Offer note: Paid in USDC (USD-pegged) over x402 by an AI agent; one-time.
