Test results

Redact Sensitive Data: PII and Secrets: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Returned the exact expected text in all 20 cases, every other character unchanged: emails, phone numbers in several national formats, IBANs with spaces, a card number, an EGN, a DNI and an SSN, IPv4 and IPv6 addresses, the password inside a connection string, a bearer token, a key in a URL, a session secret in an env file, a key in code and a temporary password. It kept order, invoice, tracking and request numbers, versions, a commit hash, the last four digits of a card, an obvious placeholder key and an empty password, and redacted as usual where the text said the data was already clean or fake.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Exact in 15 of 20. But it failed five by dropping the text before the first value it replaced: the words before an email, a From: label, the start of a Markdown link and, on the record that claimed to be fake, everything except one placeholder; once it also swallowed the label SSN into the placeholder.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
Exact redacted text (20 texts)20/2020/2015/2011/20

The same request on both sides, with the same seven placeholders; a fence around the answer is removed first. Sonnet needs no help on this test. For Haiku the skill fixed five answers, four of them cases of redacting too much and one of changed layout: the placeholder key, the last four digits of a card, a tracking number taken for a card number, the parameter name in a URL, and the line breaks of a German invoice, which it had merged into one line. It made one worse, a Markdown link where Haiku dropped the sentence before it. The other four misses are the same with and without the skill: Haiku drops the text that comes before the first value.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty texts written by us, in English, Bulgarian, German and Spanish, with every value invented and no key in a real provider's format. The expected answer is the text with each marked value replaced by its placeholder, compared character for character. The first version of the skill made Haiku think aloud in two answers; one paragraph about where the answer starts and ends was added and the side with the skill run again, the first run kept in the test folder. The side without the skill ran once. One run per model and case in each version.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Redact Sensitive Data: PII and Secrets · Card (JSON)