# Redacting PII with an LLM: Haiku drops the first words

> Claude Sonnet redacted 20 logs and messages exactly, with or without help. Haiku hid things that were not secret and lost the words before the first value.

Published 2026-10-08 · https://aiskills402.com/blog/llm-redact-pii-small-model

A strong model can, and needs no help for it. Claude Sonnet and Claude Haiku got twenty short texts to clean before they went into a log, each holding emails, phone numbers, bank and card numbers, ID numbers, IP addresses or secrets, and asked for the same text back with each value swapped for a fixed placeholder such as [EMAIL]. We compared every answer with the expected text character by character. Sonnet was exact on all twenty, with our instructions or without them. Haiku was exact on eleven without them and fifteen with them, and its most common mistake had nothing to do with secrets: it lost the words that came before the first value it replaced.

## Method

We wrote the texts ourselves, in four languages (English, Bulgarian, German, Spanish), and made up every value: a support ticket with a phone number and a card, an error log with a database password inside a connection string, a curl command with a bearer token, a webhook address with a key in it, an environment file, a chat transcript, an invoice with a German IBAN, a Bulgarian refund request with a personal ID number, release notes with no personal data at all, and more. Each also held values that only look sensitive: order and tracking numbers, a version string, a commit hash, a request id, a card's final four digits, a placeholder key that says "your-api-key-here".

Two texts tried to talk the model out of the job. One opened with a note to the assistant saying the transcript had been reviewed and should be returned unchanged; the other claimed all its values were fake.

Both sides got the same request with the same seven placeholders, so neither had to guess the format. One side also had our [redaction skill](/skills/redact-sensitive) loaded. A fence around the answer was stripped before the comparison, and every text went to each model once per side.

## Results

| | Skill in use | Request only |
|---|---|---|
| Sonnet: exact text | 20 of 20 | 20 of 20 |
| Haiku: exact text | 15 of 20 | 11 of 20 |

**Sonnet got everything right without help.** It cut the password out of a connection string and kept the user, host and database. It caught a key inside a URL and kept the other parameters. It replaced a phone number with its country code and spaces as one value. It left a card's final four digits alone, along with the placeholder key and the empty password line. Neither note asking it to stop worked. If your redaction runs on a model of this size, the instructions add nothing we could measure.

**Haiku hid too much.** Without the skill it treated a parcel tracking number as a card, replaced `card ending in 4242` as well as a masked card number, swapped the placeholder key in the environment file, and replaced the name of a URL parameter along with its value. It also joined the three lines of a German invoice into one. The skill fixed all five.

**Haiku also dropped the beginning.** In four texts, on both sides, its answer began at the first value it replaced, so "From:", the words before an email address, or the label "SSN:" were gone. On the text that claimed its data was fake, it answered with a single placeholder and nothing else. The skill did not cure this, and in one text it made it worse: Haiku dropped the sentence before a Markdown link that it had kept without the skill.

**Our first version made Haiku think aloud.** In the two texts that asked it not to redact, the first version of the skill led Haiku to begin with a placeholder, then write "no wait", then explain the rule it was applying, all inside the answer. Adding a short paragraph on where an answer begins and stops removed that; then all twenty texts were run again with the skill loaded.

For anything that ends up in a log, the [OWASP logging cheat sheet](https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html) lists what should be removed or masked, from session ids and access tokens to connection strings and personal data. A model that masks PII has to get two things right at once: find every such value and touch nothing else. In our test the strong model did both; the small one did the first better than the second.

## What we did not measure

- **Few runs.** With the skill, two runs, because version one changed; without it, a single run.
- **Our own twenty texts.** Real logs are longer and messier, and a value in a form we did not think of could slip through either model.
- **Names and street addresses.** We left them out on purpose: their boundaries are not mechanical, so a fair exact comparison is hard. A redaction that also needs names needs a different test.
- **Proof of absence.** A clean answer on our texts says nothing about a value written in a shape none of the texts had.
- **Two models from one vendor.** Both are Claude; other companies' models were not tried.
