Test results
Image Prompts in One Style: Clean Pictures for a Site: test results
Tested 2026-10-09, skill version 1.0.3 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-09
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right on all 12 briefs, read by hand: every prompt opened with the scene and closed with the house style copied exactly, gave the frame in numbers and a closed list of objects, and described clean surfaces in positive words. It turned a chalkboard of dishes, an invoice total, a chart-filled laptop, a road sign with an address and a clock at five to twelve into plain shapes, showed a hiking group from behind with hands out of frame, put the unwanted things only in the separate negative box when the tool had one, wrote three complete prompts for three services, and wrote in English for the two Bulgarian briefs.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right on 9 of 12, but it missed three: it still named the dashboard on the laptop screen, kept one negation (no markings on the kettle) in the coffee prompt where the user had asked for no text and no logo, and put a line of its own notes before one prompt although only the prompt was asked for. The order, the house style, the frame, the closed list and the people rule held in every answer.
With and without the skill
Tested 2026-10-09.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Prompts right on every check (12 prompts) | 12/12 | 0/12 | 9/12 | 0/12 |
| Prompts free of the defects alone (12 prompts) | 12/12 | 1/12 | 9/12 | 1/12 |
Same request on both sides. The second row leaves out the three structure checks the request never states (scene before style, closed list, frame in numbers) and counts only defects: negations, words that summon writing, faces, collages, language and the exact style; the price follows that stricter row. Read by hand, Sonnet without the skill wrote careful, vivid prompts and kept the house style exactly, but in 10 of 12 it answered the request for a clean picture with a list of bans (no text, no letters, no logos, no watermark), which a one-box tool reads as a request for those things; it also named the menu, the dashboard, the street name and the time on the clock, drew the smiling agent and the laughing hikers with their faces, and put the style before the scene every time. Haiku without the skill made the same moves.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twelve image briefs written by us for two house styles: ten with a trap (writing on a menu board, an invoice total, a chart-filled screen, a road sign with an address, a clock showing a time, a group of hikers, a smiling support agent, a user who asks for 'no text, no logo', three images at once, a tool with a separate negative box) and two plain controls; two briefs are in Bulgarian. The request is the same on both sides and states the tool, the house style word for word, 'premium, without the usual defects of AI images' and 'return only the prompt'. Code checks the prompt, not a picture: English, the style copied exactly, the scene before the style, a closed list, the frame in numbers, no negations in the prompt box, no words that ask for writing or name what a shape stands for, people only from behind or far away and without faces, one complete prompt per picture. Facts from the FLUX prompting guide (re-read 2026-10-09): there is no negative-prompt field, the wanted scene is described instead, a text-to-image request without references comes back square, the subject comes first. Owner-measured and NOT re-checked: a list of forbidden things in a FLUX prompt filled the picture with letters, a dial carried digits until the scene was a closed list, and ChatGPT made one collage from three scenes in one message and a square when no frame was given. Checks widened after the run, for both sides, each after a defensible answer was refused: people seen from behind or in the distance with hands out of frame, a tile grid in a bathroom, twelve dots as hour markers, and a curve shaped like a smile. The skill was also changed once after the first run (describe a shape by its look, not by the value it stands for) and its side was re-run in full. One run per model and brief. Picture probe after the test, on 9 October 2026: three trap briefs (menu board, invoice total, road sign) went to ChatGPT with the plain and the guided prompt, in two rounds (the second asked it to ignore its memory and saved preferences, because the account keeps the owner's own rules for pictures). None of the twelve pictures had garbled letters; ChatGPT wrote the requested amount and street name cleanly; stray objects appeared once in the first round only. A third round went to FLUX schnell (Cloudflare Workers AI): the plain prompts gave a board with pseudo-letters despite asking for none, a wrong amount and a sign without its number; the guided prompts gave three clean pictures. One run per picture.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.
Back to Image Prompts in One Style: Clean Pictures for a Site · Card (JSON)