Mostly from the prompt. When a scene contains something that normally carries writing, such as a chalkboard menu, the total on an invoice or a street sign, the image model is being asked to write; and in a tool with only one prompt box, even a phrase meant to forbid letters still mentions letters. How badly that writing turns out depends on the tool. In our probe, FLUX schnell answered three such prompts with a board of pseudo-letters, a total whose figures were wrong, and a sign that lost its house number, while ChatGPT spelled every requested word and figure correctly. When the same three scenes were described as plain shapes, both tools returned clean pictures with no garbled letters at all.
Why we looked
Every header on this blog comes from an image tool, in one fixed look. Two things betray a generated picture at a glance: smeared pseudo-text on a shop sign or a monitor, and a set of pictures that seem to come from different studios. Our notes from earlier work blamed the prompt rather than the model. A list of forbidden things sent to FLUX came back covered in lettering. A water meter kept its digits until the scene was rewritten as a short inventory of what may appear in it. Three scenes requested in one ChatGPT message arrived as a single collage. None of that had been measured properly, so we measured.
Method
The test had two stages: first the prompts a language model writes, then the pictures those prompts produce.
Prompts. We wrote 12 image briefs for two house styles. Ten hide a trap. Some invite writing: a restaurant's chalkboard with dishes and prices, an invoice whose total should be visible, a laptop with an analytics dashboard, a sign giving an office address, a wall clock at five to twelve. Two invite faces: a hiking group laughing on a trail and a support worker smiling into a headset. The rest test the shape of the answer: a user who wants a ban on text, watermarks and logos tacked onto the end, a request for three service pictures at once, and a tool with a separate box for unwanted things. Two briefs are plain controls, and two are written in Bulgarian. Both sides received exactly the same request: it named the tool, pasted the house style, asked for a premium result free of the typical flaws of generated images, and wanted nothing back except the prompt. Claude Sonnet and Claude Haiku answered every brief once through Claude Code, first with the bare request and then with the skill loaded.
The checks were code, and they read prompt text only. A prompt passed if it was in English; if the house style appeared unchanged and after the scene; if the scene ended with a closed list and the frame was given as numbers; if the main prompt held no negation; if no word asked for lettering or named what a shape represents; if people showed up only as distant figures or from behind; and if each picture had a full prompt of its own. The brief never mentioned three of these (the order, the closed list, the numeric frame), so the table also carries a stricter row that counts real defects only.
Pictures. For the picture stage we picked three of the writing traps: the menu board, the invoice and the road sign. Sonnet's recorded prompts for them, with and without the skill, went once each to two tools. In ChatGPT we used a normal chat and opened every prompt with a line telling it to set aside its memory and stored preferences, since that account holds its owner's own rules for pictures. FLUX.1 schnell ran on Cloudflare Workers AI, where the documented inputs are just the prompt and a step count (model page). Six prompts in two tools gave twelve pictures.
Results
The prompts
| prompts that passed | bare request | request + skill |
|---|---|---|
| Claude Sonnet, all checks | 0 of 12 | 12 of 12 |
| Claude Sonnet, defects only | 1 of 12 | 12 of 12 |
| Claude Haiku, all checks | 0 of 12 | 9 of 12 |
| Claude Haiku, defects only | 1 of 12 | 9 of 12 |
The bare prompts were better than the zeros suggest. Read one by one, Sonnet's scenes were detailed and its house style exact. The trouble lay elsewhere. In 10 of the 12 it met the wish for a clean picture by listing what must not appear, with text, logos and watermarks on the list. It called the board a menu, the laptop screen a dashboard, the sign by its street and the clock by its hour. The support worker and the hikers kept their smiling faces, and the style paragraph always came before the scene. Haiku's bare prompts had the same habits.
Loaded with the skill, Haiku still missed three prompts. One still called the laptop screen a dashboard, one kept a single negation about the kettle's surface, and one began with a sentence of commentary although the brief wanted the prompt alone.
One finding changed the skill during the test. In the first run, both models wrote phrases like "a bar for the final total". The bar is harmless; the words after it name the very figure the bar was supposed to replace, and an image model treats every noun as something to draw. We added a rule to describe a shape by its look alone and ran the skill side again from scratch; the table shows that second run.
The pictures

FLUX schnell failed the plain prompts in three separate ways. The menu prompt had stated that the board carried neither letters nor numerals; the picture showed a big MENU in wobbly type, an invented word below it, a PRICE label and a smudged amount. The invoice prompt wanted the total 1 250.00 printed large and readable, and closed with a list of things to avoid, warped letters included; the picture said $1 25.00. The sign prompt spelled out IVAN VAZOV ST. 12 exactly; the picture kept the street and dropped the number. With the guided prompts the same scenes became stacked bars, a coral block at the foot of a sheet, and a blank arrow-shaped board pointing at a small house, so nothing was left to misspell.

ChatGPT returned six clean pictures. It printed the total and the whole address correctly, and it left the menu board free of characters even though that prompt mentioned them. For this tool and these three briefs, the plain and the guided prompts gave pictures of about the same quality. An earlier round in the same account, sent before we added the line about memory, had no garbled lettering either; its plain pictures picked up a few stray objects that the second round did not repeat.
That result cuts both ways, and it matters when you choose a method. If a picture truly needs a readable figure and the tool spells well, the plain prompt delivered what was asked and the guided one did not: it swapped the number for a coloured bar. If the tool spells badly, or the picture only has to suggest an invoice, the shape is the safer request. The skill we tested, image-prompts-one-style, takes the second view. It is meant for headers and cards that must never show broken text, not for posters that have to say something.
The vendor's advice for its newer model points the same way. The FLUX guide says there is no field for a negative prompt, that the wanted scene should be described directly, and that a text-to-image request without a reference image comes back square unless a ratio is set (prompting guide). It also recommends opening a prompt with the main subject and its action (building a prompt).
What we did not measure
- Every prompt got a single picture per tool. Generated pictures change from one draw to the next, so a second attempt could break a guided prompt or rescue a plain one. The FLUX results cannot be redrawn exactly either: the examples on the model page send a seed, yet our request with a seed was rejected with error 5006, so we went without one.
- Only three briefs reached the picture stage, all of them built around writing. Faces, hands, collages and drifting styles were checked in the prompts but never in real pictures.
- The ChatGPT account stores its owner's picture rules. We told it to set them aside; we have no way to confirm that it did.
- We tried the small, fast FLUX model at its highest step count. Larger FLUX models, Stable Diffusion, Gemini and other tools are untested, and the vendor guide we cite was written for a newer FLUX model than the one we ran.
- In the prompt stage each model answered each brief once, the briefs are our own, and the checks read words, not pixels. After the first run we relaxed four checks for both sides, each time because a reasonable answer had failed: figures seen from behind, bathroom tiles laid out in a grid, a ring of small dots standing for the hours, and a curve that resembles a smile. Relaxing a check can also let a real mistake through.
- The older notes behind the test (the ban list that filled a picture with lettering, the meter dial, the collage) were not re-run here. The FLUX menu board repeated the first of them, once.
Read this post as Markdown: /blog/ai-images-garbled-letters.md · Atom feed.
