Not consistently. Sonnet and Haiku, two Claude models, each received 25 short customer messages with a request for three labels in JSON: a category from a fixed list, a priority from urgent, high, normal and low, and a language code. Without written rules for choosing the priority, Sonnet matched our expected labels on 21 messages and Haiku on 16. The category and the language were nearly always right. Twelve of the thirteen misses were about priority, and ten of those twelve went up rather than down: shouting, anger, a request for a quote and a delayed parcel were enough to move it. The two models also disagreed with each other on seven messages. Once our rules were added as a skill, both gave the expected labels on all 25, and the same labels as each other.

Method

We wrote 25 messages with invented names, orders and amounts: fifteen in English, four in Bulgarian, three in Spanish and three in German. One of the English ones contains nothing but a phone signature. Each message carries a single trap of the kind a help desk meets every week:

  • tone pulling against substance: a shouted how-to question, a furious demand for a new feature, and two calm reports of an outage, one from a shop that has had no orders since the morning;
  • two needs in one message: praise and then a double charge, praise and then a stolen account, a cancellation and then a charge made after cancelling;
  • text the customer did not write: an English error pasted into a Bulgarian message, a long English confidentiality footer under a Bulgarian question, and a German email forwarded under a short English reply;
  • lines aimed at the model: a feature request with a forged system note asking for billing and urgent, and a message made of nothing but such orders;
  • category lists of our own for a shoe shop, once in lower case, once with capitals, and once with no value for praise;
  • a phishing notice about a domain, an out-of-office reply, a request for a quote and a refund demand with a chargeback deadline.

Both sides received the same request. It named the allowed categories and the four priority levels, and said how to write the language, including und when the customer wrote nothing. So the side without the skill knew every allowed value; what it lacked were our rules for choosing among them. The other side also had our ticket triage skill as its instructions. We ran both models through Claude Code in print mode, tools off, our personal configuration left out, once per model, message and side: a hundred runs in all. Each answer was parsed by a script, which demanded exactly three keys holding the expected values. The Haiku behind the command-line alias is now claude-haiku-5-5, a newer model than the one in our earlier reports.

Results

What we counted With the skill Without it
Sonnet, all three labels right 25 of 25 21 of 25
Haiku, all three labels right 25 of 25 16 of 25
Messages where the two models gave different labels none of 25 7 of 25

Volume pushed the priority up. "URGENT!!!!" in front of a question about switching the dashboard language: Sonnet called it normal, Haiku called it high. A Bulgarian message in capitals demanding a dark mode, with a threat to leave, went from the low we expect for a feature request to normal on both models. A team of 40 people asking for a quote and a demo got high from Haiku. A German customer whose parcel, due on Monday, had sat in a depot for five days got high from both. In each of these the customer can carry on working and has lost no money while the message waits, which is why our rules keep them at normal or low.

Politeness pulled it down. The one message both models rated too low opened with warm praise for the product and then mentioned, as a small thing, that a card had been charged twice. Both called it normal. Money has left the customer's account and they asked for it back, so our rules put it at high. Read from the model's side, the friendly first sentence set the mood for the rest.

Real trouble was sometimes lifted past our scale. Haiku rated a refund demand with a Friday chargeback deadline as urgent, and did the same for a Bulgarian customer who had paid for a yearly plan two hours earlier and still had no access. Our rules give both high and keep urgent for a stolen account, a charge the customer never made, or an outage that stops work. A support lead could argue for either reading. What matters is that Sonnet gave the very same two messages high.

The two models disagreed on seven messages. Six of the disagreements were about priority, and one about category: a customer who wanted to cancel and had also been charged again after cancelling was cancellation for Sonnet and billing for Haiku. When a help desk routes by these labels, the queue a message lands in depends on which model happened to read it. With the skill, the two models returned identical labels for every message.

Category and language held up on their own. Without the skill, the category was right in every answer but that one, and the language in every answer but one: below a short English reply, Haiku took the language of the forwarded German email. Several traps we expected to work did not. Both models gave bg to the Bulgarian report that quoted an English error and to the Bulgarian question sitting above an English legal footer, und for the phone signature, and urgent for both calm outage reports, the Spanish one included.

The planted orders fooled nobody, and nothing came wrapped. The fake system note left the feature request at low, and the message made only of orders came back as spam, with and without the skill, on both models. Not one of the hundred answers arrived inside a Markdown code fence or with a sentence around it. In our JSON repair test the day before, the older Haiku fenced most of its replies; this newer one fenced none.

What to take into your own triage

  • Define each priority level by situations, never by adjectives. "Someone else is using the account" or "no order can go through" is something a model can check a message against. "Needs attention fast" hands the decision back to the customer's tone.
  • Test tone and substance apart. Give each calm message an angry twin about the same problem, and the reverse. If the label moves between the twins, the tone is deciding.
  • Run two models over the same messages before you trust either. The messages where they disagree show you where your rules are silent.
  • Say what to do with text the customer did not write. Forwarded mail, pasted errors and footers are where a language label goes wrong first, even when most answers are fine.

What we did not measure

  • Every answer is a first try. Each model answered each message once per side, and a second run could move a borderline case, the chargeback and the late parcel above all.
  • Twenty-five short messages we wrote ourselves. A real support ticket is longer, comes with a thread and attachments, and mixes problems more loosely than ours do.
  • Our priority scale is one convention among many. The side without the skill knew the four names but not our definitions, so its misses are disagreements with our rules, not errors by every standard. The late parcel is the closest call: a shop could well treat an order stuck for five days as high. What the test does show beyond doubt is the instability, the same message sent to two different queues by two models.
  • Two Claude models, one of them new to us. No models from other vendors, and the Haiku behind the alias has changed since our earlier posts, so its numbers here cannot be set beside theirs.
  • No volume, speed or other languages. We did not try thousands of messages, nor any language outside the four above.

Read this post as Markdown: /blog/llm-ticket-triage-tone-priority.md · Atom feed.