Test results

Ticket Triage: Category, Priority, Language: test results

Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.

Verdicts

Date
2026-10-08
Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
Gave all three expected labels in 25 of 25 cases, always as bare JSON with the three keys. It kept a shouted how-to question at normal, raised a politely worded double charge to high, filed a cancellation that came with an unwanted charge under billing, kept an angry demand for dark mode at low and rated a calm report of a shop that takes no orders as urgent. It read the language from the customer's own words: Bulgarian beside a pasted English error and above an English legal footer, English above a forwarded German email, und for a message holding only a phone signature. It ignored an order planted inside a feature request, labelled a message made only of such orders as spam, and used a custom category list exactly as spelled, including Returns and Other with a capital letter.
Weak · claude-haiku-5-5 (Claude Code alias "haiku")
Also 25 of 25 with the skill, with the same labels as Sonnet in every case and no code fence or text around the JSON. The rules changed its answers most on priority: without them it let capital letters, an angry tone, a sales request and a late parcel push priority up, and took the language of a forwarded German email; with them it followed the fixed levels in each of those cases.

With and without the skill

Tested 2026-10-08.

Results with and without the skill, for Sonnet and Haiku
SonnetHaiku
withwithoutwithwithout
All three labels right25/2521/2525/2516/25

One run per model and case. The side without the skill got the same category list, priority levels and language rule in the request, so it measures the written rules, not knowledge of the list. Without them, 12 of the 13 failed answers had the wrong priority: capital letters, anger, a sales request and a late delivery raised it, while praise around a double charge lowered it. Category and language were almost always right on both sides, and neither side wrapped the JSON in a fence.

Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.

Note

Twenty-five fictional messages written by us, in English, Bulgarian, Spanish and German, each built around one trap: tone against priority, a calm outage, a pasted error or a footer in another language, a forwarded email, an order planted for the model, two requests in one message, a cancellation with a wrong charge, custom category lists and a message with no customer text. The check parses the answer and requires exactly the three keys with the expected values. One run per model and case.

What was not measured

  • Models other than the two named above were not run.
  • Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
  • Full test inputs are not published here, only short excerpts of our own text.
  • Results on your own texts, languages and domains can differ.

Back to Ticket Triage: Category, Priority, Language · Card (JSON)