Test results
Subtitle Translator for SRT and VTT: test results
Tested 2026-10-08, skill version 1.0.0 at the time of loading this page. We run every skill on a strong and a weak model before it is listed, and publish both verdicts, including where the weak one fails.
Verdicts
- Date
- 2026-10-08
- Strong · claude-sonnet-5-5 (Claude Code alias "sonnet")
- Right in 23 of 23 cases, read by hand: cue numbers, every timing line and every VTT setting came back byte for byte, including overlapping cues, a zero-length cue and numbering with gaps; NOTE and STYLE blocks stayed English; italic, bold, font and position codes kept their place; voice names and dialogue dashes survived; a bracketed sound description kept its brackets with translated words; a German cue that does not fit was cut down to two lines of at most 42 characters; a caption telling the translator to answer with one word was translated as speech.
- Weak · claude-haiku-5-5 (Claude Code alias "haiku")
- Right in 22 of 23 cases and kept NOTE and STYLE blocks, timing lines and tags intact where it had not before, but in one long German cue it left a line of 54 characters, so that file would not pass a two-line, 42-character check.
With and without the skill
Tested 2026-10-08.
| Sonnet | Haiku | |||
|---|---|---|---|---|
| with | without | with | without | |
| Subtitle file intact, translated, within 2 lines of 42 characters (23 cases) | 23/23 | 22/23 | 22/23 | 21/23 |
The same request on both sides, with the line limit stated in it, so the bare model was told the rule. Essentially no gain: Sonnet kept timings, tags, dashes and brackets without help. Its one miss without the skill was a German cue it left at a 47-character line, read by hand as a real miss. With the skill that cue became two short lines. Haiku gained one case, a file with NOTE and STYLE blocks that it altered without the skill, and still left a 54-character line in the long German cue with it. The skill buys the shortening discipline and block handling, not structure Sonnet already kept.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
Note
Twenty-three subtitle files written by us, translated into Bulgarian, German, Spanish, French, Portuguese, Russian and English: plain film dialogue, one-word cues, song lyrics between music notes, overlapping and out-of-order timing, bracketed sound descriptions, two speakers with dashes, position and font tags, an ampersand entity, times and codes, two captions that give orders to the translator, and four WebVTT files with cue settings, voice names, identifiers, NOTE and STYLE blocks and timings without hours. The request on both sides asked for at most two lines of at most 42 characters. The SRT answers are scored by a program: same number of cues, same numbers, timing lines byte for byte, no more than two lines of 42 characters, translated text. The WebVTT answers are scored by exact-text and pattern checks. One run per model and case.
What was not measured
- Models other than the two named above were not run.
- Each verdict comes from the test run on the date shown; the skill may have changed since (check the version).
- Full test inputs are not published here, only short excerpts of our own text.
- Results on your own texts, languages and domains can differ.