Strong model against weak model, how we ran each test, and what we did not measure.
We ran eight SKILL.md files on a strong and a weak model. Both mostly passed our checks; the weak one failed in ways checks cannot see, and used more tokens.