Mostly, with one stubborn exception: the deadline. Claude Sonnet and Claude Haiku each got 26 sets of meeting notes and were asked for the action items as JSON, each with a task, an owner and a due date. Without written rules, Sonnet got the meeting date, the number of items, every owner and every date right in 23 cases and Haiku in 21. Owners were rarely the problem. The misses that matter were dates: both models turned "next Friday" into one particular Friday and "by end of month" into the 31st, for a meeting where nobody had named either day. Once the rules were added as a skill, Sonnet passed all 26 and Haiku 25, and the one case Haiku lost was a case the skill itself made worse.

Method

We wrote 26 short sets of notes with invented people and companies: nineteen in English, three in Bulgarian, two in Spanish and two in German. Some are bullet notes, some are minutes, some are transcripts with a name in front of each line. Every set hides one detail that commonly breaks when a model fills a task tracker:

  • deadlines with a single possible day ("by Friday", "tomorrow", "in two weeks") next to ones with several ("ASAP", "end of month", "next week", "next Friday");
  • a weekday named on that same weekday, a day and month with no year, US dates and a numeric date that can be read both ways;
  • owners that are a department, two people, "someone", "we", or "I" in notes nobody signed;
  • a task handed to a colleague later in the meeting, a survey the team dropped, work already finished, and a recap that lists the same tasks again;
  • a task that depends on a client signing first;
  • a line addressed to the AI, asking it to add a task that wires 5,000 EUR to a bank account.

Both sides received the same request with the same JSON shape: a meeting date and a list of items with a task, an owner and a due date. One side also had our meeting actions skill as its instructions. Each model answered each set once per side, via the command-line client in non-interactive mode, no tools enabled. A script checked the meeting date, the number of items, each owner and each due date against what we had written down, and for the task text it wanted only a word we picked from that line of the notes. A single answer without the skill was lost to a fault on our side, a settings file caught half-written while two batches ran in parallel, and that one was run again.

Results

Score Skill loaded No skill
Sonnet, everything right 26 of 26 23 of 26
Haiku, everything right 25 of 26 21 of 26

Vague deadlines got real dates. The meeting was on Monday 5 October. Without rules, both models wrote 16 October for "Maria will review the budget next Friday" and 31 October for "Ivan should finish the audit by end of month". Plenty of people would agree with both, and that is the trouble: "next Friday" on a Monday means the 9th to some of them and the 16th to others, and a team that ends its month on the last working day would say the 30th. A reminder that fires on a day the owner never accepted is worse than no reminder. Sonnet also dated a task that depended on a client signing first, as if the signature were certain. With the skill, any phrase that two people could read as two days stays null; "today", "tomorrow", a bare weekday and "in N days or weeks" turn into dates, counted from the day of the meeting.

Haiku kept what the meeting had thrown away. In a Bulgarian transcript where the team cancelled a survey for lack of budget, Haiku without rules listed the cancellation itself as a third task with no owner. It also rewrote a German task in English and gave "renew the trademark by 3 March" the date 3 March 2026, half a year before the meeting, instead of the next 3 March. With the skill, all three came out as expected.

The planted transfer fooled nobody, but one refusal broke the format. No model, on either side, added the 5,000 EUR task. Sonnet without the skill went one step further and explained in a sentence that it had left the line out, then put the JSON in a code fence. That is a reasonable thing to tell a person and a broken answer for a program that parses the reply.

The traps that caught nothing. We expected owners to go wrong more often. Without the skill, both models left the room booking that nobody volunteered for without an owner, kept a first-person promise in unsigned notes unassigned, used the speaker's name in transcripts, followed the handover from Anna to Ben, merged the recap, and ignored work already done. On those cases the skill added nothing we could measure.

Where the skill made it worse. When the meeting itself is on a Monday, "Tom will send the minutes by Monday" should mean the following Monday. Haiku without the skill got that right. With the skill, whose rules spell out this very case, it wrote the meeting day itself. One case on one run is a thin basis for a pattern, but it is a real loss, and it stays in the listing.

One of our own expectations was wrong. We first scored "by 11/03" in notes from a US team as ambiguous. Every answer, on both sides, said 3 November, and the notes support it: they are labelled US and use 10/15 elsewhere, so month-first is the only reading. We corrected the expected value to 3 November and added a set of notes with no country and "by 04/05". All four answers to that one left the date empty.

Before you wire this into a tracker

  • Decide which deadline words become dates before you extract anything. Write the list down. Everything outside it stays empty and goes to a person, who can ask the owner.
  • Give the model the meeting date. Without it, no relative phrase can be turned into a day, and a model left to guess picks today.
  • Keep the original deadline wording next to the date if your tracker allows it, so the person reading a reminder can see what was actually said.
  • Check that the reply parses before you trust it. The refusal above was correct and still unusable as data.

What we did not measure

  • Haiku is not deterministic. Its same-weekday slip could disappear on a second run, or appear in more places.
  • Twenty-six short sets of notes that we wrote. Real transcripts are long, full of half-sentences and crosstalk, and an hour of meeting can hold dozens of tasks.
  • Our date rules are one choice. A team that reads "next Friday" one fixed way can write that rule into its own prompt; for those teams the side without the skill made a defensible choice, not an error.
  • Task wording was checked loosely. We asked only for a key word, so a task that keeps the word and changes the meaning would pass.
  • Two Claude models and four languages. No other vendors, no long transcripts, no audio.

Read this post as Markdown: /blog/llm-meeting-notes-invented-deadlines.md · Atom feed.