BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/The Bot Took Notes and Left Out the Decision
ToolsAugust 12, 2026
Read · 5 min
ai meeting notes · meeting notes

The Bot Took Notes and Left Out the Decision

Meeting notes fail at three separate stages, and the symptom tells you which. A quality rubric, the prompt that recovers decisions, and the consent question.

Key takeaways
  • A meeting note passes through three stages that fail in different ways, and the symptom in the note tells you which one broke.
  • A decision credited to the wrong person is a diarisation failure, not a summary failure. Rewriting the prompt will not fix it.
  • Research on dialogue summarisation names omission as the dominant error, and disagreement is the thing summarisers smooth away first.
  • An independent measurement of eleven speech recognition services found accuracy varying widely by vendor and by sample, with streaming noticeably worse than batch.
  • Recording law is not uniform. Federal law in the United States sets a one party floor, about eleven states require all parties to consent, and in Europe you need a lawful basis under Article 6.

You read the note after the call and it is good. Clear paragraphs, sensible headings, an action list. Two weeks later the supplier says you agreed to the higher freight rate, and you go back to the note, and the note says the freight rate was discussed. It was discussed. It was also decided, in a sentence somebody said while somebody else was still talking, and that sentence is not in the summary and is not in the transcript either.

Notes fail in a small number of specific ways, and they are worth telling apart because they need different fixes and only one of them is about the summary.

What are the three stages?

Capture, diarisation, summarisation. Everything that goes wrong goes wrong in one of them.

Breakdown diagram of three meeting note stages: capture with microphones and accents, diarisation of who spoke, and summarisation

Capture is audio to words. Microphone quality, background noise, people talking over each other, accents the model has heard less of, the person on speakerphone in a car.

Diarisation is words to speakers. Who said which line. This is the stage nobody thinks about and where a surprising share of the damage originates.

Summarisation is speakers and words to a note. What was kept, what was dropped, what got flattened into a sentence that reads well and asserts less than what was said.

StageTypical failureWhat you see in the noteThe fix that works
CaptureCrosstalk, poor microphone, unfamiliar accentGarbled names, wrong numbers, missing sentences entirelyBetter hardware and one person speaking at a time
CaptureLive streaming transcriptionDegraded quality against the same audio processed afterwardsPrefer a recorded pass over a live one where you can
DiarisationSpeakers merged or swappedA decision credited to the wrong personNamed participants supplied in advance, separate audio channels
DiarisationShort interjections absorbed into the previous speakerObjections vanish because they were four words longCheck the transcript, not the summary, on anything contested
SummarisationOmissionA fluent note that never states what was decidedAsk for decisions explicitly rather than for a summary
SummarisationSmoothingTwo conflicting positions become one agreed directionAsk what was disagreed and left open, as its own question

The middle two rows are the ones worth internalising. If the note attributes something to the wrong person, no amount of prompt rewriting helps, because the model is faithfully summarising a transcript that already has the wrong name on the line.

How good is the transcription really?

Better than it was and less reliable than the marketing suggests, with the spread being the real story. An independent measurement of eleven common speech recognition services on higher education lectures found that accuracy ranged widely between vendors and between individual audio samples, and that streaming recognition, the kind used for live captions, measured significantly worse than the same audio processed afterwards.

The authors set that against published claims of human parity, and the mismatch they describe is exactly the one you experience. The average is fine. Your Tuesday call, with three people, one bad connection and a product name the model has never seen, is not the average.

Two practical consequences follow. Use the recorded transcription rather than the live one when the meeting matters, because the same tool usually does better with the whole file. And treat proper nouns as the weak point: supplier names, product codes and people's names carry no linguistic context to fall back on, which is why they are the words that come out wrong.

Note

Most tools let you supply a vocabulary of expected terms. Twenty minutes spent adding your product range, your suppliers and your colleagues' names is the single highest return adjustment available, and almost nobody does it.

Why does the summary leave out the decision?

Because omission is the dominant failure mode of dialogue summarisation, and a decision is structurally easy to omit. It is usually one short sentence, often said quietly, often followed by more discussion that looks like it superseded it.

The research names this directly. Work on omission in dialogue summarisation, presented at ACL 2023, introduced a dataset labelling which utterances a summary left out, on the basis that omission drives the quality of a generated note more than anything else and had barely been studied. Its finding is instructive: giving the model the omission labels produced a large improvement in the quality of the generated note, which means the information was there and the model was choosing not to keep it.

The second failure is subtler. TofuEval, a benchmark on topic focused dialogue summarisation, found that models produce more factually inconsistent summaries when asked to focus on a marginal topic, that existing models still make a considerable number of factual errors, and that larger models do not necessarily make fewer. Ask a note taker to summarise the meeting and it does reasonably. Ask it to summarise what was said about pricing, when pricing was three minutes of a forty minute call, and the error rate goes up rather than down.

Put those together and you get the pattern everyone recognises. Long meetings produce plausible notes. The two minutes that mattered are the two minutes most likely to be lost, because they were short and because asking about them specifically is the request that degrades most.

A rubric you can apply in ninety seconds

Judge a note against four questions rather than against how well it reads. Fluency is not the variable.

Card listing four things a meeting note must contain: decisions, owners, dates, and unresolved disagreements
  1. Is every decision stated as a decision? Not discussed, not considered. Decided, with what was decided.
  2. Does every action have a person's name on it? An action with no owner is a wish, and the note is where ownership was supposed to be recorded.
  3. Does every action have a date? Soon, next week and when we have time are all the absence of a date.
  4. Is every unresolved disagreement written down as unresolved? This is the one that fails almost every time and the one that costs the most later.

Run this on your last three meeting notes before changing anything. Most people find the first two mostly pass and the fourth almost never does.

The same call, two notes

Here is what smoothing looks like in practice. Take an exchange in a supplier call where one person says the new packaging will add about forty pence a unit, the other says that is fine if it holds up in transit, the first says they have not tested it in transit yet, and the second says to go ahead with a hundred units and check.

The typical generated note says: Discussed new packaging. Cost impact of around forty pence per unit noted. Team agreed to proceed.

The corrected note says: Decision: order 100 units of the new packaging as a trial, not a full run. Open risk: transit durability is untested. Owner: Sam. Check after first delivery, week of the 24th.

Everything in the second version was said out loud. The first version is not wrong so much as it is unusable: it records that a conversation happened, and the thing you needed recorded was the trial size, the untested risk and who checks.

What to ask for instead of a summary

Change the request from a genre to a set of fields. A summary is a genre and the model has strong priors about what a good one looks like, most of which involve being smooth and short. Fields have no such pull.

A prompt that recovers decisions asks for them separately and permits emptiness: list every decision made, quoting the sentence where it was made. List every action with an owner and a date, and write missing where either is absent. List every point where participants disagreed and say whether it was resolved. If a section has nothing in it, write none rather than inventing an entry.

That last instruction matters more than it looks. Without it, a decisions section will be populated with the nearest thing to a decision the model can find, which is how a discussion becomes an agreement. This is the same discipline as forcing a fixed output shape, and the reasoning behind it is set out in our piece on getting a model to hold a shape you can check.

The quoting instruction is the other half. Requiring the sentence where a decision was made turns an unverifiable claim into a pointer you can look up, and it visibly reduces confident inventions because there is nowhere to put one.

Why the transcript is the document that matters

Most people delete the transcript and keep the note, which is exactly backwards for anything that might be disputed.

The note is a derived artefact. It is one model's reading of what mattered, produced without knowing which of the forty topics would turn out to be the one you argue about in March. The transcript is closer to the source, and when a supplier says you agreed to something, the transcript is the only place the sentence either exists or does not.

Keeping it is not free, which is the tension. A full text archive of every call raises the retention question below and answers it badly if you never make a decision about it. The workable middle is short and deliberate: keep transcripts for the same window as recordings, and when a call contains a commercial commitment, promote the relevant exchange into the written record on purpose rather than relying on an archive search two months later.

There is a second reason, and it is about calibration. Reading one transcript against one generated note, for a meeting you remember, is the only way to find out what your particular setup drops. Vendor accuracy claims will not tell you, and neither will the note, because a note that lost something reads exactly like a note that did not.

Who consented to being recorded?

Nobody asks this until something goes wrong, and the answer is jurisdictional rather than a matter of etiquette.

In the United States, the Reporters Committee's guide to recording law sets out the structure. Federal law is one party consent for in person conversations, phone calls and electronic communications, so that is the national floor. About eleven states primarily require all parties to consent, and it names them: California, Delaware, Florida, Illinois, Maryland, Massachusetts, Michigan for third party recordings, Montana, New Hampshire, Pennsylvania and Washington. Four more split by conversation type, with Missouri and Oregon requiring all party consent in person but one party by phone, and Connecticut and Nevada the reverse.

Read that as a practical instruction rather than a legal briefing. If you sell across state lines, you cannot know which rule applies to a given call before it starts, so the workable policy is to announce recording every time and let people object.

In Europe the question changes shape. A recording of an identifiable person is personal data, so you need a lawful basis, and Article 6 of the GDPR lists six of them: consent, performance of a contract, a legal obligation, vital interests, a public interest task, and legitimate interests, with legitimate interests subject to being overridden by the individual's own rights. Most small business recording sits on consent or legitimate interests, and either way the person has to know it is happening.

How long should the archive live?

Shorter than the default, which is usually forever.

A searchable archive of every internal conversation is a different object from a set of meeting notes. It is a record of unfinished thinking, half positions and things people said before they knew the facts, indexed and full text searchable by anyone with a login. In a dispute, it is discoverable.

Three settings handle most of it. Set a retention window on recordings and raw transcripts, ninety days being a common choice, while keeping the finished notes for as long as you keep other business records. Decide who can search the archive, and make the answer narrower than everyone. And write down what happens when someone leaves, because their calls do not leave with them. If you are writing this into a document, the section fits naturally into an AI policy built around data classes rather than living as a separate rule nobody finds.

Questions people ask

Should I let three different bots into the same call?

No, and it is a common accident rather than a choice. Each attendee's tool joins on their behalf, so a four person call can carry three recorders, three archives and three retention policies, two of which you have never read. Agree one at the start of the call.

Does a better model fix bad audio?

Not much, and it is the wrong place to spend. A decent microphone and a rule about not talking over each other move accuracy further than any model change, because the errors originate in the signal. This is the same reason live transcription measures worse than a recorded pass: less information, made worse by having to decide immediately.

Can I trust the action items?

More than the narrative, less than you would like. Actions are the easiest thing for a summariser to extract because they are usually phrased distinctively, with a name and a verb. The failure is not usually a wrong action, it is a missing one, and specifically the one agreed in the last two minutes when everybody was already leaving.

What about voice agents that talk back?

Different problem, same stack underneath, with latency added on top. If you are considering one for customer calls rather than for internal notes, our piece on real time full duplex voice models covers what changes when the system has to respond rather than only listen.

What to change on Monday

Three things, none of which involve switching tools. Add your vocabulary of names, products and suppliers to whatever you use. Change the standing prompt from summarise this meeting to the four field version above. And read the decisions section against the recording once, this week, for a meeting you remember well, so you learn what your particular setup drops.

The last one is the one that changes behaviour, because the failure is invisible until you check it yourself. Everyone believes their notes are fine, in the same way everyone believes their spreadsheets are fine, and for the same reason: the errors do not announce themselves. If any of this touches customer conversations, our privacy policy sets out how we handle data on our own side, which is the standard worth holding a note taking vendor to as well.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building