After a client meeting wraps up, you skim the AI-generated notes and paste them straight into your team's Notion. You've probably done it again this week. While skimming, your eyes are really just checking whether the sentences read smoothly. If they do, you relax — and once you relax, you hit publish. Even if the notes contain a dollar figure the client never mentioned, or the name of someone who wasn't even in the room, a smooth sentence won't catch your eye.
This habit rests on a single belief: a human checks it at the end, so it's fine. AI scribes in medical settings run on the exact same promise — they listen to the consultation and draft the note, but nothing is final until the doctor signs off. A new study put a number on how well that promise actually holds up, and it's worth a look.
142 Consultations, Three Tools, One Audit
The researchers audited three commercially available AI medical-scribe tools under identical conditions. They assembled 142 cases — real consultations recorded in UK primary care and US outpatient clinics, plus scripted scenarios the research team wrote themselves — and collected 565 notes the three tools produced from them.
The method for finding errors is the heart of this study. The team ran twelve separate passes to cast a wide net for candidate errors, then passed only the ones that cleared a significance filter to two models from different model families. Those models weren't asked to confirm the errors — they were asked to argue against them, to find evidence that each candidate wasn't actually a mistake. Only candidates that survived that pushback were counted as confirmed errors. Finally, two physicians re-judged a sample, blind to the AI's verdict.
One in Three — and Facts Invented from Thin Air
After all that filtering, 31.3% of the notes contained a confirmed error (95% CI: 27.0%–35.6%). Tracing how the candidate count shrank down to that final figure gives a sense of just how rigorously this number was earned.
Of those, 618 errors survived. When one of the study's own physician-authors reviewed a sample, 20 of 21 held up; an independent clinician not involved in writing the paper confirmed all 12 of a separate sample.
What matters more is the type of error. The researchers report that errors clustered in three places: allergy and medication information; fabricated patient identities; and cases where a phone consultation — which by definition can't include a physical exam — got written up with exam findings anyway. Two of those three categories aren't omissions; they're details invented out of nothing. Since none of the tools were connected to actual patient records, the researchers treated identity and date errors as fields that would have been filled in correctly had a record been linked — and even after excluding those two categories, 24.8% of notes still had an error. One error type didn't fit any existing classification at all: treatment a doctor considered and then decided against got recorded as treatment that was actually given.
Also worth noting: the error rate depends as much on how you review as on which tool you use. Holding the model, evidence, and setup constant and changing only the review instructions swung the rate at which candidates got confirmed as errors from 9.3% to 79.0%. Depending on which standard you apply, the share of notes containing an error ranges anywhere from 28% to 97%.
From Skimming to Checking a List
This study is about medicine, but the structure translates directly to solo founders and planners who hand meeting notes and summaries over to AI. A skim-for-smoothness review is bad at catching omissions and worse at catching fabrications — because a fabricated sentence, by nature, reads just as smoothly as a true one.
Checking AI meeting notes for errors, then, has to mean checking against a defined list, not skimming for tone. Mapping this paper's error clusters onto a business document gives you exactly that list. The allergy-and-medication equivalent is numbers that cost you money if wrong: amounts, dates, quantities. The fabricated-identity equivalent is the names of attendees, the person responsible for a task, and company names. The phone-consultation-written-as-exam equivalent is a sentence that describes something as done when it was only discussed, or as decided when it was only reviewed. And the withdrawn-treatment case maps onto a proposal that came up, got shelved, and somehow ended up recorded as a final decision.
Another takeaway is to write your review criteria down. Without an agreed definition of what counts as an error, results vary from one reviewer to the next — sharing a written checklist for catching AI meeting-note errors across the team narrows that gap.
Before You Borrow This Number
This study looked at medical records, not meeting notes or business documents — those weren't tested directly. The 31.3% figure is specific to three tools, 142 consultations, and the review criteria the researchers themselves chose. The researchers emphasize that changing the review instructions alone moves the result substantially, so this number shouldn't be read as the error rate for AI summarization tools in general.
So what's worth taking from this paper is the method, not the percentage. The researchers published all 618 errors and the evidence behind each one, the prompts and model versions they used, and a pipeline anyone can rerun. For "a human checks it at the end" to actually function as a safeguard, what gets checked and how has to be defined first. Before you paste your next set of meeting notes into Notion, try cross-checking just three lines against the original: the numbers, the names, and the decisions.



