Hand Written Medical Record Data Extraction: The Crestor vs. Carotid Case

A 140-page patient dossier came into our platform recently. Much of it was scanned, and some of the most important pages were handwritten notes that would slow down any reader.
It’s one of the clearest examples I’ve seen of how medical record data extraction should behave when the software is uncertain. It also shows why pulling every data source into one record only helps if each data point in that record can be trusted.
What the platform read
On one page, the platform extracted “Crestor 10 mg.” The read made sense. The rest of the record showed signs of hyperlipidemia, and 10 mg is a common rosuvastatin dose.
A second entry was harder. It started with a “C,” trailed into a scrawl, and listed three years: ’19, ’21, ’22. Its shape resembled the Crestor entry. Software built to pick the most probable answer would have filed it as a second Crestor reference, maybe a refill history across three years.
The platform flagged the conflict and involved the physician, with a link to the source page. When the physician opened the original, the word was “carotid.” The three years were carotid imaging studies.

The patient was never on Crestor.
Why handwriting is the smaller problem
Transcription has always been where these errors enter the record. A study at a Swiss university hospital found documentation errors in 3.5% of prescribed agents, and 43% of patient charts had at least one. Most errors happened when handwritten prescriptions were copied into the chart.
Software will keep getting better at reading handwriting, and it will still miss. What matters for decision support is what happens after an uncertain read. Three design decisions shaped this outcome.
Safeguard 1: confidence checks in medical record data extraction
Ingestion is the step where unstructured material becomes structured data: handwritten notes, faxed scans, photographed lab printouts. Optical character recognition (OCR) and natural language processing (NLP) give each extraction a confidence score.
A smudged word, an abbreviation with two meanings, or a word shaped like a drug name all lower that score. Research on handwritten prescriptions describes the same failure: OCR confuses medicine names that differ by a letter or two.
Our standard is pre-inference validation. An extraction below the confidence threshold goes into a review queue until a physician confirms or corrects it. While it waits, it can’t populate a medication list or shape suggested interventions.
Without that gate, an ambiguous scribble becomes a medication record, and every later step treats it as fact.
Safeguard 2: escalation when two readings conflict
The second safeguard is signal-conflict detection. Here the platform had two possible readings, with some evidence for each. The hyperlipidemia context made Crestor plausible. The format, a word followed by three separate years, fit poorly with a medication line.
Software that settles conflicts on its own will usually take the more probable answer and move on. Uncertainty quantification measures how close the competing readings are. An escalation threshold sets the point where that gap is too narrow to settle automatically. Below it, the physician reviews and decides.
In our architecture, guardian agents check other agents’ outputs for this kind of disagreement. It’s the practical side of the deterministic vs. probabilistic distinction we’ve written about before. Decision support should show its uncertainty to the physician. A silent error is the expensive kind.
Safeguard 3: source-document provenance
The flag alone would have left the physician guessing. What resolved it was one click. Every field the platform extracts links to the exact spot in the original PDF it came from. The physician looked at the handwriting in context and read “carotid.”
Source-document provenance makes extraction auditable. A clinician can check any structured field against the primary record without opening another system or asking someone to pull the chart.
If a field can’t be traced to where it came from, it doesn’t belong in the dossier. It’s one of the first things worth checking when you evaluate a clinical data integration tool.

What the wrong answer would have cost
Medication reconciliation research has a name for this error: a commission, where a drug is recorded that the patient isn’t taking. In a 2026 study of admission reconciliation, commissions made up 9% of unintentional discrepancies.
Follow this one forward. A statin that was never prescribed enters the longevity protocol. The plan starts reasoning about statin-associated muscle symptoms, liver enzymes, and on-treatment lipid targets.
Meanwhile, the real signal drops out of view: three carotid imaging studies across four years. Repeat carotid imaging points to a vascular risk conversation about plaque burden, progression, and why it was being followed in the first place. That’s a different visit with a different plan.
Key takeaways
The record stayed accurate because each layer assumed the layer before it could be wrong. Three questions are worth asking any decision support vendor:
- What happens to an extraction your system isn’t confident about?
- When two interpretations conflict, who decides?
- Can I click from any data point to the original page it came from?
FAQs
What is medical record data extraction?
It’s the conversion of unstructured documents (scans, faxes, handwritten notes, PDFs) into structured fields such as medications, diagnoses, and lab values that software and physicians can work with.
How accurate is OCR on handwritten clinical notes?
It varies widely by engine and by handwriting. Treat any extraction from handwriting as provisional until it clears a confidence threshold or a physician confirms it.
What should happen when software finds two possible readings of the same entry?
It should flag the conflict, show both readings, link to the source, and leave the call to the physician.
What is source-document provenance?
Every structured field links back to the exact page and location in the original document, so a clinician can verify it in one click.
Does this slow physicians down?
Only on entries the software is unsure about. A flagged entry takes a quick look at the source page. A wrong medication that reaches the plan is much harder to catch later.



