Missing Is Not Negative

A blank space in a medical record may mean nobody asked, nobody looked, the answer lives somewhere else, or the answer was never documented. It does not mean the finding was normal.

A clinical record showing separate states for absent, unassessed, and negative information
Clinical documentation should keep absent, unassessed, and negative information separate instead of smoothing all three into the same reassuring sentence.

A generated note can sound better than the material it came from. That is part of the appeal. A scattered conversation becomes a clean history. A long chart becomes three paragraphs. The rough edges disappear.

Sometimes one of those rough edges is uncertainty.

In clinic, I may know that a patient denies fever. I may also know that fever was not discussed. Those are different facts. I may examine pedal pulses and find them absent, or I may never document a vascular examination at all. Those are different records. If a system turns both situations into “no vascular concerns,” the prose is fluent and the chart is wrong.

This is a basic problem with generated clinical language: it prefers a complete sentence even when the source is incomplete. The result can look like good documentation while quietly changing the evidence underneath it.

Three states that cannot be collapsed

Absent from the sourceThe record says nothing about the finding, and the reason for that silence is unknown.
UnassessedThe record explicitly says the question was not asked, the examination was not performed, or the assessment could not be completed.
NegativeA specific test, history question, or examination produced a documented negative result.

Even the word absent needs context. An absent symptom reported by the patient is not the same as an absent physical finding. “No drainage” may mean the patient denied drainage, the dressing was dry when removed, or the author copied a prior statement. Each has a different source and a different level of confidence.

Health-record researchers have been warning about this problem for years. EHR completeness is not one fixed property. A record may be complete enough for one purpose and inadequate for another. Missingness is often informative because what gets measured and documented depends on the patient, the clinician, the setting, and the clinical question. Treating every blank field as though it carries the same meaning is a methodological error before a language model ever sees the chart.

Silence in the source should remain silence in the draft until a person or a linked piece of evidence resolves it.

How polished prose erases the distinction

The problem rarely arrives as an outrageous fabrication. It usually arrives as a small connective phrase.

A transcript mentions that the patient is taking one medication. The generated note says the patient is “tolerating all medications without side effects,” even though tolerance was never discussed. A prior note documents an ulcer debridement. The current draft says the wound “was debrided today,” because the system blended history with the present encounter. A patient denies calf pain, and the assessment becomes “no evidence of deep venous thrombosis.” The first is a symptom report. The second is a diagnostic conclusion.

Generated prose also tends to resolve pronouns, chronology, and causality. “She stopped it” becomes a named drug. “He had surgery” becomes a specific procedure. “The dressing helped” becomes an assertion that the treatment improved the wound. Each edit makes the paragraph easier to read. Each may add a fact that was never established.

This is why a note can be grammatically improved and clinically degraded at the same time.

Where the risk becomes more than clerical

Diagnoses

A diagnosis should not appear merely because the source contains a related symptom, an old rule-out, or a test that was ordered. “Concern for osteomyelitis” is not osteomyelitis. “Biopsy pending” is not malignancy. “No shortness of breath documented” is not a negative cardiopulmonary review of systems. Once a generated diagnosis enters the assessment or problem list, it can influence later clinicians, risk adjustment, patient messages, and future model outputs.

Medications

Medication lists are full of unresolved states: prescribed, dispensed, started, stopped, held, completed, taken differently, or no longer remembered by the patient. A clean list can hide those distinctions. A safe draft should say when adherence was not assessed, when the source records a plan rather than confirmed use, and when two records conflict. “Continue” is a consequential verb. It should have a source.

Procedures

Procedure documentation needs the same discipline. A planned injection is not a completed injection. Consent discussed is not consent obtained. A template that contains anesthetic, laterality, dimensions, or tolerance language cannot establish that those details occurred today. If a required element is missing, the draft should expose the gap rather than fill it from habit or a previous note.

Medical necessity

Medical necessity is particularly vulnerable to fluent overreach because the final paragraph often has to connect findings, risk, and treatment. A model can easily produce the connection even when the record does not. That may create a persuasive explanation with no documented examination, failed conservative care, functional limitation, or risk factor behind it. The right response to a missing support element is not better rhetoric. It is a visible question for the clinician.

Clinical AI should show its work at the claim level

A source link should not mean a citation to the whole chart. It should mean that the clinician can select a sentence and see the exact transcript span, note section, result, image report, or medication event that supports it.

The FDA’s clinical decision support guidance uses a related principle: a health professional should be able to independently review the basis for a recommendation. Generated documentation deserves the same practical standard. If a sentence may affect diagnosis, treatment, coding, or follow-up, its basis should be available without a scavenger hunt.

That design changes the review task. Instead of asking a physician to reread a polished note and somehow notice what was invented, it allows the physician to inspect the claims that carry the most risk.

At minimum, I would want four visible states:

Visible uncertainty is not a failure of the interface. It is an accurate description of the chart.

Safe abstention belongs inside the note workflow

In a consumer chatbot, “I don’t know” can feel unsatisfying. In clinical documentation, it may be the safest answer available.

Abstention does not have to stop the work. The draft can preserve the unresolved language, highlight the sentence, and ask a focused question: Was the patient taking the medication? Was the procedure performed today? Was infection assessed? What finding supports medical necessity?

The important part is that the system does not answer its own question.

NIST identifies confabulation as a distinct generative-AI risk and emphasizes testing, evaluation, verification, and validation across the system lifecycle. In medicine, one of the most useful tests is whether a system becomes less certain when its source becomes weaker. A model that always produces a complete note has failed that test.

Validation has to follow the error all the way to the chart

Recent evaluations give us a better vocabulary for measuring generated clinical notes. One 2025 framework used 12,999 clinician-annotated sentences and reported both hallucination and omission rates. A 2026 health-system pilot found accidental omissions in 18% of evaluated notes, hallucinations in 11.5%, and accidental inclusions in 9.3%. Most errors were not severe, but 5.3% of the evaluated notes contained an error rated as posing serious or imminent risk if left uncorrected.

Those are useful measures. They are not the finish line.

A documentation system should be evaluated at three points: the initial draft, the clinician-approved note, and the downstream record. The final stage matters because a human review requirement does not prove that every error was caught. In the same pilot, physician editing varied widely, and 14.9% of a larger sample of AI-created notes were left entirely unedited.

I would track at least these outcomes:

The most important rate may be the one that is easiest to overlook: unsupported clinical claims that survive physician approval. That is the point where a model error becomes part of the medical record.

Preserving uncertainty is patient safety

Diagnostic uncertainty is not an unusual defect in medicine. It changes over time as symptoms evolve, records arrive, tests return, and treatment reveals more information. A useful clinical note carries that uncertainty forward so the next clinician knows what was established, what was considered, and what still needs attention.

Generated language should help organize that work without pretending it is finished.

If a patient was not asked about a symptom, say it was not assessed. If two medication lists disagree, show the conflict. If a procedure detail is missing, leave it missing and ask. If the medical-necessity argument lacks a supporting finding, do not manufacture the finding.

That may produce a less elegant draft. It will produce a more honest chart.

The safest clinical AI will not be the system that writes the smoothest note. It will be the one that knows when the record has not earned a sentence.

Sources

  1. Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research. https://pubmed.ncbi.nlm.nih.gov/22733976/
  2. Defining and measuring Completeness of electronic health records for secondary use. https://pmc.ncbi.nlm.nih.gov/articles/PMC3900159/
  3. Informative missingness in electronic health record systems: the curse of knowing. https://doi.org/10.1186/s41512-020-00077-0
  4. Defining and Measuring Diagnostic Uncertainty in Medicine: A Systematic Review. https://pmc.ncbi.nlm.nih.gov/articles/PMC5756158/
  5. A framework to assess clinical safety and hallucination rates of large language models for medical text summarisation. npj Digital Medicine. 2025. https://doi.org/10.1038/s41746-025-01670-7
  6. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics. 2026. https://medinform.jmir.org/2026/1/e86474
  7. U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff. https://www.fda.gov/media/191561/download
  8. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
Back to blog