The FDA has opened a useful conversation about generative AI in medical devices. Its new discussion paper considers a physician-inspired approach to competency assessment, using non-clinical benchmarking and clinical confirmation before a device reaches patients.1
That makes sense. If an AI system is going to summarize a chart, interpret an image, recommend a next step, or generate clinical content, it should demonstrate that it can perform the assigned task.
But medicine does not use models in isolation.
A model can answer a benchmark correctly and still fail inside a clinic. The problem may be the interface, the information available to the model, the way the answer is displayed, or the few seconds a clinician has to review it between patients. A fluent answer can hide an omitted allergy, a stale wound measurement, or a vascular study buried on page seven of an outside referral packet.
A competent model can still sit inside an unsafe clinical system.
A benchmark is not a clinic
Clinical work is full of incomplete and conflicting information. The wound photograph was taken yesterday, but the measurement came from last week. The medication list says one thing and the patient says another. A copied-forward assessment no longer matches the current exam. The home health order describes a dressing that nobody is using.
A medical AI device has to function in that environment, not in a clean prompt assembled for testing.
The FDA's proposal recognizes that premarket testing cannot stop with a technical benchmark. Clinical confirmation is part of the discussion. The agency is also asking about risk assessment, foundation models, agentic systems, and postmarket monitoring.12
The next step is to define what clinical confirmation should actually measure.
Accuracy matters, but so do omission, calibration, reviewability, and recovery. Can the clinician see which source supported the answer? Does the system make uncertainty obvious? What happens when two records conflict? Can a busy user correct the output without rebuilding the whole note? If the system takes an action, is there a visible record of what it did and why?
Those interface choices determine whether the clinician can use the tool safely.
Test the human review, too
Most medical AI products still rely on some version of human oversight. The physician remains responsible. That phrase sounds reassuring until the product is placed in a real workflow.
Clinicians do not review every output with the same attention. Review changes with time pressure, alert fatigue, interface design, staffing, and the confidence of the output. A polished paragraph may receive less scrutiny than a visibly uncertain one. A recommendation placed inside the normal order screen may feel more authoritative than the same text shown in a separate panel.
If human review is part of the safety case, then human review belongs in the evaluation.
A useful clinical test would not ask only whether the AI produced the right answer. It would ask whether clinicians using the system reached the right decision, noticed the important errors, and recovered when the model failed. It would also measure whether the system saved time or simply moved the work into another part of the day.
That is closer to the question clinicians care about. Does this tool improve care under ordinary conditions, with the ordinary friction of practice?
The model will change
Generative systems also create a lifecycle problem. Their behavior can shift when the foundation model changes, when a retrieval source is updated, when prompts are modified, or when a new tool is connected. A device that performed well at launch may behave differently six months later even if the visible interface looks the same.
The FDA discussion paper includes risk-proportionate postmarket monitoring and specifically raises foundation models and agentic AI systems.1 That matters because one-time clearance cannot answer every question about a system whose components continue to evolve.
Postmarket monitoring should look beyond average accuracy. It should identify where performance is drifting, which patient groups are affected, what kinds of errors clinicians miss, and whether updates change the balance between automation and oversight.
For a documentation tool, that might mean tracking unsupported additions, clinically meaningful omissions, and correction rates. For an imaging tool, it might mean monitoring performance across devices, sites, and patient populations. For an agent that can initiate an action, the standard should be higher: every step should be attributable, reviewable, and reversible when the clinical context allows it.
Competency should apply to the whole system
The physician analogy is useful, but it has limits. We do not judge physicians only by an examination score. We evaluate them in supervised practice, within a team, with access to records, escalation pathways, and accountability.
Medical AI deserves that broader evaluation.
A regulatory evaluation should identify the model's assigned task, test its technical performance, and then test the complete clinical system around it. That includes the data entering the model, the way results reach the clinician, the expected human review, and the monitoring that follows deployment.
The FDA is asking for public input rather than announcing a finished policy. The agency states clearly that the paper is for discussion and does not represent draft or final guidance.2 Comments are open through October 19, 2026, under docket FDA-2026-N-7874.2
Clinicians should take that invitation seriously. Device manufacturers understand model architecture. Regulators understand evidence and risk. Clinicians know where apparently reasonable systems break during an actual day of care.
That experience belongs in the regulatory record.
Testing a medical AI like a physician gives regulators one useful checkpoint. The stronger test asks whether the entire clinical system still works when the waiting room is full, the chart is messy, and the answer sounds more certain than it should.
Sources
- U.S. Food and Drug Administration. FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices. August 18, 2026.
- U.S. Food and Drug Administration. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. August 18, 2026.