The FDA has floated evaluating generative AI medical devices the way physicians are trained and evaluated, and is asking twenty-six numbered questions about whether that would work

FDA Opens Comment on Generative AI Medical Devices. The Leveraged Years regulation briefing card.

A discussion paper binds nobody and changes no obligation. What it shows is which questions the agency has not settled, and how long the comment record stays open.

The short version

Bottom line: This binds no one. The paper carries four separate disclaimers on its own first page. It is "intended for discussion purposes only and does not represent draft or final guidance", it is not intended to propose or implement policy changes, it is not intended to communicate CDRH's proposed or final regulatory expectations including expectations for supporting evidence in future marketing submissions, and it does not address whether any of it is within the FDA's existing legal authorities.

Who this affects: Physicians and clinical governance leads who may be asked to supervise generative AI device output, and regulatory affairs teams at device sponsors building on third party foundation models.

Issue date: Issued 18 August 2026. Comments are invited under docket FDA-2026-N-7874 on Regulations.gov by 19 October 2026. That is a comment deadline, not a compliance deadline.

What changed: Nothing binding. What is new is the disclosure of how CDRH is currently thinking: a possible two axis risk framework, a competency based premarket approach with an enumerated set of benchmarking elements, candidate postmarket monitoring approaches, and a voluntary Foundation Model Device Master File.

Analysis: The analogy is the substantive proposal. CDRH is considering evaluating the final user facing device rather than the model, against competencies, in a way it says is inspired by how clinicians are trained. If that survives, it changes what evidence a submission carries. The paper expressly says it is not communicating expectations for supporting evidence in future marketing submissions, so it has not survived anything yet.

Primary sources: Discussion paper (PDF), FDA · FDA press announcement, 18 August 2026 · Docket FDA-2026-N-7874

Instrument (EN)
Discussion paper and request for feedback, Considerations for the Regulation of Generative AI-Enabled Medical Devices
Authority
US Food and Drug Administration. Led by the Digital Health Center of Excellence within the Center for Devices and Radiological Health (CDRH)
Jurisdiction
United States
Status
Open for comment under docket FDA-2026-N-7874, per the FDA press announcement. The discussion paper itself announces no rulemaking
Bindingness
None. The paper states it does not represent draft or final guidance, is not intended to propose or implement policy changes, is not intended to communicate CDRH's proposed or final regulatory expectations including its expectations for supporting evidence in future marketing submissions, and does not address whether the approaches discussed are within the FDA's existing legal authorities or whether new authorities would be needed
Docket
FDA-2026-N-7874 on Regulations.gov
Issue date / next deadline
Issued 18 August 2026. Comments due 19 October 2026
Primary source
https://www.fda.gov/media/194242/download

What the FDA issued, and the four things it says this is not

On 18 August 2026 the FDA issued a discussion paper on considerations for the regulation of generative AI enabled medical devices, seeking feedback on risk assessment, premarket evaluation, postmarket monitoring and related topics. The Digital Health Center of Excellence within CDRH is leading it.

The limiting language is on the paper's first page, and it is worth reading in full because it is unusually complete. The paper is "intended for discussion purposes only and does not represent draft or final guidance." It is "not intended to propose or implement policy changes regarding how CDRH intends to regulate generative AI-enabled devices." It is "not intended to communicate CDRH's proposed (or final) regulatory expectations, including its expectations for supporting evidence in future marketing submissions."

The fourth disclaimer is the one a regulated reader should weigh most heavily: "this paper is not intended to address whether the approaches discussed below are within FDA's existing legal authorities or whether new legal authorities would be necessary." The agency is not claiming it could adopt any of this today.

CDRH Director Michelle Tarver described the exercise as a beginning: "By inviting input from the public, we are launching a transparent process to inform the development of an approach that safeguards patients and consumers, advances innovation, and serves as a potential model for regulators around the world."

Why generative devices are treated as a separate problem

The paper's own background section states the premise directly. Generative devices "have unique characteristics and behaviors that are distinct from traditional software and AI-enabled devices that FDA regulates. They may accept open-ended inputs, perform multiple subtasks, and produce variable outputs to similar inputs."

It goes further on how they change. "In some cases", the paper says, they "can also evolve over time through changes to features such as the underlying model, prompts, retrieval strategies, guardrails, orchestration logic, and/or user interface", and many are built on general-purpose foundation models developed by third parties.

That last point is the hard one. A device can change because someone who is not the sponsor changed a model. The paper's own way of putting the difficulty is that evaluation approaches developed for software with bounded inputs and fixed outputs "may not be appropriate for GenAI-enabled devices". A device that answers differently to the same question twice is the case that pressure-tests the older approach.

The two axis framework, as the paper actually defines it

Figure 1 of the paper sets out what it calls "a possible two-axis framework for thinking about risk for GenAI-enabled software functions." Both axes are defined.

The horizontal axis "represents the activity performed by the function, increasing in the degree and independence of device activity from left to right." Its four bands run from Informational: Non-Directive, through Informational: Action-Directing, to Action-Taking: HCP-Supervised, and finally Action-Taking: Fully Autonomous.

The vertical axis "represents the consequence of relying upon an incorrect output", banded Limited, Moderate and Severe. Risk increases as a function moves toward autonomous action with severe consequences.

This is a grid for calibrating expectations, not a classification rule. Nothing in it changes how a device is classified today.

Competency based premarket evaluation, and what it would actually measure

CDRH says it is considering a competency based approach to premarket evaluation consisting of non-clinical device benchmarking and clinical confirmation. The paper is explicit that what gets evaluated is "the final user-facing device, as configured and intended to be deployed for real-world use", and not, in its words, "the foundation model standing alone or other isolated subcomponent".

The paper credits the idea to published proposals rather than presenting it as an FDA invention, citing authors who have argued that generative AI is too broad, adaptive and opaque for current device frameworks and who propose regulation modelled on medical training, licensure examinations, supervised practice and periodic reevaluation.

The paper enumerates what it calls benchmarking elements, in four groups. Safety covers safety-critical recognition and escalation; scope maintenance and boundary adherence; and calibration, uncertainty communication and clinical deferral. Clinical Proficiency covers clinical knowledge and task fidelity; information gathering and clinical analysis; quantitative and measurement analysis; and communication quality and user comprehension. Generalizability covers robustness, reliability and reproducibility, and subgroup performance. A fourth group, Agentic AI Capabilities, covers additional competencies applicable only to agentic devices. CDRH adds that it "does not anticipate that all elements would necessarily apply to all devices; rather, elements would be chosen based on applicability to the device's intended use and risk profile."

The paper also says the approach "could be tailored to the device's intended use and be proportionate to its risk, potentially utilizing the two-axis risk framework", so that device activity and the consequences of an incorrect output inform the nature, rigor and amount of evidence needed. It then adds that these descriptions "are intended only to stimulate discussion and obtain stakeholder feedback."

Postmarket monitoring and the foundation model problem

Three postmarket approaches are named as examples in the paper's own comment question 19: periodic re-benchmarking, sample-based clinician review, and performance degradation monitoring. On the last, the paper describes a sponsor monitoring for degradation such as drift "that may result from changes in the input population, the data environment, or underlying model components, specifying the analyses and thresholds to detect such degradation."

Re-benchmarking is described concretely. A device benchmarked against a defined set of capabilities for premarket authorization "could potentially be re-benchmarked against the same capabilities following a modification, providing evidence to determine whether the modification has affected safety or effectiveness."

For third party models the paper floats a voluntary Foundation Model Device Master File, leveraging the existing Device Master File programme, under which foundation model developers and platform providers could voluntarily submit structured model cards or system cards. These would be held confidentially by the FDA and could be referenced by sponsors with the file holder's authorization. The paper is careful that such a submission "would not constitute authorization of the underlying model for any device intended use, and sponsors would remain responsible for independently demonstrating the safety and effectiveness of their own device". It is a proposal for comment, and voluntary in the form described.

CDRH also asks, in question 20, whether postmarket monitoring could be facilitated by machine-based supervisory agents, and what would have to be true about the reliability of the supervisory agent itself.

What a clinician or a sponsor should do with it

For a physician or clinical governance lead, the practical use is a checklist for local evaluation. Four properties the paper identifies are the ones to test before a generative device runs on your service. Does it accept open ended inputs, and how does it behave at the edge of them. Does it chain multiple subtasks, and which one fails first. Does it return different answers to the same question asked twice, and how much variation does your workflow tolerate. Is it built on a third party foundation model, in which case ask the vendor in writing what happens to your validation when that model changes without the vendor initiating the change. None of this is required by the FDA. It is what the regulator says it is currently unsure about, which makes it a defensible place to start locally.

Nothing here changes the status of a device already running on your service. No obligation has been created, altered or removed.

For a sponsor, the paper is a question set with a date on it. An agency that publishes twenty-six numbered questions before a framework is signalling that the framework is unsettled, and the comment record is where feasibility arguments get made. A sponsor that finds the competency approach unworkable has until 19 October 2026 to say so with evidence on this docket. We are not aware of a further comment opportunity having been announced, and the paper itself says nothing about comment mechanics.

To comment, file under docket FDA-2026-N-7874 on Regulations.gov. Submissions to a public docket are published, so treat anything commercially sensitive accordingly.

What we did not verify

We opened the discussion paper itself, the FDA press announcement of 18 August 2026 and the Digital Health Center of Excellence landing page. Every quotation, definition, domain list and disclaimer above is taken from those documents rather than from trade coverage.

We did not open the underlying academic papers the discussion paper cites for the competency concept, and we do not characterise their arguments beyond the paper's own one sentence summary of them.

We make no claim about whether a rulemaking will follow, on what timetable, or in what form. The paper expressly declines to address whether the approaches it discusses are within the FDA's existing legal authorities, so nothing here should be read as the agency asserting it could adopt them.

Key compliance takeaway

The FDA has put a competency based evaluation approach for generative AI devices out for comment under docket FDA-2026-N-7874, closing 19 October 2026, with both risk axes and the assessment domains actually specified. It binds nobody, it does not change any submission requirement today, and the paper expressly declines to say whether the FDA could adopt it under existing authority. The comment record is the influence point the agency has scheduled.

Source File

https://www.fda.gov/media/194242/download

Open the discussion paper PDF. Confirm the four disclaimers on the first page, the Figure 1 axis definitions in Section IV, the competency domains and the final user-facing device scope in Section V, the three postmarket mechanisms in question 19, and the voluntary Foundation Model Device Master File in Section VII.A. Then confirm docket FDA-2026-N-7874 and the 19 October 2026 date.

This paper is not intended to address whether the approaches discussed below are within FDA's existing legal authorities or whether new legal authorities would be necessary. ยท FDA discussion paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices, 18 August 2026

FAQ

Does the FDA discussion paper change any obligation for device sponsors?

No. The paper states it does not represent draft or final guidance, is not intended to propose or implement policy changes, and is not intended to communicate CDRH's proposed or final regulatory expectations, including its expectations for supporting evidence in future marketing submissions. Submissions continue to be evaluated under existing law and existing guidance.

What is the competency based approach?

A premarket evaluation approach CDRH says it is considering, consisting of non-clinical device benchmarking and clinical confirmation, applied to the final user-facing device as it would be deployed rather than to the foundation model alone. The paper groups its benchmarking elements under Safety, Clinical Proficiency, Generalizability and Agentic AI Capabilities, and says it does not anticipate that all elements would apply to all devices. It is under discussion and no reviewer applies it as a standard today.

What is the deadline and what kind of deadline is it?

Comments are due by 19 October 2026 under docket FDA-2026-N-7874 on Regulations.gov. It is a comment deadline, not a compliance deadline. Nothing happens to a device or a sponsor that does not comment, other than forgoing the chance to shape the framework.

Why does the FDA treat generative AI devices differently from other AI devices?

The paper states that these devices have unique characteristics and behaviors distinct from traditional software and AI-enabled devices, and that they may accept open-ended inputs, perform multiple subtasks, and produce variable outputs to similar inputs. It adds that they can evolve over time through changes to the underlying model, prompts, retrieval strategies, guardrails, orchestration logic or user interface, and that many are built on third-party foundation models.

Sponsored Training

Practical AI training for regulated professionals, built around verification, documentation and a defensible process. See the courses.

."}}]}