The FDA Considers Options for Regulating Generative AI

The FDA Considers Options for Regulating Generative AI

On August 18, 2026, the Food and Drug Administration (FDA) released a discussion paper on how the Center for Devices and Radiological Health should regulate generative artificial intelligence (AI)-enabled medical devices and opened a public docket (FDA-2026-N-7874) for comment through October 19, 2026. As a discussion paper, the document represents a request for feedback, not formal regulatory guidance. It is significant in that it signals the agency’s interest in building a regulatory framework for something unlike anything it has regulated before: software that generates open-ended, variable output and may change after it reaches patients.

The discussion paper comes just a little over a year after the White House published its AI Action Plan, which pressed federal agencies to move quickly to facilitate the use of AI. The action plan noted that “rigorous evaluations can be a critical tool in defining and measuring AI reliability and performance in regulated industries,” adding that “over time, regulators should explore the use of evaluations in their application of existing law to AI systems.”

The paper’s release is particularly timely: recent surveys from Gallup and Pew find that about one-quarter of Americans have turned to AI chatbots for health information or advice. A 2024 survey of U.S. hospitals found that nearly a third reported using generative AI integrated with their electronic health records. At least one state is already piloting the use of AI in prescription renewals.

Yet even as individuals and healthcare organizations embrace AI, there is growing unease. Even some of the technology’s own architects are questioning how far these systems can be trusted. Earlier this summer, OpenAI disclosed that two of its models, run in an internal test with their safeguards disabled, broke out of their sandbox and compromised parts of the AI platform Hugging Face over several days—behavior the company attributed to the models trying to solve a cybersecurity test. In an August essay, Bill Gates, a lifelong champion of the technology, warned that as AI models grow more powerful they “could begin to act against our interests and we could lose control,” and called for new guardrails to manage the risks, among them cyberattacks.

In healthcare, where the stakes are a patient’s diagnosis or prescription, such concerns quickly become questions of oversight.

But doesn’t FDA already regulate AI?

The FDA has long regulated many medical devices that incorporate AI. However, the more than 1,400 AI-enabled devices listed by the FDA are typically narrow, task-specific tools: an algorithm that flags a suspected stroke on a CT scan, one that screens retinal images for diabetic retinopathy, another that predicts a patient’s risk of sepsis. Such tools generally have bounded inputs and outputs and perform defined tasks. This allows their performance to be tested against representative cases before release. Generative AI presents a different problem: the range of possible inputs and outputs may be too large for traditional testing approaches to adequately capture. As of early 2026, no generative-AI- or large-language-model-based medical device had received FDA authorization.

The discussion paper distinguishes between such AI-enabled devices and generative AI, which the FDA abbreviates GenAI. While a conventional model classifies an image or predicts a risk score, a generative model produces new content. The paper describes a class of models that “emulate the structure and characteristics of input data in order to generate derived synthetic content,” whether text, images, audio, or video.

A GenAI-enabled device may accept open-ended inputs, produce different outputs in response to similar prompts, and change over time as its underlying model, prompts, or guardrails are updated. This effectively reduces the predictability that has allowed the FDA to test conventional devices against a representative sample of cases.

A Framework for Representing Risk

In considering the unpredictability in GenAI-enabled software functions, the discussion paper presents a two-axis framework for representing risk (figure 1). Its horizontal axis represents the range of autonomy with which the software acts: from offering information to initiating action on its own. The vertical axis represents the “consequences,” or how much harm could result from reliance on an incorrect output, from limited to severe. Risk increases on the upper right-hand corner of the graph, where a function acts on its own and has the potential for severe consequences.

Figure 1: A two-axis view of risk for generative-Al-enabled software functions, based on the framework in FDA/CDRH, “Considerations for the Regulation of Generative Al-Enabled Medical Devices” (2026), Figure 1.

The FDA is considering regulation of GenAI within a rapidly evolving landscape. It is reasonable to assume that many of the hypotheticals the discussion paper considers might be in practice before any final policy is adopted.

A Test Case in Utah

In December 2025, the Utah Department of Commerce’s Office of AI Policy (OAIP) granted a generative-AI system operated by Doctronic authority to manage prescription-medication renewals. The system operates in the sort of “sandbox” endorsed in the White House action plan, which offers temporary relief from ordinary licensing rules.

Jim O’Neill, deputy secretary of the Department of Health and Human Services (HHS), who chairs HHS’s AI Governance Board, has praised Utah’s initiative. In a post on X, he said that the state had taken “a great step forward by allowing AI to refill prescriptions” and pledged to continue to “encourage innovation, regulatory clarity, and wide adoption” of AI.

Others have been less enthusiastic about the Utah pilot. Noting that its members had been advised of the program only after it was launched, the Utah Medical Licensing Board formally asked the state to suspend the pilot.

In an April letter to the Utah Department of Commerce, the board argued that while renewing prescriptions “may seem like an innocuous task,” it is not. It observed that a prescription renewal requires reassessment, which may include dose adjustment, checking for contraindications and drug interactions, and confirming efficacy. It warned that without proper review, a patient “may remain on outdated or suboptimal therapy for months or years.”

Concern about the pilot was heightened by Axios’s reporting that the security firm Mindgard had manipulated Doctronic’s public health chatbot—which operates outside the sandbox—into tripling an OxyContin dose, mislabeling methamphetamine, and generating false vaccine claims. Doctronic and Utah’s OAIP were quick to point out that the researchers had tested a public-facing chatbot, not the sandboxed system under which each prescription renewal is currently reviewed by a licensed physician, and which does not manage controlled substances such as OxyContin. Mindgard’s researchers were unpersuaded and noted that a language model’s guardrails can be rewritten through the very sort of prompts they had used with the public chatbot.

In declining to suspend the program, OAIP and Utah’s Division of Professional Licensing observed that, in the program’s current phase, each refill still requires physician review and approval. They pointed to a phased design under which autonomy would only be expanded if safety was demonstrated. Importantly, they invited the medical board’s assistance in monitoring the program going forward.

The Utah case exemplifies many of the issues raised in the FDA’s discussion paper. A system that renews prescriptions on its own, directly for patients, would sit in the high-risk corner of the FDA’s grid. The medical board’s insistence that a “routine” refill is not routine echoes the paper’s point that risk turns on what a GenAI program does, not on what it is called.

The pilot also raises the questions for which the discussion paper is seeking answers: who is ultimately making decisions and who is responsible? Utah launched its prescription pilot program through a state innovation office. Its medical board says prescribing is the practice of medicine. The FDA paper suggests a system like this may be a medical device subject to federal review. The paper does not resolve the tension—it explicitly declines to say whether the FDA even has the authority to try.

Borrowing from Medicine

Bakul Patel and David Blumenthal, whose work is referenced in the paper, have argued that generative AI is too broad and too changeable for conventional device review. Instead, they propose that it could be overseen through a program akin to the training, examination, and continuous evaluation used in the oversight of physicians. This could entail the issuance of a credential which a developer would pursue for its marketing value. They acknowledge that the establishment of a program for the oversight of what they refer to as GAI would “require considerable research and development.”

The FDA’s paper is a careful, deliberative document about a technology that is not waiting for the guardrails the agency might ultimately impose. When, and if, the questions it poses give rise to regulations, the systems it describes will have advanced. It is also reasonable to expect that other states will have followed Utah’s lead before the federal government issues final rules. For now, neither the agency, nor the states, nor the scholars who have proposed alternatives agree on who should oversee generative AI in medicine, or how. The comment period runs through October 19. Until then, the FDA is doing the one thing a regulator can do when the answers are not yet clear: asking the questions.