FDA and Generative AI-Enabled Medical Devices: What the 2026 Discussion Paper Signals for Manufacturers

Read:

9

min

Share:

Robotic-assisted surgical procedure illustrating FDA considerations for generative AI-enabled medical devices

Generative AI medical devices are moving into a new regulatory phase. In its 2026 discussion paper, the FDA examines how risk, evidence, postmarket monitoring and third-party model dependencies may need to be assessed across the device lifecycle.

What happens when a medical device accepts open-ended inputs, generates variable outputs, evolves through changes to prompts or retrieval strategies, and depends on a foundation model that the manufacturer does not fully control?

In August 2026, the FDA’s Center for Devices and Radiological Health published a discussion paper addressing these questions. Building on the FDA’s broader regulatory framework for AI-powered medical devices, the document explores potential approaches to risk assessment, premarket evaluation, postmarket monitoring and the use of third-party foundation models in generative AI-enabled medical devices.

Although the paper does not establish new regulatory requirements, it offers an important view into the issues that may shape future FDA policy.

For manufacturers, the central message is significant: demonstrating performance at a single point in time may not be enough for a device whose behavior, technological dependencies and real-world inputs can continue to change.

How the FDA is approaching generative AI medical devices

Before examining its technical content, it is important to understand the regulatory status of the document.

The paper is intended to support discussion and collect stakeholder feedback. It does not represent draft or final guidance, does not implement policy changes and does not establish FDA expectations for evidence in future marketing submissions.

Therefore, the concepts presented should not be interpreted as requirements that manufacturers must already follow.

Nevertheless, the paper is strategically relevant because it identifies the questions the FDA considers important for the future oversight of generative AI-enabled medical devices.

The FDA is accepting feedback under docket FDA-2026-N-7874 until October 19, 2026.

Why generative AI-enabled medical devices create different regulatory challenges

Traditional medical device software is often assessed against defined inputs, expected outputs and predetermined performance criteria. Predictive AI may introduce additional complexity, but its output is generally bounded by a specific task, such as classification, detection or risk prediction.

Generative AI can behave differently.

According to the FDA discussion paper, these systems may accept open-ended inputs, perform multiple subtasks and produce different outputs in response to similar inputs. Moreover, their behavior may change when the underlying model, prompts, retrieval strategies, guardrails, orchestration logic or user interface is modified.

Many GenAI-enabled devices may also rely on general-purpose foundation models developed by third parties. As a result, the medical device manufacturer may have limited visibility into the model’s training data, architecture, evaluation methods and future updates.

These characteristics create several potential risks, including:

  • outputs that appear credible but contain inaccurate or fabricated information;
  • uncertainty about the boundaries of the device’s intended use;
  • difficulty tracing an error to the device or its underlying model;
  • performance degradation across populations or deployment environments;
  • unplanned behavioral changes following an update to a third-party model.

Importantly, the FDA does not regulate generative AI as a technology in isolation. It regulates medical devices, including device software functions enabled by generative AI, according to their intended use and risk.

Four regulatory areas manufacturers should understand

The discussion paper can be organized around four connected areas.

Regulatory areaCore questionPotential implication for manufacturers
1.Risk assessmentWhat does the function do, how independently does it act and what could happen if its output is wrong?Risk analysis may need to examine not only the intended function, but also the context, user and potential consequences of relying on an incorrect output.
2. Premarket evaluationHow can performance be demonstrated when the range of possible inputs and outputs is difficult to test exhaustively?Evidence strategies may need to combine structured benchmarking with clinically representative confirmation.
3. Postmarket monitoringHow can safety and performance remain controlled as the device and its environment evolve?Manufacturers may need defined monitoring methods, performance thresholds and reassessment triggers.
4. Third-party dependenciesHow can changes to an external foundation model be detected and controlled?Contracts, technical monitoring, change control and access to model information may become part of the regulatory strategy.

Together, these areas suggest that the regulatory evaluation of generative AI-enabled medical devices may increasingly depend on evidence generated throughout the total product lifecycle.

A new way to think about risk in FDA generative AI medical devices

One of the most relevant concepts in the paper is a possible two-axis framework for assessing risk.

The first axis considers the activity performed by the device, including how independently the function directs or takes action. The second considers the consequence of relying on an incorrect output.

Under this approach, risk increases as the device becomes more directive or autonomous and as the potential harm associated with an incorrect output becomes more severe.

Consequence of relying on an incorrect outputLower degree of device actionGreater independence or device action
Higher consequenceEven an informational output may present significant risk when users cannot independently assess its accuracyGreater concern may arise when the system diagnoses, prescribes, initiates orders or controls another clinical action
Lower consequenceNon-directive information with limited potential harm may present lower relative concernAutonomy remains relevant, although the consequence of an incorrect action continues to shape the overall risk

This table is a simplified editorial representation developed by Sobel based on the FDA concept. It is not an FDA classification tool.

Directiveness may exist on a continuum

A function does not need to issue an explicit command to influence a clinical decision.

For example, there is a meaningful difference between providing general information, relating that information to a specific patient, recommending an action and issuing a direct instruction. However, the boundaries between these behaviors may not always be clear.

Consequently, the FDA is asking whether wording, specificity, personalization and context should modify the risk associated with an informational function.

Adding a statement such as “talk to your doctor” may not necessarily make an otherwise directive output less influential.

The intended user may affect the risk

The same output may also have different implications depending on who receives it.

A healthcare professional may be better positioned to recognize limitations, compare an output with other clinical information and avoid relying on an incorrect recommendation. By contrast, a patient may have less ability to independently evaluate the basis of the output.

Still, the FDA acknowledges the potential benefits of expanding access to clinical information and patient engagement. Therefore, the discussion is not simply about restricting patient-facing functions. It is about identifying when additional safeguards may be necessary.

A similar question applies to generalist and specialist healthcare professionals. A device that provides specialist-level information may improve access to knowledge, but it may also create risks if safe interpretation depends on expertise the intended user does not possess.

Risk can evolve during a conversation

Conversational devices introduce another challenge.

A system may begin by providing non-directive information and gradually move toward an action-directing recommendation during a multi-turn interaction. For this reason, assessing isolated outputs may fail to capture the complete behavior of the device.

In practice, manufacturers may need to evaluate realistic conversational trajectories, including situations in which the user minimizes symptoms, requests information outside the intended scope or attempts to influence the system’s response.

The FDA also highlights both under-escalation and over-escalation. Failing to recommend necessary care may delay treatment. Conversely, repeatedly directing users toward unnecessary emergency care may generate anxiety, unnecessary procedures and additional burden on healthcare systems.

Could competency-based evaluation complement traditional testing?

The FDA recognizes that testing every possible input and output may be impractical for certain generative AI-enabled medical devices.

Therefore, the discussion paper presents a possible competency-based approach consisting of two components:

  1. non-clinical device benchmarking;
  2. clinical confirmation.

The approach would evaluate the final user-facing device in its intended configuration, rather than assessing the foundation model as an isolated component.

This distinction is important. Device behavior can be affected by the model, prompts, retrieval architecture, guardrails, user interface and deployment environment. As a result, strong performance by the underlying model does not automatically demonstrate the safety or effectiveness of the complete medical device.

Device benchmarking

The FDA describes benchmarking as a structured way to evaluate whether the device demonstrates the capabilities needed for its intended use.

Potential areas include:

  • safety-critical recognition and escalation;
  • maintenance of intended-use boundaries;
  • communication of uncertainty;
  • clinical knowledge and task fidelity;
  • information gathering and clinical analysis;
  • quantitative and measurement analysis;
  • communication quality and user comprehension;
  • robustness, reliability and reproducibility;
  • performance across relevant subgroups;
  • capabilities specific to agentic AI systems.

However, a benchmark would need more than a collection of test questions.

The sponsor may need to define the testing scope based on the intended use and risk profile, prespecify the evaluation methods, justify the scoring criteria and establish acceptance thresholds before testing begins.

Moreover, datasets and testing conditions should represent the clinical populations and deployment contexts in which the device is expected to operate.

Public benchmarks may also present limitations. Data contamination, repeated optimization against the same tests and limited representation of real-world conditions can make strong benchmark results less predictive of actual clinical performance.

Clinical confirmation

Even a comprehensive benchmark may not show how the device will interact with real users, workflows and patient populations.

For that reason, the FDA is considering whether clinical confirmation should complement non-clinical benchmarking.

The paper discusses several possible approaches, including:

  • retrospective evaluation using previously collected patient inputs;
  • shadow deployment in which the device operates in a live workflow but its outputs do not affect patient care;
  • standardized interactions using trained or representative participants;
  • other clinically appropriate approaches with rigor proportionate to the device’s intended use and risk.

Importantly, the FDA does not state that every generative AI-enabled device would require a prospective clinical investigation. Instead, the discussion focuses on selecting and justifying an approach that is appropriate for the function, clinical context, patient population and potential risk.

Therefore, manufacturers should avoid assuming that one evidence package will fit every GenAI-enabled device.

Why postmarket evidence may become more important

Open-ended outputs, model evolution and changes in real-world use can make it difficult for premarket testing to capture every relevant aspect of performance.

As a result, the FDA is considering whether greater premarket uncertainty could be acceptable in certain situations if it is supported by sufficiently robust postmarket monitoring.

This is not an established policy. Nevertheless, it points to a closer connection between premarket evidence and postmarket surveillance.

The paper identifies three potential monitoring approaches.

Periodic re-benchmarking

The manufacturer could reassess the device against previously established performance thresholds at defined intervals or following specific triggering events.

For example, reassessment could be triggered by a change to the underlying model, deployment architecture or other safety-relevant component.

Sample-based clinician review

Qualified and independent clinicians could review samples of real-world inputs and outputs against prospectively defined criteria.

The sampling strategy would need to represent clinically relevant presentations, populations and interaction patterns encountered during actual use.

Performance degradation monitoring

Manufacturers could monitor for drift caused by changes in the user population, data environment, model components or deployment conditions.

However, monitoring would only be meaningful if the manufacturer had already defined the applicable analyses, performance thresholds and escalation criteria.

Accordingly, postmarket monitoring should not be treated as a separate activity designed only after authorization. The premarket assessment may need to establish the baseline against which future changes and real-world performance are evaluated.

Third-party foundation models create a new change-control challenge

A generative AI-enabled medical device may depend on a foundation model supplied and updated by another organization.

That model may influence refusal behavior, content policies, output formatting, version control and safety-critical guardrails. Nevertheless, the medical device manufacturer may not control when or how the foundation model is changed.

This creates an important regulatory question: how can the manufacturer detect, assess and respond to an externally initiated change before it affects device safety or effectiveness?

The FDA is seeking feedback on mechanisms that could include:

  • contractual update-notification obligations;
  • technical methods for detecting model changes;
  • access to relevant model documentation and audit information;
  • reassessment and re-benchmarking triggers;
  • change-control procedures within the quality management system;
  • possible use of a Predetermined Change Control Plan.

The paper also discusses the potential creation of voluntary Foundation Model Device Master Files.

Under this concept, foundation model developers could submit confidential information to the FDA, such as model architecture, training data provenance, known limitations, evaluation results, safety constraints and update commitments. Device sponsors could then reference the file, with authorization from its holder, in a premarket submission.

However, a Foundation Model Device Master File would not constitute authorization of the underlying model for a medical device use.

The sponsor would remain responsible for demonstrating the safety and effectiveness of the final device.

What manufacturers should begin evaluating now

Although the FDA paper does not create new requirements, companies developing generative AI-enabled medical devices can already examine whether their regulatory and evidence strategies address the questions raised.

A practical review should include:

  1. Clearly define the intended use, intended users, clinical environment and boundaries of the device function.
  2. Map each function according to its degree of directiveness, autonomy and potential consequence if the output is incorrect.
  3. Identify every component that may affect the final device behavior, including prompts, retrieval systems, guardrails, interfaces and external models.
  4. Establish prespecified performance criteria and determine whether the proposed benchmarks represent the intended population and deployment environment.
  5. Evaluate which form of clinical confirmation would be proportionate to the intended use and risk.
  6. Define postmarket indicators, reassessment triggers and procedures for investigating performance degradation.
  7. Determine how third-party model changes will be detected, communicated, assessed and documented.
  8. Integrate these controls into the quality management system and overall change-management strategy.

Addressing these points early can help manufacturers identify evidence gaps before they become submission or postmarket problems.

Sobel Regulatory View

The most important signal in the FDA discussion paper is not a single proposed test or regulatory pathway.

It is the possibility that evidence for generative AI-enabled medical devices may need to function as a connected system throughout the product lifecycle.

Risk assessment would influence benchmarking. Benchmarking would establish a premarket baseline. Clinical confirmation would examine whether the device performs under representative conditions. Postmarket monitoring would then evaluate whether that performance remains controlled as the product, its users and its technological dependencies evolve.

In other words, the submission may become one point within a continuous evidence strategy rather than the final destination of product development.

For manufacturers, this reinforces the importance of involving regulatory, clinical, quality, software and risk-management teams before the evidence package is assembled.

What comes next

The FDA is still collecting stakeholder input, and the concepts discussed may change before they appear in future policies or guidance.

Even so, the paper provides a valuable framework for companies developing or integrating generative AI into medical device functions.

Organizations that treat GenAI as a conventional software update may underestimate the implications for intended use, evidence generation, postmarket monitoring and third-party change control.

By contrast, companies that begin connecting these elements early will be better prepared to adapt as the FDA’s regulatory approach develops.

Developing a generative AI-enabled medical device?

Sobel supports medical device companies in mapping regulatory pathways, evaluating evidence requirements and connecting clinical, risk and lifecycle considerations before submission.

Contact our regulatory team to discuss the next step for your product.

Official FDA references

Search the blog

Topics

FREE E-BOOK

ISO 10993-2025 Update

Key changes for your biocompatibility strategy.

A practical guide to understand the 2025 update and what it means for your medical device documentation.

Newsletter

Subscribe to our newsletter to receive the latest updates and insights.

You may unsubscribe at any time using the link in our newsletter.