Blog Post

Three Frameworks, Three Definitions of AI Risk

September 1, 2026

Table of Contents

Cross-mapping tables for AI evidence in life sciences already exist and are broadly right. Data integrity practice lines up against data governance requirements, software lifecycle logs against technical documentation and logging, human review checks against human oversight duties, post-market surveillance against post-market monitoring. Build one repository, present it two ways.

All of that is sound and it starts one step too late. A mapping table assumes you already know the system is in scope for both regimes. The regimes decide that using different logic, so a company can map diligently and still have built a record for the wrong systems while missing others entirely.

Two Different Definitions of Risk

The US drug regulator assesses risk from the model's role. Its draft framework rates a model on how much influence its output has over a decision, multiplied by the consequence of that decision being wrong. A model producing one input among several with human review sits lower than one determining an outcome directly, and the same architecture can occupy either position depending on how it is used.

European AI legislation assesses risk from category membership. A system is high-risk because it falls within a defined class, principally because it is a safety component of a regulated product or appears in the listed use cases. The question is what the system is rather than how much weight its output carries.

The Consequence of the Mismatch

A model with modest influence over a low-stakes decision, embedded in a regulated device, is low risk under the first system and high risk under the second. A model with decisive influence over a consequential trial decision, used purely internally and never embedded in a product, is the reverse. Neither classification is available by inference from the other, so both have to be performed.

There Are Two US Frameworks, Not One

A further complication is usually collapsed. The US regulator does not operate a single AI framework in this sector, and a company developing a therapeutic alongside an AI-enabled device is subject to both.

The drug and biologics side runs on the credibility assessment framework described above, which remains in draft. The device side runs on a separate stream, comprising final guidance on pre-authorizing bounded algorithm changes issued in August 2025, draft guidance on lifecycle management and marketing submissions for AI-enabled device software functions still pending since January 2025, and earlier transparency principles for machine learning enabled devices.

They Ask for Different Things

The drug framework centers on whether a model is credible enough for a stated context of use. The device stream centers on total product lifecycle management, covering data selection, validation, performance monitoring in the field and how changes are handled after clearance. A team that has satisfied one has not addressed the other, and the pre-authorized change envelope is the device-side mechanism worth understanding early.

Quality System Expectations Moved in February

One current obligation is easy to miss because it is not AI-specific. The amended quality system regulation took effect on 2 February 2026, withdrawing most of the previous requirements and incorporating the international medical device quality management standard by reference. Any AI lifecycle documentation now sits inside that structure rather than beside it.

The Scope Boundaries Differ Too

Coverage is not the same on either side, and the difference sits exactly where a lot of AI activity in this sector happens.

Scenario record showing qualitative and quantitative risk fields alongside the mapped regulatory frameworks and dated assessment entries
Tagging a single record with each applicable regime is what allows one assessment to answer to more than one regulator.

The US draft guidance covers nonclinical, clinical, postmarketing and manufacturing where a model's output supports a regulatory decision on safety, effectiveness or quality. It explicitly excludes early discovery and internal operational efficiency work. The European medicines regulator's reflection paper takes the opposite approach and covers discovery alongside every later stage.

Cybersecurity Falls Between Them

Worth noticing because it produces a genuine hole. The drug guidance states that cybersecurity risk sits outside its scope while recommending it be considered. European AI legislation places accuracy, robustness and cybersecurity together as requirements for high-risk systems. A sponsor working only from the first document has no cybersecurity obligation and will acquire one, and the components underneath a model are where that obligation lands first.

The Timelines Run in Opposite Directions

Sequencing matters because the two systems arrive at different moments. Expectations from the drug regulator apply to submissions now, notwithstanding the framework remaining in draft, since sponsors are being asked for credibility evidence in current interactions.

European high-risk obligations for AI embedded in regulated products arrive considerably later, with the backstop for that category falling in 2028 and stand-alone listed uses a year earlier. Transparency obligations are the exception and already apply. The practical order is therefore to satisfy the near-term evidence expectations first, provided the record built to do so is structured to serve the later obligations rather than discarded, which the revised European timeline sets out in detail.

What Both Systems Want From the Same Record

Underneath the different classification logic, the evidence overlaps substantially. Recording at attribute level rather than per document is what makes one exercise serve both.

Control assessment results against a governance framework showing average implementation maturity against target, with the weakest scores in the manage and measure functions
Assessing once against a normalized control set is what allows the result to be presented against more than one regime.

Seven attributes carry most of the weight. The specific question the model answers. Its context of use, stated precisely enough that a change is detectable. How much influence its output has over the decision. The consequence of that output being wrong. Data provenance and quality for training and evaluation. The validation performed and its results. Finally, who determined the evidence was adequate.

Both Care Who Signed the Adequacy Determination

The US framework asks that adequacy be assessed by people independent of the model development team. European rules place equivalent weight on independent review and on human oversight being exercised by someone with authority to disregard an output. The requirement converges on a familiar problem, since a reviewer who cannot evaluate the model's substance performs administration rather than challenge, and independence requires capability rather than only reporting lines.

A Change of Use Reopens Both Assessments

This is the operational trap and it catches organizations that treated classification as a one-time exercise at submission.

The US framework states that a change to the context of use, including applying a model to a new patient population or altering the oversight arrangement, requires the risk analysis to be redone. European rules treat a substantial modification, or a change in intended purpose, as reopening conformity questions and potentially moving obligations onto whoever made the change. Both trigger on a decision a project team can make without involving a regulatory function, and recording intended purpose as an inventory attribute is what makes the trigger visible.

Record Context of Use as a Field, Not a Narrative

A context of use written as a paragraph in a submission document cannot be compared against next year's deployment. Recorded as structured fields covering population, decision, oversight arrangement and model version, a change becomes detectable automatically rather than noticed by whoever happens to read both documents.

Where the Two Systems Are Converging

The picture is improving rather than fragmenting further. The two medicines regulators published joint guiding principles in January 2026, non-binding and explicitly aligned with the credibility framework, emphasizing that AI supports rather than replaces regulatory decision-making and that documentation should let a reviewer independently evaluate model behavior.

The final requirement is the useful one to build toward, because it is stricter than either framework's checklist and satisfies both. Documentation sufficient for an external reviewer to reach their own conclusion about how a model behaves will satisfy a credibility assessment and support a technical documentation file, and it is a higher bar than either states explicitly.

What Not to Wait For

Two things are frequently used as reasons to defer and neither survives examination.

  • Draft Status: The credibility framework remains in draft and sponsors are already being asked for its artifacts in submissions and meetings.
  • Absent Standards: No harmonized European standard has been cited yet, and the legislation separately requires documenting what you relied on instead.

The second is worth stating plainly for anyone planning around certification. The organization developing European standards declined to adopt the international AI management standard for the quality management requirement, writing a bespoke one instead, so a certification strategy built on that assumption needs revisiting. What to do while cited standards do not exist covers the obligation that applies in the meantime.

Classify First, Then Map

The advice to build one evidence repository and present it against several regimes is correct, and it presumes the harder question has been settled. Three frameworks reach the same estate using different logic, and one of them rates a model by how much its output influences a decision while another rates it by which category the system occupies. Establishing which systems are in scope for which regime, recording context of use as structured fields so a change is detectable, and noting where cybersecurity falls outside a framework's stated scope are the steps that decide whether the mapping exercise is pointed at the right things. Kovrr's AI compliance readiness assesses per requirement rather than per framework, which is what makes one record answer to several.

To see one assessment mapped against multiple regulatory regimes rather than repeated per framework, book a demo mapped to your own systems.

Yakir Golan

CEO

Life Sciences AI Governance FAQs

Speak to an Expert

How do FDA and EU AI Act risk classifications differ?

Do the two regimes cover the same activities?

Does the FDA have one AI framework for life sciences?

Is cybersecurity covered by the drug AI guidance?

Which attributes satisfy both regimes?

What happens if the context of use changes?

Should sponsors wait for final guidance and standards?