GAMP 5 built its reputation on a simple idea: validation effort should scale with risk and complexity, not be applied uniformly to every system regardless of what it does. That principle transfers cleanly to AI. What doesn't transfer cleanly is the assumption underneath most of GAMP 5's original tooling — that a validated system behaves the same way every time it runs the same input.
Current as of September 2026 — always confirm current expectations with your quality unit and applicable regulator before relying on the specifics below.
Traditional computerised system validation tests a fixed set of documented functions against a fixed set of expected outputs. An AI system, particularly one built on a language model, doesn't have a fixed output for a given input in the same sense — it has a distribution of plausible outputs, and "correct" is a statement about the distribution's behaviour across a representative test set, not a guarantee about any single run.
That reframes what a validation protocol has to demonstrate. Instead of "given input X, the system returns output Y," it becomes "across this representative sample of inputs within the intended use, the system's output meets this accuracy threshold, and deviations are caught by this verification layer before a human sees them." The shift is from proving determinism to proving a bounded, monitored error rate.
Every validation activity references a specific, written intended use: what the system does, on what document types or data, for what decision, within what boundaries. This matters more for AI than for conventional software because the same underlying model can plausibly be pointed at many tasks — a system validated to extract structured fields from a specific submission format is not validated to summarise clinical narratives, even if the same model sits underneath both.
Scope creep is the most common way validated AI systems quietly become unvalidated ones: a team extends a tool to a new document type or a new decision point without re-running the validation exercise against that new intended use, on the reasoning that "it's the same system." It is the same code. It is not the same validated claim.
Three categories of change matter, and only one of them is code. A model version change — a new release of the underlying language model — can shift output behaviour even with no other change to the system around it. A corpus change — the documents a retrieval-augmented system draws from — changes what evidence is available to ground an answer. A configuration change — prompts, routing rules, retrieval parameters — changes behaviour without touching a line of application code in the traditional sense. A validation plan that only watches for code deployments misses the majority of what actually changes an AI system's behaviour in production.
A validated AI system is one shown to perform within its documented boundaries and error rate — not one guaranteed to be correct every time. This is precisely why validated AI deployments in GxP contexts pair the system with defined human checkpoints rather than removing the human from the loop; the validation and the checkpoint are two parts of the same control, not a redundancy to be optimised away once confidence grows.
Where Eclypse sits: every deployment ships with a documented intended use, a defined test set and error-rate baseline, and a change-control process that treats model, corpus and configuration changes as validation-relevant events — the artefacts a quality unit needs to build its own validation package around.
See how the validation boundary interacts with human oversight in human-in-the-loop AI in healthcare, and how the same evidence supports an inspection years later in audit trail requirements for AI systems.
Its risk-based approach is generally applied to AI in GxP environments, and its second edition explicitly extends to AI/ML. It is not itself a regulation — confirm current expectations with your quality unit.
Validation has to account for performance across a representative range of inputs rather than deterministic input-output pairs, and re-validation has to be triggered by model or data changes, not only code changes.
A specific, documented statement of what the system is meant to do and on what data — the reference point every validation test is written against.
Generally yes, proportional to the change — a model version update, corpus change or configuration change can each shift performance and should trigger an impact assessment.
Yes. Validation demonstrates performance within defined boundaries and error rate, which is why validated systems are paired with human review rather than deployed unsupervised.
Our proprietary AI orchestration platform for healthcare: one engine, a registry of reusable task modules and domain agents, and a governed knowledge base — deployed inside your walls and run by your team.