Systematic literature review is one of the strongest fits for AI in evidence work, for structural reasons: the protocol is defined in advance, the volume is high, the extraction targets are specified, and correctness is checkable against the source. What does not automate is the judgement — protocol design, adjudication of borderline inclusions, quality assessment and interpretation.
This article covers where the line sits and, importantly, how to handle the recall question honestly. For the wider picture, see AI for market access and HEOR.
The clearest case. Two patterns, with very different risk profiles: prioritisation — the system ranks records by likely relevance and reviewers screen in that order — and automated exclusion — the system excludes records without human review, which introduces the possibility of missing an eligible study.
Recall is the metric that matters, not accuracy. A screening model that is 95% accurate but misses 3% of eligible studies has failed, because the missed studies are exactly the ones the review exists to find. Any deployment must measure recall specifically, against a human-screened reference set.
AI can extract the specific passages relevant to each eligibility criterion, so the reviewer assesses the criterion against the evidence rather than reading the whole paper. For structured data extraction, AI with human verification is substantially faster than manual work and, in practice, often more consistent — human extraction across multiple reviewers produces variance that is rarely measured. Numeric extraction from tables and figures warrants mandatory verification rather than sampling.
The capability that changes what is feasible rather than merely accelerating it. A living review is continuously updated as new evidence publishes, rather than a point-in-time snapshot that decays. With continuous surveillance over a governed evidence base, new records are identified, screened and extracted as they appear, flagged for reviewer adjudication. For market access teams supporting rolling submissions across many markets, this is frequently worth more than the initial review saving.
Where Eclypse sits: continuous literature surveillance runs over the same governed evidence base that generates dossiers and value stories, so a refresh becomes a re-run rather than a new project. See AI value dossier automation.
The screening pattern you choose is a methodological decision with real consequences for defensibility — worth deciding deliberately rather than inheriting from a tool's default.
It can automate abstract screening, locating relevant passages, structured extraction and continuous updating. Protocol design, adjudication and interpretation remain with the reviewer.
The relevant metric is recall, not accuracy. A model that misses eligible studies has failed regardless of overall accuracy — recall must be measured against a human-screened reference set.
A review continuously updated as new evidence publishes, rather than a point-in-time snapshot — made practical by continuous surveillance over a governed evidence base.
For structured fields, often more consistent — manual extraction across reviewers produces variance that is rarely measured. Numeric extraction needs mandatory verification either way.
Our proprietary AI orchestration platform for healthcare: one engine, a registry of reusable task modules and domain agents, and a governed knowledge base — deployed inside your walls and run by your team.