An AI agent works by looping: it plans a sequence of steps, executes one, observes what came back, and decides what to do next. Around that loop sit four supporting systems — a retrieval layer that supplies evidence, a registry of tools it is permitted to call, a router that assigns each step to an appropriate model, and a verification stage that checks the result before a person sees it.
That is the whole architecture. The engineering difficulty is not in the concept; it is in constraining the loop tightly enough to be trustworthy while leaving it loose enough to be useful. This article walks through each stage and where each one fails. For the deployment context, see enterprise AI agents.
The agent receives a request and produces a plan. Two approaches exist, and the difference matters more than any other design choice in the system. In emergent planning, the model decides the next step at each turn based on what has happened so far — flexible, and effectively impossible to review before it runs. In explicit planning, the agent produces the full plan first, as a structured, readable artefact, and then executes it.
Explicit planning costs flexibility. It buys three things regulated deployment needs: a plan a person can inspect and approve before anything executes, a record of intent separable from the record of execution, and reproducibility. Serious implementations tend to constrain planning to registered step types — the agent composes from a known vocabulary rather than inventing actions.
A model's parametric knowledge is not a defensible source for professional work — it is undated, unattributable and cannot be corrected. So the agent retrieves, pulling relevant passages from a governed corpus and carrying them into the working context. Critically, retrieved passages carry identifiers forward: when the final output makes a claim, the chain back to the source document, section and version is preserved. Without that, the output cannot be substantiated and the whole exercise fails at review.
A tool is any callable capability: a database query, a statistical routine, a document writer, an external API, another agent. Two properties distinguish an enterprise implementation. Tools are registered — a fixed, versioned, permissioned list, and the agent cannot construct new capabilities at runtime. And tools inherit the caller's permissions — the agent acts with the requesting user's access rights, not a standing service account. Skipping the second is the single most common architectural mistake in enterprise agent deployment.
After each step the agent examines what came back. Empty retrieval, tool error, result outside expected bounds, confidence below threshold — each needs a defined response, and "continue anyway" is the wrong one. An agent that responds to a failed retrieval by generating plausible content from memory has produced exactly the failure the whole architecture exists to prevent. The correct behaviour is to report the gap or escalate.
Not every step needs the same model. Extracting a table from a PDF, deciding which of forty studies is relevant, drafting a clinical summary and checking a citation are four different problems with different cost and capability profiles. A router assigns each step to a model class — a small local model for extraction and classification, a larger one for synthesis, a separate one for verification.
Before output reaches a person, an independent check runs against it, combining deterministic rule checks (do cited references exist and resolve; are numbers internally consistent) with a second model class that evaluates whether claims are supported by cited evidence. Verification produces two outputs: a filtered result, and a record that checking occurred with what outcome — the second is part of the evidence package.
Agents need to carry context across steps — the plan, results so far, retrieved evidence, decisions made. In regulated deployment this state is not an ephemeral scratchpad; it is the material the audit trail is built from, and it needs a defined retention and access policy like any other record.
Where Eclypse sits: the orchestration engine plans against registered modules, retrieves over the Validator Framework, and routes each step to the model class it actually needs — the same loop described above, run inside your environment.
The architecture is less interesting than what it is pointed at. For the definition and required components, see what are AI agents?; for whether an agent is even the right tool for a given workflow, see AI agents vs chatbots. For the permission and tool-access controls only touched on above, see enterprise AI agent security; for the ownership and approval framework, see AI agent governance.
It plans, either generating a full sequence of steps upfront or choosing each next step from results so far, then observes each outcome and adapts. In regulated deployments the plan is usually explicit and constrained to a registered vocabulary of step types.
Through retrieval over a governed evidence base rather than model memory, by carrying source citations through to the output, and by verifying claims against cited evidence before a person sees them.
Assigning different steps to different model classes based on what the step requires — small local models for extraction and classification, larger ones for synthesis, a separate one for verification.
Yes. An agent can run fully air-gapped with locally served models, a local evidence base and internal tools only.
Our proprietary AI orchestration platform for healthcare: one engine, a registry of reusable task modules and domain agents, and a governed knowledge base — deployed inside your walls and run by your team.