A small language model is one that runs on modest local hardware — typically open-weight, in the range that fits on one or a few accelerators rather than a cluster. In enterprise workflows they are frequently the better choice, not as a compromise but on the merits: for narrow, well-specified, retrieval-grounded tasks a small model is cheaper, faster, deployable where a frontier model cannot go, and often no less accurate.
The important qualifier is narrow and well-specified. This article covers where that holds, where it does not, and why model routing makes the either/or framing unhelpful. For the deployment context, see sovereign AI.
Structured extraction. Pulling defined fields from documents. The failure modes are different in a useful way: small models tend to miss things rather than invent them, which is far preferable in regulated work — a missing field is visible, a fabricated one is not.
Classification and routing. Deciding whether a document is in scope, or which category a query belongs to. Frontier models are substantially over-specified for this, and the cost difference at volume is large enough to change what is economically feasible.
Retrieval-grounded generation over short contexts. When the relevant evidence has been retrieved and placed in context, model size matters less than the quality of retrieval — the heavy lifting was already done.
Verification. Checking whether a claim is supported by cited evidence is a narrow comparison task. Using a smaller, different model for verification than for generation catches errors correlated with the generating model.
Anywhere the data cannot leave. The decisive case. In an air-gapped or on-premise deployment, frontier models are simply unavailable — the question stops being which model is better and becomes which locally deployable model is adequate. See air-gapped AI deployment.
Open-ended reasoning over long contexts — synthesising an argument across many documents. Frontier models are meaningfully better here and the gap is not closed by prompting. Tasks requiring broad world knowledge, ambiguous instructions, and novel or unusual document structures also favour larger models. First deployments where the workflow is not yet understood benefit from a more capable model absorbing the uncertainty while you learn, then substituting down once characterised.
Serious systems route: each step in a workflow goes to a model class appropriate to it. A typical evidence workflow might use a small local model for extraction across a thousand documents, a small model for relevance classification, a larger model for synthesis on the filtered set, and a separate small model for verification. Cost drops sharply, quality often improves on narrow tasks, and in restricted environments routing becomes a compliance mechanism — steps touching identifiable data pin to locally served models while de-identified summaries may use others.
The architectural requirement is that the orchestration layer treats models as substitutable. If workflows are written against one provider's interface, routing is impossible and so is model independence.
Where Eclypse sits: the orchestration engine routes each step to the model class it actually needs — see how AI agents work for where routing sits in the execution loop.
The model question is downstream of the workflow question. Once the steps are characterised, which model each needs is usually straightforward — and usually more than one.
A language model that runs on modest local hardware, typically open-weight and sized to fit on one or a few accelerators rather than a cluster.
On broad, open-ended reasoning, yes. On narrow, retrieval-grounded tasks, the gap is often small or absent, and small models tend to fail by omission rather than fabrication.
Assigning each step of a workflow to a model class suited to it — small local models for extraction and classification, larger models for synthesis, a separate model for verification.
Yes. Open-weight models can be served entirely on local hardware with no outbound connectivity, which hosted frontier models cannot.
Our proprietary AI orchestration platform for healthcare: one engine, a registry of reusable task modules and domain agents, and a governed knowledge base — deployed inside your walls and run by your team.