New · France 2027 presidential election: what the candidates propose on AI, quoted and sourced. Explore the tracker →

Last reviewed:

What is an SLM (small language model)? Definition, examples and difference from an LLM

An SLM (small language model) is a language model compact enough to run on a computer, a phone or a modest server, generally under 10 billion parameters. Less versatile than a large LLM, it is enough for many repetitive, well-defined tasks, at a much lower cost and latency.

There is no official threshold. In June 2025, NVIDIA researchers proposed a practical definition in the paper “Small Language Models are the Future of Agentic AI”: an SLM fits on a common consumer device and responds fast enough to serve one user's requests. They add that in 2025, most models under 10 billion parameters fall into this category. Examples: Mistral AI's Ministral 3B and 8B (October 2024), designed for local, privacy-first inference, Microsoft's Phi family, Hugging Face's SmolLM2 or NVIDIA's Nemotron-H models. Many are produced by distilling a large model and released with open weights. The NVIDIA paper argues a specific thesis. In an agent, most model calls are narrow, repetitive subtasks: extracting a field, calling a tool, formatting an output. A specialised SLM handles them just as well, and according to the authors, serving a 7-billion-parameter SLM is 10 to 30 times cheaper (latency, energy, compute) than a 70-to-175-billion LLM. Across three open-source agents studied, they estimate that roughly 40% to 70% of calls could move from an LLM to an SLM. The architecture they recommend is heterogeneous: SLMs for volume, a large model called only when the task requires it. This is a reasoned position, not an established consensus.

Concrete example

Illustrative case: a French accounting firm with 40 staff wants to automatically extract fields from its clients' supplier invoices (company ID, net amounts and VAT, due date). The documents are confidential and volume reaches 20,000 invoices a month. Instead of a large model via API, the firm tests an 8-billion-parameter open-weight SLM, hosted on an in-house server with a single GPU. On 300 annotated invoices, the SLM extracts the standard fields as well as the large model. Atypical invoices (foreign formats, poor scans) are routed to a large model, a small fraction of the volume. Routine data no longer leaves the firm, and the API bill disappears for most documents.

Comparison

SLM vs LLM: the differences that matter for the business
SLM (small model)Frontier LLM
Indicative sizeUnder 10 billion parametersTens to hundreds of billions of parameters
Where it runsComputer, phone, in-house single-GPU serverProvider's data centres (API)
Usage costLow, especially at high volumeHigher, billed per token
LatencyLowHigher, varies with load
VersatilityGood on a narrow taskHigh, including on open questions
CustomisationFine-tuning simpler and cheaperHeavier, often limited to prompting
ConfidentialityData kept in-house if self-hostedData sent to the provider, per contract

FAQ

What is an SLM (small language model)?

A small language model, compact enough to run on a computer, a phone or a modest server. There is no official threshold, but most models under 10 billion parameters are considered SLMs in 2025 (NVIDIA Research definition).

What is the difference between an SLM and an LLM?

Size, and what follows from it. A frontier LLM is more versatile and better on open questions, but expensive and hosted by a provider. An SLM is less versatile, but fast, cheap, and can run on your premises. For a narrow, repetitive task, it often performs on par.

What are examples of small language models?

Mistral AI's Ministral 3B and 8B (French company, October 2024), Microsoft's Phi family, Hugging Face's SmolLM2, NVIDIA's Nemotron-H models, or the distilled versions of DeepSeek-R1. Many are available with open weights.

Why are small language models called the future of agentic AI?

That is the thesis of an NVIDIA Research paper (June 2025). An agent mostly chains narrow, repetitive subtasks, which a specialised SLM handles just as well, at a serving cost 10 to 30 times lower according to the authors. They recommend mixed systems: SLMs for volume, a large model as backup.

Can an SLM work without an internet connection?

Yes, if it is installed on a company workstation or server. Mistral AI presented its Ministral models for offline assistants, on-device translation or local analytics. It is an asset for remote sites and sensitive data.

Is an SLM safer for company data?

It can be: hosted in-house, it avoids sending data to an external provider. Security then depends on your own infrastructure: access control, updates, logging. A poorly operated SLM is no safer than a well-governed API.

See also

Further reading

Small Language Models are the Future of Agentic AI, Belcak et al., NVIDIA Research, 2025 (external resource)

Sources

  1. Small Language Models are the Future of Agentic AI, Belcak, Heinrich et al., NVIDIA Research, arXiv, June 2, 2025. https://arxiv.org/abs/2506.02153 (accessed 2026-09-30)
  2. Small Language Models are the Future of Agentic AI, HTML version (WD1 definition, costs, MetaGPT, Open Operator, Cradle case studies). https://arxiv.org/html/2506.02153v1 (accessed 2026-09-30)
  3. Un Ministral, des Ministraux (Ministral 3B and 8B), Mistral AI, October 16, 2024. https://mistral.ai/news/ministraux (accessed 2026-09-30)

← Back to glossary

Address copied