Last reviewed:
What is an SLM (small language model)? Definition, examples and difference from an LLM
An SLM (small language model) is a language model compact enough to run on a computer, a phone or a modest server, generally under 10 billion parameters. Less versatile than a large LLM, it is enough for many repetitive, well-defined tasks, at a much lower cost and latency.
There is no official threshold. In June 2025, NVIDIA researchers proposed a practical definition in the paper “Small Language Models are the Future of Agentic AI”: an SLM fits on a common consumer device and responds fast enough to serve one user's requests. They add that in 2025, most models under 10 billion parameters fall into this category. Examples: Mistral AI's Ministral 3B and 8B (October 2024), designed for local, privacy-first inference, Microsoft's Phi family, Hugging Face's SmolLM2 or NVIDIA's Nemotron-H models. Many are produced by distilling a large model and released with open weights. The NVIDIA paper argues a specific thesis. In an agent, most model calls are narrow, repetitive subtasks: extracting a field, calling a tool, formatting an output. A specialised SLM handles them just as well, and according to the authors, serving a 7-billion-parameter SLM is 10 to 30 times cheaper (latency, energy, compute) than a 70-to-175-billion LLM. Across three open-source agents studied, they estimate that roughly 40% to 70% of calls could move from an LLM to an SLM. The architecture they recommend is heterogeneous: SLMs for volume, a large model called only when the task requires it. This is a reasoned position, not an established consensus.
Concrete example
Illustrative case: a French accounting firm with 40 staff wants to automatically extract fields from its clients' supplier invoices (company ID, net amounts and VAT, due date). The documents are confidential and volume reaches 20,000 invoices a month. Instead of a large model via API, the firm tests an 8-billion-parameter open-weight SLM, hosted on an in-house server with a single GPU. On 300 annotated invoices, the SLM extracts the standard fields as well as the large model. Atypical invoices (foreign formats, poor scans) are routed to a large model, a small fraction of the volume. Routine data no longer leaves the firm, and the API bill disappears for most documents.
Comparison
| SLM (small model) | Frontier LLM | |
|---|---|---|
| Indicative size | Under 10 billion parameters | Tens to hundreds of billions of parameters |
| Where it runs | Computer, phone, in-house single-GPU server | Provider's data centres (API) |
| Usage cost | Low, especially at high volume | Higher, billed per token |
| Latency | Low | Higher, varies with load |
| Versatility | Good on a narrow task | High, including on open questions |
| Customisation | Fine-tuning simpler and cheaper | Heavier, often limited to prompting |
| Confidentiality | Data kept in-house if self-hosted | Data sent to the provider, per contract |
FAQ
What is an SLM (small language model)?
A small language model, compact enough to run on a computer, a phone or a modest server. There is no official threshold, but most models under 10 billion parameters are considered SLMs in 2025 (NVIDIA Research definition).
What is the difference between an SLM and an LLM?
Size, and what follows from it. A frontier LLM is more versatile and better on open questions, but expensive and hosted by a provider. An SLM is less versatile, but fast, cheap, and can run on your premises. For a narrow, repetitive task, it often performs on par.
What are examples of small language models?
Mistral AI's Ministral 3B and 8B (French company, October 2024), Microsoft's Phi family, Hugging Face's SmolLM2, NVIDIA's Nemotron-H models, or the distilled versions of DeepSeek-R1. Many are available with open weights.
Why are small language models called the future of agentic AI?
That is the thesis of an NVIDIA Research paper (June 2025). An agent mostly chains narrow, repetitive subtasks, which a specialised SLM handles just as well, at a serving cost 10 to 30 times lower according to the authors. They recommend mixed systems: SLMs for volume, a large model as backup.
Can an SLM work without an internet connection?
Yes, if it is installed on a company workstation or server. Mistral AI presented its Ministral models for offline assistants, on-device translation or local analytics. It is an asset for remote sites and sensitive data.
Is an SLM safer for company data?
It can be: hosted in-house, it avoids sending data to an external provider. Security then depends on your own infrastructure: access control, updates, logging. A poorly operated SLM is no safer than a well-governed API.
See also
Further reading
Small Language Models are the Future of Agentic AI, Belcak et al., NVIDIA Research, 2025
Sources
- Small Language Models are the Future of Agentic AI, Belcak, Heinrich et al., NVIDIA Research, arXiv, June 2, 2025. https://arxiv.org/abs/2506.02153
- Small Language Models are the Future of Agentic AI, HTML version (WD1 definition, costs, MetaGPT, Open Operator, Cradle case studies). https://arxiv.org/html/2506.02153v1
- Un Ministral, des Ministraux (Ministral 3B and 8B), Mistral AI, October 16, 2024. https://mistral.ai/news/ministraux