Skip to main content

Small Language Model

Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/small-language-model

In short

A small language model (SLM) is a language model with far fewer parameters than the largest LLMs, cheap and fast enough to run on one GPU, a laptop or a phone.

What is a small language model?

A small language model, or SLM, works just like a large language model: it is a transformer trained to predict the next token. The difference is size. There is no official cutoff, but today the name usually covers models from a few hundred million up to roughly 10 or 15 billion parameters, while the largest LLMs have hundreds of billions or more. Examples include Microsoft's Phi, Google's Gemma, Meta's smaller Llama models and the smaller Qwen models.

Small models have improved quickly thanks to more and better-filtered training data, often including synthetic text written by larger models, and to knowledge distillation, where a small student model learns to imitate a large teacher. Microsoft's Phi-3-mini, with 3.8 billion parameters, scored close to GPT-3.5 on common benchmarks in 2024 while being small enough to run on a phone. Quantization shrinks such models further, so one with a few billion parameters fits in a few gigabytes of memory.

Teams choose SLMs for cost, speed and control. A small model responds faster, costs a fraction as much per request and can run on the user's own device, so data never leaves it and the feature keeps working offline. Fine-tuning one for a narrow job, such as sorting support tickets, pulling fields out of invoices or routing requests inside an AI agent, is cheap, and on that single job it often matches a much larger general model. It is like a pocket dictionary next to an encyclopedia: it covers less, but it goes everywhere and opens faster.

SLM vs LLM is a question of scale and trade-offs, not a different technology. A large model knows more, follows complicated instructions better and handles long, multi-step reasoning more reliably; a small one is narrower and makes more mistakes outside the tasks it was tuned for. Many products use both: a small model handles the frequent, simple requests and hands the hard ones to a large model. "Small" also keeps moving: GPT-2, a headline model in 2019, had 1.5 billion parameters, fewer than many of today's SLMs.

Key takeaways

  • A small language model works like an LLM but has far fewer parameters, roughly up to 10 or 15 billion.
  • There is no official size cutoff, and what counts as small keeps shifting.
  • Teams pick SLMs for lower cost, lower latency, privacy and offline use on the device.
  • Fine-tuned for one narrow task, a small model can match a much larger general one.
  • Large models still win on broad knowledge and complex, multi-step reasoning.

Example

Running a small model locally with Hugging Face Transformerspython
from transformers import pipeline

# About 0.5 billion parameters: small enough for an ordinary laptop CPU
generator = pipeline("text-generation", model="Qwen/Qwen2.5-0.5B-Instruct")

messages = [{
    "role": "user",
    "content": "Label this ticket as billing, bug or other: 'I was charged twice.'",
}]
result = generator(messages, max_new_tokens=10)
print(result[0]["generated_text"][-1]["content"])  # e.g. billing

Readers ask

How small is a small language model?

There is no official threshold. Today the term usually means anything from a few hundred million to about 10 or 15 billion parameters, small enough to run on a single GPU, a laptop or, at the lower end, a phone. Some researchers define it by where it can run rather than by a parameter count.

Can a small language model run on a phone?

Yes. Models of 1 to 4 billion parameters, usually quantized to 4 bits, run on recent phones, and both Android and Apple's platforms ship built-in on-device models that apps can call. They handle summarizing, rewriting and classifying text well but are less reliable for open-ended questions.

Should I use an SLM or an LLM?

Start from the task. If it is narrow and repeated, such as classification or extraction, try a small model, ideally fine-tuned, and compare it with a large one on your own evals. If it needs broad knowledge or long reasoning, a large model is usually worth the extra cost.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

More

Settings