# LLM (Large Language Model)

URL: https://softwaredictionary.org/terms/llm
Category: AI & Machine Learning
Last updated: 2026-09-29

In short: An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.

## What is an LLM?

A large language model, or LLM, is a neural network trained on a very large collection of text, such as books, websites, and source code. It learns the patterns of language well enough to answer questions, summarize documents, translate, and write code. The word large refers both to the amount of training data and to the number of parameters, which can reach billions or more.

Under the hood, an LLM works with tokens, which are small chunks of text such as words or parts of words. Given some input text, called a prompt, the model predicts a likely next token, adds it to the text, and repeats the process until the answer is complete. Most modern LLMs are based on the transformer architecture, which lets the model weigh how much each earlier token matters when predicting the next one.

Developers usually use an LLM through an API: they send a prompt and receive generated text in return. LLMs power chat assistants, coding assistants, customer support bots, and search features. A useful mental model is a very well-read autocomplete: it is excellent at producing plausible text, but it does not look facts up unless it is connected to a source of information.

An LLM is not a database or a search engine. It does not reliably store documents word for word, its knowledge stops at a training cutoff date, and it can produce confident but wrong answers, known as hallucinations. Techniques like RAG are used to ground its answers in real, up-to-date data.

## Key takeaways

- An LLM generates text by predicting the next token over and over.
- It is trained on massive text datasets and has billions of parameters.
- Its built-in knowledge is frozen at a training cutoff date.
- Output quality depends heavily on the prompt and context you provide.
- LLMs can hallucinate, so important answers should be verified.

## Example: Calling an LLM through an HTTP API

```typescript
// Send a prompt to an LLM through a generic HTTP API
const response = await fetch("https://llm.example.com/v1/generate", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    prompt: "Explain recursion in one sentence.",
    maxTokens: 60,    // limit the length of the answer
    temperature: 0.2, // lower = more predictable output
  }),
});

const data = await response.json();
console.log(data.text);
```

## Frequently asked questions

**What is the difference between an LLM and AI?**

AI is the broad field of making machines perform intelligent tasks; an LLM is one specific kind of AI model focused on understanding and generating text. Chat assistants are applications built on top of LLMs.

**What is a token in an LLM?**

A token is the unit of text an LLM reads and writes, often a whole word or part of a word. In English, one token is roughly three quarters of a word on average, and usage limits and pricing are usually measured in tokens.

**What is a context window?**

The context window is the maximum amount of text, measured in tokens, that an LLM can take into account at once, including both the prompt and its answer. Anything outside the window is invisible to the model.

## Sources

- [Vaswani et al.: Attention Is All You Need (2017)](https://arxiv.org/abs/1706.03762)
- [Brown et al.: Language Models are Few-Shot Learners (2020)](https://arxiv.org/abs/2005.14165)

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
