Learning path · Intermediate
Building with LLMs
Prompts, retrieval, tools and evaluation: what it takes to ship an AI feature.
A practical route for developers adding a language model to a product: talking to the model, giving it your own data, letting it call tools, and measuring whether the result is any good.
24 pages4 chaptersabout 46 min of reading
- AI & Machine Learning
- Security
Not started yet0/24 read
Start with LLMProgress comes from your reading history, kept only in this browser.
Chapter 1Talking to a model
- 1LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- 2TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- 3PromptAI & Machine Learning, p. 35A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- 4System PromptAI & Machine Learning, p. 44A system prompt is the instructions an app gives a language model before the conversation starts, setting its role, rules, tone and what it should know.
- 5Prompt EngineeringAI & Machine Learning, p. 36Prompt engineering is the practice of designing, testing, and refining the instructions given to an AI model so it produces accurate, consistent, useful output.
- 6Context WindowAI & Machine Learning, p. 12A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.
- 7TemperatureAI & Machine Learning, p. 45Temperature is a setting that controls how random an LLM's output is, from focused and predictable at low values to more varied and creative at high values.
- 8HallucinationAI & Machine Learning, p. 23A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
Chapter 2Giving it your data
- 9EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- 10Cosine SimilarityAI & Machine Learning, p. 13Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
- 11ChunkingAI & Machine Learning, p. 9Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
- 12Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- 13Semantic SearchAI & Machine Learning, p. 42Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
- 14RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- 15Context EngineeringAI & Machine Learning, p. 11Context engineering is the practice of choosing what an LLM sees on each call (instructions, documents, tool results, history) so it can do the task reliably.
Chapter 3Letting it act
- 16Tool CallingAI & Machine Learning, p. 47Tool calling is an LLM feature in which the model asks the application to run a specific function with structured arguments, then uses the result in its answer.
- 17Model Context ProtocolAI & Machine Learning, p. 29The Model Context Protocol is an open standard that defines how AI applications connect to external tools, data sources, and prompts through a shared interface.
- 18AI AgentAI & Machine Learning, p. 2An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
- 19Prompt InjectionSecurity, p. 29Prompt injection is an attack on LLM apps where attacker-written text is treated as instructions, so the model ignores its rules, leaks data or misuses tools.
Chapter 4Making it good
- 20EvalsAI & Machine Learning, p. 17Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
- 21Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- 22LoRAAI & Machine Learning, p. 26LoRA is a cheap way to fine-tune a large model: its weights stay frozen and only small added matrices are trained, so a new skill fits in a few megabytes.
- 23QuantizationAI & Machine Learning, p. 37Quantization is a technique that shrinks an AI model by storing its parameters in fewer bits, such as 8 or 4 instead of 16, so inference is faster and cheaper.
- 24InferenceAI & Machine Learning, p. 24Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
Along the way, compare
Pairs on this path that are easy to mix up, side by side.