# Rate Limiting

URL: https://softwaredictionary.org/terms/rate-limiting
Category: Backend & APIs
Last updated: 2026-09-30

In short: Rate limiting is a technique that caps how many requests a client can make to a server or API within a time window, protecting it from abuse and overload.

## What is rate limiting?

Rate limiting controls how often someone can call a service. A rule such as 100 requests per minute per API key, or 5 login attempts per minute per IP address, is checked on every request, and once a client goes over the limit, further requests are rejected until the time window resets. On the web, the server usually responds with the HTTP status `429 Too Many Requests`.

It works like a bouncer at a club who lets in only so many people per minute, however long the line gets. Rate limits protect servers from traffic spikes, buggy clients stuck in a loop, aggressive scraping, and brute-force password guessing, and they help share capacity fairly between customers. They are enforced in API gateways, reverse proxies, CDNs, or middleware inside the application, often with counters kept in a fast shared store such as Redis so every server sees the same count.

Common algorithms include the fixed window, which counts requests per calendar minute; the sliding window, which counts requests in the last 60 seconds; and the token bucket, where each client has a bucket that refills with tokens at a steady rate and every request spends one token. The token bucket is popular because it allows short bursts while still enforcing an average rate.

Rate limiting is closely related to throttling, and the two terms are often used interchangeably, but throttling sometimes means slowing requests down or queuing them instead of rejecting them. Well-behaved APIs tell clients about their limits with response headers such as `Retry-After` or `X-RateLimit-Remaining`, and well-behaved clients respond to a `429` by waiting and retrying with exponential backoff, meaning longer pauses after each failed attempt.

## Key takeaways

- Rate limiting caps how many requests a client can make in a period of time.
- Clients over the limit usually receive HTTP `429 Too Many Requests`.
- Limits are typically applied per API key, user, or IP address.
- Common algorithms are fixed window, sliding window, and token bucket.
- Clients should respect `Retry-After` and retry with exponential backoff.

## Example: A simple rate-limiting middleware in Express.js

```javascript
// Fixed-window limiter: 100 requests per minute per IP (in memory, one server)
const LIMIT = 100;
const WINDOW_MS = 60_000;
const hits = new Map(); // ip -> { count, start }

function rateLimit(req, res, next) {
  const now = Date.now();
  let entry = hits.get(req.ip);
  if (!entry || now - entry.start >= WINDOW_MS) {
    entry = { count: 0, start: now }; // start a new window
    hits.set(req.ip, entry);
  }
  if (++entry.count > LIMIT) return res.status(429).send("Too Many Requests");
  next();
}
```

## Frequently asked questions

**What does HTTP 429 Too Many Requests mean?**

It means you have sent more requests than the server allows in a given time. Wait before retrying, ideally for the number of seconds given in the `Retry-After` header if the response includes one.

**What is the difference between rate limiting and throttling?**

They are often used as synonyms. When a distinction is made, rate limiting rejects requests over the limit, while throttling slows them down or queues them so they are processed at a controlled pace.

**What is the token bucket algorithm?**

Each client gets a bucket that refills with tokens at a fixed rate, and every request uses one token. When the bucket is empty, requests are rejected, so short bursts are allowed but the average rate stays under the limit.

## Sources

- [RFC 6585: Additional HTTP Status Codes (429 Too Many Requests)](https://www.rfc-editor.org/rfc/rfc6585.html)

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
