# Metrics

URL: https://softwaredictionary.org/terms/metrics
Category: DevOps & Cloud
Last updated: 2026-09-30
In Turkish: Metrikler

In short: Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.

## What are metrics in software monitoring?

In software operations, metrics are numbers that describe the state or behavior of a system, recorded at regular intervals. Examples include requests per second, the percentage of failed requests, response time, memory usage, and the number of jobs waiting in a queue. Because each data point is just a name, a value, a timestamp, and a few labels, metrics are cheap to store and fast to query, even over months of history.

Most metrics fall into a few types: a counter only goes up, such as total requests served; a gauge goes up and down, such as current memory use; and a histogram sorts measurements into buckets, such as how many requests took under 100, 250, or 500 milliseconds, which lets you calculate percentiles like p95 and p99. Applications expose metrics through a library, and a monitoring system collects them every few seconds, often by scraping an HTTP endpoint, and stores them in a time-series database. Labels such as `route` or `status` let you slice a metric, but each unique combination creates a new series, so labels with unbounded values like user IDs, known as high cardinality, must be avoided.

Metrics are like the gauges on a car's dashboard: speed, fuel, and engine temperature tell you at a glance whether things are normal, without describing every event. They power dashboards, alerts, autoscaling decisions, and SLOs. A popular starting set is the four golden signals from site reliability engineering: latency, traffic, errors, and saturation, meaning how full a resource is.

Metrics are often confused with logs. A log line records one event with rich detail, while a metric aggregates many events into a number, so a metric tells you that errors jumped at 14:02 and logs or traces tell you why. Averages can also mislead: a mean response time of 200 milliseconds can hide that 1% of users wait five seconds, which is why teams track percentiles.

## Key takeaways

- Metrics are numeric measurements recorded over time and stored as time series.
- The common types are counters, gauges, and histograms.
- Percentiles such as p95 and p99 reveal slow requests that averages hide.
- Labels add dimensions, but high-cardinality labels like user IDs cause cost problems.
- Metrics show that something changed; logs and traces help explain why.

## Example: Exposing a counter and a histogram in Python

```python
from prometheus_client import Counter, Histogram, start_http_server

# A counter only goes up; labels let you slice it by route and status
REQUESTS = Counter("http_requests_total", "Requests served", ["route", "status"])
# A histogram sorts durations into buckets so percentiles can be calculated
LATENCY = Histogram("http_request_duration_seconds", "Request duration", ["route"])

start_http_server(9100)  # serves the metrics for the monitoring system to scrape

def handle_checkout(request):
    with LATENCY.labels(route="/checkout").time():
        response = process(request)
    REQUESTS.labels(route="/checkout", status=response.status).inc()
    return response
```

## Frequently asked questions

**What is the difference between a counter and a gauge?**

A counter is a value that only increases, such as the total number of requests, and is usually viewed as a rate per second. A gauge is a value that can go up or down, such as current memory usage or the number of active connections.

**What are the four golden signals?**

They are latency, traffic, errors, and saturation, a set of metrics recommended in site reliability engineering for monitoring any user-facing service. Together they show how fast the service is, how much it is used, how often it fails, and how close it is to its limits.

**What does p99 latency mean?**

p99 latency is the response time that 99% of requests are faster than, so only the slowest 1% take longer. It shows what users in the long tail experience, which an average hides.

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
