Regression
- In Turkish
- Regresyon
In short
Regression is the kind of supervised learning where a model predicts a number, such as a price, a temperature or a delivery time, from the input it is given.
What is regression in machine learning?
In machine learning, regression means predicting a continuous number rather than a category. A model learns from past examples where the answer is known, such as flats with their size, location and sale price, and then estimates that number for new inputs. Forecasting demand, estimating how long a delivery will take and predicting energy use are typical regression tasks.
The simplest and oldest method is linear regression, which fits a straight line, or a flat plane when there are several inputs, through the data. It picks the line that makes the squared differences between its predictions and the real values as small as possible, a method called least squares that dates back to the early 1800s. The fitted weights are easy to read, for example how much the price rises per extra square meter, which is why linear regression is also a standard tool in statistics. Decision trees, gradient-boosted trees and neural networks handle patterns that don't follow a straight line.
Regression models are judged by how far their predictions land from the truth, usually as the mean absolute error (MAE), the root mean squared error (RMSE), which punishes big misses more, or R², the share of the variation the model explains. The name comes from Francis Galton, who in the 1880s noticed that the children of very tall parents tend to be closer to average height, as if their height "regressed" toward the mean.
Regression vs classification is the core distinction in supervised learning: regression outputs a number, classification a category. Logistic regression, despite its name, is a classification method. And regression in machine learning has nothing to do with a regression in software testing, a bug that breaks something that used to work, which regression testing is meant to catch.
Key takeaways
- Regression predicts a number; classification predicts a category.
- Linear regression fits a straight line by least squares and is the classic starting point.
- Errors are measured with MAE, RMSE or R².
- Logistic regression is a classifier, despite its name.
- Regression in machine learning is unrelated to a regression bug in software testing.
Example
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error
# Flat size in square meters (X) and sale price in thousands (y)
X = [[50], [65], [80], [100], [120]]
y = [150, 190, 235, 290, 345]
model = LinearRegression().fit(X, y)
print(model.coef_[0], model.intercept_) # price added per square meter, and the base
print(model.predict([[90]])) # estimated price of a 90 square meter flat
print(mean_absolute_error(y, model.predict(X))) # average miss, in thousandsReaders ask
What is the difference between regression and classification?
Regression predicts a number, such as a price or a temperature; classification predicts a category, such as spam or not spam. A quick test: if it makes sense to ask how far off a prediction was, it is regression.
Is logistic regression a regression model?
No, despite the name it is used for classification. It computes a probability between 0 and 1 that an input belongs to a class, and a threshold turns that probability into a predicted class.
Is regression in machine learning the same as a regression in testing?
No. In testing, a regression is a bug where something that used to work breaks after a change, and regression tests rerun old checks to catch it. The two meanings of the word are unrelated.
See also
- Supervised LearningAI & Machine Learning, p. 47Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.
- ClassificationAI & Machine Learning, p. 10Classification is the kind of supervised learning where a model predicts a category for each input, such as spam or not spam, or which digit a photo shows.
- Machine LearningAI & Machine Learning, p. 29Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Gradient DescentAI & Machine Learning, p. 23Gradient descent is an optimization algorithm that trains machine learning models by repeatedly nudging their parameters in the direction that reduces error.
- UnderfittingAI & Machine Learning, p. 55Underfitting happens when a machine learning model is too simple or too briefly trained to learn the real pattern, so it does poorly even on its training data.
- Regression TestingTesting & Quality, p. 22Regression testing is the practice of re-running existing tests after a code change to make sure that features which used to work have not broken.
Sources
Spotted a mistake or something missing on this page?Suggest an edit