Skip to main content

Side by side

ClassificationvsRegression

What is the difference between classification and regression?

Updated 3 min read7 differences

In short

Classification predicts a category, such as spam or not spam; regression predicts a number, such as a price or a delivery time. Both are supervised learning.

Classification

Classification is the kind of supervised learning where a model predicts a category for each input, such as spam or not spam, or which digit a photo shows.

Read the page on Classification

Regression

Regression is the kind of supervised learning where a model predicts a number, such as a price, a temperature or a delivery time, from the input it is given.

Read the page on Regression

Classification and Regression compared

AspectClassificationRegression
PredictsA category, called a classA continuous number
Example questionsIs this email spam? Which digit is in this image?What will this flat sell for? How long will delivery take?
OutputA probability per class, turned into a label by a thresholdA single value, such as 235 or 4.7
Measured withAccuracy, precision, recall and the confusion matrixMAE, RMSE and R²
A wrong answerPicks the wrong classMisses the true value by some amount
Classic first modelLogistic regressionLinear regression
Common pitfallAccuracy misleads when one class is rareOutliers pull the predictions and inflate RMSE

The difference, explained

Classification and regression are the two main kinds of supervised learning: in both, a model learns from labeled examples, inputs paired with the correct answer, and then predicts that answer for new inputs. In classification the answer is a category, called a class, such as fraud or not fraud, or which topic a support ticket is about. In regression it is a continuous number, such as a flat's sale price, tomorrow's temperature or how long a delivery will take.

The type of answer changes how a model is judged. A classifier usually outputs a probability for each class and turns it into a decision with a threshold, and it is measured by accuracy, precision, recall and a confusion matrix. A regression model outputs a number and is measured by how far it misses, with the mean absolute error (MAE), the root mean squared error (RMSE) or R².

Many algorithms come in both versions, such as decision trees, gradient-boosted trees and neural networks, so the choice depends on the question, not the tool. A quick test: if it makes sense to ask how far off a prediction was, it is regression; if the prediction is simply right or wrong, it is classification. A number can also become a class with a cut-off, for example predicting a delivery time and then flagging orders that will be late.

A common source of confusion is logistic regression, which despite its name is a classification model: it computes a probability between 0 and 1, and a threshold turns it into a class. Another is treating labels that happen to be numbers, such as product codes or the digits 0 to 9 in an image, as a regression target; they are names, not quantities, so they call for classification.

Which one should you use?

Choose Classification when…

  • The answer is one of a fixed set of labels, such as spam or not spam.
  • You need a yes-or-no decision, such as approving or flagging a transaction.
  • Items can carry several tags at once, such as the topics of an article.
  • The labels are names that happen to be numbers, such as product codes.

Choose Regression when…

  • The answer is a quantity, such as a price, a duration or a temperature.
  • Being close counts: a small miss is better than a big one.
  • You forecast values such as demand, sales or energy use.
  • You want to know how much each input moves the result, as linear regression shows.

The same inputs predicting a class and a number (scikit-learn)

Classificationpython
from sklearn.tree import DecisionTreeClassifier

# Flat size in m² and distance to the center in km -> sold within a month?
X = [[50, 2], [65, 8], [80, 3], [100, 12], [120, 5]]
y = ["yes", "no", "yes", "no", "yes"]  # categories

clf = DecisionTreeClassifier(random_state=0).fit(X, y)
print(clf.predict([[90, 4]]))        # ['yes']: a class
print(clf.predict_proba([[90, 4]]))  # a probability for each class
Regressionpython
from sklearn.tree import DecisionTreeRegressor

# The same inputs -> sale price in thousands
X = [[50, 2], [65, 8], [80, 3], [100, 12], [120, 5]]
y = [150, 190, 235, 290, 345]  # numbers

reg = DecisionTreeRegressor(random_state=0).fit(X, y)
print(reg.predict([[90, 4]]))  # an estimated price: a number

Readers ask

Is logistic regression classification or regression?

Classification. Despite its name, it computes the probability that an input belongs to a class, and a threshold, often 0.5, turns that probability into a predicted class.

Can the same algorithm do both classification and regression?

Yes. Decision trees, random forests, gradient-boosted trees, k-nearest neighbors and neural networks all come in a classification and a regression version; what changes is the output and the error the model is trained to reduce.

Is predicting a rating from 1 to 5 classification or regression?

Either can work. Treating the stars as classes ignores that 4 is closer to 5 than to 1, while regression keeps that order but may predict values such as 3.6; ordinal regression is a middle ground made for ordered categories.

More

Settings