# Classification

URL: https://softwaredictionary.org/terms/classification
Category: AI & Machine Learning
Last updated: 2026-10-06
In Turkish: Sınıflandırma

In short: Classification is the kind of supervised learning where a model predicts a category for each input, such as spam or not spam, or which digit a photo shows.

## What is classification in machine learning?

Classification is one of the two main kinds of supervised learning; the other is regression, which predicts numbers. A classification model learns from labeled examples, inputs paired with the correct category, and then assigns one of those categories, called classes, to new inputs. A spam filter, a fraud detector and a model that tags support tickets as billing, bug or feature request are all classifiers.

There are three common setups. Binary classification chooses between two classes, such as fraud or not fraud; multiclass classification picks exactly one of several, such as which digit from 0 to 9 an image shows; multi-label classification can attach several tags at once, such as the topics of an article. Most classifiers output a probability or score for each class, and the application turns it into a decision with a threshold, often 0.5 by default, which can be moved to catch more cases or to raise fewer false alarms. Logistic regression, decision trees, gradient-boosted trees and neural networks are all common classification models.

Accuracy, the share of correct predictions, is the obvious score but can mislead. If only 1 in 100 transactions is fraud, a model that always answers "not fraud" is 99 percent accurate and useless. That is why classifiers are also judged by precision (how many flagged cases were real), recall (how many real cases were caught) and a confusion matrix, a table that counts each kind of right and wrong answer.

Classification vs regression is the key distinction: classification predicts a category, regression a number, such as a price or a delivery time. Confusingly, logistic regression is a classification model; the name comes from the logistic function it uses to turn a score into a probability. Classification is also different from clustering, an unsupervised technique that groups unlabeled data without being told what the groups are.

## Key takeaways

- Classification predicts a category, called a class, for each input.
- It is supervised: the model learns from examples labeled with the right class.
- Binary, multiclass and multi-label classification are the three common setups.
- Accuracy can mislead on imbalanced data; precision and recall tell more.
- Regression predicts numbers instead, and logistic regression is in fact a classifier.

## Example: Training and evaluating a classifier with scikit-learn

```python
from sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
from sklearn.model_selection import train_test_split

# Labeled data: tumor measurements (X) and the class, 0 = malignant, 1 = benign (y)
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

# Despite its name, logistic regression is a classification model
clf = LogisticRegression(max_iter=5000).fit(X_train, y_train)

print(clf.predict_proba(X_test[:1]))  # probability of each class
print(classification_report(y_test, clf.predict(X_test)))  # precision, recall, F1
```

## Frequently asked questions

**What is the difference between classification and regression?**

Classification predicts a category, such as spam or not spam; regression predicts a number, such as tomorrow's temperature. Both are supervised learning, and some algorithms, such as decision trees, come in both versions.

**What is a confusion matrix?**

A table that compares a classifier's predictions with the true labels. For two classes it counts true positives, false positives, true negatives and false negatives, which shows exactly what kind of mistakes the model makes.

**Can LLMs do classification?**

Yes. A language model can sort text into categories from a prompt alone, sometimes with a few examples, which is handy when there is little labeled data. At high volumes, a small trained classifier is usually cheaper, faster and more consistent.

## Sources

- [scikit-learn documentation: Classification metrics](https://scikit-learn.org/stable/modules/model_evaluation.html#classification-metrics)
- [scikit-learn documentation: Logistic regression](https://scikit-learn.org/stable/modules/linear_model.html#logistic-regression)

---

Software Dictionary: https://softwaredictionary.org/ · https://softwaredictionary.org/llms.txt
