Side by side
OverfittingvsUnderfitting
What is the difference between overfitting and underfitting?
Updated 3 min read8 differences
In short
An overfit model memorizes its training data, noise included, and fails on new data; an underfit model is too simple to learn the pattern and fails on both.
Overfitting
Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.
Read the page on OverfittingUnderfitting
Underfitting happens when a machine learning model is too simple or too briefly trained to learn the real pattern, so it does poorly even on its training data.
Read the page on UnderfittingOverfitting and Underfitting compared
| Aspect | Overfitting | Underfitting |
|---|---|---|
| What goes wrong | The model memorizes training examples, noise included | The model misses the real pattern |
| Training score | Very high | Low |
| Score on new data | Much worse than on the training data | About as low as on the training data |
| Model complexity | Too flexible for the amount of data | Too simple for the problem |
| Bias and variance | Low bias, high variance | High bias, low variance |
| Typical causes | Few examples, a very large model, training for too long | Weak features, heavy regularization, training stopped too early |
| Fixes | More data, a simpler model, regularization, early stopping | More capacity, better features, less regularization, longer training |
| Does more data help? | Yes, it is the classic remedy | Rarely, if it is more of the same data |
The difference, explained
A machine learning model should learn general patterns that also hold for data it has never seen. Overfitting is learning too much: the model memorizes its training examples, including their noise and quirks, so it scores very well on training data and noticeably worse on new data. Underfitting is learning too little: the model misses the real relationship between inputs and outputs, so it does poorly on its training data and just as poorly on new data.
The two are opposite ends of the bias-variance trade-off. An underfit model has high bias: its built-in assumptions are too rigid, like a straight line fitted to a curved pattern. An overfit model has high variance: it is so flexible that it changes with the particular examples it saw, like a large neural network trained on a few hundred examples, or trained for too long. The goal is the point in between, where the score on held-out data is best.
Comparing training and validation scores shows which problem you have. If both are low and close together, the model is underfitting; if training is excellent but validation is much worse, it is overfitting. The fixes run in opposite directions: underfitting calls for a more flexible model, better features, weaker regularization and longer training, while overfitting calls for more and more varied data, a simpler model, stronger regularization such as dropout or weight decay, and early stopping.
A common misconception is that more data always helps. It is the classic remedy for overfitting, but an underfit model can't use the information it already has, so more of the same data rarely fixes it. Another is that a high training score means a good model: on its own it can't tell real learning from memorizing, which is why performance must always be measured on data the model never trained on.
Where each one is a risk
Overfitting is a risk when…
- The model is large or flexible compared with the amount of training data.
- Training runs for many epochs without anyone watching the validation score.
- A large model is fine-tuned on a small dataset for too many steps.
- Features carry noise, IDs or details that won't repeat in new data.
Underfitting is a risk when…
- A simple model, such as a straight line, faces a curved or complex pattern.
- The inputs lack the information needed to predict the target.
- Regularization is too strong, or training stops too early.
- A fine-tuning run is too short for the model to pick up the new format.
One curve fitted too closely and too loosely (scikit-learn)
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
rng = np.random.default_rng(0)
X = rng.uniform(-3, 3, size=(30, 1)); y = X[:, 0] ** 2 + rng.normal(0, 1, 30)
X_new = rng.uniform(-3, 3, size=(30, 1)); y_new = X_new[:, 0] ** 2 + rng.normal(0, 1, 30)
wiggly = make_pipeline(PolynomialFeatures(15), LinearRegression()).fit(X, y)
print(wiggly.score(X, y)) # high: the curve bends through the noisy points
print(wiggly.score(X_new, y_new)) # much lower on new data: overfitting# The same data as in the overfitting example
line = LinearRegression().fit(X, y) # a straight line can't follow a curve
print(line.score(X, y)) # near 0, even on its own training data
print(line.score(X_new, y_new)) # just as low on new data: underfitting
# In between: degree 2 matches the real pattern
good = make_pipeline(PolynomialFeatures(2), LinearRegression()).fit(X, y)
print(good.score(X_new, y_new)) # the best of the three on new dataReaders ask
How do you tell overfitting from underfitting?
Compare the score on the training data with the score on a validation set. A high training score with a much lower validation score means overfitting; low scores on both mean underfitting.
What is the bias-variance trade-off?
The tension between a model that is too rigid, with high bias and underfitting, and one that is too sensitive to its training examples, with high variance and overfitting. Making a model more flexible lowers bias but raises variance, so the best model sits in between.
Do large language models overfit or underfit?
Both can happen during fine-tuning. Too many steps on a small dataset can make a model repeat training examples word for word and lose general abilities; too few leave it without the new format or style. Watching the validation loss helps find the right length.