Skip to main content

Side by side

OverfittingvsUnderfitting

What is the difference between overfitting and underfitting?

Updated 3 min read8 differences

In short

An overfit model memorizes its training data, noise included, and fails on new data; an underfit model is too simple to learn the pattern and fails on both.

Overfitting

Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.

Read the page on Overfitting

Underfitting

Underfitting happens when a machine learning model is too simple or too briefly trained to learn the real pattern, so it does poorly even on its training data.

Read the page on Underfitting

Overfitting and Underfitting compared

AspectOverfittingUnderfitting
What goes wrongThe model memorizes training examples, noise includedThe model misses the real pattern
Training scoreVery highLow
Score on new dataMuch worse than on the training dataAbout as low as on the training data
Model complexityToo flexible for the amount of dataToo simple for the problem
Bias and varianceLow bias, high varianceHigh bias, low variance
Typical causesFew examples, a very large model, training for too longWeak features, heavy regularization, training stopped too early
FixesMore data, a simpler model, regularization, early stoppingMore capacity, better features, less regularization, longer training
Does more data help?Yes, it is the classic remedyRarely, if it is more of the same data

The difference, explained

A machine learning model should learn general patterns that also hold for data it has never seen. Overfitting is learning too much: the model memorizes its training examples, including their noise and quirks, so it scores very well on training data and noticeably worse on new data. Underfitting is learning too little: the model misses the real relationship between inputs and outputs, so it does poorly on its training data and just as poorly on new data.

The two are opposite ends of the bias-variance trade-off. An underfit model has high bias: its built-in assumptions are too rigid, like a straight line fitted to a curved pattern. An overfit model has high variance: it is so flexible that it changes with the particular examples it saw, like a large neural network trained on a few hundred examples, or trained for too long. The goal is the point in between, where the score on held-out data is best.

Comparing training and validation scores shows which problem you have. If both are low and close together, the model is underfitting; if training is excellent but validation is much worse, it is overfitting. The fixes run in opposite directions: underfitting calls for a more flexible model, better features, weaker regularization and longer training, while overfitting calls for more and more varied data, a simpler model, stronger regularization such as dropout or weight decay, and early stopping.

A common misconception is that more data always helps. It is the classic remedy for overfitting, but an underfit model can't use the information it already has, so more of the same data rarely fixes it. Another is that a high training score means a good model: on its own it can't tell real learning from memorizing, which is why performance must always be measured on data the model never trained on.

Where each one is a risk

Overfitting is a risk when…

  • The model is large or flexible compared with the amount of training data.
  • Training runs for many epochs without anyone watching the validation score.
  • A large model is fine-tuned on a small dataset for too many steps.
  • Features carry noise, IDs or details that won't repeat in new data.

Underfitting is a risk when…

  • A simple model, such as a straight line, faces a curved or complex pattern.
  • The inputs lack the information needed to predict the target.
  • Regularization is too strong, or training stops too early.
  • A fine-tuning run is too short for the model to pick up the new format.

One curve fitted too closely and too loosely (scikit-learn)

Overfittingpython
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures

rng = np.random.default_rng(0)
X = rng.uniform(-3, 3, size=(30, 1)); y = X[:, 0] ** 2 + rng.normal(0, 1, 30)
X_new = rng.uniform(-3, 3, size=(30, 1)); y_new = X_new[:, 0] ** 2 + rng.normal(0, 1, 30)

wiggly = make_pipeline(PolynomialFeatures(15), LinearRegression()).fit(X, y)
print(wiggly.score(X, y))          # high: the curve bends through the noisy points
print(wiggly.score(X_new, y_new))  # much lower on new data: overfitting
Underfittingpython
# The same data as in the overfitting example
line = LinearRegression().fit(X, y)  # a straight line can't follow a curve
print(line.score(X, y))              # near 0, even on its own training data
print(line.score(X_new, y_new))      # just as low on new data: underfitting

# In between: degree 2 matches the real pattern
good = make_pipeline(PolynomialFeatures(2), LinearRegression()).fit(X, y)
print(good.score(X_new, y_new))      # the best of the three on new data

Readers ask

How do you tell overfitting from underfitting?

Compare the score on the training data with the score on a validation set. A high training score with a much lower validation score means overfitting; low scores on both mean underfitting.

What is the bias-variance trade-off?

The tension between a model that is too rigid, with high bias and underfitting, and one that is too sensitive to its training examples, with high variance and overfitting. Making a model more flexible lowers bias but raises variance, so the best model sits in between.

Do large language models overfit or underfit?

Both can happen during fine-tuning. Too many steps on a small dataset can make a model repeat training examples word for word and lose general abilities; too few leave it without the new format or style. Watching the validation loss helps find the right length.

More

Settings