Underfitting
- In Turkish
- Yetersiz Öğrenme
In short
Underfitting happens when a machine learning model is too simple or too briefly trained to learn the real pattern, so it does poorly even on its training data.
What is underfitting?
Underfitting is the opposite of memorizing: the model hasn't learned enough. It misses the real relationship between inputs and outputs, so its predictions are poor on the data it was trained on and just as poor on new data. The telltale sign is that training and validation scores are both low and close together.
Common causes are a model that is too simple for the problem, such as a straight line fitted to a curved pattern, inputs that don't carry the information needed, regularization that is too strong, and training that stops too early. The fixes go the other way: a more flexible model, better or additional features, weaker regularization and longer training. Unlike with overfitting, collecting more of the same data rarely helps, because the model can't yet use the information it already has.
Picture a student who only skimmed the chapter titles before an exam. They fail the practice questions and the real exam alike, not because the questions changed but because they never learned the material. With large language models, a similar effect shows up when fine-tuning runs for too few steps: the model hasn't yet picked up the new format or style.
Underfitting vs overfitting: an underfit model is too simple and does badly everywhere, while an overfit model is too complex, memorizes noise and does well only on its training data. The tension between the two is called the bias-variance trade-off. Underfitting means high bias, a model whose built-in assumptions are too rigid; overfitting means high variance, a model that changes too much with the particular training examples. The goal is the point in between, where the score on held-out data is best.
Key takeaways
- An underfit model does poorly on both its training data and new data.
- Too simple a model, weak features, heavy regularization or too little training cause it.
- More capacity, better features and longer training fix it; more of the same data rarely does.
- Overfitting is the opposite: great on training data, poor on new data.
- Comparing training and validation scores shows which problem you have.
Example
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures
rng = np.random.default_rng(0)
X = rng.uniform(-3, 3, size=(200, 1))
y = X[:, 0] ** 2 + rng.normal(0, 0.5, 200) # a curved pattern with some noise
line = LinearRegression().fit(X, y) # a straight line can't follow a curve
print(line.score(X, y)) # R² near 0, even on its own training data
curve = make_pipeline(PolynomialFeatures(2), LinearRegression()).fit(X, y)
print(curve.score(X, y)) # close to 1: now flexible enoughReaders ask
How do you know if a model is underfitting?
Check its score on the training data itself. If it is poor there and the validation score is about as poor, the model is underfitting; if training is excellent but validation is much worse, it is overfitting.
What is the difference between underfitting and overfitting?
An underfit model is too simple and misses the pattern, so it fails on all data. An overfit model is too complex and learns noise, so it does well only on the examples it trained on.
Does adding more data fix underfitting?
Usually not. A model that can't capture the pattern in the data it has won't capture it in more of the same; it needs more capacity, better features or more training. More data is the classic fix for overfitting instead.
See also
- OverfittingAI & Machine Learning, p. 36Overfitting happens when a machine learning model learns its training data so closely, including its noise, that it performs poorly on new, unseen data.
- Supervised LearningAI & Machine Learning, p. 47Supervised learning is machine learning where a model learns from labeled examples, inputs paired with correct answers, to predict outputs for new data.
- Training DataAI & Machine Learning, p. 53Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
- Model ParametersAI & Machine Learning, p. 32Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
- RegressionAI & Machine Learning, p. 42Regression is the kind of supervised learning where a model predicts a number, such as a price, a temperature or a delivery time, from the input it is given.
- Machine LearningAI & Machine Learning, p. 29Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
Sources
- scikit-learn documentation: Underfitting vs. Overfittingscikit-learn.org(opens in a new tab)
- scikit-learn documentation: Validation curves: plotting scores to evaluate modelsscikit-learn.org(opens in a new tab)
- Goodfellow, Bengio and Courville: Deep Learning, Chapter 5: Machine Learning Basics (2016)deeplearningbook.org(opens in a new tab)
Spotted a mistake or something missing on this page?Suggest an edit