Table of Contents
Machine learning models recommend videos, flag fraud, read medical scans and power chatbots, and every one of them sometimes gets things wrong. Often the mistakes are not random. They follow from how the model learned, and once you know the patterns, you can predict when a model is likely to fail, and avoid building one that will.
Rather than describe these failures in the abstract, we ran five small experiments with Python and scikit-learn, one for each common cause. Every number below comes from actually running the code, with fixed random seeds so anyone can reproduce it. You do not need to follow the code to follow the ideas, but it is there for students who want to try it.
How a model learns, in one paragraph
A machine learning model is a mathematical function with adjustable settings. Training means showing it many examples, inputs with the right answers, and adjusting the settings until its outputs match those answers as closely as possible. The hope is that it has learned the underlying pattern, so it will also be right about new examples it has never seen. Almost every mistake comes from that hope not quite coming true.
1. Overfitting: memorising instead of learning
Imagine a student who memorises the answers to last year's paper instead of understanding the topic. They score perfectly on that paper and badly on the real exam. Models do the same thing. Here, fifteen noisy measurements of a smooth wave are fitted with three models of increasing flexibility:
import numpy as np
from numpy.polynomial import Polynomial
rng = np.random.default_rng(0)
def truth(x):
return np.sin(x) # the real pattern
x_train = np.sort(rng.uniform(0, 6, 15))
y_train = truth(x_train) + rng.normal(0, 0.25, 15) # 15 noisy measurements
x_test = np.linspace(0.2, 5.8, 200)
y_test = truth(x_test) + rng.normal(0, 0.25, 200) # new data the model never saw
def rmse(a, b):
return float(np.sqrt(np.mean((a - b) ** 2)))
for degree in (1, 4, 12):
model = Polynomial.fit(x_train, y_train, degree)
print(f"degree {degree:>2}: error on training data {rmse(model(x_train), y_train):.2f}"
f" error on new data {rmse(model(x_test), y_test):.2f}")
degree 1: error on training data 0.43 error on new data 0.65
degree 4: error on training data 0.13 error on new data 0.28
degree 12: error on training data 0.04 error on new data 9.61
The straight line is too simple to follow the wave, so it is wrong everywhere, a problem called underfitting. The most flexible model has the lowest error on the training data, because it bends to pass close to every noisy point, but its error on new data is by far the worst of the three: between and beyond the training points it swings wildly. It learned the noise, not the wave. That is overfitting, and it is the single most common reason a model that looked brilliant in testing disappoints in real use.
The rule every data scientist follows
Never judge a model on the data it was trained on. Always hold back data it has never seen and measure it there. A model that is much better on training data than on new data is overfitting.
2. Not enough data
A model can only learn patterns that appear in its examples. With too few, it learns accidents of the particular examples it happened to see. Here is the same model trained on more and more examples, and always tested on the same two thousand new ones:
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
X, y = make_classification(n_samples=6000, n_features=10, n_informative=5, random_state=1)
X_pool, X_test, y_pool, y_test = train_test_split(X, y, test_size=2000, random_state=1)
for n in (20, 100, 500, 4000):
model = DecisionTreeClassifier(max_depth=5, random_state=1).fit(X_pool[:n], y_pool[:n])
print(f"trained on {n:>4} examples: accuracy on new data {model.score(X_test, y_test):.0%}")
trained on 20 examples: accuracy on new data 71%
trained on 100 examples: accuracy on new data 76%
trained on 500 examples: accuracy on new data 85%
trained on 4000 examples: accuracy on new data 86%
Nothing about the model changed. Only the amount of data did. This is why companies value data so highly, and why a model trained on a small, narrow set of examples should always be treated with suspicion.
3. A high score that hides failure
Sometimes the model is fine but the way it is being judged is not. Suppose 5% of transactions are fraudulent. A "model" that simply says "not fraud" every time will be right about 95% of the time, and completely useless:
import numpy as np
from sklearn.metrics import accuracy_score, recall_score
rng = np.random.default_rng(3)
y_true = (rng.random(1000) < 0.05).astype(int) # 5% of 1,000 cases are real positives
lazy_prediction = np.zeros(1000, dtype=int) # a "model" that always says no
print("positives in the data:", y_true.sum())
print(f"accuracy: {accuracy_score(y_true, lazy_prediction):.1%}")
print(f"positives it caught (recall): {recall_score(y_true, lazy_prediction):.0%}")
positives in the data: 47
accuracy: 95.3%
positives it caught (recall): 0%
This is called the accuracy paradox. For rare events such as fraud, disease or equipment failure, you have to measure how many of the real cases the model catches (called recall) and how many of its alarms are real (called precision). A headline accuracy number on its own can hide a model that never does its actual job.
4. Questions outside its experience
Models are generally reliable only for inputs similar to the ones they trained on. Ask about something outside that range and they extrapolate, often confidently and wrongly. Here a simple model learns how ice cream sales rise with temperature, using only spring days between 10 and 25 degrees:
import numpy as np
from sklearn.linear_model import LinearRegression
# ice cream sales rise with temperature, but not in a straight line forever
def sales(temp):
return 20 * temp - 0.4 * temp ** 2
train_temps = np.arange(10, 26).reshape(-1, 1) # trained on a mild spring, 10 to 25 degrees
model = LinearRegression().fit(train_temps, sales(train_temps.ravel()))
for t in (15, 25, 35, 45):
predicted = model.predict([[t]])[0]
print(f"{t} degrees: predicted {predicted:6.0f} actual {sales(t):6.0f}"
+ (" (inside training range)" if 10 <= t <= 25 else " (outside training range)"))
15 degrees: predicted 204 actual 210 (inside training range)
25 degrees: predicted 264 actual 250 (inside training range)
35 degrees: predicted 324 actual 210 (outside training range)
45 degrees: predicted 384 actual 90 (outside training range)
Inside its training range, the model is close. Far outside it, it predicts sales will keep climbing, when in reality they level off and fall as it gets too hot to go out. The model was never shown a heatwave, so it had no way to know. The same thing happens when a model trained before a big change, a new product, a pandemic, a different country, is used after it. Data scientists call this distribution shift.
5. Unbalanced training data
If one group dominates the training data, the model learns the pattern that fits that group, and can do noticeably worse for everyone else, even when nobody intended it. Here, 95% of the training examples come from group A and 5% from group B, whose pattern is slightly different:
import numpy as np
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(7)
def group(n, shift):
X = rng.normal(0, 1, (n, 2))
y = (X[:, 0] + shift * X[:, 1] > 0).astype(int) # the right rule differs a little by group
return X, y
XA, yA = group(1900, 0.2) # group A: 95% of the training data
XB, yB = group(100, -1.5) # group B: only 5%
model = LogisticRegression().fit(np.vstack([XA, XB]), np.concatenate([yA, yB]))
XA_new, yA_new = group(1000, 0.2)
XB_new, yB_new = group(1000, -1.5)
print(f"accuracy for group A: {model.score(XA_new, yA_new):.0%}")
print(f"accuracy for group B: {model.score(XB_new, yB_new):.0%}")
accuracy for group A: 99%
accuracy for group B: 64%
The same model is right 99% of the time for the group it saw most, and 64% for the group it barely saw. In real systems this is how face recognition, speech recognition or medical tools can work well for some people and poorly for others. It is also why testing a model separately for different groups, not just overall, matters so much.
What this means for chatbots and everyday AI
Large language models such as ChatGPT are far bigger than these examples, but related problems appear in them too. They can state wrong facts confidently, especially about topics that were rare in their training data, and they can do worse on questions unlike anything they were trained on. That is why checking AI answers matters, and why our guide on using ChatGPT to study without cheating spends a whole section on catching confident wrong answers. For how language models work underneath, see how AI actually works.
A model is only as good as the examples it learned from and the test it was judged by.
How we teach it
In our AI and machine learning classes, students run experiments like these themselves before learning the theory, in line with the learning-by-building principle on our how we teach page. Seeing a model overfit on your own screen teaches more than a definition. See AI and machine learning with Python, taught live, one to one or in small groups of 5 to 10.
Frequently asked questions
Common reasons are overfitting (memorising training examples instead of learning the pattern), too little data, being judged by a misleading score, being asked about situations unlike their training data, and training data that under-represents some groups. Each of these can be measured and reduced.
Overfitting is when a model memorises its training examples, including their noise, instead of learning the general pattern. It performs very well on the data it was trained on and worse on new data, like a student who memorised last year's answers.
Compare its performance on training data with its performance on data it has never seen. If it is much better on the training data, it is overfitting. That is why data is always split into training and test sets.
When one outcome is rare, a model can achieve high accuracy by always predicting the common outcome. With 5% fraud, always saying "not fraud" is 95% accurate but catches nothing. Precision and recall show whether the model finds the rare cases.
Often it helps a lot, as the experiment in this post shows, but not always. More data does not fix a misleading score, and more of the same biased data does not fix unbalanced data. The data has to cover the situations the model will face.
Usually because some groups were under-represented in the training data, so the model learned patterns that fit the majority better. Testing performance separately for each group, and collecting more balanced data, are the main ways to find and reduce this.
No. The ideas in this post need only an intuitive understanding of patterns and averages. Building and tuning models well does use more maths, especially statistics and linear algebra, which can be learned step by step.