Artificial Intelligence

Why Do Machine Learning Models Make Mistakes?

Five common causes, each demonstrated with a small experiment you can reproduce: overfitting, too little data, misleading scores, unfamiliar situations and unbalanced data.

Modern Age Coders Team
Modern Age Coders Team September 28, 2026
8 min read
Three machine learning models fitted to the same points: too simple, about right, and overfitted to noise

Machine learning models recommend videos, flag fraud, read medical scans and power chatbots, and every one of them sometimes gets things wrong. Often the mistakes are not random. They follow from how the model learned, and once you know the patterns, you can predict when a model is likely to fail, and avoid building one that will.

Rather than describe these failures in the abstract, we ran five small experiments with Python and scikit-learn, one for each common cause. Every number below comes from actually running the code, with fixed random seeds so anyone can reproduce it. You do not need to follow the code to follow the ideas, but it is there for students who want to try it.

How a model learns, in one paragraph

A machine learning model is a mathematical function with adjustable settings. Training means showing it many examples, inputs with the right answers, and adjusting the settings until its outputs match those answers as closely as possible. The hope is that it has learned the underlying pattern, so it will also be right about new examples it has never seen. Almost every mistake comes from that hope not quite coming true.

1. Overfitting: memorising instead of learning

Imagine a student who memorises the answers to last year's paper instead of understanding the topic. They score perfectly on that paper and badly on the real exam. Models do the same thing. Here, fifteen noisy measurements of a smooth wave are fitted with three models of increasing flexibility:

overfitting.py
import numpy as np
from numpy.polynomial import Polynomial

rng = np.random.default_rng(0)
def truth(x):
    return np.sin(x)                                  # the real pattern

x_train = np.sort(rng.uniform(0, 6, 15))
y_train = truth(x_train) + rng.normal(0, 0.25, 15)    # 15 noisy measurements
x_test = np.linspace(0.2, 5.8, 200)
y_test = truth(x_test) + rng.normal(0, 0.25, 200)     # new data the model never saw

def rmse(a, b):
    return float(np.sqrt(np.mean((a - b) ** 2)))

for degree in (1, 4, 12):
    model = Polynomial.fit(x_train, y_train, degree)
    print(f"degree {degree:>2}: error on training data {rmse(model(x_train), y_train):.2f}"
          f"   error on new data {rmse(model(x_test), y_test):.2f}")
Output
degree  1: error on training data 0.43   error on new data 0.65
degree  4: error on training data 0.13   error on new data 0.28
degree 12: error on training data 0.04   error on new data 9.61
Training error and new-data error for polynomial models of degree 1, 4 and 12, showing the degree 12 model fits training data best but new data worse
The most flexible model wins on old data and loses on new data.

The straight line is too simple to follow the wave, so it is wrong everywhere, a problem called underfitting. The most flexible model has the lowest error on the training data, because it bends to pass close to every noisy point, but its error on new data is by far the worst of the three: between and beyond the training points it swings wildly. It learned the noise, not the wave. That is overfitting, and it is the single most common reason a model that looked brilliant in testing disappoints in real use.

๐Ÿ’ก

The rule every data scientist follows

Never judge a model on the data it was trained on. Always hold back data it has never seen and measure it there. A model that is much better on training data than on new data is overfitting.

2. Not enough data

A model can only learn patterns that appear in its examples. With too few, it learns accidents of the particular examples it happened to see. Here is the same model trained on more and more examples, and always tested on the same two thousand new ones:

data_size.py
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier

X, y = make_classification(n_samples=6000, n_features=10, n_informative=5, random_state=1)
X_pool, X_test, y_pool, y_test = train_test_split(X, y, test_size=2000, random_state=1)

for n in (20, 100, 500, 4000):
    model = DecisionTreeClassifier(max_depth=5, random_state=1).fit(X_pool[:n], y_pool[:n])
    print(f"trained on {n:>4} examples: accuracy on new data {model.score(X_test, y_test):.0%}")
Output
trained on   20 examples: accuracy on new data 71%
trained on  100 examples: accuracy on new data 76%
trained on  500 examples: accuracy on new data 85%
trained on 4000 examples: accuracy on new data 86%
Accuracy on new data rising as a decision tree is trained on 20, 100, 500 and 4,000 examples
Nothing changed but the amount of data.

Nothing about the model changed. Only the amount of data did. This is why companies value data so highly, and why a model trained on a small, narrow set of examples should always be treated with suspicion.

3. A high score that hides failure

Sometimes the model is fine but the way it is being judged is not. Suppose 5% of transactions are fraudulent. A "model" that simply says "not fraud" every time will be right about 95% of the time, and completely useless:

accuracy_paradox.py
import numpy as np
from sklearn.metrics import accuracy_score, recall_score

rng = np.random.default_rng(3)
y_true = (rng.random(1000) < 0.05).astype(int)       # 5% of 1,000 cases are real positives
lazy_prediction = np.zeros(1000, dtype=int)          # a "model" that always says no

print("positives in the data:", y_true.sum())
print(f"accuracy: {accuracy_score(y_true, lazy_prediction):.1%}")
print(f"positives it caught (recall): {recall_score(y_true, lazy_prediction):.0%}")
Output
positives in the data: 47
accuracy: 95.3%
positives it caught (recall): 0%
A model that always predicts no reaching 95.3% accuracy while catching none of the real positive cases
Accuracy sounds great and means nothing here.

This is called the accuracy paradox. For rare events such as fraud, disease or equipment failure, you have to measure how many of the real cases the model catches (called recall) and how many of its alarms are real (called precision). A headline accuracy number on its own can hide a model that never does its actual job.

4. Questions outside its experience

Models are generally reliable only for inputs similar to the ones they trained on. Ask about something outside that range and they extrapolate, often confidently and wrongly. Here a simple model learns how ice cream sales rise with temperature, using only spring days between 10 and 25 degrees:

distribution_shift.py
import numpy as np
from sklearn.linear_model import LinearRegression

# ice cream sales rise with temperature, but not in a straight line forever
def sales(temp):
    return 20 * temp - 0.4 * temp ** 2

train_temps = np.arange(10, 26).reshape(-1, 1)        # trained on a mild spring, 10 to 25 degrees
model = LinearRegression().fit(train_temps, sales(train_temps.ravel()))

for t in (15, 25, 35, 45):
    predicted = model.predict([[t]])[0]
    print(f"{t} degrees: predicted {predicted:6.0f}   actual {sales(t):6.0f}"
          + ("   (inside training range)" if 10 <= t <= 25 else "   (outside training range)"))
Output
15 degrees: predicted    204   actual    210   (inside training range)
25 degrees: predicted    264   actual    250   (inside training range)
35 degrees: predicted    324   actual    210   (outside training range)
45 degrees: predicted    384   actual     90   (outside training range)

Inside its training range, the model is close. Far outside it, it predicts sales will keep climbing, when in reality they level off and fall as it gets too hot to go out. The model was never shown a heatwave, so it had no way to know. The same thing happens when a model trained before a big change, a new product, a pandemic, a different country, is used after it. Data scientists call this distribution shift.

5. Unbalanced training data

If one group dominates the training data, the model learns the pattern that fits that group, and can do noticeably worse for everyone else, even when nobody intended it. Here, 95% of the training examples come from group A and 5% from group B, whose pattern is slightly different:

unbalanced_data.py
import numpy as np
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(7)
def group(n, shift):
    X = rng.normal(0, 1, (n, 2))
    y = (X[:, 0] + shift * X[:, 1] > 0).astype(int)  # the right rule differs a little by group
    return X, y

XA, yA = group(1900, 0.2)       # group A: 95% of the training data
XB, yB = group(100, -1.5)       # group B: only 5%
model = LogisticRegression().fit(np.vstack([XA, XB]), np.concatenate([yA, yB]))

XA_new, yA_new = group(1000, 0.2)
XB_new, yB_new = group(1000, -1.5)
print(f"accuracy for group A: {model.score(XA_new, yA_new):.0%}")
print(f"accuracy for group B: {model.score(XB_new, yB_new):.0%}")
Output
accuracy for group A: 99%
accuracy for group B: 64%

The same model is right 99% of the time for the group it saw most, and 64% for the group it barely saw. In real systems this is how face recognition, speech recognition or medical tools can work well for some people and poorly for others. It is also why testing a model separately for different groups, not just overall, matters so much.

Five reasons machine learning models make mistakes, each measured in this post: overfitting, too little data, misleading scores, new situations, and unbalanced data
Every one of these is visible in the experiments above.

What this means for chatbots and everyday AI

Large language models such as ChatGPT are far bigger than these examples, but related problems appear in them too. They can state wrong facts confidently, especially about topics that were rare in their training data, and they can do worse on questions unlike anything they were trained on. That is why checking AI answers matters, and why our guide on using ChatGPT to study without cheating spends a whole section on catching confident wrong answers. For how language models work underneath, see how AI actually works.

A model is only as good as the examples it learned from and the test it was judged by.

How we teach it

In our AI and machine learning classes, students run experiments like these themselves before learning the theory, in line with the learning-by-building principle on our how we teach page. Seeing a model overfit on your own screen teaches more than a definition. See AI and machine learning with Python, taught live, one to one or in small groups of 5 to 10.

Frequently asked questions

Common reasons are overfitting (memorising training examples instead of learning the pattern), too little data, being judged by a misleading score, being asked about situations unlike their training data, and training data that under-represents some groups. Each of these can be measured and reduced.

Overfitting is when a model memorises its training examples, including their noise, instead of learning the general pattern. It performs very well on the data it was trained on and worse on new data, like a student who memorised last year's answers.

Compare its performance on training data with its performance on data it has never seen. If it is much better on the training data, it is overfitting. That is why data is always split into training and test sets.

When one outcome is rare, a model can achieve high accuracy by always predicting the common outcome. With 5% fraud, always saying "not fraud" is 95% accurate but catches nothing. Precision and recall show whether the model finds the rare cases.

Often it helps a lot, as the experiment in this post shows, but not always. More data does not fix a misleading score, and more of the same biased data does not fix unbalanced data. The data has to cover the situations the model will face.

Usually because some groups were under-represented in the training data, so the model learned patterns that fit the majority better. Testing performance separately for each group, and collecting more balanced data, are the main ways to find and reduce this.

No. The ideas in this post need only an intuitive understanding of patterns and averages. Building and tuning models well does use more maths, especially statistics and linear algebra, which can be learned step by step.

Modern Age Coders Team

About Modern Age Coders Team

Expert educators making coding and maths clear for ages 6 to 67.

Keep exploring Modern Age Coders

More from the blog

Learn more

Free resources

From the blog

Start here

Ask Misti AI
Chat with us
Enroll Watch Class Priority Demo Enrol Book a Demo Watch Class WhatsApp Book demo today