New · France 2027 presidential election: what the candidates propose on AI, quoted and sourced. Explore the tracker →

Last reviewed:

What is overfitting? Definition and difference with underfitting

Overfitting describes an AI model that fits its training data so closely that it memorises it instead of learning general rules: it scores very well in internal tests, then makes mistakes on new cases. It is one of the most common reasons why a model that shines in a demo disappoints in production.

In its machine learning course, Google defines overfitting as creating a model that matches its training set so closely that it fails to make correct predictions on new data. The desired opposite is called generalisation: predicting well on cases never seen before. Underfitting is the reverse defect: a model too simple, which fails even on its training data. Two causes dominate. Training data that poorly represents reality (too little, too old, from a single site or a single season). A model too complex for the amount of data available, with enough parameters to learn the noise by heart. The classic sign: error keeps falling on the training data while it rises on validation data set aside. The basic safeguard is to split the data into three sets (training, validation, test) and to judge the model only on data it has never seen. Regularisation techniques come on top: in 2012, the AlexNet network, with its 60 million parameters, already used a dedicated method to limit overfitting. LLMs face a variant of the problem, contamination: if a test's questions appear in the training data, the score measures the model's memory rather than its ability to reason.

Concrete example

Real case, documented in the journal Science in March 2014: Google Flu Trends, launched in 2008, estimated flu activity in the United States from Google searches. The initial method searched among 50 million search terms for those that best matched 1,152 official data points. The authors of the article see overfitting in this: seasonal terms unrelated to flu, such as high school basketball, followed the same curve. In February 2013, the tool predicted more than double the doctor visits for influenza-like illness reported by the US Centers for Disease Control and Prevention (CDC), and it had overestimated flu in 100 out of 108 weeks starting in August 2011.

Illustrative case: a French mid-sized services company predicts customer churn. In the demo, the model shows 96% correct predictions. In production, it falls below 65%. The audit shows it had been tested on the same customers as those used for training.

Comparison

Underfitting, good fit and overfitting
UnderfittingGood fitOverfitting
On training dataPoorGoodExcellent, sometimes perfect
On new dataPoorClose to the training scoreClearly worse
Typical causeModel too simple, missing variablesModel and data well sizedModel too complex, too little or unrepresentative data
What the executive seesThe model fails to convince from the demoPilot results hold up in productionBrilliant demo, disappointment in production
RemedyRicher model, better variablesRegular performance monitoringMore real data, testing on unseen data, regularisation

FAQ

What is overfitting in AI?

It is the defect of a model that has memorised its training data instead of learning general rules. It performs very well on examples it has already seen and makes mistakes on new ones. It is said to generalise poorly.

What is the difference between overfitting and underfitting?

Overfitting is a model too closely fitted to its data: excellent in training, poor on new cases. Underfitting is a model too simple: poor everywhere, including on its training data. The goal is the middle ground, a model that generalises.

How do you detect overfitting?

By comparing performance on the training data with performance on held-out data the model has never seen. If the gap is large, or if validation error rises while training error falls, the model is overfitting.

How do you prevent overfitting?

Use more data that represents production, split into training, validation and test sets, choose a model suited to the amount of data, apply regularisation techniques and stop training when validation error stops falling.

What is an example of overfitting?

Google Flu Trends: according to a 2014 article in Science, the tool selected search terms among 50 million to match 1,152 data points, including seasonal terms unrelated to flu. In February 2013, it predicted more than double the doctor visits measured by US health authorities.

Can LLMs like ChatGPT overfit?

Yes, in a particular form. An LLM can recite passages from its training data, or pass a public test because it saw the questions during training (contamination). That is why a serious evaluation uses your own, unpublished cases.

See also

Further reading

Overfitting, Machine Learning Crash Course, Google for Developers (external resource)

Sources

  1. Overfitting, Machine Learning Crash Course, Google for Developers. https://developers.google.com/machine-learning/crash-course/overfitting/overfitting (accessed 2026-09-30)
  2. The Parable of Google Flu: Traps in Big Data Analysis, Lazer, Kennedy, King and Vespignani, Science, vol. 343, no. 6176, March 14, 2014. https://www.science.org/doi/10.1126/science.1248506 (accessed 2026-09-30)
  3. ImageNet Classification with Deep Convolutional Neural Networks, Krizhevsky, Sutskever and Hinton, NeurIPS (NIPS) 2012. https://proceedings.neurips.cc/paper_files/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html (accessed 2026-09-30)

← Back to glossary

Address copied