Last reviewed:
What is overfitting? Definition and difference with underfitting
Overfitting describes an AI model that fits its training data so closely that it memorises it instead of learning general rules: it scores very well in internal tests, then makes mistakes on new cases. It is one of the most common reasons why a model that shines in a demo disappoints in production.
In its machine learning course, Google defines overfitting as creating a model that matches its training set so closely that it fails to make correct predictions on new data. The desired opposite is called generalisation: predicting well on cases never seen before. Underfitting is the reverse defect: a model too simple, which fails even on its training data. Two causes dominate. Training data that poorly represents reality (too little, too old, from a single site or a single season). A model too complex for the amount of data available, with enough parameters to learn the noise by heart. The classic sign: error keeps falling on the training data while it rises on validation data set aside. The basic safeguard is to split the data into three sets (training, validation, test) and to judge the model only on data it has never seen. Regularisation techniques come on top: in 2012, the AlexNet network, with its 60 million parameters, already used a dedicated method to limit overfitting. LLMs face a variant of the problem, contamination: if a test's questions appear in the training data, the score measures the model's memory rather than its ability to reason.
Concrete example
Real case, documented in the journal Science in March 2014: Google Flu Trends, launched in 2008, estimated flu activity in the United States from Google searches. The initial method searched among 50 million search terms for those that best matched 1,152 official data points. The authors of the article see overfitting in this: seasonal terms unrelated to flu, such as high school basketball, followed the same curve. In February 2013, the tool predicted more than double the doctor visits for influenza-like illness reported by the US Centers for Disease Control and Prevention (CDC), and it had overestimated flu in 100 out of 108 weeks starting in August 2011.
Illustrative case: a French mid-sized services company predicts customer churn. In the demo, the model shows 96% correct predictions. In production, it falls below 65%. The audit shows it had been tested on the same customers as those used for training.
Comparison
| Underfitting | Good fit | Overfitting | |
|---|---|---|---|
| On training data | Poor | Good | Excellent, sometimes perfect |
| On new data | Poor | Close to the training score | Clearly worse |
| Typical cause | Model too simple, missing variables | Model and data well sized | Model too complex, too little or unrepresentative data |
| What the executive sees | The model fails to convince from the demo | Pilot results hold up in production | Brilliant demo, disappointment in production |
| Remedy | Richer model, better variables | Regular performance monitoring | More real data, testing on unseen data, regularisation |
FAQ
What is overfitting in AI?
It is the defect of a model that has memorised its training data instead of learning general rules. It performs very well on examples it has already seen and makes mistakes on new ones. It is said to generalise poorly.
What is the difference between overfitting and underfitting?
Overfitting is a model too closely fitted to its data: excellent in training, poor on new cases. Underfitting is a model too simple: poor everywhere, including on its training data. The goal is the middle ground, a model that generalises.
How do you detect overfitting?
By comparing performance on the training data with performance on held-out data the model has never seen. If the gap is large, or if validation error rises while training error falls, the model is overfitting.
How do you prevent overfitting?
Use more data that represents production, split into training, validation and test sets, choose a model suited to the amount of data, apply regularisation techniques and stop training when validation error stops falling.
What is an example of overfitting?
Google Flu Trends: according to a 2014 article in Science, the tool selected search terms among 50 million to match 1,152 data points, including seasonal terms unrelated to flu. In February 2013, it predicted more than double the doctor visits measured by US health authorities.
Can LLMs like ChatGPT overfit?
Yes, in a particular form. An LLM can recite passages from its training data, or pass a public test because it saw the questions during training (contamination). That is why a serious evaluation uses your own, unpublished cases.
See also
Further reading
Overfitting, Machine Learning Crash Course, Google for Developers
Sources
- Overfitting, Machine Learning Crash Course, Google for Developers. https://developers.google.com/machine-learning/crash-course/overfitting/overfitting
- The Parable of Google Flu: Traps in Big Data Analysis, Lazer, Kennedy, King and Vespignani, Science, vol. 343, no. 6176, March 14, 2014. https://www.science.org/doi/10.1126/science.1248506
- ImageNet Classification with Deep Convolutional Neural Networks, Krizhevsky, Sutskever and Hinton, NeurIPS (NIPS) 2012. https://proceedings.neurips.cc/paper_files/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html