Last reviewed:
What is unsupervised learning? Definition, examples and uses
Unsupervised learning is a machine learning method that looks for structure in raw data, without any expected answer provided in advance: groups of similar items, unusual behaviour, variables that go together. It is mainly used to explore and segment (customers, products, transactions), and its results must be interpreted by people who know the business.
According to the CNIL, the French data protection authority, unsupervised learning is a machine learning process in which the algorithm uses a raw dataset and obtains a result by detecting similarities between some of these data. Nobody tells it the right answer: there is no label. That is the difference from supervised learning, which learns from annotated examples. Three uses dominate. Clustering groups items that resemble each other: Google presents it as the technique many unsupervised models rely on. Anomaly detection spots what deviates from usual behaviour: a transaction, a sensor reading, an expense claim. Dimensionality reduction summarises dozens of variables into a few readable axes, often to visualise data. Unsupervised learning is also at the heart of today's generative AI. Embeddings, numerical representations in which texts with close meanings have close coordinates, make it possible to group thousands of customer reviews or documents by theme automatically, without a predefined grid. Its limit is the lack of ground truth. A clustering algorithm always finds groups, even when they have no useful meaning, and it is often a human who sets how many. Evaluation therefore relies on business judgement: are the segments stable, understandable, actionable? In practice, approaches are often combined: unsupervised grouping helps define categories, which are then labelled to train a supervised model.
Concrete example
Illustrative cases (fictitious companies). A French online home decor retailer with 40 employees has three years of purchase history for about 60,000 customers. A clustering algorithm, applied to purchase frequency, average basket and categories bought, reveals five groups. The marketing team recognises four; the fifth, customers who only buy during sales, was not being tracked. The team adapts its email calendar for this segment. In a services SME, management control applies anomaly detection to expense claims: the tool flags expenses that are unusual compared with each team's habits, and a person checks every alert before any action is taken.
Comparison
| Criterion | Supervised | Unsupervised |
|---|---|---|
| Data required | History annotated with the right answer | Raw data, no annotation |
| Goal | Predict a known category or value | Discover groups, anomalies, structures |
| Evaluation | Comparison with real answers on unseen cases | Business judgement: stability, clarity, usefulness |
| Examples | Scoring, known fraud detection, ticket classification | Customer segmentation, anomalies, grouping of comments |
| Role of business teams | Provide and check the labels | Name and validate the groups obtained |
FAQ
Unsupervised learning: what is a simple definition?
It is a method in which the AI receives data without any expected answer and looks on its own for similarities, groupings or unusual cases. It is used to explore data that nobody has classified yet.
What is the difference between supervised and unsupervised learning?
Supervised learning learns from labelled examples (this customer left, that one stayed) to predict a known answer. Unsupervised learning has no labels: it discovers structures, such as groups of customers with similar behaviour. The first is checked by comparing its predictions with the real answers; the second is judged mainly on the business usefulness of what it reveals.
What are examples of unsupervised learning?
Customer segmentation, detection of abnormal transactions or sensor readings, grouping customer reviews or tickets by theme, recommendations of products often bought together, spotting duplicates in a supplier database.
What is clustering?
It is the best-known unsupervised learning technique: it divides items into groups (clusters) so that items in the same group resemble each other more than those in other groups. The k-means algorithm, which requires the number of groups to be set in advance, is the classic example.
How does it relate to embeddings and generative AI?
Embeddings turn a text, an image or a product into a series of numbers that captures its meaning. Applying clustering to these embeddings makes it possible to group thousands of documents or comments by theme without a predefined grid. They are also the foundation of the vector databases used in RAG.
Where can I take a course on unsupervised learning?
For a free first approach, Google offers an online introductory machine learning course that presents the three paradigms, including clustering. For an executive, the issue is less the technique than the ability to ask the right questions: which variables, how many groups, who validates the meaning of the results.
See also
Further reading
What is Machine Learning?, Google for Developers (free introductory course)
Sources
- Apprentissage non supervisé (unsupervised learning), definition, CNIL (French data protection authority). https://www.cnil.fr/fr/definition/apprentissage-non-supervise
- What is Machine Learning?, Introduction to Machine Learning course, Google for Developers. https://developers.google.com/machine-learning/intro-to-ml/what-is-ml
- Apprentissage supervisé (supervised learning), definition, CNIL (French data protection authority). https://www.cnil.fr/fr/definition/apprentissage-supervise