library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]
# Visualize the raw data
ggplot(df, aes(x, y)) +
geom_point() +
theme_minimal()
Find arbitrarily-shaped clusters and flag outliers as noise — without choosing k
A practical guide to DBSCAN in R: see why k-means fails on non-spherical shapes, choose the eps radius from a k-nearest-neighbor distance plot, run dbscan() with eps and MinPts, and visualize the clusters and noise points with factoextra’s fviz_cluster(). Worked on the multishapes data set.
June 24, 2026
July 7, 2026
eps (the neighborhood radius) and MinPts (the minimum points to form a dense region).eps from a k-nearest-neighbor distance plot (dbscan::kNNdistplot()) — the “knee” (sharp bend) of the curve is a good value.dbscan(df, eps, minPts) from the dbscan package, then visualize with fviz_cluster() from factoextra — noise points are drawn in black.DBSCAN — Density-Based Spatial Clustering of Applications with Noise (Ester et al., 1996) — is a density-based clustering algorithm that can identify clusters of any shape in data that contains noise and outliers. The intuition is human: look at a scatter of points and you naturally see clusters as dense regions separated by sparser gaps, with a few stray points belonging to nothing. DBSCAN formalizes exactly that.
This is what makes it different from the methods you’ve seen so far. Partitioning methods (k-means, PAM) and hierarchical clustering work well only for compact, well-separated, roughly spherical clusters, and they are easily thrown off by outliers — every point must land in some cluster. DBSCAN has two advantages those methods lack:
This lesson is part of the Cluster Analysis in R series. We’ll first show why k-means fails on non-spherical data, then build up DBSCAN’s parameters, choose eps properly, and read the result.
Real data is rarely a tidy set of round blobs. To make the point, we’ll use the simulated multishapes data set from the factoextra package. It contains five clusters of genuinely different shapes — plus scattered outliers — exactly the situation that defeats k-means.
Load it and look at the raw points first (we use the first two columns, the x–y coordinates):

By eye there are 5 clusters plus outliers: two oval clusters (the concentric rings on the left), two linear clusters (the two parallel bands), and one compact round cluster — with a sprinkle of noise. A human sees these instantly because of the density differences.
Now hand the same data to k-means and ask it for 5 clusters. fviz_cluster() from factoextra draws the result:

We know there are 5 clusters, but k-means gets them wrong. It carves the rings and the linear bands into wedge-shaped pieces, because it assumes spherical, equal-size clusters and assigns every point to the nearest center. The shapes here are non-convex, so no set of centers can describe them — and the outliers are forced into a cluster instead of being recognized as noise.
This is precisely where DBSCAN shines.
DBSCAN measures density as the number of points close to a given point. “Close” and “enough” are set by two parameters:
eps (epsilon, \(\epsilon\)) — the radius of the neighborhood around a point. Points within eps of point x form its \(\epsilon\)-neighborhood.MinPts — the minimum number of points that must fall inside that radius for a region to count as dense.With those, every point in the data is one of three types:
MinPts neighbors within eps (it sits inside a dense region);MinPts neighbors, but lies within the eps-neighborhood of some core point (it’s on the edge of a cluster);A cluster is then a group of density-connected points: start from a core point and keep absorbing every point reachable through chains of overlapping dense neighborhoods. Points that no chain reaches are left as noise. That single mechanism is why DBSCAN follows arbitrary shapes — it grows clusters along the data’s own density, not toward fixed centers — and why it needs no k.
k to specify — the number of clusters falls out of the density.Free to read, ads included. Go Pro to remove them.
MinPts is the easier parameter: the larger and noisier the data, the larger it should be, with a floor of about 3 (a common default for 2-D data is MinPts = 5). The sensitive one is eps — too small and dense clusters fracture into noise; too large and separate clusters merge.
The standard recipe for eps is a k-nearest-neighbor distance plot. The idea: for every point, compute the distance to its k-th nearest neighbor (with k = MinPts), sort those distances in ascending order, and plot them. Points inside a cluster have small k-NN distances; points in sparse regions have large ones, so the curve stays low and then bends sharply upward at the boundary between cluster and noise. That knee is a good value for eps.
kNNdistplot() from the dbscan package draws it:

The curve is flat for most points, then turns sharply upward. The knee sits at a distance of about 0.15 — so we’ll use eps = 0.15 with MinPts = 5.
We’ll compute DBSCAN with the dbscan package, which is a fast, modern re-implementation of the algorithm, and visualize with factoextra. Install them once if needed:
The original SThDA tutorial used fpc::dbscan(). We use the dbscan package’s dbscan() here — it is faster and the maintained choice today, and it gives the same clusters. The only spelling difference is the argument name: dbscan() takes minPts (lowercase m), where fpc::dbscan() took MinPts.
The call is one line — dbscan(data, eps, minPts):
DBSCAN clustering for 1100 objects.
Parameters: eps = 0.15, minPts = 5
Using euclidean distances and borderpoints = TRUE
The clustering contains 5 cluster(s) and 31 noise points.
0 1 2 3 4 5
31 410 405 104 99 51
Available fields: cluster, eps, minPts, metric, borderPoints
The print-out tells the whole story: DBSCAN found 5 clusters and 31 noise points — exactly the five shapes we could see by eye, with the scattered outliers correctly set aside. No k was supplied; the number of clusters emerged from the density. The cluster sizes confirm it:
0 1 2 3 4 5
31 410 405 104 99 51
The membership is stored in db$cluster, an integer vector. Cluster 0 is reserved for noise (the 31 outliers); clusters 1–5 are the real groups. A quick look at a random sample:
Now the signature figure. fviz_cluster() plots the points colored by cluster — and crucially draws the noise points in black, so you can see what DBSCAN chose to ignore. We pass stand = FALSE (the coordinates are already on the right scale) and geom = "point":
library(dbscan)
library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]
set.seed(123)
db <- dbscan::dbscan(df, eps = 0.15, minPts = 5)
# Black points are noise/outliers
fviz_cluster(db, data = df, stand = FALSE,
ellipse = FALSE, show.clust.cent = FALSE,
geom = "point", palette = "jco",
ggtheme = theme_minimal())
This is the result k-means could not produce. Each of the five shapes — the two concentric rings, the two parallel bands, and the compact blob — is recovered as its own cluster, following the data’s true geometry. The black points are the noise: the scattered outliers DBSCAN correctly refused to assign. fviz_cluster() also uses slightly different symbols for core (seed) and border points within each cluster.
Reading a DBSCAN result comes down to three things:
eps/MinPts are off, not your k.0 is noise. Its size (31 here) tells you how many points were too sparse to belong anywhere — a built-in outlier report.eps is the lever. Too small fragments dense clusters into noise; too large fuses distinct clusters. DBSCAN is sensitive to it, especially when clusters have different densities — a single eps may not suit them all (see Common issues).So which method should you reach for?
| Use DBSCAN when… | Use k-means when… |
|---|---|
| clusters are non-spherical (rings, bands, S-shapes) | clusters are compact and roughly spherical |
| the data has outliers/noise you want flagged | the data is clean, every point should be assigned |
| you don’t know the number of clusters | you can choose / estimate k |
| densities are similar across clusters | clusters are similar in size |
In short: DBSCAN trades the “pick k” problem for a “pick eps” problem, and in exchange it handles arbitrary shapes and outliers that break centroid methods.
Edit and run the code below in your browser — no installation needed. Try changing eps (e.g. to 0.1 or 0.3) or minPts, then re-run and watch how the number of clusters and noise points changes.
Not sure whether your data is a job for DBSCAN or k-means? Ask Prova “does my data have non-spherical clusters or outliers, and what eps should I use?” — it answers with R code you can run on your own data set, drawing the kNN distance plot and reading the knee for you. The runtime is the judge. Ask Prova →
Go Pro for an ad-free Datanovia.
eps and MinPts. Don’t guess eps. Draw the kNN distance plot (kNNdistplot(df, k = MinPts)) and read the knee. For MinPts, start at 5 for 2-D data and raise it for larger or noisier sets (a rule of thumb is MinPts ≥ dimensions + 1, at least 3). Remember the argument is minPts in the dbscan package, MinPts in fpc::dbscan().eps problem. If almost all points come back as cluster 0 (noise), eps is too small — increase it. If the whole data collapses into a single cluster, eps is too large — decrease it. The kNN-distance knee is the corrective.eps, so when some clusters are much denser than others, no single value fits all: a small eps splits the sparse clusters into noise, a large one merges the dense ones. If your clusters have very different densities, consider HDBSCAN (dbscan::hdbscan()), which adapts the density threshold per cluster.eps separates poorly. Reduce dimensions first with PCA, or scale your variables so no one feature dominates the distance.DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is an unsupervised algorithm that groups together points packed in dense regions and marks points in sparse regions as noise (outliers). It can find clusters of arbitrary shape and, unlike k-means, does not require you to choose the number of clusters in advance — they emerge from the density of the data.
Use a k-nearest-neighbor distance plot: dbscan::kNNdistplot(df, k = MinPts). It sorts every point’s distance to its k-th nearest neighbor and plots them. The curve bends sharply (the “knee”) where points stop being inside clusters and start being noise — that distance is a good value for eps. On the multishapes data the knee sits near 0.15.
K-means needs you to fix the number of clusters k, assigns every point to the nearest center, and assumes roughly spherical, equal-size clusters — so it fails on rings, bands, and outliers. DBSCAN discovers the number of clusters from the density, follows arbitrary shapes, and labels outliers as noise. Use k-means for compact spherical groups; use DBSCAN for non-spherical clusters or noisy data.
A noise point is an outlier that DBSCAN assigns to no cluster: it is neither a core point (≥ MinPts neighbors within eps) nor a border point (within a core point’s neighborhood). In the result, noise points are coded as cluster 0 (db$cluster == 0) and fviz_cluster() draws them in black.
Usually yes, when your variables are on different scales — DBSCAN uses distances and a single global eps, so a high-magnitude variable would dominate. Standardize with scale() first. The multishapes x–y coordinates here are already comparable, which is why we pass stand = FALSE to fviz_cluster().
Run DBSCAN on the multishapes data with a larger radius, eps = 0.3 (keep minPts = 5). Print the result and the cluster sizes. How many clusters does it find now, and how many noise points, compared to eps = 0.15? Then draw the cluster plot.
Load dbscan and factoextra, build df <- multishapes[, 1:2], then call dbscan::dbscan(df, eps = 0.3, minPts = 5). Use table(db$cluster) for the sizes (cluster 0 = noise) and fviz_cluster() to plot.
library(dbscan)
library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]
set.seed(123)
db <- dbscan::dbscan(df, eps = 0.3, minPts = 5)
db
table(db$cluster)
fviz_cluster(db, data = df, stand = FALSE,
ellipse = FALSE, show.clust.cent = FALSE,
geom = "point", palette = "jco",
ggtheme = theme_minimal())A larger eps merges clusters and absorbs noise: you get fewer clusters and far fewer noise points than with eps = 0.15, because a wider radius links groups that were separate and pulls outliers into them. This is exactly why the kNN-distance knee matters — it keeps eps from being too large.
Quick check. A colleague runs DBSCAN and almost every point comes back in cluster 0. What is cluster 0, and what is the one parameter to adjust?
Cluster 0 is noise (outliers assigned to no cluster). If nearly everything is noise, eps is too small — the neighborhoods don’t contain MinPts points, so no dense regions form. Increase eps (read it off the kNNdistplot() knee), or lower MinPts.
You saw why k-means fails on the non-spherical multishapes data, learned DBSCAN’s two parameters (eps and MinPts) and its core/border/noise point types, chose eps from the kNN distance plot’s knee, ran dbscan(df, eps = 0.15, minPts = 5) to recover 5 clusters and 31 noise points, and visualized them with fviz_cluster(). DBSCAN is the method to reach for when clusters have arbitrary shapes or the data carries outliers — the cases that defeat centroid-based clustering. When clusters have very different densities, look at HDBSCAN; when they’re compact and spherical, k-means is faster and simpler.
k when you do use a partitioning method. · Model-based clustering — another way to handle non-spherical groups, via mixture models. · Cluster Analysis in R — the full series.Prefer a book? Practical Guide to Cluster Analysis in R is available as a downloadable PDF — every lesson in this series, offline and yours to keep.
Prove you can do it. Master the whole Cluster Analysis in R series — track your path, build projects, and earn a certificate.
Go Pro — unlimited Prova on your own data and a verifiable certificate that proves the skill.
Cancel anytime · 30-day full refund
✓ You're Pro — keep going. The runtime is the judge.
Ready to level up?
Get new R & Python lessons by email
Practical, reproducible, no spam. Unsubscribe anytime.
Double opt-in. We never share your email.
This lesson is reproducible: every figure was produced by the code shown — edit any block and Run, and the sandbox + quiz re-run live in your browser. The runtime is the judge.
@online{2026,
author = {},
title = {DBSCAN: {Density-Based} {Clustering} in {R}},
date = {2026-06-24},
url = {https://www.datanovia.com/learn/machine-learning/clustering/dbscan},
langid = {en}
}