DBSCAN: Density-Based Clustering in R

Find arbitrarily-shaped clusters and flag outliers as noise — without choosing k

Machine Learning

A practical guide to DBSCAN in R: see why k-means fails on non-spherical shapes, choose the eps radius from a k-nearest-neighbor distance plot, run dbscan() with eps and MinPts, and visualize the clusters and noise points with factoextra’s fviz_cluster(). Worked on the multishapes data set.

Published

June 24, 2026

Modified

July 7, 2026

TipKey takeaways
  • DBSCAN (Density-Based Spatial Clustering of Applications with Noise) finds clusters of any shape by grouping dense regions, and labels low-density points as noise (outliers).
  • Unlike k-means, you don’t pre-specify the number of clusters — DBSCAN discovers them from the density of the data.
  • It needs two parameters: eps (the neighborhood radius) and MinPts (the minimum points to form a dense region).
  • Choose eps from a k-nearest-neighbor distance plot (dbscan::kNNdistplot()) — the “knee” (sharp bend) of the curve is a good value.
  • Run it with dbscan(df, eps, minPts) from the dbscan package, then visualize with fviz_cluster() from factoextra — noise points are drawn in black.
  • DBSCAN beats k-means on non-spherical clusters and noisy data; k-means is faster and better for compact, roughly spherical groups.
Get the book — Practical Guide to Cluster Analysis in R (PDF)

Introduction

DBSCAN — Density-Based Spatial Clustering of Applications with Noise (Ester et al., 1996) — is a density-based clustering algorithm that can identify clusters of any shape in data that contains noise and outliers. The intuition is human: look at a scatter of points and you naturally see clusters as dense regions separated by sparser gaps, with a few stray points belonging to nothing. DBSCAN formalizes exactly that.

This is what makes it different from the methods you’ve seen so far. Partitioning methods (k-means, PAM) and hierarchical clustering work well only for compact, well-separated, roughly spherical clusters, and they are easily thrown off by outliers — every point must land in some cluster. DBSCAN has two advantages those methods lack:

  • you don’t choose the number of clusters in advance — it finds them from the density;
  • it can recover arbitrary shapes (linear, oval, S-shaped) and explicitly flags outliers as noise.

This lesson is part of the Cluster Analysis in R series. We’ll first show why k-means fails on non-spherical data, then build up DBSCAN’s parameters, choose eps properly, and read the result.

Why k-means fails on arbitrary shapes

Real data is rarely a tidy set of round blobs. To make the point, we’ll use the simulated multishapes data set from the factoextra package. It contains five clusters of genuinely different shapes — plus scattered outliers — exactly the situation that defeats k-means.

Load it and look at the raw points first (we use the first two columns, the x–y coordinates):

library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]

# Visualize the raw data
ggplot(df, aes(x, y)) +
  geom_point() +
  theme_minimal()

Scatter plot of the multishapes data set showing two concentric oval clusters, two parallel linear clusters, one compact round cluster, and scattered outlier points.

By eye there are 5 clusters plus outliers: two oval clusters (the concentric rings on the left), two linear clusters (the two parallel bands), and one compact round cluster — with a sprinkle of noise. A human sees these instantly because of the density differences.

Now hand the same data to k-means and ask it for 5 clusters. fviz_cluster() from factoextra draws the result:

library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]

set.seed(123)
km.res <- kmeans(df, centers = 5, nstart = 25)

fviz_cluster(km.res, df, geom = "point",
             ellipse = FALSE, show.clust.cent = FALSE,
             palette = "jco", ggtheme = theme_minimal())

factoextra cluster plot of k-means with five clusters on the multishapes data, slicing the rings and bands into wedge-shaped groups that ignore the true shapes.

We know there are 5 clusters, but k-means gets them wrong. It carves the rings and the linear bands into wedge-shaped pieces, because it assumes spherical, equal-size clusters and assigns every point to the nearest center. The shapes here are non-convex, so no set of centers can describe them — and the outliers are forced into a cluster instead of being recognized as noise.

This is precisely where DBSCAN shines.

The concepts: eps, MinPts, and three kinds of point

DBSCAN measures density as the number of points close to a given point. “Close” and “enough” are set by two parameters:

  • eps (epsilon, \(\epsilon\)) — the radius of the neighborhood around a point. Points within eps of point x form its \(\epsilon\)-neighborhood.
  • MinPts — the minimum number of points that must fall inside that radius for a region to count as dense.

With those, every point in the data is one of three types:

  • a core point — has at least MinPts neighbors within eps (it sits inside a dense region);
  • a border point — has fewer than MinPts neighbors, but lies within the eps-neighborhood of some core point (it’s on the edge of a cluster);
  • a noise point (outlier) — neither core nor border; it sits in a sparse region and belongs to no cluster.

A cluster is then a group of density-connected points: start from a core point and keep absorbing every point reachable through chains of overlapping dense neighborhoods. Points that no chain reaches are left as noise. That single mechanism is why DBSCAN follows arbitrary shapes — it grows clusters along the data’s own density, not toward fixed centers — and why it needs no k.

NoteDBSCAN’s three advantages over k-means
  1. No k to specify — the number of clusters falls out of the density.
  2. Any shape — clusters need not be circular or convex.
  3. Outliers are first-class — sparse points are labeled noise, not forced into a group.
Advertisement

Free to read, ads included. Go Pro to remove them.

Choosing eps with the k-NN distance plot

MinPts is the easier parameter: the larger and noisier the data, the larger it should be, with a floor of about 3 (a common default for 2-D data is MinPts = 5). The sensitive one is eps — too small and dense clusters fracture into noise; too large and separate clusters merge.

The standard recipe for eps is a k-nearest-neighbor distance plot. The idea: for every point, compute the distance to its k-th nearest neighbor (with k = MinPts), sort those distances in ascending order, and plot them. Points inside a cluster have small k-NN distances; points in sparse regions have large ones, so the curve stays low and then bends sharply upward at the boundary between cluster and noise. That knee is a good value for eps.

kNNdistplot() from the dbscan package draws it:

library(dbscan)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]

# k = MinPts; the knee marks a good eps
kNNdistplot(df, k = 5)
abline(h = 0.15, lty = 2)

k-nearest-neighbor distance plot for the multishapes data with k equals 5, showing a flat curve that bends sharply upward near a distance of 0.15, marked by a dashed horizontal line.

The curve is flat for most points, then turns sharply upward. The knee sits at a distance of about 0.15 — so we’ll use eps = 0.15 with MinPts = 5.

Running DBSCAN

We’ll compute DBSCAN with the dbscan package, which is a fast, modern re-implementation of the algorithm, and visualize with factoextra. Install them once if needed:

install.packages("dbscan")
install.packages("factoextra")
NoteA note on which package

The original SThDA tutorial used fpc::dbscan(). We use the dbscan package’s dbscan() here — it is faster and the maintained choice today, and it gives the same clusters. The only spelling difference is the argument name: dbscan() takes minPts (lowercase m), where fpc::dbscan() took MinPts.

The call is one line — dbscan(data, eps, minPts):

library(dbscan)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]

# Compute DBSCAN with the eps from the kNN plot
set.seed(123)
db <- dbscan::dbscan(df, eps = 0.15, minPts = 5)

# Print a summary of the result
db
DBSCAN clustering for 1100 objects.
Parameters: eps = 0.15, minPts = 5
Using euclidean distances and borderpoints = TRUE
The clustering contains 5 cluster(s) and 31 noise points.

  0   1   2   3   4   5 
 31 410 405 104  99  51 

Available fields: cluster, eps, minPts, metric, borderPoints

The print-out tells the whole story: DBSCAN found 5 clusters and 31 noise points — exactly the five shapes we could see by eye, with the scattered outliers correctly set aside. No k was supplied; the number of clusters emerged from the density. The cluster sizes confirm it:

library(dbscan)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]
set.seed(123)
db <- dbscan::dbscan(df, eps = 0.15, minPts = 5)

# Cluster membership counts: cluster 0 = noise/outliers
table(db$cluster)

  0   1   2   3   4   5 
 31 410 405 104  99  51 

The membership is stored in db$cluster, an integer vector. Cluster 0 is reserved for noise (the 31 outliers); clusters 1–5 are the real groups. A quick look at a random sample:

library(dbscan)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]
set.seed(123)
db <- dbscan::dbscan(df, eps = 0.15, minPts = 5)

# Cluster of a random subset of points (0 = noise)
set.seed(42)
db$cluster[sample(seq_len(nrow(df)), 20)]
 [1] 2 1 5 2 1 1 1 1 2 2 3 2 4 4 2 3 1 2 5 1

Visualizing the clusters

Now the signature figure. fviz_cluster() plots the points colored by cluster — and crucially draws the noise points in black, so you can see what DBSCAN chose to ignore. We pass stand = FALSE (the coordinates are already on the right scale) and geom = "point":

library(dbscan)
library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]
set.seed(123)
db <- dbscan::dbscan(df, eps = 0.15, minPts = 5)

# Black points are noise/outliers
fviz_cluster(db, data = df, stand = FALSE,
             ellipse = FALSE, show.clust.cent = FALSE,
             geom = "point", palette = "jco",
             ggtheme = theme_minimal())

factoextra cluster plot of DBSCAN on the multishapes data: the two oval rings, two linear bands and one compact cluster are each recovered as a distinct colored cluster, with outlier noise points drawn in black.

This is the result k-means could not produce. Each of the five shapes — the two concentric rings, the two parallel bands, and the compact blob — is recovered as its own cluster, following the data’s true geometry. The black points are the noise: the scattered outliers DBSCAN correctly refused to assign. fviz_cluster() also uses slightly different symbols for core (seed) and border points within each cluster.

Interpreting the result, and DBSCAN vs k-means

Reading a DBSCAN result comes down to three things:

  • The number of clusters is an output, not an input — here, 5. If it doesn’t match what you expect, your eps/MinPts are off, not your k.
  • Cluster 0 is noise. Its size (31 here) tells you how many points were too sparse to belong anywhere — a built-in outlier report.
  • eps is the lever. Too small fragments dense clusters into noise; too large fuses distinct clusters. DBSCAN is sensitive to it, especially when clusters have different densities — a single eps may not suit them all (see Common issues).

So which method should you reach for?

Use DBSCAN when… Use k-means when…
clusters are non-spherical (rings, bands, S-shapes) clusters are compact and roughly spherical
the data has outliers/noise you want flagged the data is clean, every point should be assigned
you don’t know the number of clusters you can choose / estimate k
densities are similar across clusters clusters are similar in size

In short: DBSCAN trades the “pick k” problem for a “pick eps” problem, and in exchange it handles arbitrary shapes and outliers that break centroid methods.

Try it live

Edit and run the code below in your browser — no installation needed. Try changing eps (e.g. to 0.1 or 0.3) or minPts, then re-run and watch how the number of clusters and noise points changes.

🟢 With an AI agent

Not sure whether your data is a job for DBSCAN or k-means? Ask Prova “does my data have non-spherical clusters or outliers, and what eps should I use?” — it answers with R code you can run on your own data set, drawing the kNN distance plot and reading the knee for you. The runtime is the judge. Ask Prova →

Advertisement

Go Pro for an ad-free Datanovia.

Common issues

  • Choosing eps and MinPts. Don’t guess eps. Draw the kNN distance plot (kNNdistplot(df, k = MinPts)) and read the knee. For MinPts, start at 5 for 2-D data and raise it for larger or noisier sets (a rule of thumb is MinPts ≥ dimensions + 1, at least 3). Remember the argument is minPts in the dbscan package, MinPts in fpc::dbscan().
  • Everything is noise, or everything is one cluster. That’s an eps problem. If almost all points come back as cluster 0 (noise), eps is too small — increase it. If the whole data collapses into a single cluster, eps is too large — decrease it. The kNN-distance knee is the corrective.
  • Clusters of varying density. DBSCAN uses one global eps, so when some clusters are much denser than others, no single value fits all: a small eps splits the sparse clusters into noise, a large one merges the dense ones. If your clusters have very different densities, consider HDBSCAN (dbscan::hdbscan()), which adapts the density threshold per cluster.
  • High-dimensional data. Distances become less meaningful as dimensions grow (the curse of dimensionality), so a single eps separates poorly. Reduce dimensions first with PCA, or scale your variables so no one feature dominates the distance.

Frequently asked questions

DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is an unsupervised algorithm that groups together points packed in dense regions and marks points in sparse regions as noise (outliers). It can find clusters of arbitrary shape and, unlike k-means, does not require you to choose the number of clusters in advance — they emerge from the density of the data.

Use a k-nearest-neighbor distance plot: dbscan::kNNdistplot(df, k = MinPts). It sorts every point’s distance to its k-th nearest neighbor and plots them. The curve bends sharply (the “knee”) where points stop being inside clusters and start being noise — that distance is a good value for eps. On the multishapes data the knee sits near 0.15.

K-means needs you to fix the number of clusters k, assigns every point to the nearest center, and assumes roughly spherical, equal-size clusters — so it fails on rings, bands, and outliers. DBSCAN discovers the number of clusters from the density, follows arbitrary shapes, and labels outliers as noise. Use k-means for compact spherical groups; use DBSCAN for non-spherical clusters or noisy data.

A noise point is an outlier that DBSCAN assigns to no cluster: it is neither a core point (≥ MinPts neighbors within eps) nor a border point (within a core point’s neighborhood). In the result, noise points are coded as cluster 0 (db$cluster == 0) and fviz_cluster() draws them in black.

Usually yes, when your variables are on different scales — DBSCAN uses distances and a single global eps, so a high-magnitude variable would dominate. Standardize with scale() first. The multishapes x–y coordinates here are already comparable, which is why we pass stand = FALSE to fviz_cluster().

Test your understanding

Run DBSCAN on the multishapes data with a larger radius, eps = 0.3 (keep minPts = 5). Print the result and the cluster sizes. How many clusters does it find now, and how many noise points, compared to eps = 0.15? Then draw the cluster plot.

Load dbscan and factoextra, build df <- multishapes[, 1:2], then call dbscan::dbscan(df, eps = 0.3, minPts = 5). Use table(db$cluster) for the sizes (cluster 0 = noise) and fviz_cluster() to plot.

library(dbscan)
library(factoextra)
data("multishapes", package = "factoextra")
df <- multishapes[, 1:2]

set.seed(123)
db <- dbscan::dbscan(df, eps = 0.3, minPts = 5)
db
table(db$cluster)

fviz_cluster(db, data = df, stand = FALSE,
             ellipse = FALSE, show.clust.cent = FALSE,
             geom = "point", palette = "jco",
             ggtheme = theme_minimal())

A larger eps merges clusters and absorbs noise: you get fewer clusters and far fewer noise points than with eps = 0.15, because a wider radius links groups that were separate and pulls outliers into them. This is exactly why the kNN-distance knee matters — it keeps eps from being too large.

Quick check. A colleague runs DBSCAN and almost every point comes back in cluster 0. What is cluster 0, and what is the one parameter to adjust?

Cluster 0 is noise (outliers assigned to no cluster). If nearly everything is noise, eps is too small — the neighborhoods don’t contain MinPts points, so no dense regions form. Increase eps (read it off the kNNdistplot() knee), or lower MinPts.

Conclusion

You saw why k-means fails on the non-spherical multishapes data, learned DBSCAN’s two parameters (eps and MinPts) and its core/border/noise point types, chose eps from the kNN distance plot’s knee, ran dbscan(df, eps = 0.15, minPts = 5) to recover 5 clusters and 31 noise points, and visualized them with fviz_cluster(). DBSCAN is the method to reach for when clusters have arbitrary shapes or the data carries outliers — the cases that defeat centroid-based clustering. When clusters have very different densities, look at HDBSCAN; when they’re compact and spherical, k-means is faster and simpler.

References

  • Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining (KDD-96), 226–231.

Reuse

Citation

BibTeX citation:
@online{2026,
  author = {},
  title = {DBSCAN: {Density-Based} {Clustering} in {R}},
  date = {2026-06-24},
  url = {https://www.datanovia.com/learn/machine-learning/clustering/dbscan},
  langid = {en}
}
For attribution, please cite this work as:
“DBSCAN: Density-Based Clustering in R.” 2026. June 24. https://www.datanovia.com/learn/machine-learning/clustering/dbscan.