Davies–Bouldin Index Calculator
Paste your data points and cluster labels to get the Davies–Bouldin Index, each cluster's dispersion and centroid, and the full pairwise Rᵢⱼ matrix. Lower is better. Matches scikit-learn davies_bouldin_score. No signup, nothing uploaded.
How it works
The Davies–Bouldin Index (DBI) measures how well a set of clusters is separated relative to how spread out each one is. Like the silhouette score it is an internal validity index — it needs no ground-truth labels, so it is a standard way to judge a k-means, hierarchical, or DBSCAN result and to compare different choices of k. It was defined by David Davies and Donald Bouldin in 1979 and is what scikit-learn returns from davies_bouldin_score.
For a clustering with k clusters (k ≥ 2), the tool computes:
- Centroid cᵢ. The coordinate-wise mean of the points in cluster i.
- Dispersion Sᵢ. The mean Euclidean distance from each point in cluster i to its centroid:
Sᵢ = (1 / |Cᵢ|) · Σ ‖x − cᵢ‖. This is scikit-learn's p = q = 1 choice. Smaller means a tighter cluster. - Separation Mᵢⱼ. The Euclidean distance between the centroids of clusters i and j:
Mᵢⱼ = ‖cᵢ − cⱼ‖. Larger means the clusters sit further apart. - Similarity ratio Rᵢⱼ.
Rᵢⱼ = (Sᵢ + Sⱼ) / Mᵢⱼ. It is high when two clusters are both loose and close — that is, when they overlap. - Worst rival Dᵢ.
Dᵢ = max over j ≠ i of Rᵢⱼ. Each cluster is judged by its most similar competing cluster.
The index is the plain average of Dᵢ over all clusters: DB = (1 / k) · Σ Dᵢ. Distance is Euclidean throughout, matching scikit-learn's default. The minimum is 0 — reached only by perfectly tight, perfectly separated clusters — and there is no fixed upper bound, so lower is betterand DBI values are only comparable across clusterings of the same data. Two edge cases are guarded exactly as scikit-learn does: if every dispersion is zero or every centroid coincides, the index is reported as 0; and if two distinct clusters happen to share a centroid, that pair's ratio is treated as 0 rather than a division by zero. As a credibility check the tool also recomputes the index a second way — an independent reference pass that rebuilds every centroid, dispersion, and centroid distance from scratch — and confirms the two agree.
Worked examples
Frequently asked questions
Sources & references
- scikit-learn — sklearn.metrics.davies_bouldin_score (the average-Euclidean-distance dispersion, the mean-over-clusters form, the 0-floor, and the coincident-centroid handling this tool matches)
- Davies, D.L. & Bouldin, D.W. (1979) — A Cluster Separation Measure, IEEE TPAMI (the original definition of Rᵢⱼ, Dᵢ, and the index)
The formulas on this page were last cross-checked against these sources on 2026-07-12. The Davies–Bouldin Index is a stable mathematical definition, so this tool needs no rate or schedule updates — only the worked examples are periodically re-reconciled against scikit-learn.
Related tools
Comments & feedback
Spotted a bug or want an improvement? Tell us — our team reviews every comment, and good ideas get built. Comments are public and anonymous.
Found a bug, edge case, or want to suggest an improvement?
Email me at [email protected] — most fixes ship within 24 hours.