目录

DBSCAN采用居于中心的密度定义,居于密度的聚类算法是根据密度而不是距离来计算样本相似度,将样本中的高密度区域划分为簇。
DBSCAN算法根据密度可达关系求出所有密度相连样本的最大集合,将这些样本点作为同一个簇。
(1)从样本中选择一点,给定半径epsilon和圆内的最小近邻点数min_points。
(2)如果该点满足在其半径为epsilon的邻域圆内至少有min_points个近邻点,则将圆心转移到下一样本点。
(3)若一样本点不满足上述条件,则重新选择样本点。按照设定的半径epsilon和min_points进行迭代聚类。
- # -*- coding: utf-8 -*-
- """
- ===================================
- Demo of DBSCAN clustering algorithm
- ===================================
-
- Finds core samples of high density and expands clusters from them.
-
- """
- print(__doc__)
-
- import numpy as np
- from sklearn.cluster import DBSCAN
- from sklearn import metrics
- #from sklearn.datasets.samples_generator import make_blobs
- from sklearn.datasets import make_blobs
- from sklearn.preprocessing import StandardScaler
- import matplotlib.pyplot as plt
- plt.rcParams['font.sans-serif']=['SimHei'] #用来正常显示中文标签
- plt.rcParams['axes.unicode_minus']=False #用来正常显示负号
-
-
- # #############################################################################
- # Generate sample data
- centers = [[1, 1], [-1, -1], [1, -1]]
- X, labels_true = make_blobs(n_samples=750, centers=centers, cluster_std=0.4, random_state=0)
- X = StandardScaler().fit_transform(X)
- # Compute DBSCAN
- db = DBSCAN(eps=0.3, min_samples=10).fit(X)
- core_samples_mask = np.zeros_like(db.labels_, dtype=bool)
- core_samples_mask[db.core_sample_indices_] = True
- labels = db.labels_
- n_clusters_ = len(set(labels)) - (1 if -1 in labels else 0)
-
- print('Estimated number of clusters: %d' % n_clusters_)
- print("Homogeneity: %0.3f" % metrics.homogeneity_score(labels_true, labels))
- print("Completeness: %0.3f" % metrics.completeness_score(labels_true, labels))
- print("V-measure: %0.3f" % metrics.v_measure_score(labels_true, labels))
- print("Adjusted Rand Index: %0.3f"
- % metrics.adjusted_rand_score(labels_true, labels))
- print("Adjusted Mutual Information: %0.3f"
- % metrics.adjusted_mutual_info_score(labels_true, labels))
- print("Silhouette Coefficient: %0.3f"
- % metrics.silhouette_score(X, labels))
-
- # #############################################################################
-
- # Black removed and is used for noise instead.
- unique_labels = set(labels)
- colors = [plt.cm.Spectral(each)
- for each in np.linspace(0, 1, len(unique_labels))]
- for k, col in zip(unique_labels, colors):
- if k == -1:
- # Black used for noise.
- col = [0, 0, 0, 1]
-
- class_member_mask = (labels == k)
-
- xy = X[class_member_mask & core_samples_mask]
- plt.plot(xy[:, 0], xy[:, 1], 'o', markerfacecolor=tuple(col),
- markeredgecolor='k', markersize=14)
-
- xy = X[class_member_mask & ~core_samples_mask]
- plt.plot(xy[:, 0], xy[:, 1], 'o', markerfacecolor=tuple(col),
- markeredgecolor='k', markersize=6)
-
- plt.title('估计类的数量: %d' % n_clusters_)
- plt.show()

五、实验总结 DBSCAN算法的优缺点,它的优点是可以适用非凸数据集,能够发现异常点;它的缺点体现在应对密度不均匀、样本距离相差很大的数据集效果不好;样本集规模较大时,聚类时间较长。