• 【生物信息学】使用HSIC LASSO方法进行特征选择


    目录

    一、实验介绍

    二、实验环境

    1. 配置虚拟环境

    2. 库版本介绍

    3. IDE

    三、实验内容

    0. 导入必要的工具

    1. 读取数据

    2. 划分训练集和测试集

    3. 进行HSIC LASSO特征选择

    4. 特征提取

    5. 使用随机森林进行分类(使用所有特征)

    6. 使用随机森林进行分类(使用HSIC选择的特征):

    7. 代码整合


    一、实验介绍

            本实验实现了HSIC LASSO(Hilbert-Schmidt independence criterion LASSO)方法进行特征选择,并使用随机森林分类器对选择的特征子集进行分类。

            特征选择是机器学习中的重要任务之一,它可以提高模型的效果、减少计算开销,并帮助我们理解数据的关键特征。

            HSIC LASSO是一种基于核的独立性度量方法,用于寻找对输出值具有强统计依赖性的非冗余特征。

    二、实验环境

        本系列实验使用了PyTorch深度学习框架,相关操作如下(基于深度学习系列文章的环境):

    1. 配置虚拟环境

    深度学习系列文章的环境

    conda create -n DL python=3.7 
    conda activate DL
    pip install torch==1.8.1+cu102 torchvision==0.9.1+cu102 torchaudio==0.8.1 -f https://download.pytorch.org/whl/torch_stable.html
    
    conda install matplotlib
    conda install scikit-learn

    新增加

    conda install pandas
    conda install seaborn
    conda install networkx
    conda install statsmodels
    pip install pyHSICLasso

    注:本人的实验环境按照上述顺序安装各种库,若想尝试一起安装(天知道会不会出问题)

    2. 库版本介绍

    软件包本实验版本目前最新版
    matplotlib3.5.33.8.0
    numpy1.21.61.26.0
    python3.7.16
    scikit-learn0.22.11.3.0
    torch1.8.1+cu1022.0.1
    torchaudio0.8.12.0.2
    torchvision0.9.1+cu1020.15.2

    新增

    networkx2.6.33.1
    pandas1.2.32.1.1
    pyHSICLasso1.4.21.4.2
    seaborn0.12.20.13.0
    statsmodels0.13.50.14.0

    3. IDE

            建议使用Pycharm(其中,pyHSICLasso库在VScode出错,尚未找到解决办法……)

    win11 安装 Anaconda(2022.10)+pycharm(2022.3/2023.1.4)+配置虚拟环境_QomolangmaH的博客-CSDN博客https://blog.csdn.net/m0_63834988/article/details/128693741icon-default.png?t=N7T8https://blog.csdn.net/m0_63834988/article/details/128693741

    三、实验内容

    0. 导入必要的工具

    1. import random
    2. import pandas as pd
    3. from pyHSICLasso import HSICLasso
    4. from sklearn.preprocessing import LabelEncoder
    5. from sklearn.model_selection import train_test_split
    6. from sklearn.metrics import confusion_matrix, classification_report
    7. from sklearn.ensemble import RandomForestClassifier

    1. 读取数据

    1. data = pd.read_csv("cancer_subtype.csv")
    2. x = data.iloc[:, :-1]
    3. y = data.iloc[:, -1]

    2. 划分训练集和测试集

    X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.3, random_state=10)

            将数据集划分为训练集(X_trainy_train)和测试集(X_testy_test)。其中,测试集占总数据的30%。

    3. 进行HSIC LASSO特征选择

    1. random.seed(1)
    2. le = LabelEncoder()
    3. y_hsic = le.fit_transform(y_train)
    4. x_hsic, fea_n = X_train.to_numpy(), X_train.columns.tolist()
    5. hsic.input(x_hsic, y_hsic, featname=fea_n)
    6. hsic.classification(200)
    7. genes = hsic.get_features()
    8. score = hsic.get_index_score()
    9. res = pd.DataFrame([genes, score]).T
    • 设置随机种子,确保随机过程的可重复性
    • 使用LabelEncoder将目标变量进行标签编码,将其转换为数值形式。
    • 通过将训练集数据X_train和标签y_hsic输入HSIC LASSO模型进行特征选择。
      • hsic.input用于设置输入数据和特征名称
      • hsic.classification用于运行HSIC LASSO算法进行特征选择
        • 选择的特征保存在genes中;
        • 对应的特征得分保存在score中;
      • genes、score存储在DataFrame res中。

    4. 特征提取

    1. hsic_x_train = X_train[res[0]]
    2. hsic_x_test = X_test[res[0]]

            根据HSIC LASSO选择的特征索引,从原始的训练集X_train和测试集X_test中提取相应的特征子集,分别保存在hsic_x_trainhsic_x_test中。

    5. 使用随机森林进行分类(使用所有特征)

    1. rf_model = RandomForestClassifier(20)
    2. rf_model.fit(X_train, y_train)
    3. rf_pred = rf_model.predict(X_test)
    4. print("RF all feature")
    5. print(confusion_matrix(y_test, rf_pred))
    6. print(classification_report(y_test, rf_pred, digits=5))

            使用随机森林分类器(RandomForestClassifier)对具有所有特征的训练集进行训练,并在测试集上进行预测。预测结果存储在rf_pred中,并输出混淆矩阵(confusion matrix)和分类报告(classification report)。

    6. 使用随机森林进行分类(使用HSIC选择的特征):

    1. rf_hsic_model = RandomForestClassifier(20)
    2. rf_hsic_model.fit(hsic_x_train, y_train)
    3. rf_hsic_pred = rf_hsic_model.predict(hsic_x_test)
    4. print("RF HSIC feature")
    5. print(confusion_matrix(y_test, rf_hsic_pred))
    6. print(classification_report(y_test, rf_hsic_pred, digits=5))

            使用随机森林分类器对使用HSIC LASSO选择的特征子集hsic_x_train进行训练,并在测试集的相应特征子集hsic_x_test上进行预测。预测结果存储在rf_hsic_pred中,并输出混淆矩阵和分类报告。

    7. 代码整合

    1. # HSIC LASSO
    2. # HSIC全称“Hilbert-Schmidt independence criterion”,“希尔伯特-施密特独立性指标”,跟互信息一样,它也可以用来衡量两个变量之间的独立性
    3. # 核函数的特定选择,可以在基于核的独立性度量(如Hilbert-Schmidt独立性准则(HSIC))中找到对输出值具有很强统计依赖性的非冗余特征
    4. # CIN 107 EBV 23 GS 50 MSI 47 normal 33
    5. import random
    6. import pandas as pd
    7. from pyHSICLasso import HSICLasso
    8. from sklearn.preprocessing import LabelEncoder
    9. from sklearn.model_selection import train_test_split
    10. from sklearn.metrics import confusion_matrix, classification_report
    11. from sklearn.ensemble import RandomForestClassifier
    12. data = pd.read_csv("cancer_subtype.csv")
    13. x = data.iloc[:, :-1]
    14. y = data.iloc[:, -1]
    15. X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.3, random_state=10)
    16. random.seed(1)
    17. le = LabelEncoder()
    18. hsic = HSICLasso()
    19. y_hsic = le.fit_transform(y_train)
    20. x_hsic, fea_n = X_train.to_numpy(), X_train.columns.tolist()
    21. hsic.input(x_hsic, y_hsic, featname=fea_n)
    22. hsic.classification(200)
    23. genes = hsic.get_features()
    24. score = hsic.get_index_score()
    25. res = pd.DataFrame([genes, score]).T
    26. hsic_x_train = X_train[res[0]]
    27. hsic_x_test = X_test[res[0]]
    28. rf_model = RandomForestClassifier(20)
    29. rf_model.fit(X_train, y_train)
    30. rf_pred = rf_model.predict(X_test)
    31. print("RF all feature")
    32. print(confusion_matrix(y_test, rf_pred))
    33. print(classification_report(y_test, rf_pred, digits=5))
    34. rf_hsic_model = RandomForestClassifier(20)
    35. rf_hsic_model.fit(hsic_x_train, y_train)
    36. rf_hsic_pred = rf_hsic_model.predict(hsic_x_test)
    37. print("RF HSIC feature")
    38. print(confusion_matrix(y_test, rf_hsic_pred))
    39. print(classification_report(y_test, rf_hsic_pred, digits=5))

  • 相关阅读:
    大数据之Spark(二)
    一文读懂PostgreSQL中的索引
    【一起学Rust | 设计模式】习惯语法——使用借用类型作为参数、格式化拼接字符串、构造函数
    Mininet玩转流表
    八、MySql表的复合查询
    K线形态识别_大阴线
    机器学习之偏差与方差的区别
    数据结构 | 图
    关于汽车html网页设计完整版,10个以汽车为主题的网页设计与实现
    hashMap索引原理
  • 原文地址:https://blog.csdn.net/m0_63834988/article/details/133443975