• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
杜晓菲, 韦必忠, 王佳慧. 基于超图网络的多组学数据融合分类J. 桂林电子科技大学学报, xxxx, x(x): 1-7. DOI: 10.16725/j.1673-808X.202488
引用本文: 杜晓菲, 韦必忠, 王佳慧. 基于超图网络的多组学数据融合分类J. 桂林电子科技大学学报, xxxx, x(x): 1-7. DOI: 10.16725/j.1673-808X.202488
DU Xiaofei, WEI Bizhong, WANG Jiahui. The fusion of multi-omics data and complex disease classification based on hypergraph networksJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-7. DOI: 10.16725/j.1673-808X.202488
Citation: DU Xiaofei, WEI Bizhong, WANG Jiahui. The fusion of multi-omics data and complex disease classification based on hypergraph networksJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-7. DOI: 10.16725/j.1673-808X.202488

基于超图网络的多组学数据融合分类

The fusion of multi-omics data and complex disease classification based on hypergraph networks

  • 摘要: 针对多组学数据维度差异性在特征级层面融合效果欠佳,导致模型分类准确性不高的问题,提出了一种基于超图的多组学数据融合分类方法,该分类方法对于维度较低的数据特征融合,采用矩阵拼接方式融合不同组学数据,对于维度较高的数据特征融合,采用融合距离矩阵融合多组学数据,经过融合得到的数据包含了组学之间的关联信息,有利于提高模型的学习性能。基于超图神经网络设计了一种带有权重表示的权重超图神经网络模型WHGNN,该模型在卷积过程中引入顶点权重矩阵和超边权重矩阵表示节点和超边的重要性。超图结构可以实现节点的高阶邻域信息传递和超边结构特征学习,权重矩阵可以减少过拟合现象,加强模型对重要性较高的节点和超边特征的学习,提高模型的分类性能。实验结果表明,在公开数据集阿尔茨海默症、乳腺癌细胞、胶质母瘤细胞上的分类精准度分别为81.9%、87%、89%,均优于KNN、RF、SVM等经典分类算法。

     

    Abstract: To address the problem of poor fusion effectiveness at the feature level due to the dimensional disparity among multi-omics data, resulting in low model classification accuracy, this paper proposes a multi-omics data fusion classification method based on hypergraphs. For datasets with lower-dimensional features, a matrix concatenation approach is employed to fuse different omics data, while for datasets with higher-dimensional features, a fusion distance matrix is utilized to integrate multiple omics data. The fused data contain correlation information between omics, facilitating enhanced model learning performance. A Weighted Hypergraph Neural Network model (WHGNN) with weight representation is designed based on the hypergraph neural network. During the convolution process, vertex weight matrices and hyperedge weight matrices are introduced to represent the importance of nodes and hyperedges, respectively. The hypergraph structure enables the transmission of high-order neighborhood information of nodes and the learning of hyperedge structure features, while the weight matrices help mitigate overfitting and enhance the learning of features of important nodes and hyperedges, thereby improving the model's classification performance. Experimental results demonstrate classification accuracies of 81.9%, 87%, and 89% on publicly available datasets of Alzheimer's disease, breast cancer cells, and glioblastoma cells, respectively, outperforming classical classification algorithms such as KNN, RF, and SVM.

     

/

返回文章
返回