A Knowledge Distillation Framework for Hypergraph Neural Networks with Fast Inference
-
摘要: 超图神经网络(HGNN)凭借高阶关系建模能力被广泛关注,但其计算量大、推理效率低,难以在实际场景部署。现有HGNN向多层感知机(MLP)知识蒸馏方法存在可解释性差、精度偏低等问题,主要原因在于:MLP激活函数单一且层间解耦、Softmax软标签存在信息丢失、未关注节点可靠性差异。为此,该文提出面向快速推理的超图神经网络知识蒸馏框架(DH2KAN)。该方法以柯尔莫哥洛夫阿诺德网络(KAN)为学生模型,用可学习单变量样条函数提升表征能力;设计表征相似性蒸馏,以节点表征监督KAN训练;提出高可靠节点感知蒸馏,基于信息熵量化节点可靠性并利用可靠节点软标签强化指导。在11个真实数据集上的实验表明,DH2KAN在精度与HGNN相当甚至更优的前提下,推理速度提升75倍;相比KAN在速度相近的情况下,准确率提升10.53%,适用于大规模低延迟场景。
-
关键词:
- 超图神经网络 /
- 柯尔莫哥洛夫-阿诺德网络 /
- 知识蒸馏 /
- 快速推理
Abstract:Objective HyperGraph Neural Networks (HGNNs) have received increasing attention because of their ability to model higher-order relationships among entities. However, their high computational cost and low inference efficiency limit their deployment in large-scale applications. Existing knowledge distillation methods that transfer knowledge from HGNNs to Multi-Layer Perceptrons (MLPs) have limited interpretability and accuracy. They also suffer from information loss in softmax-based soft labels and do not account for differences in node reliability. To address these limitations, a hypergraph knowledge distillation framework, Distill Hypergraph neural network to Kolmogorov-Arnold Network (DH2KAN), is proposed for fast inference. Methods DH2KAN consists of three main modules ( Fig. 2 ). First, a Kolmogorov-Arnold Network (KAN) is adopted as the student model instead of an MLP. Learnable spline-based univariate functions replace fixed activation functions and linear weights to improve the fitting ability and interpretability of the student model. Second, a representation similarity distillation mechanism is introduced to directly align the pre-logit representations of the HGNN and KAN. This alignment reduces the information loss caused by softmax normalization and preserves higher-order structural knowledge in the hypergraph. Third, a high-reliability node-aware distillation method is introduced (Fig. 3 ). Node reliability is quantified from changes in information entropy under noise perturbations. High-reliability nodes that are less sensitive to noise are selected, and their soft labels are used as additional supervision to improve the quality of distilled knowledge.Results and Discussions DH2KAN is evaluated under both transductive learning ( Table 2 ) and production learning (Table 3 ) settings. Under transductive learning, its average accuracy is 10.53 percentage points higher than that of the baseline KAN student model, 1.73 percentage points higher than that of the teacher HGNN, and 1.20 percentage points higher than that of the MLP-based distillation baseline. These results confirm the effectiveness of knowledge transfer from the HGNN to the KAN and show that the proposed method achieves higher inference accuracy than conventional MLP-based distillation methods. In addition, DH2KAN is particularly well suited to feature-dominated hypergraphs and remains robust on hypergraphs with sparse structures and noisy nodes.Conclusions DH2KAN is proposed to accelerate HGNN inference in large-scale, low-latency applications. Knowledge distillation narrows the performance gap between the KAN and HGNN and removes dependence on the hypergraph structure during inference, thereby enabling fast, low-complexity prediction. Representation similarity distillation and high-reliability node-aware distillation use the teacher model’s pre-logit representations and the soft labels of high-reliability nodes as additional supervision, respectively. These mechanisms enable effective transfer of task-relevant knowledge to the student model and support the practical deployment of DH2KAN in large-scale, low-latency scenarios. -
表 1 训练和推理期间的时间复杂度比较
KANs HGNN DH2KAN 训练 O(LNF2 (G+K)) O(LN2F+LNF2) O(NMC+LNF2
(G+K))推理 O(LNF2(G+K)) O(LN2F+LNF2) O(LNF2(G+K)) 表 2 直推式学习下在8个超图数据集的实验结果(%)
数据集 KAN HGNN LHGNN DH2KAN △KAN △HGNN △LHGN News20 67.4±2.58 75.02±1.42 75.06±1.03 77.54±1.19 10.14 2.52 2.48 CA-Cora 57.26±4.78 63.85±3.66 65.96±4.76 67.91±5.53 10.65 4.06 1.95 CC-Cora 56.77±4.74 61.93±4.41 63.52±4.23 67.66±3.15 9.89 5.73 4.14 CC-Citeseer 59.29±1.3 59.91±1.94 61.45±1.77 65.19±2.15 5.9 5.28 3.74 DBLP-Conf 67.32±0.57 93.21±0.6 91.46±0.46 85.97±0.63 18.65 –7.24 –5.49 DBLP-Paper 67.32±0.57 70.85±2.03 72.68±1.66 75.42±0.76 8.1 4.57 2.74 DBLP-Term 67.32±0.57 80.19±2.22 77.68±2.39 78.36±2.31 11.04 –1.83 0.68 IMDB-AW 40.88±1.28 49.01±2.25 50.4±2.04 49.72±2.46 8.84 0.71 –0.68 平均排名/平均值 4 2.5 2 1.5 10.53 1.73 1.2 表 3 生产式学习模式下在8个超图数据集的实验结果(%)
数据集 设置 KAN HGNN LHGNN DH2KAN △KAN △HGNN △LHGNN News20 prod 66.68±2.16 75.65±0.70 76.80±0.85 77.52±0.45 10.84 1.87 0.72 Ind 66.60±2.17 75.58±0.63 76.74±0.83 77.17±0.37 10.57 1.59 0.43 trans 67.44±2.26 76.30±1.72 77.33±1.27 78.17±1.08 10.73 1.87 0.84 CA-Cora prod 57.70±2.68 70.50±2.54 67.86±2.05 71.30±3.56 13.60 0.80 3.44 Ind 57.75±2.61 70.64±2.58 65.08±2.04 67.75±3.68 10.00 –2.89 2.67 trans 57.22±4.10 70.81±3.98 69.11±1.93 72.15±2.98 14.93 1.34 3.04 CC-Cora prod 58.03±2.32 67.17±2.90 61.32±1.88 63.64±2.46 5.61 –3.53 2.32 Ind 58.11±2.22 67.16±2.91 60.75±1.8 63.17±2.58 5.06 –3.99 2.42 trans 57.33±3.76 63.64±2.89 64.39±3.34 66.44±3.25 9.11 2.80 2.05 CC-Citeseer prod 59.10±1.28 62.62±0.71 59.87±1.64 60.63±4.85 1.53 –1.99 0.76 Ind 59.19±1.10 62.74±0.72 59.59±1.69 60.48±4.71 1.29 –2.26 0.89 trans 58.38±4.42 62.46±3.32 62.39±4.07 60.36±4.88 1.98 –2.60 –2.03 DBLP-Conf prod 67.12±0.75 93.96±0.31 76.17±0.69 78.53±1.41 11.41 –15.43 2.36 Ind 67.07±0.74 93.87±0.28 74.18±0.79 76.51±1.16 9.44 –17.36 2.33 trans 67.60±3.08 94.19±1.23 94.02±1.86 90.48±3.6 22.88 –3.71 –0.94 DBLP-Paper prod 67.12±0.75 69.62±1.29 71.15±1.06 73.26±1.3 6.14 3.64 2.11 Ind 67.07±0.74 69.67±1.26 71.07±1.09 73.42±0.95 6.35 3.75 2.35 trans 67.60±3.08 68.88±2.25 71.83±3.18 73.66±4.30 6.06 4.78 1.83 DBLP-Term prod 67.12±0.75 81.48±1.71 75.32±0.52 75.73±1.03 8.61 –5.75 0.41 Ind 67.07±0.74 81.40±1.71 74.74±0.61 75.37±1.02 8.30 –6.03 0.63 trans 67.60±3.08 81.17±1.72 80.45±1.76 78.41±2.41 10.81 –2.76 –2.04 IMDB-AW prod 41.29±1.59 46.28±2.04 49.61±1.15 49.90±1.65 8.61 3.62 0.29 Ind 41.32±1.53 46.25±1.96 48.24±1.10 48.94±1.62 7.55 2.69 0.70 trans 41.07±3.89 44.49±3.68 50.29±3.53 50.36±3.08 9.29 2.87 0.07 平均排名/平均值 4 1.92 2.45 1.63 8.78 –1.53 1.15 表 4 图和超图数据集上的实验结果(%)
Category Model Graph Datasets Hypergraph Datasets Avg.Rank Cora Citeseer Pubmed CA-Cora dblp-paper IMDB-AW MLPs MLP 49.99±1.20 52.45±2.53 65.92±1.81 51.74±1.58 63.02±1.43 40.69±1.51 12 KANs eff-kan 55.80±0.67 60.32±2.33 68.88±1.57 54.85±3.36 66.39±1.15 41.64±1.04 11 GNNs GCN 79.67±1.68 69.63±1.75 76.95±1.59 70.73±1.64 69.09±1.51 46.31±1.34 6 GAT 78.46±2.17 69.41±2.42 76.86±1.48 70.62±1.57 69.48±1.31 45.78±1.69 7 HGNNs HGNN 76.57±2.92 65.55±2.16 75.27±2.87 70.50±2.54 69.82±1.29 46.28±2.04 7.83 HGNN+ 75.65±1.86 65.74±1.97 75.14±1.85 71.93±2.15 70.75±2.04 47.86±2.95 6.17 GNNs-to-MLP GLNN 80.89±1.78 69.75±1.52 78.24±1.86 71.18±3.79 69.51±1.59 46.38±1.62 3.83 KRD 79.53±1.62 69.70±2.36 78.67±1.79 70.83±3.61 69.61±3.46 45.35±2.34 5 NOSMOG 80.22±1.15 70.79±1.23 80.39±1.25 68.89±3.42 69.42±1.98 46.16±1.59 4.83 HGNNs-to-MLP LHGNN 77.92±4.16 66.35±2.57 78.60±2.86 68.66±3.62 72.30±1.71 50.32±1.70 5.08 LHGNN+ 77.92±4.16 65.59±1.78 78.09±2.93 67.53±4.69 72.60±1.57 49.76±2.19 6.08 DH2KAN 78.05±2.99 69.78±0.95 77.71±2.45 71.30±6.13 74.40±1.04 49.90±1.59 3.17 表 5 消融实验对比结果(%)
数据集 w/o KAN w/o RSD w/o HKD DH2KAN △KAN △RSD △HKD News20 76.26±1.03 76.94±1.09 76.76±1.25 77.54±1.19 ↑1.68 ↑0.78 ↑1.02 CA-Cora 66.96±3.75 67.21±5.78 67.30±5.58 67.91±5.53 ↑1.42 ↑1.04 ↑0.91 CC-Cora 66.72±3.42 67.36±3.24 67.46±4.37 67.66±3.15 ↑1.41 ↑0.45 ↑0.30 CC-Citeseer 63.55±1.57 65.80±2.09 65.33±3.12 65.19±2.15 ↑2.58 ↓–0.93 ↓–0.21 DBLP-Conf 88.26±0.34 84.97±0.54 85.67±0.45 85.97±0.63 ↓–2.59 ↑1.18 ↑0.35 DBLP-Paper 73.88±1.45 75.10±0.80 75.24±1.82 75.42±0.76 ↑2.08 ↑0.43 ↑0.24 DBLP-Term 78.28±2.19 77.67±2.22 77.81±1.80 78.36±2.31 ↑0.10 ↑0.89 ↑0.71 IMDB-AW 48.40±2.24 48.98±2.65 48.53±2.51 49.72±2.46 ↑2.73 ↑1.51 ↑2.45 平均值 70.29 70.50 70.59 70.97 ↑1.18 ↑0.67 ↑0.72 -
[1] 谢丽霞, 史镜琛, 杨宏宇, 等. 基于图神经网络模型校准的成员推理攻击[J]. 电子与信息学报, 2025, 47(3): 780–791. doi: 10.11999/JEIT240477.XIE Lixia, SHI Jingchen, YANG Hongyu, et al. Membership inference attacks based on graph neural network model calibration[J]. Journal of Electronics & Information Technology, 2025, 47(3): 780–791. doi: 10.11999/JEIT240477. [2] 吴翼腾, 刘伟, 于洪涛, 等. 基于局部影响分析模型的图神经网络对抗攻击[J]. 电子与信息学报, 2022, 44(7): 2576–2583. doi: 10.11999/JEIT210448.WU Yiteng, LIU Wei, YU Hongtao, et al. Adversarial attacks on graph neural network based on local influence analysis model[J]. Journal of Electronics & Information Technology, 2022, 44(7): 2576–2583. doi: 10.11999/JEIT210448. [3] GAO Yue, ZHANG Zizhao, LIN Haojie, et al. Hypergraph learning: Methods and practices[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(5): 2548–2566. doi: 10.1109/TPAMI.2020.3039374. [4] TIAN Yijun, PEI Shichao, ZHANG Xiangliang, et al. Knowledge distillation on graphs: A survey[J]. ACM Computing Surveys, 2025, 57(8): 189. doi: 10.1145/3711121. [5] FENG Yifan, LUO Yihe, YING Shihui, et al. LightHGNN: Distilling hypergraph neural networks into MLPs for 100x faster inference[C]. The Twelfth International Conference on Learning Representations, Vienna, Austria, 2024. [6] LIU Ziming, WANG Yixuan, VAIDYA S, et al. KAN: Kolmogorov–Arnold networks[C]. The Thirteenth International Conference on Learning Representations, Singapore, Singapore, 2025: 70367–70413. [7] LIAO Sihao, XIE Liang, DU Yuanchuang, et al. Stock trend prediction based on dynamic hypergraph spatio-temporal network[J]. Applied Soft Computing, 2024, 154: 111329. doi: 10.1016/J.ASOC.2024.111329. [8] KIM S, LEE S Y, GAO Yue, et al. A survey on hypergraph neural networks: An in-depth and step-by-step guide[C]. The 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain, 2024: 6534–6544. [9] BERGE C. Hypergraphs: Combinatorics of Finite Sets[M]. Amsterdam: Elsevier, 1989. [10] FENG Yifan, YOU Haoxuan, ZHANG Zizhao, et al. Hypergraph neural networks[C]. The 33rd AAAI Conference on Artificial Intelligence, Honolulu, USA, 2019: 3558–3565. [11] GAO Yue, FENG Yifan, JI Shuyi, et al. HGNN+: General hypergraph neural networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3181–3199. doi: 10.1109/TPAMI.2022.3182052. [12] YADATI N, NIMISHAKAVI M, YADAV P, et al. HyperGCN: A new method of training graph convolutional networks on hypergraphs[C]. The 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, 2019: 135. [13] DONG Yihe, SAWIN W, and BENGIO Y. HNHN: Hypergraph networks with hyperedge neurons[J]. arXiv: 2006.12278, 2020. doi: 10.48550/arXiv.2006.12278. [14] JIANG Jianwen, WEI Yuxuan, FENG Yifan, et al. Dynamic hypergraph neural networks[C]. The Twenty-Eighth International Joint Conference on Artificial Intelligence, Macao, China, 2019: 2635–2641. doi: 10.24963/ijcai.2019/366. [15] BAI Song, ZHANG Feihu, and TORR P H S. Hypergraph convolution and hypergraph attention[J]. Pattern Recognition, 2021, 110: 107637. doi: 10.1016/j.patcog.2020.107637. [16] JO J, BAEK J, LEE S, et al. Edge representation learning with hypergraphs[C]. The 35th International Conference on Neural Information Processing Systems, 2021: 577. [17] HUANG Jing and YANG Jie. UniGNN: A unified framework for graph and hypergraph neural networks[C]. The Thirtieth International Joint Conference on Artificial Intelligence, Montreal, Canada, 2021: 2563–2569. doi: 10.24963/IJCAI.2021/353. [18] CHIEN E, PAN Chao, PENG Jianhao, et al. You are AllSet: A multiset function framework for hypergraph neural networks[C]. The Tenth International Conference on Learning Representations, 2022. [19] HINTON G, VINYALS O, and DEAN J. Distilling the knowledge in a neural network[J]. arXiv: 1503.02531, 2015. doi: 10.48550/arXiv.1503.02531. [20] 罗一畅, 齐析屿, 张博锐, 等. 分割一切模型的轻量化研究综述[J]. 电子与信息学报, 2026, 48(2): 713–731. doi: 10.11999/JEIT250894.LUO Yichang, QI Xiyu, ZHANG Borui, et al. A survey of lightweight techniques for segment anything model[J]. Journal of Electronics & Information Technology, 2026, 48(2): 713–731. doi: 10.11999/JEIT250894. [21] ZHANG Shichang, LIU Yozen, SUN Yizhou, et al. Graph-less neural networks: Teaching old MLPs new tricks via distillation[C]. The Tenth International Conference on Learning Representations, 2022. [22] YANG Cheng, LIU Jiawei, and SHI Chuan. Extract the knowledge of graph neural networks and go beyond it: An effective knowledge distillation framework[C]. The Web Conference 2021, Ljubljana, Slovenia, 2021: 1227–1237. doi: 10.1145/3442381.3450068. [23] WU Lirong, LIN Haitao, HUANG Yufei, et al. Quantifying the knowledge in GNNs for reliable distillation into MLPs[C]. The 40th International Conference on Machine Learning, Honolulu, USA, 2023: 37571–37581. [24] LIU Ziming, MA Pingchuan, WANG Yixuan, et al. KAN 2.0: Kolmogorov-Arnold networks meet science[J]. arXiv: 2408.10205, 2024. doi: 10.48550/arxiv.2408.10205. [25] 郑庆河, 刘方霖, 余礼苏, 等. 基于改进Kolmogorov-Arnold混合卷积神经网络的调制识别方法[J]. 电子与信息学报, 2025, 47(8): 2584–2597. doi: 10.11999/JEIT250161.ZHENG Qinghe, LIU Fanglin, YU Lisu, et al. An improved modulation recognition method based on hybrid Kolmogorov-Arnold convolutional neural network[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2584–2597. doi: 10.11999/JEIT250161. [26] Blealtan. An efficient implementation of Kolmogorov-Arnold network[EB/OL]. https://github.com/Blealtan/efficient-kan, 2024. [27] ZHANG Fan and ZHANG Xin. GraphKAN: Enhancing feature extraction with graph Kolmogorov Arnold networks[J]. arXiv: 2406.13597, 2024. doi: 10.48550/arXiv.2406.13597. [28] GAL Y and GHAHRAMANI Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning[C]. The 33nd International Conference on Machine Learning, New York, USA, 2016: 1050–1059. [29] SZEGEDY C, ZAREMBA W, SUTSKEVER I, et al. Intriguing properties of neural networks[C]. 2nd International Conference on Learning Representations, Banff, Canada, 2014. [30] PHAM H, DAI Zihang, XIE Qizhe, et al. Meta pseudo labels[C]. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 11552–11563. doi: 10.1109/CVPR46437.2021.01139. [31] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]. The 31st International Conference on Neural Information Processing Systems, Long Beach, USA, 2017: 6000–6010. -
下载: