A knowledge distillation framework for hypergraph neural networks with rapid inference capabilities
-
摘要: 超图神经网络(HGNN)凭借高阶关系建模能力被广泛关注,但其计算量大、推理效率低,难以在实际场景部署。现有HGNN向多层感知机(MLP)知识蒸馏方法存在可解释性差、精度偏低等问题,主要原因在于:MLP激活函数单一且层间解耦、Softmax软标签存在信息丢失、未关注节点可靠性差异。为此,本文提出面向快速推理的超图神经网络知识蒸馏框架DH2KAN。该方法以柯尔莫哥洛夫阿诺德网络(KAN)为学生模型,用可学习单变量样条函数提升表征能力;设计表征相似性蒸馏,以节点表征监督KAN训练;提出高可靠节点感知蒸馏,基于信息熵量化节点可靠性并利用可靠节点软标签强化指导。在11个真实数据集上的实验表明,DH2KAN在精度与HGNN相当甚至更优的前提下,推理速度提升75倍;相比KAN在速度相近的情况下,准确率提升10.53%,适用于大规模低延迟场景。本文代码已开源至GitHub:https://github.com/ljzology/DH2KAN。
-
关键词:
- 超图神经网络 /
- 柯尔莫哥洛夫-阿诺德网络 /
- 知识蒸馏 /
- 快速推理
Abstract:Objective Hypergraph Neural Networks (HGNNs) have gained widespread attention for their strong ability to model high-order correlations among entities, but their computational complexity and memory consumption grow exponentially as the hypergraph scale expands, severely restricting their deployment in large-scale industrial scenarios. Existing knowledge distillation methods that distill HGNNs into Multi-Layer Perceptrons (MLPs) are troubled by poor interpretability, low accuracy, severe information loss caused by Softmax-based soft labels, and neglect of node reliability heterogeneity. To address these critical challenges, this paper proposes a novel hypergraph knowledge distillation framework named DH2KAN (Distill Hypergraph Neural Network to Kolmogorov-Arnold Network) for fast inference, which breaks through the bottlenecks of traditional hypergraph knowledge distillation. Methods This paper designs a three-module knowledge distillation framework DH2KAN(图2). Firstly, we replace the traditional MLP with KAN as the student model, which uses learnable spline-based univariate functions instead of fixed activation functions and linear weights to improve the fitting ability and interpretability of the student model. Secondly, we propose a representation similarity distillation mechanism, which directly aligns the pre-logits representations of HGNN and KAN to avoid information loss caused by the Softmax normalization layer and completely retain the high-order structural knowledge of hypergraphs. Thirdly, we introduce a high-reliable node-aware distillation method(图3), which quantifies the node reliability by information entropy variation, screens out robust nodes with strong anti-noise ability, and takes their soft labels as the core supervision signal to improve the purity of distilled knowledge. Results and Discussions The DH2KAN algorithm achieves remarkable performance under both transductive learning (表2) and production learning (表3) settings. Quantitative experimental results reveal that DH2KAN obtains an accuracy improvement of approximately 10.53% over the vanilla KAN student model, 1.73% over the teacher HGNN model, and 1.2% over existing MLP-based distillation methods. Such results verify the effectiveness of knowledge transfer from HGNN to KAN, and demonstrate that the proposed method outperforms conventional MLP-oriented distillation schemes in inference performance. In addition, DH2KAN achieves optimal performance on feature-dominated hypergraph datasets, and possesses strong robustness when handling sparse structures and noisy node samples. Conclusions This paper proposes DH2KAN to accelerate HGNN inference for large-scale low-latency applications. Via knowledge distillation, it bridges the performance gap between KAN and HGNN and eliminates structural dependence for efficient reasoning. With representation similarity and reliable node-aware distillation, it transfers effective task knowledge via pre-logit features and soft labels, showing great practical application potential. -
表 1 训练和推理期间的时间复杂度比较
KANs HGNN DH2KAN Training O(LNF2 (G+K)) O(LN2 F+LNF2) O(NMC+LNF2
(G+K))Inference O(LNF2(G+K)) O(LN2 F+LNF2) O(LNF2(G+K)) 表 2 直推式学习下在8个超图数据集的实验结果(%)
Datasets KAN HGNN LHGNN DH2KAN △KAN △HGNN △LHGN News20 67.4±2.58 75.02±1.42 75.06±1.03 77.54±1.19 10.14 2.52 2.48 CA-Cora 57.26±4.78 63.85±3.66 65.96±4.76 67.91±5.53 10.65 4.06 1.95 CC-Cora 56.77±4.74 61.93±4.41 63.52±4.23 67.66±3.15 9.89 5.73 4.14 CC-Citeseer 59.29±1.3 59.91±1.94 61.45±1.77 65.19±2.15 5.9 5.28 3.74 DBLP-Conf 67.32±0.57 93.21±0.6 91.46±0.46 85.97±0.63 18.65 –7.24 –5.49 DBLP-Paper 67.32±0.57 70.85±2.03 72.68±1.66 75.42±0.76 8.1 4.57 2.74 DBLP-Term 67.32±0.57 80.19±2.22 77.68±2.39 78.36±2.31 11.04 –1.83 0.68 IMDB-AW 40.88±1.28 49.01±2.25 50.4±2.04 49.72±2.46 8.84 0.71 –0.68 Avg.Rank/Avg 4 2.5 2 1.5 10.53 1.73 1.2 表 3 生产式学习模式下在8个超图数据集的实验结果(%)
Dataset Setting KAN HGNN LHGNN DH2KAN △KAN △HGNN △LHGNN News20 prod 66.68±2.16 75.65±0.70 76.80±0.85 77.52±0.45 10.84 1.87 0.72 Ind 66.60±2.17 75.58±0.63 76.74±0.83 77.17±0.37 10.57 1.59 0.43 trans 67.44±2.26 76.30±1.72 77.33±1.27 78.17±1.08 10.73 1.87 0.84 CA-Cora prod 57.70±2.68 70.50±2.54 67.86±2.05 71.30±3.56 13.60 0.80 3.44 Ind 57.75±2.61 70.64±2.58 65.08±2.04 67.75±3.68 10.00 –2.89 2.67 trans 57.22±4.10 70.81±3.98 69.11±1.93 72.15±2.98 14.93 1.34 3.04 CC-Cora prod 58.03±2.32 67.17±2.90 61.32±1.88 63.64±2.46 5.61 –3.53 2.32 Ind 58.11±2.22 67.16±2.91 60.75±1.8 63.17±2.58 5.06 –3.99 2.42 trans 57.33±3.76 63.64±2.89 64.39±3.34 66.44±3.25 9.11 2.80 2.05 CC-Citeseer prod 59.10±1.28 62.62±0.71 59.87±1.64 60.63±4.85 1.53 –1.99 0.76 Ind 59.19±1.10 62.74±0.72 59.59±1.69 60.48±4.71 1.29 –2.26 0.89 trans 58.38±4.42 62.46±3.32 62.39±4.07 60.36±4.88 1.98 –2.60 –2.03 DBLP-Conf prod 67.12±0.75 93.96±0.31 76.17±0.69 78.53±1.41 11.41 –15.43 2.36 Ind 67.07±0.74 93.87±0.28 74.18±0.79 76.51±1.16 9.44 –17.36 2.33 trans 67.60±3.08 94.19±1.23 94.02±1.86 90.48±3.6 22.88 –3.71 –0.94 DBLP-Paper prod 67.12±0.75 69.62±1.29 71.15±1.06 73.26±1.3 6.14 3.64 2.11 Ind 67.07±0.74 69.67±1.26 71.07±1.09 73.42±0.95 6.35 3.75 2.35 trans 67.60±3.08 68.88±2.25 71.83±3.18 73.66±4.30 6.06 4.78 1.83 DBLP-Term prod 67.12±0.75 81.48±1.71 75.32±0.52 75.73±1.03 8.61 –5.75 0.41 Ind 67.07±0.74 81.40±1.71 74.74±0.61 75.37±1.02 8.30 –6.03 0.63 trans 67.60±3.08 81.17±1.72 80.45±1.76 78.41±2.41 10.81 –2.76 –2.04 IMDB-AW prod 41.29±1.59 46.28±2.04 49.61±1.15 49.90±1.65 8.61 3.62 0.29 Ind 41.32±1.53 46.25±1.96 48.24±1.10 48.94±1.62 7.55 2.69 0.70 trans 41.07±3.89 44.49±3.68 50.29±3.53 50.36±3.08 9.29 2.87 0.07 Avg.Rank/Avg. 4 1.92 2.45 1.63 8.78 –1.53 1.15 表 4 图和超图数据集上的实验结果(%)
Category Model Graph Datasets Hypergraph Datasets Avg.Rank Cora Citeseer Pubmed CA-Cora dblp-paper IMDB-AW MLPs MLP 49.99±1.20 52.45±2.53 65.92±1.81 51.74±1.58 63.02±1.43 40.69±1.51 12 KANs eff-kan 55.80±0.67 60.32±2.33 68.88±1.57 54.85±3.36 66.39±1.15 41.64±1.04 11 GNNs GCN 79.67±1.68 69.63±1.75 76.95±1.59 70.73±1.64 69.09±1.51 46.31±1.34 6 GAT 78.46±2.17 69.41±2.42 76.86±1.48 70.62±1.57 69.48±1.31 45.78±1.69 7 HGNNs HGNN 76.57±2.92 65.55±2.16 75.27±2.87 70.50±2.54 69.82±1.29 46.28±2.04 7.83 HGNN+ 75.65±1.86 65.74±1.97 75.14±1.85 71.93±2.15 70.75±2.04 47.86±2.95 6.17 GNNs-to-MLP GLNN 80.89±1.78 69.75±1.52 78.24±1.86 71.18±3.79 69.51±1.59 46.38±1.62 3.83 KRD 79.53±1.62 69.70±2.36 78.67±1.79 70.83±3.61 69.61±3.46 45.35±2.34 5 NOSMOG 80.22±1.15 70.79±1.23 80.39±1.25 68.89±3.42 69.42±1.98 46.16±1.59 4.83 HGNNs-to-MLP LHGNN 77.92±4.16 66.35±2.57 78.60±2.86 68.66±3.62 72.30±1.71 50.32±1.70 5.08 LHGNN+ 77.92±4.16 65.59±1.78 78.09±2.93 67.53±4.69 72.60±1.57 49.76±2.19 6.08 DH2KAN 78.05±2.99 69.78±0.95 77.71±2.45 71.30±6.13 74.40±1.04 49.90±1.59 3.17 表 5 消融实验对比结果
Datasets w/o KAN w/o RSD w/o HKD DH2KAN △KAN △RSD △HKD News20 76.26±1.03 76.94±1.09 76.76±1.25 77.54±1.19 ↑1.68% ↑0.78% ↑1.02% CA-Cora 66.96±3.75 67.21±5.78 67.30±5.58 67.91±5.53 ↑1.42% ↑1.04% ↑0.91% CC-Cora 66.72±3.42 67.36±3.24 67.46±4.37 67.66±3.15 ↑1.41% ↑0.45% ↑0.30% CC-Citeseer 63.55±1.57 65.80±2.09 65.33±3.12 65.19±2.15 ↑2.58% ↓-0.93% ↓-0.21% DBLP-Conf 88.26±0.34 84.97±0.54 85.67±0.45 85.97±0.63 ↓-2.59% ↑1.18% ↑0.35% DBLP-Paper 73.88±1.45 75.10±0.80 75.24±1.82 75.42±0.76 ↑2.08% ↑0.43% ↑0.24% DBLP-Term 78.28±2.19 77.67±2.22 77.81±1.80 78.36±2.31 ↑0.10% ↑0.89% ↑0.71% IMDB-AW 48.40±2.24 48.98±2.65 48.53±2.51 49.72±2.46 ↑2.73% ↑1.51% ↑2.45% Avg.Rank/Avg 70.29 70.50 70.59 70.97 ↑1.18% ↑0.67% ↑0.72% -
[1] 谢丽霞, 史镜琛, 杨宏宇, 等. 基于图神经网络模型校准的成员推理攻击[J]. 电子与信息学报, 2025, 47(3): 780–791. doi: 10.11999/JEIT240477.XIE Lixia, SHI Jingchen, YANG Hongyu, et al. Membership inference attacks based on graph neural network model calibration[J]. Journal of Electronics & Information Technology, 2025, 47(3): 780–791. doi: 10.11999/JEIT240477. [2] 吴翼腾, 刘伟, 于洪涛, 等. 基于局部影响分析模型的图神经网络对抗攻击[J]. 电子与信息学报, 2022, 44(7): 2576–2583. doi: 10.11999/JEIT210448.WU Yiteng, LIU Wei, YU Hongtao, et al. Adversarial attacks on graph neural network based on local influence analysis model[J]. Journal of Electronics & Information Technology, 2022, 44(7): 2576–2583. doi: 10.11999/JEIT210448. [3] GAO Yue, ZHANG Zizhao, LIN Haojie, et al. Hypergraph learning: Methods and practices[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(5): 2548–2566. doi: 10.1109/TPAMI.2020.3039374. [4] TIAN Yijun, PEI Shichao, ZHANG Xiangliang, et al. Knowledge distillation on graphs: A survey[J]. ACM Computing Surveys, 2025, 57(8): 189. doi: 10.1145/3711121. [5] FENG Yifan, LUO Yihe, YING Shihui, et al. LightHGNN: Distilling hypergraph neural networks into MLPs for 100x faster inference[C]. The Twelfth International Conference on Learning Representations, Vienna, Austria, 2024. [6] LIU Ziming, WANG Yixuan, VAIDYA S, et al. KAN: Kolmogorov–Arnold networks[C]. The Thirteenth International Conference on Learning Representations, Singapore, Singapore, 2025: 70367–70413. [7] LIAO Sihao, XIE Liang, DU Yuanchuang, et al. Stock trend prediction based on dynamic hypergraph spatio-temporal network[J]. Applied Soft Computing, 2024, 154: 111329. doi: 10.1016/J.ASOC.2024.111329. [8] KIM S, LEE S Y, GAO Yue, et al. A survey on hypergraph neural networks: An in-depth and step-by-step guide[C]. Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain, 2024: 6534–6544. [9] BERGE C. Hypergraphs: Combinatorics of Finite Sets[M]. Amsterdam: Elsevier, 1989. (查阅网上资料, 请补充引用页码并核对出版年信息). [10] FENG Yifan, YOU Haoxuan, ZHANG Zizhao, et al. Hypergraph neural networks[C]. Proceedings of the 33rd AAAI Conference on Artificial Intelligence, Honolulu, USA, 2019: 3558–3565. [11] GAO Yue, FENG Yifan, JI Shuyi, et al. HGNN+: General hypergraph neural networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3181–3199. doi: 10.1109/TPAMI.2022.3182052. [12] YADATI N, NIMISHAKAVI M, YADAV P, et al. HyperGCN: A new method of training graph convolutional networks on hypergraphs[C]. Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, 2019: 135. [13] DONG Yihe, SAWIN W, and BENGIO Y. HNHN: Hypergraph networks with hyperedge neurons[J]. arXiv: 2006.12278, 2020. doi: 10.48550/arXiv.2006.12278. [14] JIANG Jianwen, WEI Yuxuan, FENG Yifan, et al. Dynamic hypergraph neural networks[C]. Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, Macao, China, 2019: 2635–2641. doi: 10.24963/ijcai.2019/366. [15] BAI Song, ZHANG Feihu, and TORR P H S. Hypergraph convolution and hypergraph attention[J]. Pattern Recognition, 2021, 110: 107637. doi: 10.1016/j.patcog.2020.107637. [16] JO J, BAEK J, LEE S, et al. Edge representation learning with hypergraphs[C]. Proceedings of the 35th International Conference on Neural Information Processing Systems, 2021: 577. (查阅网上资料, 未找到本条文献出版地信息, 请确认并补充). [17] HUANG Jing and YANG Jie. UniGNN: A unified framework for graph and hypergraph neural networks[C]. Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, Montreal, Canada, 2021: 2563–2569. doi: 10.24963/IJCAI.2021/353. [18] CHIEN E, PAN Chao, PENG Jianhao, et al. You are AllSet: A multiset function framework for hypergraph neural networks[C]. The Tenth International Conference on Learning Representations, 2022. (查阅网上资料, 未找到本条文献出版地信息, 请确认并补充). [19] HINTON G, VINYALS O, and DEAN J. Distilling the knowledge in a neural network[J]. arXiv: 1503.02531, 2015. doi: 10.48550/arXiv.1503.02531. [20] 罗一畅, 齐析屿, 张博锐, 等. 分割一切模型的轻量化研究综述[J]. 电子与信息学报, 2026, 48(2): 713–731. doi: 10.11999/JEIT250894.LUO Yichang, QI Xiyu, ZHANG Borui, et al. A survey of lightweight techniques for segment anything model[J]. Journal of Electronics & Information Technology, 2026, 48(2): 713–731. doi: 10.11999/JEIT250894. [21] ZHANG Shichang, LIU Yozen, SUN Yizhou, et al. Graph-less neural networks: Teaching old MLPs new tricks via distillation[C]. The Tenth International Conference on Learning Representations, 2022. (查阅网上资料, 未找到本条文献出版地信息, 请确认并补充). [22] YANG Cheng, LIU Jiawei, and SHI Chuan. Extract the knowledge of graph neural networks and go beyond it: An effective knowledge distillation framework[C]. Proceedings of the Web Conference 2021, Ljubljana, Slovenia, 2021: 1227–1237. doi: 10.1145/3442381.3450068. [23] WU Lirong, LIN Haitao, HUANG Yufei, et al. Quantifying the knowledge in GNNs for reliable distillation into MLPs[C]. Proceedings of the 40th International Conference on Machine Learning, Honolulu, USA, 2023: 37571–37581. [24] LIU Ziming, MA Pingchuan, WANG Yixuan, et al. KAN 2.0: Kolmogorov-Arnold networks meet science[J]. arXiv: 2408.10205, 2024. doi: 10.48550/arxiv.2408.10205. [25] 郑庆河, 刘方霖, 余礼苏, 等. 基于改进Kolmogorov-Arnold混合卷积神经网络的调制识别方法[J]. 电子与信息学报, 2025, 47(8): 2584–2597. doi: 10.11999/JEIT250161.ZHENG Qinghe, LIU Fanglin, YU Lisu, et al. An improved modulation recognition method based on hybrid Kolmogorov-Arnold convolutional neural network[J]. Journal of Electronics & Information Technology, 2025, 47(8): 2584–2597. doi: 10.11999/JEIT250161. [26] Blealtan. An efficient implementation of Kolmogorov-Arnold network[EB/OL]. https://github.com/Blealtan/efficient-kan, 2024. [27] ZHANG Fan and ZHANG Xin. GraphKAN: Enhancing feature extraction with graph Kolmogorov Arnold networks[J]. arXiv: 2406.13597, 2024. doi: 10.48550/arXiv.2406.13597. [28] GAL Y and GHAHRAMANI Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning[C]. Proceedings of the 33nd International Conference on Machine Learning, New York, USA, 2016: 1050–1059. [29] SZEGEDY C, ZAREMBA W, SUTSKEVER I, et al. Intriguing properties of neural networks[C]. 2nd International Conference on Learning Representations, Banff, Canada, 2014. [30] PHAM H, DAI Zihang, XIE Qizhe, et al. Meta pseudo labels[C]. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 11552–11563. doi: 10.1109/CVPR46437.2021.01139. [31] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, USA, 2017: 6000–6010. -
下载: