Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement
-
摘要: 词义消歧(WSD)是生物医学文本理解与信息挖掘的核心技术,广泛应用于医学文献分析、临床数据处理等领域。为解决现有模型在生物医学场景中面临的术语语义易受噪声干扰、细粒度语义类区分难、小样本泛化不足等问题,该文提出了一种将生物医学术语中的词形、词性、语义类作为核心消歧特征,输入至Electra、mDeBERTa与T5 3路模型中来获取领域适配的动态词向量,使用Focal loss和Margin loss组合的复合损失函数3路并行架构。该文引入卡方统计量引导的注意力机制,筛选高代表性特征,有效抑制冗余噪声干扰。在模型训练阶段,采用两阶段困难样本挖掘策略:第1阶段精准挖掘困难样本,第2阶段针对性攻克样本学习难题。同时引入基于生物医学术语裁剪的对比学习方法,构建语义一致的双视图进行对比训练,显著增强特征的语义鲁棒性与区分能力。实验结果证明,该方法能够显著提升模型对近义语义类的区分能力,并增强模型在噪声高、少样本场景下的鲁棒性与泛化能力,表现出更优秀的生物医学词义消歧性能。Abstract:
Objective Biomedical Word Sense Disambiguation (WSD) is important for biomedical text mining and clinical data analysis. However, biomedical terms are susceptible to semantic noise, fine-grained semantic classes are difficult to distinguish, and model generalization is limited in small-sample settings. A three-branch parallel WSD framework with contrastive learning is proposed to address these problems. Electra, mDeBERTa, and Flan-T5 (FT5) are integrated to obtain complementary semantic features. Chi-square attention, a Focal Loss and Margin Loss hybrid loss, and hard sample mining are further incorporated to improve feature discrimination and model robustness. Methods The proposed framework uses Electra, mDeBERTa, and FT5 to extract complementary semantic features from biomedical terms. A chi-square attention module is designed to identify representative biomedical terms and guide token-level attention allocation. Focal Loss and Margin Loss are combined to improve the learning of hard samples and class boundaries under class imbalance. A two-stage hard sample mining strategy is developed by jointly considering training loss and prediction uncertainty. In addition, a constrained contrastive learning mechanism based on core disambiguation tokens is introduced. Random cropping is used to generate semantically equivalent augmented views, and the NT-Xent loss is applied to optimize the representation space. Results and Discussions Experiments are conducted on the MSH WSD dataset, which contains 203 ambiguous biomedical terms. The proposed model achieves an accuracy of 95.27%, a precision of 90.33%, a recall of 85.49%, and an F1 score of 87.84%. It outperforms the Neural Concept Embeddings model, which achieves an accuracy of 94.34%, by 0.93 percentage points. Ablation experiments show that the successive addition of Word2Vec, FT5, hard sample mining, chi-square attention, the hybrid loss, and contrastive learning improves model performance. Among the evaluated contrastive learning settings, NT-Xent with a temperature of 0.10 and a contrastive weight of 1.00 achieves the best performance. The proposed hard sample mining strategy, which combines training loss and prediction uncertainty, also outperforms alternative sample selection methods. Conclusions A three-branch parallel contrastive learning framework is proposed for biomedical WSD. Electra, mDeBERTa, and FT5 are integrated to capture complementary semantic features. Chi-square attention is used to strengthen representative feature extraction, whereas the Focal Loss and Margin Loss hybrid loss improves learning of hard samples and class boundaries. Two-stage hard sample mining and constrained contrastive learning further enhance the discrimination of fine-grained semantic classes. Experimental results demonstrate that the proposed method improves biomedical WSD performance and reduces confusion among semantically similar biomedical concepts. The framework is evaluated on English biomedical texts from the MSH WSD dataset and can be further extended to multilingual biomedical corpora and domain-specific knowledge graphs. -
Key words:
- Word sense disambiguation /
- Biomedicine /
- Contrastive learning /
- Disambiguation feature /
- Loss function
-
表 1 对比学习算法对模型性能的影响(%)
算法 准确度 精确度 召回率 F1 Negative samples 95.27 90.33 85.49 87.84 Cluster 94.49 87.61 84.36 85.95 Asymmetric Network 94.14 85.71 84.87 85.29 Memory Bank 91.19 83.92 69.24 75.87 Adversarial Examples 93.37 85.87 80.04 82.85 注:加粗数值表示性能最优值 表 2 对比损失对模型性能的影响(%)
对比损失 准确度 精确度 召回率 F1 NT-Xent 95.27 90.33 85.49 87.84 Triplet 94.61 88.66 83.74 86.13 SupCon 94.42 86.77 85.07 85.91 Barlow Twins 93.20 87.49 76.99 81.91 Circle 94.01 87.38 81.89 84.55 注:加粗数值表示性能最优值。 表 3 第2视图对模型性能的影响(%)
方法 准确度 精确度 召回率 F1 synonym 94.16 87.31 82.82 85.01 mask 92.90 85.58 77.57 81.38 crop 95.27 90.33 85.49 87.84 drop-modality 93.83 87.92 80.14 83.85 表 4 困难样本集构建方法对模型性能的影响(%)
所基于的方法 准确度 精确度 召回率 F1 训练损失 94.77 88.56 84.77 86.63 不确定性 94.12 87.86 81.94 84.79 训练损失及不确定性 95.27 90.33 85.49 87.84 自步学习 94.20 88.98 81.00 84.81 在线困难样本挖掘 94.50 87.11 85.10 86.09 随机样本选取 93.93 86.75 82.19 84.41 “注:加粗数值表示性能最优值。 表 5 α对模型性能的影响(%)
α 准确度 精确度 召回率 F1 0 93.67 87.39 79.88 83.46 0.2 94.11 89.02 80.49 84.54 0.4 94.22 88.30 81.94 85.00 0.6 95.27 90.33 85.49 87.84 0.8 94.19 88.89 81.11 84.82 1.0 94.27 88.50 81.99 85.12 注:加粗数值表示性能最优值。 表 6 消融实验(%)
模型 准确度 精确度 召回率 F1 EM 85.20 65.13 55.92 60.18 EMw 88.39 72.17 68.26 70.16 EMwT 91.56 85.04 70.16 76.89 EMwT-H 92.47 83.89 77.16 80.39 EMwT-HC 93.95 87.33 81.58 84.36 EMwT-HCL 94.16 87.31 82.82 85.01 EMwT-HCL-Cl 95.27 90.33 85.49 87.84 注:加粗数值表示性能最优值。 表 7 对比实验(%)
模型 准确度 Electra 70.89 mDeBERTa 69.90 FT5 72.70 Bert 78.27 Word-Concept Model 89.10 Attention Neural Network 91.38 Attention cct-T 93.94 Neural Concept Embeddings 94.34 EMwT-HCL-Cl 95.27 注:加粗数值表示性能最优值。 表 8 参数量及时间复杂度
模型 参数量 (M) 时间复杂度 (Big-O) Electra 110 $ O(L\cdot H_\text{E}^{2}) $ mDeBERTa 180 $ O(L\cdot H_\text{M}^{2}) $ FT5 250 $ O(L\cdot H_\text{T}^{2}) $ 卡方注意力 0.77 $ O(L\cdot H\cdot D) $ 困难样本挖掘 0 $ O(N\cdot {f}_{\text{forward}}) $ 对比学习 0.34 $ O(F\cdot H) $ EMwT-HCL-Cl ≈541.1 $ O(L(H_\text{E}^{2}+H_\text{M}^{2}+H_\text{T}^{2})+LDH+FH) $ -
[1] 李俊辉, 侯兴松. 基于伪监督注意力短期记忆与多尺度去伪影网络的图像分块压缩感知[J]. 电子与信息学报, 2024, 46(2): 472–480. doi: 10.11999/JEIT231069.LI Junhui and HOU Xingsong. Pseudo supervised attention short-term memory and multi-scale deartifacting network based on image block compressed sensing[J]. Journal of Electronics & Information Technology, 2024, 46(2): 472–480. doi: 10.11999/JEIT231069. [2] 杨春玲, 梁梓文. 静态与动态域先验增强的两阶段视频压缩感知重构网络[J]. 电子与信息学报, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.YANG Chunling and LIANG Ziwen. Static and dynamic-domain prior enhancement two-stage video compressed sensing reconstruction network[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295. [3] SUNG S F, HU Yahan, and CHEN Chongyan. Disambiguating clinical abbreviations by one-to-all classification: Algorithm development and validation study[J]. JMIR Medical Informatics, 2024, 12: e56955. doi: 10.2196/56955. [4] KADDOURA S and NASSAR R. EnhancedBERT: A feature-rich ensemble model for Arabic word sense disambiguation with statistical analysis and optimized data collection[J]. Journal of King Saud University-Computer and Information Sciences, 2024, 36(1): 101911. doi: 10.1016/j.jksuci.2023.101911. [5] CAO Yukun, JIN Chengkun, TANG Yijia, et al. Word sense disambiguation combining knowledge graph and text hierarchical structure[J]. ACM Transactions on Asian and Low-Resource Language Information Processing, 2024, 23(12): 161. doi: 10.1145/3677524. [6] 张春祥, 孙颖, 高可心, 等. 结合预训练模型的双向门控图卷积对抗词义消歧[J]. 电子与信息学报, 2025, 47(11): 4549–4559. doi: 10.11999/JEIT250386.ZHANG Chunxiang, SUN Ying, GAO Kexin, et al. Combine the pre-trained model with bidirectional gated recurrent units and graph convolutional network for adversarial word sense disambiguation[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4549–4559. doi: 10.11999/JEIT250386. [7] ION R, PĂIȘ V, MITITELU V B, et al. Unsupervised word sense disambiguation using transformer’s attention mechanism[J]. Machine Learning and Knowledge Extraction, 2025, 7(1): 10. doi: 10.3390/make7010010. [8] RANSING R and GULATI A. Unsupervised word sense disambiguation for Marathi language using word embeddings[J]. International Journal of Intelligent Systems and Applications in Engineering, 2024, 12(3): 1374–1380. [9] PADWAD H, KESWANI G, BISEN W, et al. Leveraging contextual factors for word sense disambiguation in Hindi language[J]. International Journal of Intelligent Systems and Applications in Engineering, 2024, 12(12s): 129–136. [10] RANSING R and GULATI A. Marathi word sense disambiguation through unsupervised k-means clustering[J]. Engineering, Technology & Applied Science Research, 2025, 15(3): 22837–22843. doi: 10.48084/etasr.9975. [11] MARTIN D I, BERRY M W, and MARTIN J C. Semantic unsupervised learning for word sense disambiguation[M]. BERRY M W, MOHAMED A, and YAP B W. Supervised and Unsupervised Learning for Data Science. Cham, Switzerland: Springer, 2019: 101–120. doi: 10.1007/978-3-030-22475-2_6. [12] RAHMANI S, FAKHRAHMAD S M, and SADREDDINI M H. Co-occurrence graph-based context adaptation: A new unsupervised approach to word sense disambiguation[J]. Digital Scholarship in the Humanities, 2021, 36(2): 449–471. doi: 10.1093/llc/fqz048. [13] LU Wenpeng, MENG Fanqing, WANG Shoujin, et al. Graph-based Chinese word sense disambiguation with multi-knowledge integration[J]. CMC-Computers, Materials & Continua, 2019, 61(1): 197–212. doi: 10.32604/cmc.2019.06068. [14] BLEVINS T and ZETTLEMOYER L. Moving down the long tail of word sense disambiguation with Gloss Informed Bi-encoders[C].The 58th Annual Meeting of the Association for Computational Linguistics, 2020: 1006–1017. doi: 10.18653/v1/2020.acl-main.95. [15] ANTUNES R and MATOS S. Supervised learning and knowledge-based approaches applied to biomedical word sense disambiguation[J]. Journal of Integrative Bioinformatics, 2017, 14(4): 20170051. doi: 10.1515/jib-2017-0051. [16] ALSAEEDAN W, MENAI M E B, and AL-AHMADI S. A hybrid genetic-ant colony optimization algorithm for the word sense disambiguation problem[J]. Information Sciences, 2017, 417: 20–38. doi: 10.1016/j.ins.2017.07.002. [17] SARROUTI M and OUATIK EL ALAOUI S. SemBioNLQA: A semantic biomedical question answering system for retrieving exact and ideal answers to natural language questions[J]. Artificial Intelligence in Medicine, 2020, 102: 101767. doi: 10.1016/j.artmed.2019.101767. [18] CHEN Fangyi, ZHANG Gongbo, CHEN Si, et al. Clinical note structural knowledge improves word sense disambiguation[J]. AMIA Joint Summits on Translational Science Proceedings, 2024, 2024: 515–524. [19] RAIS M and LACHKAR A. An empirical study of word sense disambiguation for biomedical information retrieval system[C]. The 6th International Conference on Bioinformatics and Biomedical Engineering, Granada, Spain, 2018: 314–326. doi: 10.1007/978-3-319-78723-7_27. [20] PASHUK A V, GURINOVICH A B, VOLOROVA N A, et al. Analysis of the methods of word sense disambiguation in the biomedical domain[J]. Doklady BGUIR, 2019, 75(5): 60–65. doi: 10.35596/1729-7648-2019-123-5-60-65. [21] YAP B P, KOH A, and CHNG E S. Adapting BERT for word sense disambiguation with gloss selection objective and example sentences[C]. Findings of the Association for Computational Linguistics, 2020: 41–46. doi: 10.18653/v1/2020.findings-emnlp.4. [22] KWON S, OH D, and KO Y. Word sense disambiguation based on context selection using knowledge-based word similarity[J]. Information Processing & Management, 2021, 58(4): 102551. doi: 10.1016/j.ipm.2021.102551. [23] ALMOUSA M, BENLAMRI R, and KHOURY R. A novel word sense disambiguation approach using WordNet knowledge graph[J]. Computer Speech & Language, 2022, 74: 101337. doi: 10.1016/j.csl.2021.101337. [24] CHEN Ting, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations[C]. 37th International Conference on Machine Learning, 2020: 1597–1607. [25] CARON M, MISRA I, MAIRAL J, et al. Unsupervised learning of visual features by contrasting cluster assignments[C].The 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2020: 831. [26] GRILL J B, STRUB F, ALTCHÉ F, et al. Bootstrap your own latent a new approach to self-supervised learning[C]. Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2020: 1786. [27] WU Zhirong, XIONG Yuanjun, YU S X, et al. Unsupervised feature learning via non-parametric instance discrimination[C].2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 3733–3742. doi: 10.1109/CVPR.2018.00393. [28] HU Qianjiang, WANG Xiao, HU Wei, et al. AdCo: Adversarial contrast for efficient learning of unsupervised representations from self-trained negative adversaries[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 1074–1083. doi: 10.1109/CVPR46437.2021.00113. [29] YEPES A J and BERLANGA R. Knowledge based word-concept model estimation and refinement for biomedical text mining[J]. Journal of Biomedical Informatics, 2015, 53: 300–307. doi: 10.1016/j.jbi.2014.11.015. [30] ZHANG Linfeng, CHEN Xin, ZHANG Junbo, et al. Contrastive deep supervision[C]. 17th European Conference on Computer Vision, Tel Aviv, Israel, 2022: 1–19. doi: 10.1007/978-3-031-19809-0_1. [31] SABBIR A K M, JIMENO-YEPES A, and KAVULURU R. Knowledge-based biomedical word sense disambiguation with neural concept embeddings[C]. 2017 IEEE 17th International Conference on Bioinformatics and Bioengineering, Washington, USA, 2017: 163–170. doi: 10.1109/BIBE.2017.00-61. [32] ZHANG Canlin, BIŚ D, LIU Xiuwen, et al. Biomedical word sense disambiguation with bidirectional long short-term memory and attention-based neural networks[J]. BMC Bioinformatics, 2019, 20(S16): 502. doi: 10.1186/s12859-019-3079-8. [33] CHUNG H W, HOU Le, LONGPRE S, et al. Scaling instruction-finetuned language models[J]. The Journal of Machine Learning Research, 2024, 25(1): 70. -
下载: