Advanced Search
Turn off MathJax
Article Contents
ZHANG Chunxiang, ZHANG Huibin, GAO Xueyao. Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260061
Citation: ZHANG Chunxiang, ZHANG Huibin, GAO Xueyao. Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260061

Biomedical Word Sense Disambiguation via Contrastive Learning with Feature and Sample Enhancement

doi: 10.11999/JEIT260061 cstr: 32379.14.JEIT260061
Funds:  The National Natural Science Foundation of China (61502124, 60903082), China Postdoctoral Science Foundation (2014M560249), Heilongjiang Provincial Natural Science Foundation of China (PL2025F017)
  • Received Date: 2026-01-16
  • Accepted Date: 2026-08-10
  • Rev Recd Date: 2026-08-10
  • Available Online: 2026-08-18
  •   Objective  Biomedical Word Sense Disambiguation (WSD) is important for biomedical text mining and clinical data analysis. However, biomedical terms are susceptible to semantic noise, fine-grained semantic classes are difficult to distinguish, and model generalization is limited in small-sample settings. A three-branch parallel WSD framework with contrastive learning is proposed to address these problems. Electra, mDeBERTa, and Flan-T5 (FT5) are integrated to obtain complementary semantic features. Chi-square attention, a Focal Loss and Margin Loss hybrid loss, and hard sample mining are further incorporated to improve feature discrimination and model robustness.  Methods  The proposed framework uses Electra, mDeBERTa, and FT5 to extract complementary semantic features from biomedical terms. A chi-square attention module is designed to identify representative biomedical terms and guide token-level attention allocation. Focal Loss and Margin Loss are combined to improve the learning of hard samples and class boundaries under class imbalance. A two-stage hard sample mining strategy is developed by jointly considering training loss and prediction uncertainty. In addition, a constrained contrastive learning mechanism based on core disambiguation tokens is introduced. Random cropping is used to generate semantically equivalent augmented views, and the NT-Xent loss is applied to optimize the representation space.  Results and Discussions  Experiments are conducted on the MSH WSD dataset, which contains 203 ambiguous biomedical terms. The proposed model achieves an accuracy of 95.27%, a precision of 90.33%, a recall of 85.49%, and an F1 score of 87.84%. It outperforms the Neural Concept Embeddings model, which achieves an accuracy of 94.34%, by 0.93 percentage points. Ablation experiments show that the successive addition of Word2Vec, FT5, hard sample mining, chi-square attention, the hybrid loss, and contrastive learning improves model performance. Among the evaluated contrastive learning settings, NT-Xent with a temperature of 0.10 and a contrastive weight of 1.00 achieves the best performance. The proposed hard sample mining strategy, which combines training loss and prediction uncertainty, also outperforms alternative sample selection methods.  Conclusions  A three-branch parallel contrastive learning framework is proposed for biomedical WSD. Electra, mDeBERTa, and FT5 are integrated to capture complementary semantic features. Chi-square attention is used to strengthen representative feature extraction, whereas the Focal Loss and Margin Loss hybrid loss improves learning of hard samples and class boundaries. Two-stage hard sample mining and constrained contrastive learning further enhance the discrimination of fine-grained semantic classes. Experimental results demonstrate that the proposed method improves biomedical WSD performance and reduces confusion among semantically similar biomedical concepts. The framework is evaluated on English biomedical texts from the MSH WSD dataset and can be further extended to multilingual biomedical corpora and domain-specific knowledge graphs.
  • loading
  • [1]
    李俊辉, 侯兴松. 基于伪监督注意力短期记忆与多尺度去伪影网络的图像分块压缩感知[J]. 电子与信息学报, 2024, 46(2): 472–480. doi: 10.11999/JEIT231069.

    LI Junhui and HOU Xingsong. Pseudo supervised attention short-term memory and multi-scale deartifacting network based on image block compressed sensing[J]. Journal of Electronics & Information Technology, 2024, 46(2): 472–480. doi: 10.11999/JEIT231069.
    [2]
    杨春玲, 梁梓文. 静态与动态域先验增强的两阶段视频压缩感知重构网络[J]. 电子与信息学报, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.

    YANG Chunling and LIANG Ziwen. Static and dynamic-domain prior enhancement two-stage video compressed sensing reconstruction network[J]. Journal of Electronics & Information Technology, 2024, 46(11): 4247–4258. doi: 10.11999/JEIT240295.
    [3]
    SUNG S F, HU Yahan, and CHEN Chongyan. Disambiguating clinical abbreviations by one-to-all classification: Algorithm development and validation study[J]. JMIR Medical Informatics, 2024, 12: e56955. doi: 10.2196/56955.
    [4]
    KADDOURA S and NASSAR R. EnhancedBERT: A feature-rich ensemble model for Arabic word sense disambiguation with statistical analysis and optimized data collection[J]. Journal of King Saud University-Computer and Information Sciences, 2024, 36(1): 101911. doi: 10.1016/j.jksuci.2023.101911.
    [5]
    CAO Yukun, JIN Chengkun, TANG Yijia, et al. Word sense disambiguation combining knowledge graph and text hierarchical structure[J]. ACM Transactions on Asian and Low-Resource Language Information Processing, 2024, 23(12): 161. doi: 10.1145/3677524.
    [6]
    张春祥, 孙颖, 高可心, 等. 结合预训练模型的双向门控图卷积对抗词义消歧[J]. 电子与信息学报, 2025, 47(11): 4549–4559. doi: 10.11999/JEIT250386.

    ZHANG Chunxiang, SUN Ying, GAO Kexin, et al. Combine the pre-trained model with bidirectional gated recurrent units and graph convolutional network for adversarial word sense disambiguation[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4549–4559. doi: 10.11999/JEIT250386.
    [7]
    ION R, PĂIȘ V, MITITELU V B, et al. Unsupervised word sense disambiguation using transformer’s attention mechanism[J]. Machine Learning and Knowledge Extraction, 2025, 7(1): 10. doi: 10.3390/make7010010.
    [8]
    RANSING R and GULATI A. Unsupervised word sense disambiguation for Marathi language using word embeddings[J]. International Journal of Intelligent Systems and Applications in Engineering, 2024, 12(3): 1374–1380.
    [9]
    PADWAD H, KESWANI G, BISEN W, et al. Leveraging contextual factors for word sense disambiguation in Hindi language[J]. International Journal of Intelligent Systems and Applications in Engineering, 2024, 12(12s): 129–136.
    [10]
    RANSING R and GULATI A. Marathi word sense disambiguation through unsupervised k-means clustering[J]. Engineering, Technology & Applied Science Research, 2025, 15(3): 22837–22843. doi: 10.48084/etasr.9975.
    [11]
    MARTIN D I, BERRY M W, and MARTIN J C. Semantic unsupervised learning for word sense disambiguation[M]. BERRY M W, MOHAMED A, and YAP B W. Supervised and Unsupervised Learning for Data Science. Cham, Switzerland: Springer, 2019: 101–120. doi: 10.1007/978-3-030-22475-2_6.
    [12]
    RAHMANI S, FAKHRAHMAD S M, and SADREDDINI M H. Co-occurrence graph-based context adaptation: A new unsupervised approach to word sense disambiguation[J]. Digital Scholarship in the Humanities, 2021, 36(2): 449–471. doi: 10.1093/llc/fqz048.
    [13]
    LU Wenpeng, MENG Fanqing, WANG Shoujin, et al. Graph-based Chinese word sense disambiguation with multi-knowledge integration[J]. CMC-Computers, Materials & Continua, 2019, 61(1): 197–212. doi: 10.32604/cmc.2019.06068.
    [14]
    BLEVINS T and ZETTLEMOYER L. Moving down the long tail of word sense disambiguation with Gloss Informed Bi-encoders[C].The 58th Annual Meeting of the Association for Computational Linguistics, 2020: 1006–1017. doi: 10.18653/v1/2020.acl-main.95.
    [15]
    ANTUNES R and MATOS S. Supervised learning and knowledge-based approaches applied to biomedical word sense disambiguation[J]. Journal of Integrative Bioinformatics, 2017, 14(4): 20170051. doi: 10.1515/jib-2017-0051.
    [16]
    ALSAEEDAN W, MENAI M E B, and AL-AHMADI S. A hybrid genetic-ant colony optimization algorithm for the word sense disambiguation problem[J]. Information Sciences, 2017, 417: 20–38. doi: 10.1016/j.ins.2017.07.002.
    [17]
    SARROUTI M and OUATIK EL ALAOUI S. SemBioNLQA: A semantic biomedical question answering system for retrieving exact and ideal answers to natural language questions[J]. Artificial Intelligence in Medicine, 2020, 102: 101767. doi: 10.1016/j.artmed.2019.101767.
    [18]
    CHEN Fangyi, ZHANG Gongbo, CHEN Si, et al. Clinical note structural knowledge improves word sense disambiguation[J]. AMIA Joint Summits on Translational Science Proceedings, 2024, 2024: 515–524.
    [19]
    RAIS M and LACHKAR A. An empirical study of word sense disambiguation for biomedical information retrieval system[C]. The 6th International Conference on Bioinformatics and Biomedical Engineering, Granada, Spain, 2018: 314–326. doi: 10.1007/978-3-319-78723-7_27.
    [20]
    PASHUK A V, GURINOVICH A B, VOLOROVA N A, et al. Analysis of the methods of word sense disambiguation in the biomedical domain[J]. Doklady BGUIR, 2019, 75(5): 60–65. doi: 10.35596/1729-7648-2019-123-5-60-65.
    [21]
    YAP B P, KOH A, and CHNG E S. Adapting BERT for word sense disambiguation with gloss selection objective and example sentences[C]. Findings of the Association for Computational Linguistics, 2020: 41–46. doi: 10.18653/v1/2020.findings-emnlp.4.
    [22]
    KWON S, OH D, and KO Y. Word sense disambiguation based on context selection using knowledge-based word similarity[J]. Information Processing & Management, 2021, 58(4): 102551. doi: 10.1016/j.ipm.2021.102551.
    [23]
    ALMOUSA M, BENLAMRI R, and KHOURY R. A novel word sense disambiguation approach using WordNet knowledge graph[J]. Computer Speech & Language, 2022, 74: 101337. doi: 10.1016/j.csl.2021.101337.
    [24]
    CHEN Ting, KORNBLITH S, NOROUZI M, et al. A simple framework for contrastive learning of visual representations[C]. 37th International Conference on Machine Learning, 2020: 1597–1607.
    [25]
    CARON M, MISRA I, MAIRAL J, et al. Unsupervised learning of visual features by contrasting cluster assignments[C].The 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2020: 831.
    [26]
    GRILL J B, STRUB F, ALTCHÉ F, et al. Bootstrap your own latent a new approach to self-supervised learning[C]. Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2020: 1786.
    [27]
    WU Zhirong, XIONG Yuanjun, YU S X, et al. Unsupervised feature learning via non-parametric instance discrimination[C].2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 3733–3742. doi: 10.1109/CVPR.2018.00393.
    [28]
    HU Qianjiang, WANG Xiao, HU Wei, et al. AdCo: Adversarial contrast for efficient learning of unsupervised representations from self-trained negative adversaries[C]. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 2021: 1074–1083. doi: 10.1109/CVPR46437.2021.00113.
    [29]
    YEPES A J and BERLANGA R. Knowledge based word-concept model estimation and refinement for biomedical text mining[J]. Journal of Biomedical Informatics, 2015, 53: 300–307. doi: 10.1016/j.jbi.2014.11.015.
    [30]
    ZHANG Linfeng, CHEN Xin, ZHANG Junbo, et al. Contrastive deep supervision[C]. 17th European Conference on Computer Vision, Tel Aviv, Israel, 2022: 1–19. doi: 10.1007/978-3-031-19809-0_1.
    [31]
    SABBIR A K M, JIMENO-YEPES A, and KAVULURU R. Knowledge-based biomedical word sense disambiguation with neural concept embeddings[C]. 2017 IEEE 17th International Conference on Bioinformatics and Bioengineering, Washington, USA, 2017: 163–170. doi: 10.1109/BIBE.2017.00-61.
    [32]
    ZHANG Canlin, BIŚ D, LIU Xiuwen, et al. Biomedical word sense disambiguation with bidirectional long short-term memory and attention-based neural networks[J]. BMC Bioinformatics, 2019, 20(S16): 502. doi: 10.1186/s12859-019-3079-8.
    [33]
    CHUNG H W, HOU Le, LONGPRE S, et al. Scaling instruction-finetuned language models[J]. The Journal of Machine Learning Research, 2024, 25(1): 70.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(5)  / Tables(8)

    Article Metrics

    Article views (144) PDF downloads(14) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return