Advanced Search
Turn off MathJax
Article Contents
HU Tianwei, ZHANG Xiangrui, CHEN Jian, DUAN Haodong, JIA Jie. A Structure-Preserving Semantic Transmission Method for Low-Bandwidth Networks[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260525
Citation: HU Tianwei, ZHANG Xiangrui, CHEN Jian, DUAN Haodong, JIA Jie. A Structure-Preserving Semantic Transmission Method for Low-Bandwidth Networks[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260525

A Structure-Preserving Semantic Transmission Method for Low-Bandwidth Networks

doi: 10.11999/JEIT260525 cstr: 32379.14.JEIT260525
Funds:  The National Natural Science Foundation of China (No.62132004, No.62572109)
  • Accepted Date: 2026-08-17
  • Rev Recd Date: 2026-08-17
  • Available Online: 2026-08-25
  •   Objective  Low-bandwidth visual transmission is essential for edge-intelligence applications such as disaster inspection, underwater exploration, and remote assistance. These scenarios require visual communication systems to simultaneously achieve low bitrate, high structural fidelity, stable color reconstruction, and low end-to-end latency. However, conventional image coding methods suffer from severe quality degradation at extremely low bitrates due to the digital-cliff effect and compression artifacts. Although semantic communication provides a promising solution by transmitting task-relevant representations rather than pixel-level information, existing approaches still face challenges in balancing bitrate efficiency, reconstruction fidelity, and decoding complexity. Therefore, a structure-preserving semantic transmission method, termed Edge-Link, is proposed for low-bandwidth networks.  Methods  Edge-Link adopts a multimodal decoupled representation framework that separates image information into three complementary streams: a semantic stream, a structural stream, and a low-frequency appearance stream. The semantic stream is extracted using a frozen CLIP encoder to provide global semantic guidance, while the structural stream is obtained from Canny edge information to explicitly preserve object boundaries. A low-resolution color map is introduced as the appearance stream to maintain global color distribution and illumination characteristics with minimal transmission overhead. Furthermore, a FastSAM-based semantic gateway is developed to distinguish foreground objects from background regions, and an object-aware bitrate allocation strategy is designed to prioritize important semantic regions under bandwidth constraints. At the receiver, a deterministic dual-stream generation network based on SPADE is proposed, where structural information provides spatial constraints and semantic features guide the reconstruction process, avoiding the high latency caused by iterative diffusion sampling. A real LoRa communication prototype based on GNU Radio and USRP is also implemented to validate transmission feasibility and robustness under practical wireless conditions.  Results and Discussions  Experiments are conducted on Set14 and DIV2K to evaluate transmission efficiency, reconstruction quality, perceptual naturalness, and latency. The proposed method achieves an average payload of 10.96 KB for 512×512 images, corresponding to about 0.31 bpp, while keeping the end-to-end inference latency below 60 ms; under a LoRa narrowband link of about 8 kbps, the airtime of a single frame is about 11 s, verifying its feasibility in bandwidth-limited environments (Table 1). The ablation study shows that the object-aware coding strategy improves the bitrate-quality trade-off by reducing the average payload on DIV2K from 126.09 KB to 93.55 KB while preserving better foreground reconstruction quality (Table 2). Compared with JPEG, the proposed method also shows better structural preservation and perceptual quality at low bitrates; for example, on image 0843 with a payload of 15.54 KB, the global SSIM is improved from 0.783 to 0.963, and the subject-region LPIPS is reduced from 0.561 to 0.307 (Table 4). Under the same bitrate constraint, the proposed method further outperforms VQGAN and ControlNet+Canny, achieving PSNR, SSIM, LPIPS, and NIQE values of 23.738, 0.622, 0.223, and 4.570, respectively, indicating a better balance between fidelity and perceptual quality in low-bandwidth semantic reconstruction (Table 5).  Conclusions  A structure-preserving semantic transmission framework for low-bandwidth networks is presented. By combining multimodal decoupled representation, object-aware bitrate allocation, and deterministic dual-stream reconstruction, the framework balances transmission efficiency, structural fidelity, perceptual quality, and decoding latency. The reported results show that the proposed method is well suited to high-reliability edge visual communication scenarios in which accurate contours, stable color appearance, and efficient inference are simultaneously required. The real-link validation on a LoRa prototype further suggests its practical potential for bandwidth-constrained edge networks.
  • loading
  • [1]
    JERNBERG C, SANDIN J, ZIEMKE T, et al. The effect of latency, speed and task on remote operation of vehicles[J]. Transportation Research Interdisciplinary Perspectives, 2024, 26: 101152. doi: 10.1016/j.trip.2024.101152.
    [2]
    BOURTSOULATZE E, KURKA D B, and GÜNDÜZ D. Deep joint source-channel coding for wireless image transmission[J]. IEEE Transactions on Cognitive Communications and Networking, 2019, 5(3): 567–579. doi: 10.1109/TCCN.2019.2919300.
    [3]
    DONG Chao, DENG Yubin, LOY C C, et al. Compression artifacts reduction by a deep convolutional network[C]. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2015: 576–584. doi: 10.1109/ICCV.2015.73.
    [4]
    LIU Yating, WANG Xiaojie, NING Zhaolong, et al. A survey on semantic communications: Technologies, solutions, applications and challenges[J]. Digital Communications and Networks, 2024, 10(3): 528–545. doi: 10.1016/j.dcan.2023.05.010.
    [5]
    CHAI Jingxuan, XIAO Yong, and SHI Guangming. On the rate-distortion-complexity tradeoff for semantic communication[J]. IEEE Internet of Things Journal, 2026, 13(14): 31768–31781. doi: 10.1109/JIOT.2026.3689652.
    [6]
    陈阳, 马欢, 姬智, 等. 面向图像恢复任务的语义通信网络能耗优化[J]. 电子与信息学报, 2026, 48(1): 183–190. doi: 10.11999/JEIT250915.

    CHEN Yang, MA Huan, JI Zhi, et al. Optimization of energy consumption in semantic communication networks for image recovery tasks[J]. Journal of Electronics & Information Technology, 2026, 48(1): 183–190. doi: 10.11999/JEIT250915.
    [7]
    AITHAL S K, MAINI P, LIPTON Z, et al. Understanding hallucinations in diffusion models through mode interpolation[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 134614–134644. doi: 10.52202/079017-4278.
    [8]
    JIA Zhaoyang, LI Jiahao, LI Bin, et al. Generative latent coding for ultra-low bitrate image compression[C]. Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, 2024: 26088–26098. doi: 10.1109/CVPR52733.2024.02465.
    [9]
    PEZONE F, MUSA O, CAIRE G, et al. Semantic-preserving image coding based on conditional diffusion models[C]. ICASSP 2024–2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Korea, 2024: 13501–13505. doi: 10.1109/ICASSP48485.2024.10447279.
    [10]
    ZHANG Chaoning, CHO J, PUSPITASARI F D, et al. A survey on Segment Anything Model (SAM): Vision foundation model meets prompt engineering[EB/OL]. arXiv preprint arXiv: 2306.06211. https://arxiv.org/abs/2306.06211, 2023.
    [11]
    KINGMA D P and WELLING M. Auto-encoding variational Bayes[EB/OL]. arXiv preprint arXiv: 1312.6114. https://arxiv.org/abs/1312.6114, 2013.
    [12]
    BURGESS C P, HIGGINS I, PAL A, et al. Understanding disentangling in β-VAE[EB/OL]. arXiv preprint arXiv: 1804.03599. https://arxiv.org/abs/1804.03599?context=cs.LG#1, 2018.
    [13]
    罗一畅, 齐析屿, 张博锐, 等. 分割一切模型的轻量化研究综述[J]. 电子与信息学报, 2026, 48(2): 713–731. doi: 10.11999/JEIT250894.

    LUO Yichang, QI Xiyu, ZHANG Borui, et al. A survey of lightweight techniques for segment anything model[J]. Journal of Electronics & Information Technology, 2026, 48(2): 713–731. doi: 10.11999/JEIT250894.
    [14]
    TAPPAREL J, AFISIADIS O, MAYORAZ P, et al. An open-source LoRa physical layer prototype on GNU radio[C]. Proceedings of the 2020 IEEE 21st International Workshop on Signal Processing Advances in Wireless Communications (SPAWC), Atlanta, USA, 2020: 1–5. doi: 10.1109/SPAWC48557.2020.9154273.
    [15]
    SUN Xiaorui, LIU Jun, SHEN Hengtao, et al. On efficient variants of Segment Anything Model: A survey[J]. International Journal of Computer Vision, 2025, 133(10): 7406–7436. doi: 10.1007/s11263-025-02539-8.
    [16]
    RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]. Proceedings of the 38th International Conference on Machine Learning, 2021: 8748–8763. (查阅网上资料, 未找见本条文献出版地信息, 请确认).
    [17]
    PARK T, LIU Mingyu, WANG Tingchun, et al. Semantic image synthesis with spatially-adaptive normalization[C]. Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, USA, 2019: 2332–2341. doi: 10.1109/CVPR.2019.00244.
    [18]
    BLAU Y and MICHAELI T. The perception-distortion tradeoff[C]. Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 6228–6237. doi: 10.1109/CVPR.2018.00652.
    [19]
    JIANG Wei, ZHAI Yongqi, LI Hangyu, et al. Learned image compression with ROI-weighted distortion and bit allocation[EB/OL]. CoRR. https://arxiv.org/html/2401.08154v2, 2024.
    [20]
    MAO Qi, YANG Tinghan, ZHANG Yinuo, et al. Extreme image compression using fine-tuned VQGANs[C]. Proceedings of the 2024 Data Compression Conference (DCC), Snowbird, USA, 2024: 203–212. doi: 10.1109/DCC58796.2024.00028.
    [21]
    CHEN Weilong, XU Wenxuan, CHEN Haoran, et al. Semantic communication based on large language model for underwater image transmission[J]. IEEE Transactions on Mobile Computing, 2026, 25(2): 2060–2075. doi: 10.1109/TMC.2025.3607717.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(10)  / Tables(7)

    Article Metrics

    Article views (79) PDF downloads(9) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return