高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

互补衰减学习的图像分类网络

袁姮 田雯月 张晟翀

袁姮, 田雯月, 张晟翀. 互补衰减学习的图像分类网络[J]. 电子与信息学报. doi: 10.11999/JEIT260751
引用本文: 袁姮, 田雯月, 张晟翀. 互补衰减学习的图像分类网络[J]. 电子与信息学报. doi: 10.11999/JEIT260751
YUAN Heng, TIAN Wenyue, ZHANG Shengchong. Image Classification Network Based on Complementary Decay Learning[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260751
Citation: YUAN Heng, TIAN Wenyue, ZHANG Shengchong. Image Classification Network Based on Complementary Decay Learning[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260751

互补衰减学习的图像分类网络

doi: 10.11999/JEIT260751 cstr: 32379.14.JEIT260751
基金项目: 国家自然科学基金(61601213),辽宁省自然科学基金(20170540426),辽宁省教育厅重点基金(LJYL049)
详细信息
    作者简介:

    袁姮:女,博士,副教授,研究方向为图像与视觉信息计算、模式识别与人工智能,邮箱 lntuyuanheng@163.com

    田雯月:女,在读研究生,研究方向为图像与视觉信息计算、模式识别与人工智能,邮箱 15642274116@163.com

    张晟翀:男,高级工程师,研究方向为数字信号处理、模式识别与人工智能

    通讯作者:

    袁姮 lntuyuanheng@163.com

  • 中图分类号: TP391

Image Classification Network Based on Complementary Decay Learning

Funds: National Natural Science Foundation of China (61601213),, Natural Science Foundation of Liaoning Province (20170540426), Key Foundation of Education Department of Liaoning Province (LJYL049)
  • 摘要: 针对卷积神经网络在复杂场景中过度依赖正相关特征响应,导致特征关注由目标主体向局部区域偏移的问题,该文提出互补衰减学习的图像分类网络(CDLNet)。首先,提出空间互补衰减(SCD)模块,将特征划分为正响应、负响应与全局响应三类分支,并分别进行幅值衰减调节,使前景细节、背景结构与整体信息在空间上得到保留;然后,设计通道互补衰减(CCD)模块,通过对通道幅值施加指数衰减处理,降低少数高响应通道的主导作用,保留中低幅值通道信息;随后,设计互补衰减注意力(CDA)模块,并嵌入残差块中,实现空间结构与通道语义的协同调节。在CIFAR-10、CIFAR-100、SVHN、Imagenette、Imagewoof和ImageNet数据集上分别取得96.68%、81.81%、97.38%、92.45%、85.46%和62.00%的分类准确率。实验结果表明,CDLNet能够有效缓解高响应特征主导引起的特征关注漂移问题,增强低幅值与负响应信息的表达能力,提升图像分类能力。
  • 图  1  正负响应与边缘结构对应关系示意图

    Figure  1.  Schematic diagram of the correspondence between positive and negative responses and edge structures

    图  2  SCD模块结构图

    Figure  2.  SCD module structure diagram

    图  3  CCD模块结构图

    Figure  3.  Structure diagram of CCD module

    图  4  CDA模块结构图

    Figure  4.  Structure diagram of CDA module

    图  5  加权融合前后特征图像对比

    Figure  5.  Comparison of feature images before and after weighted fusion

    图  6  两种残差块

    Figure  6.  Two types of residual blocks

    图  7  CDLNet总体架构

    Figure  7.  Overall architecture of CDLNet

    图  8  SCD通道分配对分类准确率的影响

    Figure  8.  The impact of SCD channel allocation on classification

    图  9  CDA内部组合结构图

    Figure  9.  Internal combination structure diagram of CDA

    图  10  CDA内部组合方式对分类准确率的影响

    Figure  10.  The impact of CDA internal combination methods on classification accuracy

    图  11  ResNet-34和CDLNet在4个数据集上的混淆矩阵

    Figure  11.  Confusion matrices of ResNet-34 and CDLNet on four datasets

    图  12  注意力机制热力图对比

    Figure  12.  Comparison of heatmaps of attention mechanisms

    表  1  实验数据集

    Table  1.   Experimental datasets

    名称尺寸类别训练样本测试样本
    CIFAR-1032×32105000010000
    CIFAR-10032×321005000010000
    SVHN32×32107325726032
    Imagenette224×2241094693925
    Imagewoof224×2241090253929
    STL-1096×961050008000
    ImageNet32×321000128116750000
    下载: 导出CSV

    表  2  复杂度和训练时间对比

    Table  2.   Complexity and training time comparison

    网络分辨率Params(M)FLOPs(G)Time(h)
    ResNet-3432×3221.331.162.22
    CDLNet32×3232.921.944.54
    ResNet-34224×22421.293.681.41
    CDLNet224×22432.936.053.31
    下载: 导出CSV

    表  3  β值对CCD模块性能的影响

    Table  3.   Impact of β value on CCD module performance


    β
    分类准确率/%
    CIFAR-10CIFAR-100SVHNImagenetteImagewoof
    195.5180.1596.7592.4085.43
    295.5079.7197.1192.4085.28
    395.7679.9696.7392.0784.70
    495.5079.7996.9491.2684.85
    596.6881.8197.3892.4585.46
    695.7279.4396.8191.9785.23
    795.6979.7396.5692.3684.55
    895.8479.5196.7791.7785.05
    995.5579.7996.8392.4384.55
    下载: 导出CSV

    表  4  CDLNet的消融实验结果

    Table  4.   Ablation experiment results of CDLNet


    网络
    分类准确率/%
    CIFAR-10CIFAR-100SVHNImagenetteImagewoof
    CDLNet96.6881.8297.3892.4885.47
    Net193.8775.2196.3989.2683.29
    Net293.8774.2496.4289.6283.06
    Net387.2168.7491.60
    下载: 导出CSV

    表  5  不同网络在ImageNet数据集上的分类结果

    Table  5.   Classification results of different networks on ImageNet

    网络Params/MFLOPs/GImageNet/%
    ResNet-3421.331.1660.88
    WRN-28-10[17]37.105.2459.04
    GLPool[26]21.831.4957.84
    DSRNet[27]23.201.5458.22
    ff-EBMs[28]21.40.6546.00
    iGPT-L[29]1362154660.32
    BH-HRNet32[30]12.360.1656.00
    EM-Softmax[31]10.30.9449.22
    hEP[32]4.140.1636.50
    DP[33]43.990.3441.48
    CDLNet32.931.9462.88
    下载: 导出CSV

    表  6  ImageNet数据集上的分类结果

    Table  6.   Classification results on the ImageNet dataset

    网络Val-Acc%Top-1/%Top-5/%
    ResNet-3459.9760.8883.27
    ResNet-34+SCD60.7461.4683.86
    ResNet-34+CCD61.1861.9284.18
    CDLNet62.0062.8884.67
    下载: 导出CSV

    表  7  各网络在5个数据集上的分类准确率

    Table  7.   Classification accuracies of different networks on five datasets

    网络 CIFAR-10% CIFAR-100% SVHN% Imagenette% Imagewoof%
    ResNet-34[2] 88.10 72.35 91.74 86.75 77.82
    CAPR-DenseNet[3] 94.24 78.84 94.95 87.72 77.91
    EfficientNets[13] 94.01 75.96 93.32 88.01 77.93
    GhostNet[14] 94.92 77.15 93.86 87.83 78.22
    QKFormer[15] 96.18 80.26 97.13 88.32 81.65
    SSLLNet[16] 95.51 79.23 96.91 87.93 80.89
    WideResnet-28-10[17] 95.83 79.50 95.21 88.34 78.71
    TLENet[18] 95.46 78.42 96.83 87.62 80.57
    RepViT-M0.9[19] 95.88 80.21 97.02 90.91 83.47
    DCDENet[20] 96.66 80.08 96.35 91.87 84.52
    Couplformer[21] 93.54 73.92 94.26 87.91 77.89
    ATONet[22] 94.51 78.54 95.21 86.67 80.19
    FDPRNet[23] 95.70 80.01 96.96 90.83 83.41
    SCAM-Net[24] 96.28 80.40 97.25 91.62 84.18
    FasterNet-T2[25] 95.62 79.84 96.94 90.38 82.76
    CDLNet 96.68±0.08 81.81±0.09 97.38±0.04 92.45±0.10 85.46±0.11
    下载: 导出CSV
  • [1] KRIZHEVSKY A, SUTSKEVER I, and HINTON G E. ImageNet classification with deep convolutional neural networks[C]. NIPS, Lake Tahoe, USA, 2012: 1097–1105. doi: 10.1145/3065386. (查阅网上资料,未找到本条文献母体文献、出版地、页码和doi,请确认).
    [2] HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition[C]. 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 770–778. doi: 10.1109/CVPR.2016.90.
    [3] ZHANG Ke, GUO Yurong, WANG Xinsheng, et al. Channel-wise and feature-points reweights DenseNet for image classification[C]. 2019 IEEE International Conference on Image Processing, Taipei, China, 2019: 410–414. doi: 10.1109/ICIP.2019.8802982.
    [4] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: Transformers for image recognition at scale[C]. International Conference on Learning Representations, Vienna, Austria, 2021.
    [5] LIU Ze, LIN Yutong, CAO Yue, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 9992–10002. doi: 10.1109/ICCV48922.2021.00986.
    [6] HU Jie, SHEN Li, and SUN Gang. Squeeze-and-excitation networks[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7132–7141. doi: 10.1109/CVPR.2018.00745.
    [7] WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: Efficient channel attention for deep convolutional neural networks[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 11531–11539. doi: 10.1109/CVPR42600.2020.01155.
    [8] 郭礼华, 王广飞. 基于任务感知关系网络的少样本图像分类[J]. 电子与信息学报, 2024, 46(3): 977–985. doi: 10.11999/JEIT230162.

    GUO Lihua and WANG Guangfei. Few-shot image classification based on task-aware relation network[J]. Journal of Electronics & Information Technology, 2024, 46(3): 977–985. doi: 10.11999/JEIT230162.
    [9] 赵凤, 耿苗苗, 刘汉强, 等. 卷积神经网络与视觉Transformer联合驱动的跨层多尺度融合网络高光谱图像分类方法[J]. 电子与信息学报, 2024, 46(5): 2237–2248. doi: 10.11999/JEIT231209.

    ZHAO Feng, GENG Miaomiao, LIU Hanqiang, et al. Convolutional neural network and vision transformer-driven cross-layer multi-scale fusion network for hyperspectral image classification[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2237–2248. doi: 10.11999/JEIT231209.
    [10] 文泓力, 胡庆浩, 黄立威, 等. 基于参数高效ViT与多模态导引的遥感图像小样本分类方法[J]. 电子与信息学报, 2025, 47(12): 4689–4703. doi: 10.11999/JEIT250996.

    WEN Hongli, HU Qinghao, HUANG Liwei, et al. Few-shot remote sensing image classification based on parameter-efficient vision transformer and multimodal guidance[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4689–4703. doi: 10.11999/JEIT250996.
    [11] GORIS R L T, COEN-CAGLI R, MILLER K D, et al. Response sub-additivity and variability quenching in visual cortex[J]. Nature Reviews Neuroscience, 2024, 25(4): 237–252. doi: 10.1038/s41583-024-00795-0.
    [12] JIANG Wentao, YUAN Heng, and LIU Wanjun. Neuron signal attenuation activation mechanism for deep learning[J]. Patterns, 2025, 6(1): 101117. doi: 10.1016/j.patter.2024.101117.
    [13] TAN Mingxing and LE Q V. EfficientNet: Rethinking model scaling for convolutional neural networks[C]. Proceedings of the 36th International Conference on Machine Learning, Long Beach, USA, 2019: 6105–6114.
    [14] HAN Kai, WANG Yunhe, TIAN Qi, et al. GhostNet: More features from cheap operations[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 1577–1586. doi: 10.1109/CVPR42600.2020.00165.
    [15] ZHOU Chenlin, ZHANG Han, ZHOU Zhaokun, et al. QKFormer: Hierarchical spiking transformer using Q-K attention[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 416.
    [16] MA Chenxiang, WU Jibin, SI Chenyang, et al. Scaling supervised local learning with augmented auxiliary networks[C]. The 12th International Conference on Learning Representations, Vienna, Austria, 2024.
    [17] CHRABASZCZ P, LOSHCHILOV I, and HUTTER F. A downsampled variant of ImageNet as an alternative to the CIFAR datasets[EB/OL]. https://arxiv.org/abs/1707.08819, 2017.
    [18] SHIN H and CHOI D W. Teacher as a lenient expert: Teacher-agnostic data-free knowledge distillation[C]. Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, 2024: 14991–14999. doi: 10.1609/aaai.v38i13.29420.
    [19] WANG Ao, CHEN Hui, LIN Zijia, et al. Rep ViT: Revisiting mobile CNN from ViT perspective[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 15909–15920. doi: 10.1109/CVPR52733.2024.01506.
    [20] 袁姮, 范桐桐, 高原. 双频通道差异增强的图像分类网络[J]. 计算机工程与应用, 2026, 62(11): 259–271. doi: 10.3778/j.issn.1002-8331.2503-0175.

    YUAN Heng, FAN Tongtong, and GAO Yuan. Dual-frequency channel difference enhancement for image classification[J]. Computer Engineering and Applications, 2026, 62(11): 259–271. doi: 10.3778/j.issn.1002-8331.2503-0175.
    [21] LAN Hai, WANG Xihao, SHEN Hao, et al. Couplformer: Rethinking vision transformer with coupling attention[C]. 2023 IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, 2023: 6464–6473. doi: 10.1109/WACV56688.2023.00641.
    [22] WU Xidong, GAO Shangqian, ZHANG Zeyu, et al. Auto-train-once: Controller network guided automatic network pruning from scratch[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 16163–16173. doi: 10.1109/CVPR52733.2024.01530.
    [23] 袁姮, 刘杰, 姜文涛, 等. 特征重排列注意力机制的双池化残差分类网络[J]. 中国图象图形学报, 2025, 30(1): 110–129. doi: 10.11834/jig.240061.

    YUAN Heng, LIU Jie, JIANG Wentao, et al. Double-pooling residual classification network based on feature reordering attention mechanism[J]. Journal of Image and Graphics, 2025, 30(1): 110–129. doi: 10.11834/jig.240061.
    [24] 姜文涛, 王鑫杰, 张晟翀. 空间约束注意力机制的图像分类网络[J]. 智能系统学报, 2025, 20(6): 1444–1460. doi: 10.11992/tis.202505025.

    JIANG Wentao, WANG Xinjie, and ZHANG Shengchong. Spatially constrained attention mechanism for image classification network[J]. CAAI Transactions on Intelligent Systems, 2025, 20(6): 1444–1460. doi: 10.11992/tis.202505025.
    [25] CHEN Jierun, KAO S H, HE Hao, et al. Run, don't walk: Chasing higher FLOPS for faster neural networks[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 12021–12031. doi: 10.1109/CVPR52729.2023.01157.
    [26] ZHANG Xingpeng and ZHANG Xiaohong. Global learnable pooling with enhancing distinctive feature for image classification[J]. IEEE Access, 2020, 8: 98539–98547. doi: 10.1109/ACCESS.2020.2997078.
    [27] ZHANG Xingpeng and ZHANG Xiaohong. Feature recalibration in deep learning via depthwise squeeze and refinement operations[J]. IEEE Access, 2020, 8: 79046–79055. doi: 10.1109/ACCESS.2020.2990658.
    [28] NEST T and ERNOULT M. Towards training digitally-tied analog blocks via hybrid gradient computation[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 2666.
    [29] CHEN M, RADFORD A, CHILD R, et al. Generative pretraining from pixels[C]. Proceedings of the 37th International Conference on Machine Learning, 2020: 158. (查阅网上资料, 未找到本条文献出版地, 请确认).
    [30] NING Jinlai, GUAN Haoyan, and SPRATLING M W. Rethinking the backbone architecture for tiny object detection[C]. Proceedings of the 18th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Lisbon, Portugal, 2023: 103–114.
    [31] WANG Xiaobo, ZHANG Shifeng, LEI Zhen, et al. Ensemble soft-margin softmax loss for image classification[C]. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden, 2018: 992–998.
    [32] LABORIEUX A and ZENKE F. Holomorphic equilibrium propagation computes exact gradients through finite size oscillations[C]. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, USA, 2022: 941.
    [33] HØIER R, STAUDT D, and ZACH C. Dual propagation: Accelerating contrastive Hebbian learning with dyadic neurons[C]. Proceedings of the 40th International Conference on Machine Learning, Honolulu, USA, 2023: 533.
    [34] WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1.
    [35] ZHANG H, ZHANG S, WANG X, et al. Polarized self-attention: Towards high-quality pixel-wise regression[C]. IEEE Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2021: 123–132. (查阅网上资料, 未找到本条文献信息, 请确认).
  • 加载中
图(12) / 表(7)
计量
  • 文章访问数:  7
  • HTML全文浏览量:  0
  • PDF下载量:  1
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-04-10
  • 修回日期:  2026-08-26
  • 录用日期:  2026-08-26
  • 网络出版日期:  2026-09-02

目录

    /

    返回文章
    返回