Image Classification Network Based on Complementary Decay Learning
-
摘要: 针对卷积神经网络在复杂场景中过度依赖正相关特征响应,导致特征关注由目标主体向局部区域偏移的问题,该文提出互补衰减学习的图像分类网络(CDLNet)。首先,提出空间互补衰减(SCD)模块,将特征划分为正响应、负响应与全局响应三类分支,并分别进行幅值衰减调节,使前景细节、背景结构与整体信息在空间上得到保留;然后,设计通道互补衰减(CCD)模块,通过对通道幅值施加指数衰减处理,降低少数高响应通道的主导作用,保留中低幅值通道信息;随后,设计互补衰减注意力(CDA)模块,并嵌入残差块中,实现空间结构与通道语义的协同调节。在CIFAR-10、CIFAR-100、SVHN、Imagenette、Imagewoof和ImageNet数据集上分别取得96.68%、81.81%、97.38%、92.45%、85.46%和62.00%的分类准确率。实验结果表明,CDLNet能够有效缓解高响应特征主导引起的特征关注漂移问题,增强低幅值与负响应信息的表达能力,提升图像分类能力。Abstract:
Objective Image classification depends on complete and discriminative feature representations. Existing convolutional neural networks usually enhance positive and high-amplitude responses through activation functions and attention mechanisms. However, excessive reliance on dominant positive responses may shift attention from the whole object to local salient regions, while negative and low-amplitude responses containing edge, texture, and foreground-background transition information are often weakened. To address this problem, a Complementary Decay Learning Network (CDLNet) is proposed to suppress high-response dominance and preserve complementary feature information. Methods Inspired by the signal attenuation mechanism of the biological visual system, CDLNet introduces complementary decay learning into a residual network. The Spatial Complementary Decay (SCD) module divides features into positive-response, negative-response, and global-response branches, and applies differentiated decay to preserve salient regions, boundary details, and contextual information ( Fig.2 ). The Channel Complementary Decay (CCD) module attenuates high-response channels while retaining middle- and low-response channels, thereby reducing channel dominance and promoting cooperative channel representation (Fig.3 ). The Complementary Decay Attention (CDA) module integrates SCD and CCD in parallel and is embedded into ResNet residual blocks to jointly regulate spatial structures and channel semantics (Fig.4 ,Fig.6 ).Results and Discussions Experiments are conducted on CIFAR-10, CIFAR-100, SVHN, Imagenette, Imagewoof, and ImageNet datasets. CDLNet achieves classification accuracies of 96.68%, 81.81%, 97.38%, 92.45%, 85.46%, and 62.88%, respectively. Compared with ResNet-34 and representative classification networks, CDLNet obtains higher accuracy on multiple datasets ( Table 5 ,Table 6 ). Ablation experiments demonstrate that removing either SCD or CCD reduces classification accuracy, indicating that spatial and channel complementary decay both contribute to feature regulation (Table 4 ,Table 5 ). Visualization results show that CDLNet can enhance target regions, preserve structural details, and suppress irrelevant background responses (Fig.1 ,Fig.12 ). Although CDA increases parameters and computation, the accuracy improvement shows a reasonable balance between performance and complexity (Table 3 ).Conclusions CDLNet introduces spatial and channel complementary decay mechanisms into residual networks. By suppressing excessive high-response dominance and preserving negative-response and low-amplitude information, the proposed method improves the completeness and balance of feature representations, alleviates attention drift, and enhances image classification performance. Future work will further optimize the decay strategy and model complexity to improve efficiency and generalization. -
表 1 实验数据集
Table 1. Experimental datasets
名称 尺寸 类别 训练样本 测试样本 CIFAR-10 32×32 10 50000 10000 CIFAR-100 32×32 100 50000 10000 SVHN 32×32 10 73257 26032 Imagenette 224×224 10 9469 3925 Imagewoof 224×224 10 9025 3929 STL-10 96×96 10 5000 8000 ImageNet 32×32 1000 1281167 50000 表 2 复杂度和训练时间对比
Table 2. Complexity and training time comparison
网络 分辨率 Params(M) FLOPs(G) Time(h) ResNet-34 32×32 21.33 1.16 2.22 CDLNet 32×32 32.92 1.94 4.54 ResNet-34 224×224 21.29 3.68 1.41 CDLNet 224×224 32.93 6.05 3.31 表 3 β值对CCD模块性能的影响
Table 3. Impact of β value on CCD module performance
β分类准确率/% CIFAR-10 CIFAR-100 SVHN Imagenette Imagewoof 1 95.51 80.15 96.75 92.40 85.43 2 95.50 79.71 97.11 92.40 85.28 3 95.76 79.96 96.73 92.07 84.70 4 95.50 79.79 96.94 91.26 84.85 5 96.68 81.81 97.38 92.45 85.46 6 95.72 79.43 96.81 91.97 85.23 7 95.69 79.73 96.56 92.36 84.55 8 95.84 79.51 96.77 91.77 85.05 9 95.55 79.79 96.83 92.43 84.55 表 4 CDLNet的消融实验结果
Table 4. Ablation experiment results of CDLNet
网络分类准确率/% CIFAR-10 CIFAR-100 SVHN Imagenette Imagewoof CDLNet 96.68 81.82 97.38 92.48 85.47 Net1 93.87 75.21 96.39 89.26 83.29 Net2 93.87 74.24 96.42 89.62 83.06 Net3 87.21 68.74 91.60 ─ ─ 表 5 不同网络在ImageNet数据集上的分类结果
Table 5. Classification results of different networks on ImageNet
表 6 ImageNet数据集上的分类结果
Table 6. Classification results on the ImageNet dataset
网络 Val-Acc% Top-1/% Top-5/% ResNet-34 59.97 60.88 83.27 ResNet-34+SCD 60.74 61.46 83.86 ResNet-34+CCD 61.18 61.92 84.18 CDLNet 62.00 62.88 84.67 表 7 各网络在5个数据集上的分类准确率
Table 7. Classification accuracies of different networks on five datasets
网络 CIFAR-10% CIFAR-100% SVHN% Imagenette% Imagewoof% ResNet-34[2] 88.10 72.35 91.74 86.75 77.82 CAPR-DenseNet[3] 94.24 78.84 94.95 87.72 77.91 EfficientNets[13] 94.01 75.96 93.32 88.01 77.93 GhostNet[14] 94.92 77.15 93.86 87.83 78.22 QKFormer[15] 96.18 80.26 97.13 88.32 81.65 SSLLNet[16] 95.51 79.23 96.91 87.93 80.89 WideResnet-28-10[17] 95.83 79.50 95.21 88.34 78.71 TLENet[18] 95.46 78.42 96.83 87.62 80.57 RepViT-M0.9[19] 95.88 80.21 97.02 90.91 83.47 DCDENet[20] 96.66 80.08 96.35 91.87 84.52 Couplformer[21] 93.54 73.92 94.26 87.91 77.89 ATONet[22] 94.51 78.54 95.21 86.67 80.19 FDPRNet[23] 95.70 80.01 96.96 90.83 83.41 SCAM-Net[24] 96.28 80.40 97.25 91.62 84.18 FasterNet-T2[25] 95.62 79.84 96.94 90.38 82.76 CDLNet 96.68±0.08 81.81±0.09 97.38±0.04 92.45±0.10 85.46±0.11 -
[1] KRIZHEVSKY A, SUTSKEVER I, and HINTON G E. ImageNet classification with deep convolutional neural networks[C]. NIPS, Lake Tahoe, USA, 2012: 1097–1105. doi: 10.1145/3065386. (查阅网上资料,未找到本条文献母体文献、出版地、页码和doi,请确认). [2] HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition[C]. 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 2016: 770–778. doi: 10.1109/CVPR.2016.90. [3] ZHANG Ke, GUO Yurong, WANG Xinsheng, et al. Channel-wise and feature-points reweights DenseNet for image classification[C]. 2019 IEEE International Conference on Image Processing, Taipei, China, 2019: 410–414. doi: 10.1109/ICIP.2019.8802982. [4] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: Transformers for image recognition at scale[C]. International Conference on Learning Representations, Vienna, Austria, 2021. [5] LIU Ze, LIN Yutong, CAO Yue, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C]. 2021 IEEE/CVF International Conference on Computer Vision, Montreal, Canada, 2021: 9992–10002. doi: 10.1109/ICCV48922.2021.00986. [6] HU Jie, SHEN Li, and SUN Gang. Squeeze-and-excitation networks[C]. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, 2018: 7132–7141. doi: 10.1109/CVPR.2018.00745. [7] WANG Qilong, WU Banggu, ZHU Pengfei, et al. ECA-Net: Efficient channel attention for deep convolutional neural networks[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 11531–11539. doi: 10.1109/CVPR42600.2020.01155. [8] 郭礼华, 王广飞. 基于任务感知关系网络的少样本图像分类[J]. 电子与信息学报, 2024, 46(3): 977–985. doi: 10.11999/JEIT230162.GUO Lihua and WANG Guangfei. Few-shot image classification based on task-aware relation network[J]. Journal of Electronics & Information Technology, 2024, 46(3): 977–985. doi: 10.11999/JEIT230162. [9] 赵凤, 耿苗苗, 刘汉强, 等. 卷积神经网络与视觉Transformer联合驱动的跨层多尺度融合网络高光谱图像分类方法[J]. 电子与信息学报, 2024, 46(5): 2237–2248. doi: 10.11999/JEIT231209.ZHAO Feng, GENG Miaomiao, LIU Hanqiang, et al. Convolutional neural network and vision transformer-driven cross-layer multi-scale fusion network for hyperspectral image classification[J]. Journal of Electronics & Information Technology, 2024, 46(5): 2237–2248. doi: 10.11999/JEIT231209. [10] 文泓力, 胡庆浩, 黄立威, 等. 基于参数高效ViT与多模态导引的遥感图像小样本分类方法[J]. 电子与信息学报, 2025, 47(12): 4689–4703. doi: 10.11999/JEIT250996.WEN Hongli, HU Qinghao, HUANG Liwei, et al. Few-shot remote sensing image classification based on parameter-efficient vision transformer and multimodal guidance[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4689–4703. doi: 10.11999/JEIT250996. [11] GORIS R L T, COEN-CAGLI R, MILLER K D, et al. Response sub-additivity and variability quenching in visual cortex[J]. Nature Reviews Neuroscience, 2024, 25(4): 237–252. doi: 10.1038/s41583-024-00795-0. [12] JIANG Wentao, YUAN Heng, and LIU Wanjun. Neuron signal attenuation activation mechanism for deep learning[J]. Patterns, 2025, 6(1): 101117. doi: 10.1016/j.patter.2024.101117. [13] TAN Mingxing and LE Q V. EfficientNet: Rethinking model scaling for convolutional neural networks[C]. Proceedings of the 36th International Conference on Machine Learning, Long Beach, USA, 2019: 6105–6114. [14] HAN Kai, WANG Yunhe, TIAN Qi, et al. GhostNet: More features from cheap operations[C]. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 1577–1586. doi: 10.1109/CVPR42600.2020.00165. [15] ZHOU Chenlin, ZHANG Han, ZHOU Zhaokun, et al. QKFormer: Hierarchical spiking transformer using Q-K attention[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 416. [16] MA Chenxiang, WU Jibin, SI Chenyang, et al. Scaling supervised local learning with augmented auxiliary networks[C]. The 12th International Conference on Learning Representations, Vienna, Austria, 2024. [17] CHRABASZCZ P, LOSHCHILOV I, and HUTTER F. A downsampled variant of ImageNet as an alternative to the CIFAR datasets[EB/OL]. https://arxiv.org/abs/1707.08819, 2017. [18] SHIN H and CHOI D W. Teacher as a lenient expert: Teacher-agnostic data-free knowledge distillation[C]. Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, 2024: 14991–14999. doi: 10.1609/aaai.v38i13.29420. [19] WANG Ao, CHEN Hui, LIN Zijia, et al. Rep ViT: Revisiting mobile CNN from ViT perspective[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 15909–15920. doi: 10.1109/CVPR52733.2024.01506. [20] 袁姮, 范桐桐, 高原. 双频通道差异增强的图像分类网络[J]. 计算机工程与应用, 2026, 62(11): 259–271. doi: 10.3778/j.issn.1002-8331.2503-0175.YUAN Heng, FAN Tongtong, and GAO Yuan. Dual-frequency channel difference enhancement for image classification[J]. Computer Engineering and Applications, 2026, 62(11): 259–271. doi: 10.3778/j.issn.1002-8331.2503-0175. [21] LAN Hai, WANG Xihao, SHEN Hao, et al. Couplformer: Rethinking vision transformer with coupling attention[C]. 2023 IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, 2023: 6464–6473. doi: 10.1109/WACV56688.2023.00641. [22] WU Xidong, GAO Shangqian, ZHANG Zeyu, et al. Auto-train-once: Controller network guided automatic network pruning from scratch[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 16163–16173. doi: 10.1109/CVPR52733.2024.01530. [23] 袁姮, 刘杰, 姜文涛, 等. 特征重排列注意力机制的双池化残差分类网络[J]. 中国图象图形学报, 2025, 30(1): 110–129. doi: 10.11834/jig.240061.YUAN Heng, LIU Jie, JIANG Wentao, et al. Double-pooling residual classification network based on feature reordering attention mechanism[J]. Journal of Image and Graphics, 2025, 30(1): 110–129. doi: 10.11834/jig.240061. [24] 姜文涛, 王鑫杰, 张晟翀. 空间约束注意力机制的图像分类网络[J]. 智能系统学报, 2025, 20(6): 1444–1460. doi: 10.11992/tis.202505025.JIANG Wentao, WANG Xinjie, and ZHANG Shengchong. Spatially constrained attention mechanism for image classification network[J]. CAAI Transactions on Intelligent Systems, 2025, 20(6): 1444–1460. doi: 10.11992/tis.202505025. [25] CHEN Jierun, KAO S H, HE Hao, et al. Run, don't walk: Chasing higher FLOPS for faster neural networks[C]. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 2023: 12021–12031. doi: 10.1109/CVPR52729.2023.01157. [26] ZHANG Xingpeng and ZHANG Xiaohong. Global learnable pooling with enhancing distinctive feature for image classification[J]. IEEE Access, 2020, 8: 98539–98547. doi: 10.1109/ACCESS.2020.2997078. [27] ZHANG Xingpeng and ZHANG Xiaohong. Feature recalibration in deep learning via depthwise squeeze and refinement operations[J]. IEEE Access, 2020, 8: 79046–79055. doi: 10.1109/ACCESS.2020.2990658. [28] NEST T and ERNOULT M. Towards training digitally-tied analog blocks via hybrid gradient computation[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 2666. [29] CHEN M, RADFORD A, CHILD R, et al. Generative pretraining from pixels[C]. Proceedings of the 37th International Conference on Machine Learning, 2020: 158. (查阅网上资料, 未找到本条文献出版地, 请确认). [30] NING Jinlai, GUAN Haoyan, and SPRATLING M W. Rethinking the backbone architecture for tiny object detection[C]. Proceedings of the 18th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Lisbon, Portugal, 2023: 103–114. [31] WANG Xiaobo, ZHANG Shifeng, LEI Zhen, et al. Ensemble soft-margin softmax loss for image classification[C]. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden, 2018: 992–998. [32] LABORIEUX A and ZENKE F. Holomorphic equilibrium propagation computes exact gradients through finite size oscillations[C]. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, USA, 2022: 941. [33] HØIER R, STAUDT D, and ZACH C. Dual propagation: Accelerating contrastive Hebbian learning with dyadic neurons[C]. Proceedings of the 40th International Conference on Machine Learning, Honolulu, USA, 2023: 533. [34] WOO S, PARK J, LEE J Y, et al. CBAM: Convolutional block attention module[C]. 15th European Conference on Computer Vision, Munich, Germany, 2018: 3–19. doi: 10.1007/978-3-030-01234-2_1. [35] ZHANG H, ZHANG S, WANG X, et al. Polarized self-attention: Towards high-quality pixel-wise regression[C]. IEEE Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2021: 123–132. (查阅网上资料, 未找到本条文献信息, 请确认). -
下载: