Advanced Search
Turn off MathJax
Article Contents
NAN Longmei, WANG Haoyu, DU Yiran, LI Wei, CHEN Tao. A Reconfigurable Parallelized Coprocessor Design for the RISC-V-based Grain Cryptographic Algorithm[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260391
Citation: NAN Longmei, WANG Haoyu, DU Yiran, LI Wei, CHEN Tao. A Reconfigurable Parallelized Coprocessor Design for the RISC-V-based Grain Cryptographic Algorithm[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260391

A Reconfigurable Parallelized Coprocessor Design for the RISC-V-based Grain Cryptographic Algorithm

doi: 10.11999/JEIT260391 cstr: 32379.14.JEIT260391
  • Received Date: 2026-04-02
  • Accepted Date: 2026-07-28
  • Rev Recd Date: 2026-07-23
  • Available Online: 2026-08-07
  •   Objective  To address the performance limitations of Grain cryptographic algorithms on General-Purpose Processors (GPPs), as well as the inflexibility and high hardware overhead of Application-Specific Integrated Circuit (ASIC) implementations, a dedicated cryptographic hardware accelerator is integrated into an RISC-V coprocessor through a custom instruction extension mechanism. A reconfigurable parallelized architecture is proposed for the Grain algorithm family based on the RISC-V coprocessor interface. Corresponding custom instructions are designed to support the flexible and efficient execution of Grain-80, Grain-128, Grain-128a, and Grain-128AEAD on a unified hardware platform. The proposed architecture provides a favorable balance between processing efficiency, design flexibility, and hardware resource overhead, making it suitable for resource-constrained embedded systems.  Methods  A unified shift-register architecture is adopted to support flexible switching among Grain-80, Grain-128, Grain-128a, and Grain-128AEAD, with a configurable parallelization degree of 1 to 8. To implement the custom instructions for the Grain cryptographic algorithms, the software and hardware functions are analyzed, and the cryptographic process is divided between the processor and coprocessor to achieve efficient execution. The proposed coprocessor uses a streamlined architecture that reuses common hardware resources across the four algorithms. Combined with a configurable feedback tap selection network and feedback logic tailored to a predefined set of algorithms, the architecture enables reconfigurable parallel execution with limited additional hardware resource overhead.  Results and Discussions  The proposed custom instructions enable flexible implementation of four Grain cryptographic algorithms on the same hardware platform. Compared with implementations without instruction extensions, the proposed custom instructions reduce the number of clock cycles and instructions required for cryptographic processing while improving throughput (Table 5). On the Hummingbird E203 platform with a parallelization degree of 4, the proposed implementation requires 105, 129, and 183 clock cycles for Grain-80, Grain-128/Grain-128a, and Grain-128AEAD, respectively, with corresponding throughputs of 220.67, 179.74, and 126.70 Mbit/(s·Hz). Compared with purely software-based implementations, the proposed approach substantially reduces both the number of executed instructions and the number of clock cycles (Table 5). Synthesis results further demonstrate that the coprocessor occupies 7 252.55 μm² in a 65 nm process and achieves a throughput of 1.56 Gbit/(s·Hz) at a parallelization degree of 4. Although the reconfigurable architecture requires slightly more area than a dedicated single-algorithm implementation, it supports four Grain algorithms on a unified hardware platform and provides improved hardware resource reuse and design flexibility.  Conclusions  A hardware-software cooperative reconfigurable parallelization scheme is designed to accelerate the Grain cryptographic algorithm family in lightweight embedded systems. The scheme exploits the RISC-V custom instruction extension mechanism and combines a unified shift-register architecture, a configurable feedback tap selection network, and feedback logic tailored to a predefined set of algorithms. Four algorithms, namely Grain-80, Grain-128, Grain-128a, and Grain-128AEAD, can therefore be flexibly selected and processed in parallel on a single hardware platform. This design improves processing efficiency while maintaining design flexibility and limiting hardware resource overhead. The present work focuses on reconfigurable parallelization for the Grain algorithm family. Future research will investigate more general reconfigurable architectures for nonlinear Boolean functions by incorporating configurable units, such as LookUp Tables (LUTs) or programmable logic arrays. Such architectures may further improve compatibility with multiple stream cipher algorithms while maintaining high throughput and low hardware resource overhead.
  • loading
  • [1]
    ZHU Yufei, XING Zuocheng, XUE Jinhui, et al. Area-efficient parallel reconfigurable stream processor for symmetric cryptograph[J]. IEEE Access, 2021, 9: 28377–28392. doi: 10.1109/ACCESS.2021.3057866.
    [2]
    陈艺文. 基于Grain类算法结构的流密码设计与分析[D]. [硕士论文], 桂林电子科技大学, 2023. doi: 10.27049/d.cnki.ggldc.2023.001518.

    CHEN Yiwen. Design and analysis of stream ciphers based on grain-like algorithm structure[D]. [Master dissertation], Guilin University of Electronic Technology, 2023. doi: 10.27049/d.cnki.ggldc.2023.001518.
    [3]
    MANSOURI S S and DUBROVA E. An improved hardware implementation of the grain stream cipher[C]. The 2010 13th Euromicro Conference on Digital System Design: Architectures, Methods and Tools, Lille, France, 2010: 433–440. doi: 10.1109/DSD.2010.49.
    [4]
    LI Wei, ZENG Xiaoyang, DAI Zibin, et al. A high energy-efficient reconfigurable VLIW symmetric cryptographic processor with loop buffer structure and chain processing mechanism[J]. Chinese Journal of Electronics, 2017, 26(6): 1161–1167. doi: 10.1049/cje.2017.06.010.
    [5]
    ANANTH R, RAO P R M V, and RAMAIAH N S. An efficient Grain-80 stream cipher with unrolling features to enhance the throughput on hardware platform[J]. Indonesian Journal of Electrical Engineering and Computer Science, 2024, 33(1): 218–226. doi: 10.11591/ijeecs.v33.i1.pp218-226.
    [6]
    GILL V K, CHENNA R T, KANDULA N K, et al. High throughput efficient implementation of grain v1, lizard and plantlet stream ciphers for resource constraint devices[J]. International Journal of Computing and Digital Systems, 2022, 12(1): 1051–1061. doi: 10.12785/ijcds/120184.
    [7]
    LI Bohan, ZHANG Hailong, and LIN Dongdai. Efficient (masked) hardware implementation of grain-128AEADv2[J]. Security and Communication Networks, 2023, 2023(1): 8044164. doi: 10.1155/2023/8044164.
    [8]
    MANSOURI S S and DUBROVA E. An improved hardware implementation of the grain-128a stream cipher[C]. International Conference on Information Security and Cryptology, Seoul, Korea, 2013: 278–292. doi: 10.1007/978-3-642-37682-5_20.
    [9]
    NAN Longmei, YANG Xuan, ZENG Xiaoyang, et al. A VLIW architecture stream cryptographic processor for information security[J]. China Communications, 2019, 16(6): 185–199. doi: 10.23919/jcc.2019.06.015.
    [10]
    刘小罗, 林洪怡, 刘盼. RISC-V指令集架构及其应用综述[J]. 中国集成电路, 2025, 34(3): 16–20,49. doi: 10.3969/j.issn.1681-5289.2025.03.003.

    LIU Xiaoluo, LIN Hongyi, and LIU Pan. An overview of the RISC-V instruction set architecture and its applications[J]. China Integrated Circuit, 2025, 34(3): 16–20,49. doi: 10.3969/j.issn.1681-5289.2025.03.003.
    [11]
    李伟, 别梦妮, 陈韬, 等. RISCV密码专用处理器能效概率模型与体系结构研究[J]. 电子与信息学报, 2021, 43(6): 1541–1549. doi: 10.11999/JEIT210004.

    LI Wei, BIE Mengni, CHEN Tao, et al. Research on energy efficiency probability model and architecture of RISCV cryptographic processor[J]. Journal of Electronics & Information Technology, 2021, 43(6): 1541–1549. doi: 10.11999/JEIT210004.
    [12]
    于斌, 闵玉新, 张自豪, 等. 基于RISC-V指令扩展的双线性对协处理器设计[J]. 电子与信息学报, 2025, 47(9): 3137–3145. doi: 10.11999/JEIT250367.

    YU Bin, MIN Yuxin, ZHANG Zihao, et al. Design of a bilinear pairing coprocessor based on RISC-V instruction extension[J]. Journal of Electronics & Information Technology, 2025, 47(9): 3137–3145. doi: 10.11999/JEIT250367.
    [13]
    王明登, 严迎建, 郭朋飞, 等. 基于RISC-V指令扩展方式的国密算法SM2、SM3和SM4的高效实现[J]. 电子学报, 2024, 52(8): 2850–2865. doi: 10.12263/DZXB.20230391.

    WANG Mingdeng, YAN Yingjian, GUO Pengfei, et al. Efficient implementation of national security algorithms SM2, SM3, and SM4 based on RISC-V instruction extension method[J]. Acta Electronica Sinica, 2024, 52(8): 2850–2865. doi: 10.12263/DZXB.20230391.
    [14]
    李伟, 陈億, 陈韬, 等. 面向边缘计算的可重构CNN协处理器研究与设计[J]. 电子与信息学报, 2024, 46(4): 1499–1512. doi: 10.11999/JEIT230509.

    LI Wei, CHEN Yi, CHEN Tao, et al. A research and design of reconfigurable CNN co-processor for edge computing[J]. Journal of Electronics & Information Technology, 2024, 46(4): 1499–1512. doi: 10.11999/JEIT230509.
    [15]
    HELL M, JOHANSSON T, MAXIMOV A, et al. Grain-128AEADv2: Strengthening the initialization against key reconstruction[C]. 20th International Conference on Cryptology and Network Security, Vienna, Austria, 2021: 24–41. doi: 10.1007/978-3-030-92548-2_2.
    [16]
    陈韬, 赵旺鹏, 别梦妮, 等. 格基后量子密码双域可重构多项式乘法运算单元架构研究[J]. 电子与信息学报, 2026, 48(4): 1646–1658. doi: 10.11999/JEIT250929.

    CHEN Tao, ZHAO Wangpeng, BIE Mengni, et al. Research on the architecture of dual-field reconfigurable polynomial multiplication unit for lattice-based post-quantum cryptography[J]. Journal of Electronics & Information Technology, 2026, 48(4): 1646–1658. doi: 10.11999/JEIT250929.
    [17]
    李伟. 面向序列密码的反馈移位寄存器可重构并行化设计技术研究[D]. [硕士论文], 解放军信息工程大学, 2009.

    LI Wei. Research on technology of reconfigurable parallel feedback shift register targeted at stream ciphers[D]. [Master dissertation], PLA Information Engineering University, 2009.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(9)  / Tables(6)

    Article Metrics

    Article views (302) PDF downloads(13) Cited by()
    Proportional views
    Related

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return