高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

QuadPipe:面向FPGA大模型推理的低误差高吞吐Softmax加速器

余裕鑫 周海洋 郭涛 王硕

余裕鑫, 周海洋, 郭涛, 王硕. QuadPipe:面向FPGA大模型推理的低误差高吞吐Softmax加速器[J]. 电子与信息学报. doi: 10.11999/JEIT260610
引用本文: 余裕鑫, 周海洋, 郭涛, 王硕. QuadPipe:面向FPGA大模型推理的低误差高吞吐Softmax加速器[J]. 电子与信息学报. doi: 10.11999/JEIT260610
YU Yuxin, ZHOU Haiyang, GUO Tao, WANG Shuo. QuadPipe: A Low-Error and High-Throughput Softmax Accelerator for Large Model Inference on FPGA[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260610
Citation: YU Yuxin, ZHOU Haiyang, GUO Tao, WANG Shuo. QuadPipe: A Low-Error and High-Throughput Softmax Accelerator for Large Model Inference on FPGA[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260610

QuadPipe:面向FPGA大模型推理的低误差高吞吐Softmax加速器

doi: 10.11999/JEIT260610 cstr: 32379.14.JEIT260610
详细信息
    作者简介:

    余裕鑫:男,硕士,研究方向为神经网络加速、计算机体系结构、大模型辅助设计,邮箱 15910878756@163.com

    周海洋:男,博士,研究方向为计算机体系结构、复杂嵌入式系统设计、人工智能计算,邮箱 zhouhy@mxtronics.com

    郭涛:男,博士生,研究方向为嵌入式软件

    王硕:男,博士,研究方向为可重构芯片架构、电子设计自动化

    通讯作者:

    周海洋 zhouhy@mxtronics.com

  • 中图分类号: TP301

QuadPipe: A Low-Error and High-Throughput Softmax Accelerator for Large Model Inference on FPGA

  • 摘要: 随着大模型推理在云端与边缘场景的规模化部署,FPGA凭借可重构计算、高资源效率与低推理时延等优势,已成为补充GPU的异构加速方案。然而,在FPGA上实现高效Softmax仍面临三重瓶颈:长序列分母累加引发舍入误差积累与数值漂移,固定并行度流水线在变长向量尾拍对齐时利用率显著下降,以及访存带宽与计算密度失衡制约系统吞吐。针对上述问题,本文提出QuadPipe Softmax加速器,以三项设计协同应对:(1)Q48.47格式96位扩展精度累加器抑制长序列分母误差传播;(2)基于负无穷填充的硬件级边界无关处理,消除变长向量的软件预补齐开销;(3)X–Y双维批处理命令调度与多级反压闭环,统一计算核心与AXI接口的吞吐节拍。在Xilinx Kintex-7 325T开发板上,QuadPipe以仅占器件总逻辑4.72%的LUT与8个DSP的轻量级资源占用,实现了59周期流水线深度、每周期4个结果的稳态吞吐与150 MHz下600 M elements/s的峰值吞吐。板级实测MAE维持在10-11–10-10量级,Q48.47扩展累加使长序列分母误差较标准FP32累加降低87.59%;硬件边界无关处理使非对齐向量有效吞吐较软件预补齐最高提升6.49%。该设计在统一流水线框架内协同应对了长序列数值稳定性、变长输入处理与系统吞吐匹配三项工程挑战,为FPGA平台上大模型推理的Softmax加速提供了一套高效的系统方案。(QuadPipe开源地址:https://github.com/Yuxin-Yu/QuadPipe.git)
  • 图  3  计算核心流水线阶段与时延分配示意图

    图  1  顶层分层与模块职责示意图

    图  2  两遍式数据流与存储复用示意图

    图  4  反压链路与控制信号传播示意图

    表  1  FPGA Softmax 加速器横向对比

    工作年份目标器件精度(w)并行度nfmax / MHzLUTFFFOMDSPBRAM方法特征来源
    Zhu2020Zc706定点可调(16)15003954988.96*10600精度可调分段[11]
    Gao2020KC705定点16115422292241.00*106//Taylor 近似[4]
    Koca2023ZCU102定点1684769093334.91*1070/adder/shifter 近似[15]
    HyFT2024Xc7z030FP168645117213613.26*107//混合数值格式[12]
    DB-Attn2025Alveo U280DBFP(m8,e5)10244551087237435.10*108//动态块浮点+DH-LUT[28]
    QuadPipe2026Kintex-7 325TINT8→FP324150945574531.14*106814两遍式 + Q48.47 累加/
    下载: 导出CSV

    表  2  不同累加路径在长序列上相对 FP64 参考的平均绝对误差(MAE)

    向量长度N RTL MAE均值 RTL MAE最大值 FP32 MAE均值 FP32 MAE最大值 acc96 proxy均值 acc96 proxy最大值 RTL是否通过10–6
    128 $ 2.92\times {10}^{-10} $ $ 3.56\times {10}^{-10} $ $ 5.94\times {10}^{-10} $ $ 1.34\times {10}^{-9} $ $ 8.33\times {10}^{-11} $ $ 1.00\times {10}^{-10} $
    512 $ 1.14\times {10}^{-10} $ $ 1.56\times {10}^{-10} $ $ 2.18\times {10}^{-10} $ $ 3.49\times {10}^{-10} $ $ 2.58\times {10}^{-11} $ $ 2.94\times {10}^{-11} $
    1024 $ 4.32\times {10}^{-11} $ $ 5.61\times {10}^{-11} $ $ 2.93\times {10}^{-10} $ $ 5.57\times {10}^{-10} $ $ 1.08\times {10}^{-11} $ $ 1.10\times {10}^{-11} $
    2048 $ 2.47\times {10}^{-11} $ $ 3.34\times {10}^{-11} $ $ 1.99\times {10}^{-10} $ $ 3.63\times {10}^{-10} $ $ 6.49\times {10}^{-12} $ $ 7.02\times {10}^{-12} $
    下载: 导出CSV

    表  3  变长序列场景下 hw_auto_pad 与 sw_pre_pad 的系统性能对比(

    长度N 模式 启动长度 额外填充元素 批处理时延/ns 注入间隙/ns 有效元素吞吐/elem·ns–1 对齐后吞吐/elem·ns–1
    127 hw_auto_pad 127 0 35240198 2264013 5.80×10–5 5.80×10–5
    127 sw_pre_pad 128 1 33452189 2632016 6.10×10–5 6.10×10–5
    511 hw_auto_pad 511 0 35402200 4038026 2.31×10–4 2.31×10–4
    511 sw_pre_pad 512 1 37452211 5684035 2.18×10–4 2.19×10–4
    1025 hw_auto_pad 1025 0 40040225 7226039 4.10×10–4 4.10×10–4
    1025 sw_pre_pad 1028 3 42588240 7170043 3.85×10–4 3.86×10–4
    2049 hw_auto_pad 2049 0 49498279 14624083 6.62×10–4 6.62×10–4
    2049 sw_pre_pad 2052 3 49896282 14902086 6.57×10–4 6.58×10–4
    下载: 导出CSV
  • [1] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[C]. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, USA, 2017: 5998–6008.
    [2] YUAN Bo. Efficient hardware architecture of softmax layer in deep neural network[C]. Proceedings of the 29th IEEE International System-on-Chip Conference, Seattle, USA, 2016: 323–326. doi: 10.1109/SOCC.2016.7905501.
    [3] WANG Meiqi, LU Siyuan, ZHU Danyang, et al. A high-speed and low-complexity architecture for softmax function in deep learning[C]. Proceedings of 2018 IEEE Asia Pacific Conference on Circuits and Systems, Chengdu, China, 2018: 223–226. doi: 10.1109/APCCAS.2018.8605654.
    [4] GAO Yue, LIU Weiqiang, and LOMBARDI F. Design and implementation of an approximate softmax layer for deep neural networks[C]. Proceedings of 2020 IEEE International Symposium on Circuits and Systems, Seville, Spain, 2020: 1–5. doi: 10.1109/ISCAS45731.2020.9180870.
    [5] KIM J, KIM S, CHOI K, et al. Hardware-efficient softmax architecture with bit-wise exponentiation and reciprocal calculation[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2024, 71(10): 4574–4585. doi: 10.1109/TCSI.2024.3443270.
    [6] TARASOV I and POTEKHIN D. Calculation of activation functions in FPGA-based neuroprocessors using the CORDIC algorithm[C]. Proceedings of the 11th International Conference on High-Performance Computing Systems and Technologies in Scientific Research, Automation of Control and Production, Barnaul, Russia, 2022: 16–30. doi: 10.1007/978-3-030-94141-3_2.
    [7] SUDHA J, HANUMANTHARAJU M C, VENKATESWARULU V, et al. A novel method for computing exponential function using CORDIC algorithm[J]. Procedia Engineering, 2012, 30: 519–528. doi: 10.1016/j.proeng.2012.01.893.
    [8] STEVENS J R, VENKATESAN R, DAI S, et al. Softermax: Hardware/software co-design of an efficient softmax for transformers[C]. Proceedings of the 58th ACM/IEEE Design Automation Conference, San Francisco, USA, 2021: 469–474. doi: 10.1109/DAC18074.2021.9586134.
    [9] ZHANG Huan, ZHANG Yonggang, PENG Lele, et al. Base-2 softmax function: Suitability for training and efficient hardware implementation[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2022, 69(9): 3605–3618. doi: 10.1109/TCSI.2022.3175534.
    [10] SPAGNOLO F, PERRI S, and CORSONELLO P. Aggressive approximation of the softmax function for power-efficient hardware implementations[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2022, 69(3): 1652–1656. doi: 10.1109/TCSII.2021.3120495.
    [11] ZHU Danyang, LU Siyuan, WANG Meiqi, et al. Efficient precision-adjustable architecture for softmax function in deep learning[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2020, 67(12): 3382–3386. doi: 10.1109/TCSII.2020.3002564.
    [12] XIA Tianhua and ZHANG Saiqian. Hyft: A reconfigurable softmax accelerator with hybrid numeric format for both training and inference[C]. Proceedings of the 29th ACM/IEEE International Symposium on Low Power Electronics and Design, Newport Beach, USA, 2024. doi: 10.1145/3665314.3670816.
    [13] WU Yuanchen, XIE Zhiheng, PAN Hongbing, et al. MBS: A high-precision approximation method for softmax and efficient hardware implementation[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2025, 72(7): 3366–3375. doi: 10.1109/TCSI.2025.3559069.
    [14] CUI Zhanhao, MA Yuhao, CHENG Yuan, et al. TEA-SPS: A tiny and efficient architecture for softmax with parallelism and sparsity adaptability[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2026, 73(2): 1245–1258. doi: 10.1109/TCSI.2025.3610210.
    [15] KOCA N A, DO A T, and CHANG C H. Hardware-efficient softmax approximation for self-attention networks[C]. Proceedings of 2023 IEEE International Symposium on Circuits and Systems, Monterey, USA, 2023. doi: 10.1109/ISCAS46773.2023.10181465.
    [16] DAO T, FU D Y, ERMON S, et al. FLASHATTENTION: Fast and memory-efficient exact attention with IO-awareness[C]. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, USA, 2022: 1189.
    [17] YE Wenhua, ZHOU Xu, ZHOU J, et al. Accelerating attention mechanism on FPGAs based on efficient reconfigurable systolic array[J]. ACM Transactions on Embedded Computing Systems, 2023, 22(6): 93. doi: 10.1145/3549937.
    [18] KABIR E, KABIR A, DOWNEY A R J, et al. FAMOUS: Flexible accelerator for the attention mechanism of transformer on Ultrascale+ FPGAs[C]. Proceedings of 2024 International Conference on Field Programmable Technology, Sydney, Australia, 2024. doi: 10.1109/ICFPT64416.2024.11113430.
    [19] ISLAMOGLU G, SCHERER M, PAULIN G, et al. ITA: An energy-efficient attention and softmax accelerator for quantized transformers[C]. Proceedings of 2023 IEEE/ACM International Symposium on Low Power Electronics and Design, Vienna, Austria, 2023. doi: 10.1109/ISLPED58423.2023.10244348.
    [20] LI Tianyang, ZHANG Fan, FAN Xitian, et al. Unified accelerator for attention and convolution in inference based on FPGA[C]. Proceedings of 2023 IEEE International Symposium on Circuits and Systems, Monterey, USA, 2023. doi: 10.1109/ISCAS46773.2023.10182145.
    [21] 李天阳, 张帆, 王松, 等. 基于FPGA的卷积神经网络和视觉Transformer通用加速器[J]. 电子与信息学报, 2024, 46(6): 2663–2672. doi: 10.11999/JEIT230713.

    LI Tianyang, ZHANG Fan, WANG Song, et al. FPGA-based unified accelerator for convolutional neural network and vision Transformer[J]. Journal of Electronics & Information Technology, 2024, 46(6): 2663–2672. doi: 10.11999/JEIT230713.
    [22] 姜小波, 邓晗珂, 莫志杰, 等. 规则压缩模型和灵活架构的Transformer加速器设计[J]. 电子与信息学报, 2024, 46(3): 1079–1088. doi: 10.11999/JEIT230188.

    JIANG Xiaobo, DENG Hanke, MO Zhijie, et al. Design of Transformer accelerator with regular compression model and flexible architecture[J]. Journal of Electronics & Information Technology, 2024, 46(3): 1079–1088. doi: 10.11999/JEIT230188.
    [23] ZENG Shulin, LIU Jun, DAI Guohao, et al. FlightLLM: Efficient large language model inference with a complete mapping flow on FPGAs[C]. Proceedings of 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays, Monterey, USA, 2024. doi: 10.1145/3626202.3637562.
    [24] CHEN Hongzheng, ZHANG Jiahao, DU Yixiao, et al. Understanding the potential of FPGA-based spatial acceleration for large language model inference[J]. ACM Transactions on Reconfigurable Technology and Systems, 2025, 18(1): 5. doi: 10.1145/3656177.
    [25] LI Cong, YIN Yihan, WU Xintong, et al. H2-LLM: Hardware-dataflow co-exploration for heterogeneous hybrid-bonding-based low-batch LLM inference[C]. Proceedings of the 52nd Annual International Symposium on Computer Architecture, Tokyo, Japan, 2025. doi: 10.1145/3695053.3731008.
    [26] 王泽昊, 朱振华, 谢童欣, 等. 混合专家大语言模型的系统与架构优化技术综述[J]. 电子与信息学报, 2025, 47(11): 4055–4078. doi: 10.11999/JEIT250407.

    WANG Zehao, ZHU Zhenhua, XIE Tongxin, et al. A survey on system and architecture optimization techniques for mixture-of-experts large language models[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4055–4078. doi: 10.11999/JEIT250407.
    [27] 林志平, 肖亮, 陈宏毅, 等. 面向大语言模型的抗干扰协同推断技术[J]. 电子与信息学报, 2025, 47(11): 4572–4582. doi: 10.11999/JEIT250675.

    LIN Zhiping, XIAO Liang, CHEN Hongyi, et al. Collaborative inference for large language models against jamming attacks[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4572–4582. doi: 10.11999/JEIT250675.
    [28] WANG Hui, CHENG Yuan, HAN Xiaomeng, et al. Pushing the limits of BFP on narrow precision LLM inference[C]. Proceedings of the 39th AAAI Conference on Artificial Intelligence, Philadelphia, USA, 2025: 21099–21107. doi: 10.1609/aaai.v39i20.35407.
  • 加载中
图(4) / 表(3)
计量
  • 文章访问数:  19
  • HTML全文浏览量:  4
  • PDF下载量:  1
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-05-14
  • 修回日期:  2026-09-14
  • 录用日期:  2026-09-14
  • 网络出版日期:  2026-09-21

目录

    /

    返回文章
    返回