面向大规模MIMO信号检测的深度学习模型混合精度量化方法

Mixed-Precision Quantization Method for Deep Learning Models in Massive MIMO Signal Detection

  • 摘要: 在大规模多输入多输出(Multiple-Input Multiple-Output,MIMO)系统的信号检测中,为使深度学习模型在资源受限的硬件平台上部署,通常需要优化其存储开销、计算复杂度和能耗。模型量化是提高部署效率的一种可行途径。本文提出一种基于量化单元敏感度的混合精度模型量化方法。首先,根据深度学习模型中各层的位置与功能对其进行量化单元划分,将具有相似结构特征的层归为同一量化单元,把层级比特分配问题转化为量化单元级分配问题以缩减搜索空间;其次,设计融合检测性能与能耗开销的量化单元敏感度指标,依据敏感度差异进行比特分配,在有限位宽范围内搜索量化配置;最后,建立包含计算、权重传输及激活值传输在内的能耗模型,用于敏感度指标计算和量化策略评估。仿真结果表明,与全精度深度学习模型、固定精度量化方案和传统线性检测算法相比,所提方法的检测性能接近全精度模型,优于相同总位宽的固定精度量化和传统线性检测算法,同时其能耗相较全精度模型显著降低。

     

    Abstract: In the signal detection of massive multiple-input multiple-output (MIMO) systems, deploying deep learning models on resource-constrained hardware typically requires optimizing storage overhead, computational complexity, and energy consumption. Model quantization provides a feasible approach to improve deployment efficiency. This paper proposes a mixed-precision model quantization method based on quantization-unit sensitivity. First, according to the position and functional role of each layer in the deep learning model, we partitioned the network into quantization units by grouping layers with similar structural characteristics; then, we transformed the hierarchical bit-allocation problem into a quantization-unit-level allocation problem to reduce the search space. Second, we designed a quantization-unit sensitivity metric that jointly accounts for detection performance and energy overhead and assigned bit widths based on sensitivity differences; within a limited bit-width range, we searched for an appropriate quantization configuration. Meanwhile, an energy-consumption model was constructed to include computation, weight transmission, and activation-value transmission, which is used to compute the sensitivity metric and evaluate quantization strategies. Simulation results showed that, compared with the full-precision model, fixed-precision quantization schemes, and conventional linear detection algorithms, the proposed method achieved a detection performance close to that of the full-precision model and outperformed fixed-precision quantization under the same total bit-width constraint and conventional linear detection. In addition, its energy consumption was significantly lower than that of the full-precision model.

     

/

返回文章
返回