一种基于隐状态空间激活的免训练安全检测方法
A Training-Free Security Detection Method Based on Hidden State Space Activation
-
摘要: 大语言模型(Large Language Models,LLMs)在重塑人机交互范式的同时,正面临日益严峻的越狱攻击威胁。现有的防御方案多依赖于计算成本高昂的安全微调或泛化性受限的外部拦截机制,难以深度表征模型内部隐状态空间的几何演化特性。针对上述挑战,本文提出一种基于隐状态空间激活的免训练安全检测方法(Training-free Security Detection based on Hidden State Space Activation,TSSA)。首先,TSSA揭示了安全与风险查询在模型深层激活流形中遵循显著可分离的统计分布规律,为构建内生防御机制提供了理论支撑;其次,通过关键层段自适应选择机制,精准定位对安全语义最具判别力的神经元层级,并进一步开发安全方向向量校准算法,利用马氏空间度量消除特征空间的各向异性,从而提取最优判别轴;最后,引入自适应阈值校准策略,通过量化查询冲突强度动态确立判定边界,实现对深度语义伪装攻击的精准拦截。在DeepSeek-R1、Qwen2.5等主流模型上的实验结果表明,TSSA在不引入显著推理时延的前提下,显著提升了防御成功率与跨模型泛化能力,为理解大模型内部安全边界及构建轻量化实时防御体系提供了新视角。Abstract: Large Language Models (LLMs) are transforming the paradigm of human-computer interaction, but these systems are also facing increasingly serious and evolving threats from jailbreak attacks. Existing defense solutions predominantly rely on computationally expensive security-oriented fine-tuning strategies or external interception mechanisms with limited generalization capability, making it difficult to thoroughly and systematically characterize the intrinsic geometric evolution patterns within the model’s internal hidden state space. To address these limitations, this study proposes a novel, training-free security detection method based on Hidden State Space Activation (TSSA). First, TSSA demonstrates that security and risk queries follow a highly separable statistical distribution pattern in the model’s deep activation manifold, providing a theoretical foundation for constructing an intrinsic security defense mechanism. Second, through a key-segment-aware adaptive selection strategy, the most discriminative neural layer for security semantics is precisely identified. Furthermore, a security direction vector calibration algorithm is developed, utilizing a Mahalanobis distance-based metric to eliminate feature-space anisotropy, thereby extracting the optimal discriminative direction. Finally, an adaptive threshold calibration strategy is introduced to dynamically establish the decision boundary by quantitatively measuring the query conflict intensity, achieving robust and precise detection of deep semantic spoofing attacks. Experimental results on mainstream models such as DeepSeek-R1 and Qwen2.5 demonstrate that TSSA substantially enhances the defense success rate and cross-model generalization ability without introducing significant inference latency, providing a new perspective for understanding the internal security boundary of large models and enabling the development of a lightweight, real-time defense framework.
下载: