Abstract:
With the rapid evolution and widespread deployment of large language models (LLMs), the risks they introduce have become increasingly complex, rendering traditional artificial intelligence safety assessment approaches insufficient. Existing studies on LLM risk evaluation tend to focus on isolated risk factors or static indicators and often overlook the dynamic relationships between a model's intrinsic capabilities and systemic risks. Moreover, current assessment procedures usually rely on expert judgment with direct weighting, without adequately correcting for subjective bias among experts. To address these limitations, this paper proposes a multidimensional, hierarchical, and dynamic risk assessment framework grounded in the objective capabilities of LLMs. The framework first constructs a comprehensive risk indicator system across four layers—model, individual, organization, and environment—covering 13 risk dimensions in total. Building on a structured capability evaluation scheme for LLMs, we further design a capability–risk dynamic association matrix to characterize the nonlinear mappings and threshold effects between model capabilities and risk manifestation, enabling quantitative tracing from capability signals to risk outcomes. For risk dimension weighting, we improve the conventional Analytic Hierarchy Process (AHP) by evaluating both the distributional consistency of expert judgments and inter-dimensional agreement, thereby dynamically calibrating expert influence and producing a more objective and scientifically grounded risk weight matrix. In a case study on risk prediction within the legal domain, the proposed framework successfully identified the high-risk profile of the involved model under a scenario characterized by "high factual accuracy requirements + high legal liability constraints + low fault-tolerant decision chains." It uncovered a cross-level propagation chain—"insufficient reliability capability → elevated algorithmic risk → misleading professional judgment → operational risk → manifestation of legal risk"—and elucidated its underlying mechanism. At the methodological level, the constructed risk quantification mapping algorithm enabled automated and explainable computation from objective capabilities to an overall risk score. The improved dynamic weighted AHP enhanced the scientific rigor and accuracy of the risk weight matrix. Experimental and case-based validations demonstrate that the proposed framework can accurately assess composite risks triggered by LLM capabilities. More importantly, through systematic analysis of potential risk propagation processes, it provides a practical and quantitative decision-support tool for risk prediction, traceability-based governance, and proactive prevention in large-scale systems.