Statistical Intelligence for Detecting AI Model Aging and Concept Drift in Adaptive Cybersecurity Systems
Ashish Kumar Routray
Department of Statistics, Ravenshaw University, Cuttack, India.
Thambe Sai Pavan *
Department of Data Science, Lakshya Institute of Technology, Bhubaneswar, India.
Nirmal Kumar Behera
Department of Computer Science, Lakshya Institute of Technology, Bhubaneswar, India.
Abhibhav Mohanty
Department of Data Science, Lakshya Institute of Technology, Bhubaneswar, India.
Soumyajeet Chakra
Department of Data Science, Lakshya Institute of Technology, Bhubaneswar, India.
Anshu Sharma
Department of Data Science, Lakshya Institute of Technology, Bhubaneswar, India.
*Author to whom correspondence should be addressed.
Abstract
Machine-learning intrusion detectors are typically validated once at deployment and then trusted indefinitely. This is unsafe because the traffic distribution used for training differs from what deployed detectors encounter, with the gap widening as adversaries introduce unseen techniques. We formalise this degradation as model aging and develop a statistical decision framework that identifies, without ground-truth labels, when a cybersecurity model becomes unreliable and when retraining is economically justified.
The framework combines the Population Stability Index, Kullback–Leibler and Jensen–Shannon divergences, Wasserstein-1 distance, the two-sample Kolmogorov–Smirnov statistic, and sequential ADWIN, Page–Hinkley, and CUSUM procedures with two bounded state variables: the Model Aging Index (MAI), which accumulates null-calibrated distributional stress under exponential forgetting, and the Cybersecurity Drift Index (CDI), which incorporates the asymmetric costs of missed intrusions and false alerts. We prove that both indices are bounded, monotone, and Lipschitz, while their stationary exceedance probability admits an explicit bound for certified false-alarm control. The underlying marginal transport statistic also lower-bounds the joint Wasserstein radius, establishing alarm soundness.
Retraining is formulated as an optimal-stopping problem, yielding a control-limit policy analogous to an inventory replenishment threshold. Experiments on NSL-KDD introduce 17 genuinely novel attack types across 100 batches, five drift regimes, and seven architectures. Novel-family drift reduces attack recall from 0.99 to 0.57–0.66. The CDI controller cuts expected cost by 44.3% versus no retraining and 17.3% versus periodic retraining, while closing 79% of the oracle gap. Notably, periodic retraining can be harmful under adversarial drift.
Keywords: Concept drift, model aging, intrusion detection, cybersecurity analytics, Wasserstein distance, statistical process control, distribution shift, optimal stopping, adaptive retraining, adversarial evasion