Attackers Strike Back? Not Anymore -- An Ensemble of RL Defenders Awakens for APT Detection

  • 2025-08-26 14:29:10
  • Sidahmed Benabderrahmane, Talal Rahwan
  • 0

Abstract

Advanced Persistent Threats (APTs) represent a growing menace to moderndigital infrastructure. Unlike traditional cyberattacks, APTs are stealthy,adaptive, and long-lasting, often bypassing signature-based detection systems.This paper introduces a novel framework for APT detection that unites deeplearning, reinforcement learning (RL), and active learning into a cohesive,adaptive defense system. Our system combines auto-encoders for latentbehavioral encoding with a multi-agent ensemble of RL-based defenders, eachtrained to distinguish between benign and malicious process behaviors. Weidentify a critical challenge in existing detection systems: their staticnature and inability to adapt to evolving attack strategies. To this end, ourarchitecture includes multiple RL agents (Q-Learning, PPO, DQN, adversarialdefenders), each analyzing latent vectors generated by an auto-encoder. Whenany agent is uncertain about its decision, the system triggers an activelearning loop to simulate expert feedback, thus refining decision boundaries.An ensemble voting mechanism, weighted by each agent's performance, ensuresrobust final predictions.

 

Quick Read (beta)

loading the full paper ...