This paper addresses the challenge of operating multi‑UAV networks in threat‑prone environments by proposing a deployment framework that jointly maximizes global energy efficiency (EE) and ensures safety. The framework proceeds in three steps:
In the first step, a threat‑aware K‑means (TAKM) algorithm determines the minimum number of UAVs required and computes safe initial cluster centroids based on the spatial threat distribution.
The second step performs optimal matching, assigning physical UAVs to the centroids so that flight and communication energy consumption are minimized.
The third step introduces a threat‑aware multi‑agent twin delayed deep deterministic policy gradient (MATD3) algorithm, which simultaneously optimizes UAV trajectories, transmit power, and user association. Safety constraints are embedded in the reward function, enabling the agents to learn policies that avoid threat zones autonomously.
Simulation results show that the proposed framework achieves zero safety violations while delivering superior EE and faster convergence compared with other learning baselines and non‑clustering approaches. Compared with heuristic methods, it outperforms greedy particle swarm optimization (GPSO) and matches the performance of optimized PSO (OPSO) with far lower online computational complexity. Additional experiments demonstrate strong generalization to unseen user distributions, larger UAV fleets, and varied threat geometries, all while maintaining zero violations.
Review: By tightly integrating threat‑aware clustering with multi‑agent reinforcement learning, the study offers a practical, energy‑optimal, and safety‑guaranteed solution for dynamic UAV network deployment in hostile settings.